Blog
Which LLM Should Your Business Actually Pay For?
Sam Carr

There isn't one. For most UK small businesses the sensible default is ChatGPT, which Layer3Labs calls "the best default choice for many small businesses" because it is flexible, familiar and useful across a lot of teams; Claude suits careful document and writing work, Gemini suits Google Workspace shops, and Copilot wins if your team lives in Word, Excel, Outlook and Teams. Costs are close enough that they rarely decide it: ChatGPT Business is $25 per user per month billed monthly or $20 annually, and Claude Pro is $20 a month monthly or $17 a month billed annually. What should decide it is your own test on your own work, plus the privacy setting: OpenAI says that by default it does not use data from ChatGPT Business, Enterprise, Edu or its API platform for training, while on Plus, Pro or Free in a personal workspace data sharing is on by default.
An LLM is a text tool that drafts, summarises and rewrites, and that is most of what it does for a business
An LLM, or large language model, is a text tool. You give it some text and it gives you text back. That is the whole shape of it.
Underneath, the mechanism is simpler than the marketing suggests. CSET's explainer describes it plainly: LLMs work by receiving an input or prompt, calculating what is most likely to come next, and then producing an output or completion. There is no lookup of a fact and no reasoning about your business. There is a very good guess at what text usually follows text like yours.
One small correction to the usual description. The model does not predict the next word. As Towards Data Science puts it, "It does not predict the next word, ChatGPT predicts the next token." A token is a chunk of text, often a word or part of a word. This matters mainly because most vendors charge you by the token, so tokens are the unit your bill is measured in.
The guessing gets good because the models are large. OpenAI's GPT-2, released in 2019, had 1.6 billion parameters and was trained on 100 billion tokens, and models since then have grown well beyond that. Parameters are the internal settings the model learns during training.
So what does this buy a business? Anything that is text in, text out. Drafting a first version of a quote, an email or a job advert. Summarising a long thread or a set of notes into something a colleague can read in a minute. Rewriting a rambling paragraph into plain English. Turning messy free text into a tidy list.
What it is not is a database, a calculator or a member of staff. Because the output is a prediction of likely text, it can produce something that reads perfectly and is wrong. Every LLM task in a business needs a human who checks the output before it goes anywhere.
In practice the choice is between four names: ChatGPT, Claude, Gemini and Copilot
Ignore the leaderboards for a moment. For a UK small business without a technical team, the realistic shortlist is four products: ChatGPT, Claude, Gemini and Microsoft Copilot.
The simplest way to tell them apart is by the job each one is built around. Layer3Labs sums it up neatly: "ChatGPT is the broad assistant. Claude is the careful document and writing assistant. Gemini is the Google Workspace assistant."
ChatGPT is the general purpose one. Layer3Labs calls it "the best default choice for many small businesses", on the grounds that it is flexible, familiar and useful across a lot of teams. If you have no strong reason to pick something else, this is the safe starting point.
Claude is the one to reach for when the output is a document. Drafting, editing, long written work: that is where it is positioned.
Gemini is the Google option. If your business runs on Gmail, Docs and Drive, it sits inside the tools your staff already open every morning.
Microsoft Copilot is the same argument in reverse. If you are already on Microsoft 365, Dreams AI notes that Copilot is built into Word, Excel, Outlook and Teams, and calls the integration unmatched. Worth knowing that Copilot is no longer tied to a single model behind the scenes. Alphabold reports that Wave 3 makes Microsoft 365 Copilot model-diverse by design, with users in the Frontier program able to select Anthropic's Claude models directly within it.
So the practical question is rarely "which model is cleverest". It is which of these four already lives in the software your team uses all day.
There is no single best LLM, the winner changes with the task
No. The honest answer is that the best model depends on what you are asking it to do, and on what you are willing to pay per task.
The pricing spread alone makes a single winner unlikely. One 2026 benchmarking round-up of more than 30 models found prices ranging from $0.05 to $50 per million tokens, and concluded that the practical comparison is capability per dollar for your task tier rather than a single leaderboard rank. A token is roughly three quarters of a word, so a million tokens is a lot of text, but the gap between the cheap end and the expensive end is a factor of a thousand. Paying top rates to summarise routine emails is money burnt.
Head to head tests show the same thing on quality. In AI Magicx's 2026 comparison, Claude 4 gave the deepest analysis on a data task, spotting a correlation between particular product categories and refund rates that the other models missed. Gemini also performed well on that same data task, and the reviewers singled out scalability as its strength: their test used 500 rows, but Gemini's context window handles far more. Context window means how much text the model can hold in mind at once.
So the sensible way to choose is by job, not by brand. Buyer comparisons in 2026 tend to lay out OpenAI, Anthropic, Google, Meta's open-weight models and Mistral side by side on strengths and tradeoffs, precisely because no column wins every row. Work out your two or three most common tasks first. Then test the shortlist on those, with your own data.
Expect roughly £17 to £25 per user per month, and here is what each tier buys
Most business plans are priced the same way: per user, per month. OpenAI says its paid plans (Plus, Pro, Business and Enterprise) are priced per user per month, with monthly billing for Plus, Pro and Business and annual options as well. Anthropic prices Claude the same way. So the sum you care about is headcount multiplied by a fairly small monthly figure, not a big one-off licence.
The published numbers sit in a narrow band. ChatGPT Business is $25 per user per month billed monthly, or $20 per user per month if you pay annually, for most countries. Claude Pro is $20 a month billed monthly, or $17 a month billed annually, which is $200 paid up front.
| Plan | Monthly billing | Annual billing | Notes |
|---|---|---|---|
| ChatGPT Business | $25 per user per month | $20 per user per month | Price given for most countries |
| Claude Pro | $20 per month | $17 per month | Annual is $200 paid up front |
Two warnings before you budget. First, those figures are quoted in US dollars, so your actual bill in sterling depends on the exchange rate on the day. Second, pricing varies by region and excludes taxes, so add VAT on top when you model the cost.
The practical read: for a five person team, you are looking at roughly $85 to $125 a month depending on which product you pick and whether you commit for a year. Paying annually is the cheaper route on both products, but it means handing over the full year up front on Claude Pro.
The free versions are fine for trying things out, but they fall down on limits and data terms
Free tiers are a good way to find out whether an LLM helps you at all. Spend a week pasting in your usual work and see what comes back. What they are not good for is anything confidential, and that is a terms question rather than a quality one.
The default on free consumer plans is that your chats can train the model. OpenAI's help centre states that on a ChatGPT Plus, Pro or Free plan in a personal workspace, data sharing is enabled by default, though there is a setting you can change. The paid business tiers work the other way round. OpenAI says that with ChatGPT Business, Enterprise, Edu and its API platform, it does not use provided inputs and outputs to train its models by default.
Google is blunter still on the free developer tier. Its Gemini API additional terms say: "Do not submit sensitive, confidential, or personal information to the Unpaid Services." That is not a warning about quality. It is an instruction, and it rules out client records, staff data and anything you hold under a duty of confidence.
Retention matters too. Google's Gemini Apps privacy hub says reviewed feedback, the conversations attached to it and related data are kept for up to three years, disconnected from your Google account. So switching off history later does not necessarily pull back what a human reviewer has already seen.
The practical rule for a small UK business: use the free tier for drafting, summarising public documents and learning what the tool is like. The moment real customer or employee data is involved, move to a paid business plan where training is off by default, or don't put it in at all.
Yes, data location and training settings matter, and business plans handle them differently to consumer ones
The question splits in two: who can read your data, and whether your data is used to train the model. The second one has a clear published answer, and it turns on which plan you are on.
On OpenAI's side, the consumer products are the risk. OpenAI states that when you use its services for individuals, such as ChatGPT, Sora or Operator, it may use your content to train its models. Its business products are treated differently. OpenAI says that by default it does not use data from ChatGPT Enterprise, ChatGPT Business, ChatGPT Edu, ChatGPT for Healthcare, ChatGPT for Teachers or its API platform for training.
Anthropic draws the same line. Its commercial plans, covering the API, Claude for Work, Claude Team and Claude Enterprise, are not used to train its models, and this is described as a contractual guarantee rather than a setting buried in a menu.
The consumer side of Claude changed. After 8 October 2025, users have to make a choice. Leave the toggle off and your conversations are never used for training. Turn it on and they are.
What this means in practice for a small business. If your team is pasting client details, quotes, contracts or staff information into a free or personal account, assume it may be used for training unless you have actively turned that off. A paid business or team plan is not just about extra features. It is the version where not training on your work is the default rather than something an employee has to remember to switch. If you handle client data, that difference is worth the seat cost on its own.
One caveat. Training settings are only part of the picture. Where the data physically sits, and what your contract says about it, is a separate question you should put to the vendor before you commit.
Standardising on one keeps control simple, but a second tool earns its place in specific cases
Start with one. A single main tool means one bill, one set of admin controls, one place where your data goes, and one thing to train people on. That is most of the value for most small firms.
The alternative is not really "two tools". It is whatever staff pick on their own. Worklytics reports that 78% of AI users bring their own AI tools, which it calls BYOAI, and notes this skews self-reported usage numbers. In practice that means work being pasted into accounts you do not control and cannot audit. Choosing one tool and paying for it properly is the cheapest way to stop that.
Tool sprawl also creeps in quietly. Inside Consulting describes it as a symptom of decentralised decision-making, fast growth and the low friction of monthly subscriptions, rather than a procurement failure. Nobody signs off on six AI subscriptions. They just appear, one card payment at a time.
There is a switching cost too. ActivTrak found focus efficiency, the share of total work time spent in focused, uninterrupted work, fell to 60%, a three-year low. Every extra app is another tab and another context switch.
When does a second tool earn its place? Three cases. First, a genuine capability gap: your main tool cannot do something a specific team needs every week, not once a quarter. Second, a data restriction: one client or contract requires processing that your main provider's terms do not cover. Third, a hedge on a critical process: if one system going down would stop invoicing or customer replies, a tested fallback is reasonable.
Set a rule for the second tool. Name the owner, name the use case, and review it in six months. If it has not been used for the stated purpose, cancel it.
On usage, more is not automatically better. ActivTrak found employees spending 7 to 10% of their total work hours in AI tools had the highest productivity rates, at 95%, of any usage tier. That is roughly three to four hours a week in a full-time role. Aim for depth in one tool rather than shallow use of five.
Run a two week trial on your own real tasks before you buy anything
Benchmark tables tell you how a model performs in general. They do not tell you how it performs on your quotes, your customer emails or your supplier chasers. Turian's guidance on evaluating LLMs makes the same point: benchmarks show general capability, while your own metrics, such as the success rate in answering your customers' top 100 questions, tell you something specific to your business.
So run a short trial on real work before you commit to anything.
Pick tasks you already do. Take a job that eats time every week. Drafting replies to enquiries. Summarising site notes. Turning a phone call into a written quote. Use real inputs, not tidy examples you made up for the test.
Weight the trial towards the things that would hurt. Advice from projectsupply.in is to focus evaluation effort on the edge cases and high-stakes queries that pose the greatest risk to your business if the model fails. For most small firms that means anything touching price, legal wording, safety or a customer complaint. A model that handles routine work well and mangles a refund policy is not a model you can leave alone.
Test it where it will actually live. Turian distinguishes system evaluation, which looks at the model's performance once it is integrated, from testing the model in isolation. Real-world assessment exposes it to the unpredictability of actual usage. A model that reads well in a chat window can behave differently once it is sitting inside your inbox or your helpdesk.
Score it simply. Count how many outputs you could send with a light edit, how many needed a rewrite, and how many were wrong in a way you would not have caught. Do the same set of tasks on a second model. The comparison is the point.
Log usage, then price it. The AI Consultancy recommends running a one-month pilot on a Teams tier with usage logged, then using that usage profile to model the total cost of an Enterprise plan. The same logic works at any size. Two weeks of real usage tells you how many seats you need and how heavily people lean on it, which is a far better basis for a budget than a per-seat price on a website.
At the end of the trial you should be able to say which tasks the tool does well enough to keep, which need a human check every time, and what it will cost you a month. If you cannot answer those three, run it another two weeks before you sign anything.
An LLM is the wrong tool when the answer must be exact, auditable or legally accountable
An LLM predicts likely text. It does not calculate, and it does not check. So the moment your answer has to be exact, traceable to a document, or defensible to a regulator, the model stops being the right tool and becomes a risk.
Arithmetic and precise reasoning. The ARB benchmark, which tests models on advanced reasoning problems, scores failures by category: misreading the problem, taking the wrong approach, logical error or hallucination, and arithmetic mistake. Those are separate ways to get to a confident wrong answer. If the output is a VAT figure, a payroll total or a quoted price, put it in a spreadsheet or your accounting software, not a chat window.
Anything that must match the source document. Research on document-based queries found that hallucinations cluster: when a response contains any hallucination, it usually contains several, with Gemini averaging four per affected response. That matters because a plausible summary of a contract or a policy is worse than no summary. One invented clause is easy to spot. Four, wrapped in otherwise accurate text, are not.
Decisions about people. Under UK data protection rules, solely automated decisions with significant effects on individuals are now permitted for most personal data, but only subject to the Article 22C safeguards. And the ICO is clear that a Data Protection Impact Assessment, a written risk assessment you have to do before you start, is always required for systematic and extensive profiling or other automated evaluation of personal data used for decisions with legal or similar effects. So if you are thinking of letting a model screen job applicants, set credit terms or decide who gets a refund, the compliance work is the project. The model is the easy part.
A practical line to draw: use an LLM for drafting, sorting, summarising for your own eyes and first-pass triage. Keep a human in the loop wherever the output is exact, auditable or legally accountable.
Common questions
What is a token, and why does it keep coming up?
A token is a chunk of text, usually a word or part of a word. As Towards Data Science puts it, "It does not predict the next word, ChatGPT predicts the next token." It matters mainly for pay-as-you-go API pricing, where one 2026 round-up of over 30 models found prices ranging from $0.05 to $50 per million tokens.
If we already pay for Microsoft Copilot, are we locked into one model?
Not necessarily. Alphabold reports that Wave 3 makes Microsoft 365 Copilot model-diverse by design, with users in the Frontier program able to select Anthropic's Claude models directly within it. So the tool you buy and the model underneath it are increasingly separate decisions.
How much should staff actually be using these tools?
Less than you might think. ActivTrak found that employees spending 7 to 10% of their total work hours in AI tools had the highest productivity rates, at 95%, of any usage tier. More time in the tool is not automatically better.
What does a wrong answer actually look like?
Errors are not usually one small slip. Research on document-based queries found that when a response contains any hallucination it typically contains several, with Gemini averaging four per affected response. Benchmarks like ARB sort failures into categories such as misreading the problem, taking the wrong approach, logical errors and arithmetic mistakes.
Do we need a formal assessment if we use an LLM in decisions about people?
Probably yes. The ICO is clear that a Data Protection Impact Assessment is always required for systematic and extensive profiling or other automated evaluation of personal data used for decisions. Solely automated decisions with significant effects are now permitted for most personal data, but only subject to the Article 22C safeguards.
Drafted by our blog writer. Read, checked and published by Sam.
Sources
- The Surprising Power of Next Word Prediction: Large Language Models Explained, Part 1 | Center for Security and Emerging Technology
CSET (Georgetown)
- Neural Network Internals: How Next-Token Prediction Really Works | by Saipriya Damarapati | Medium
Medium (Saipriya Damarapati)
- Unleashing the ChatGPT Tokenizer | Towards Data Science
Towards Data Science
- The Surprising Power of Next Word Prediction: Large Language Models Explained, Part 1 | Center for Security and Emerging Technology
CSET (Georgetown)
- Best LLM 2026: Which One I Actually Pay For | Dreams AI
Dreams AI
- ChatGPT vs Claude vs Gemini (2026): Verdict by Business Use Case
Layer3Labs
- ChatGPT vs Claude vs Gemini (2026): Verdict by Business Use Case
Layer3Labs
- Guide to Microsoft Copilot Pricing & Licensing
Alphabold
- LLM Comparison 2026: 30+ Models Benchmarked & Ranked
Iternal.ai
- AI Model Comparison 2026: GPT-4o vs Claude 4 vs Gemini 2.0 vs Mistral — Which Should You Use? | AI Magicx Blog | AI Magicx
AI Magicx Blog
- AI Model Comparison 2026: GPT-4o vs Claude 4 vs Gemini 2.0 vs Mistral — Which Should You Use? | AI Magicx Blog | AI Magicx
AI Magicx Blog
- Best AI Models in 2026: GPT, Claude, Gemini, and More Compared | AIntelligenceHub
AIntelligenceHub
- ChatGPT Business - Overview | OpenAI Help Center
OpenAI Help Center
- Claude Team: Pricing, Features & Benefits for Remote Teams
Tactiq
- Pricing | ChatGPT
OpenAI
- Claude pricing in 2026: every plan, API rate, and what it actually costs
CloudZero
- What if I want to keep my history on but disable model training? | OpenAI Help Center
OpenAI
- Gemini API Additional Terms of Service | Google AI for Developers
Google
- Gemini Apps Privacy Hub - Gemini Apps Help
Google
- What if I want to keep my history on but disable model training? | OpenAI Help Center
OpenAI
- Business data privacy, security, and compliance | OpenAI
OpenAI
- Does Anthropic Train on My Data? Clear Answer 2026
formation-claude-ia.fr
- How your data is used to improve model performance | OpenAI
OpenAI
- Claude privacy: How Anthropic handles your data | Anonyome
Anonyome
- The Hidden Cost of Too Many Tools: A SaaS Vendor Consolidation Guide - Inside Consulting
Inside Consulting
- 2026 State of the Workplace: AI Adoption and Workforce Performance Benchmarks – ActivTrak
ActivTrak
- 2026 State of the Workplace: AI Adoption and Workforce Performance Benchmarks – ActivTrak
ActivTrak
- Benchmarking Employee AI Adoption: Closing the Gap Between Daily Usage and Leadership Estimates in 2025 | Worklytics
Worklytics
- How To Evaluate LLM Models: Approaches, Metrics and Tips | Turian Blog
Turian Blog
- How to Evaluate an LLM for a Specific Business Use Case in 2026 — The Framework - Blog
projectsupply.in
- Claude vs ChatGPT Enterprise: UK Anthropic Consulting view
The AI Consultancy
- How To Evaluate LLM Models: Approaches, Metrics and Tips | Turian Blog
Turian Blog
- Not Wrong, But Untrue: LLM Overconfidence in Document-Based Queries
arXiv (Not Wrong, But Untrue)
- Legal framework | ICO
ICO
- ARB: Advanced Reasoning Benchmark for Large Language Models
arXiv (ARB benchmark)
- Automated Decision-Making | UK GDPR and DUAA 2025
bratby.law