AI · Automation · Engineering

LLM Costs for UK Teams: API vs GPU vs Hybrid

By Lazar MilicevicOctober 8, 20268 min read
Rows of GPU server racks in a data center, illustrating LLM infrastructure costs for UK teams

The first thing I do on any LLM cost review is ignore the per-token price. The number that decides the bill is almost never what the model costs, it is how many tokens your architecture wastes, and whether a GPU would sit idle at 3 a.m. while you pay for it. If you are a UK founder or CTO comparing API pricing, self-hosted GPUs and hybrids, this is how I think about it.

One caveat up front: I am not publishing client invoices here, and I will not invent them. What follows is the cost structure I use on real deployments, with public list prices from the sources I checked on 8 October 2026. Vendor prices move constantly, so verify against the vendor's own pricing page before you commit a budget.

What the APIs cost per million tokens

List prices are cheap enough that for most workloads the API is the correct default, and the spread between model tiers matters far more than the spread between vendors. The figures below come from third-party trackers, not the vendors' own pricing pages, so treat them as a snapshot.

Model Input / MTok Output / MTok Source
Claude Haiku 4.5 $1 $5 BenchLM
Claude Sonnet 5 $2 $10 BenchLM
Claude Opus 5 $5 $25 BenchLM
GPT-5 $1.25 $10 OpenAI model page

A few details change the maths more than the headline rate:

  • Batch is half price. Anthropic's Batch API for Haiku 4.5 is $0.50 input and $2.50 output per MTok, per Maximem's tracker. OpenAI batch also halves every rate, per Morph.
  • Caching is a bigger lever than model choice. Haiku 4.5 cache reads are listed at $0.10/MTok, a tenth of the input price. GPT-5 cached input is $0.125 on OpenAI's own page.
  • Tiers within one vendor span a huge range. BenchLM puts OpenAI's range at roughly 600x between the cheapest and most expensive models (source).

The practical consequence: pick the cheapest model that passes your eval, not the best model you can afford. I treat "which tier" as an eval question, not a procurement question.

Where self-hosted GPUs actually land

A GPU only beats the API when you keep it busy, and the break-even is about utilization, not about price per hour. The AWS p5.48xlarge (8x H100) lists at $55.04/hr on-demand, roughly $6.88 per H100 GPU-hour, according to Spheron's write-up of AWS list prices. That is third-party reporting, and sources I found disagree on the exact figure, so check AWS directly. Across providers including spot and marketplace rates, CloudZero reports H100 rental from about $1.49 to $6.98 per GPU-hour.

Here is an example with round numbers, clearly illustrative and not a quote. Take a single H100 at $3 per GPU-hour, running around the clock:

  • 24 hours x 30 days = 720 hours
  • 720 x $3 = about $2,160 per month, before engineering time, storage, egress and monitoring

That $2,160 buys you a fixed amount of capacity whether you use 5% of it or 95%. The same money on a mid-tier API buys a lot of tokens, and you pay nothing for the idle hours. For a team sending a few thousand requests a day, the API wins by a wide margin. For a team with sustained, predictable, high-volume traffic, or with hard data residency requirements, the calculation can flip.

The costs people forget when they quote "$3 per GPU-hour":

  • Someone has to run it. Inference servers, autoscaling, model upgrades, GPU driver issues, on-call. In my experience this is the line item that kills the business case for small teams.
  • Quality gap. Open-weight models you can host are not the same as the frontier APIs. You need an eval harness to know whether the gap matters for your task.
  • Idle capacity. Overnight and weekend troughs. Reserved pricing helps the rate but locks in the idle time.

I run local models through Ollama for RAG pipelines where data must not leave the machine, and for development. That is a privacy and experimentation choice. I would not pretend it is a cost win at small scale.

The hybrid setup I reach for most

Route cheap, high-volume, low-risk work to a small model or local model, and send only the hard cases to a frontier model. This is not exotic. It is the pattern behind most of the production systems I build, including the multi-agent content machine behind BizFlowAI ContentStudio, where different stages have very different quality requirements.

The shape looks like this:

request
  -> cache lookup (exact + semantic)       # free-ish
  -> small model: classify / extract / draft
  -> confidence or validator check
       pass -> return
       fail -> escalate to frontier model
  -> async/offline work -> batch API (50% off)

What makes it work in production:

  1. A cheap, honest validator. Schema checks, retrieval-grounding checks, or a small judge. Without it, you cannot tell which requests deserve escalation.
  2. Batch everything that is not user-facing. Scheduled workers, nightly enrichment, backfills. Half price for tolerating latency is the easiest saving there is.
  3. Prompt caching by design. Put the stable part of the prompt (system instructions, tool definitions, retrieved reference material) first so the cache actually hits. I have seen teams lose most of the discount by interpolating a timestamp at the top of the prompt.
  4. Log cost per request, per route. If you cannot see which path costs what, you cannot tune the router.

Most of the savings I have delivered on automation work came from this kind of plumbing, not from negotiating a better rate. It is the same discipline as the analytics migration that saved EUR 30-60k a year: find the waste in the architecture first.

UK tax and finance notes for founders

Two things matter most: cloud and compute costs can qualify for R&D relief, and the corporation tax rate shapes how much a given spend really costs you. I am an engineer, not an accountant, so use this to know what to ask, then ask a qualified adviser.

  • R&D tax relief covers cloud computing and datasets. The government's reform extended qualifying expenditure to include data and cloud computing costs, with new definitions added to section 1125 of CTA 2009. Source: GOV.UK, R&D tax relief reform changes. That is relevant if your LLM spend sits inside genuine R&D work (experimental, resolving technical uncertainty). Whether your particular GPU or API spend qualifies depends on the project, so get a specialist to look at it. Do not assume routine production inference counts.
  • Corporation Tax main rate is 25% on profits over £250,000, with a 19% small profits rate under £50,000 and a taper in between, per the House of Commons Library, which confirms these from 1 April 2023 and unchanged for 2024/25. For current-year rates, check GOV.UK's "Corporation Tax rates and allowances" page rather than relying on me.
  • VAT: overseas API vendors and UK cloud providers can have different VAT treatment on invoices, and your registration status changes what you can reclaim. I did not verify current thresholds from an official source, so confirm with HMRC guidance or your accountant.

The structural point for founders: GPU spend is capital-like and committed, API spend is variable and easier to attribute to features. That difference shows up in cash flow and in how cleanly you can document R&D costs. Keep per-project cost logging from day one. It makes both the engineering and the paperwork easier.

What I'd do

If I were advising a UK team starting a new LLM product today:

  1. Start on an API. Always. Use a mid-tier model, build the eval set first, and measure real token usage for a few weeks. Do not buy GPUs against a forecast.
  2. Turn on caching and batch before anything else. They are the cheapest wins and need no new infrastructure.
  3. Add routing once you have data. A small model in front, a frontier model behind it, a validator in between. Only then ask whether the small tier should be self-hosted.
  4. Self-host for a reason other than price. Data residency, latency guarantees, offline operation, or truly sustained high volume. If the reason is "it should be cheaper," run the utilization maths honestly first.
  5. Budget engineering time as a real cost. A $2,000 GPU bill plus a part-time person keeping it healthy is not a $2,000 solution.
  6. Re-check prices quarterly. In this snapshot, sources disagreed on whether Sonnet 5 is still at its introductory price or has already moved to a higher rate, so check Anthropic's own pricing page before you budget (one tracker, Layer3 Labs, covers the disagreement). Prices and model lineups shift fast. Build so that swapping the model is a config change, not a rewrite.
  7. Talk to an accountant about R&D relief before the year-end, not after, so the cost records exist when you need them.

None of this is glamorous. The teams that keep LLM bills sane are the ones with boring cost logging and a router, not the ones with the biggest cluster.

If you are sizing an LLM system and want a second pair of eyes on the architecture or the cost model, you can reach me at lazar-milicevic.com/#contact. Otherwise, the rest of the blog goes deeper on RAG, agent workflows and production AI automation.

Frequently asked questions

Is it cheaper to use an LLM API or host my own GPU?

For most teams the API is cheaper, because you pay only for the tokens you use and nothing for idle hours. A GPU gives you a fixed amount of capacity whether you use 5% or 95% of it. As an illustration, one H100 at $3 per GPU-hour running 24/7 is about $2,160 per month before engineering time, storage and monitoring. Self-hosting only tends to win with sustained, predictable, high-volume traffic or hard data residency requirements.

What is the break-even point for self-hosting an LLM versus using an API?

The break-even depends on GPU utilization more than on the hourly price. A rented GPU costs the same whether it is busy or idle, so a team sending only a few thousand requests a day will usually pay more than on an API. The calculation can flip when traffic is steady and high, or when data cannot leave your infrastructure. I also factor in engineering time, on-call, and the quality gap of open-weight models against frontier APIs.

How much does an H100 GPU cost per hour in the cloud?

Prices vary widely by provider and pricing type. Third-party reporting puts the AWS p5.48xlarge (8x H100) at $55.04/hr on-demand, roughly $6.88 per H100 GPU-hour, while CloudZero reports H100 rental from about $1.49 to $6.98 per GPU-hour across providers including spot and marketplace rates. Sources disagree on exact figures, so check the vendor's own pricing page before budgeting.

What is a hybrid LLM architecture and how does it cut costs?

A hybrid setup routes cheap, high-volume, low-risk work to a small or local model and escalates only the hard cases to a frontier model. A typical flow is a cache lookup, then a small model, then a validator check, with failures escalated and non-urgent work sent to a batch API at half price. It only works if you have a cheap, honest validator and log cost per request by route. In my experience this plumbing delivers more savings than negotiating a better per-token rate.

How can I reduce LLM API costs without changing models?

Use batch processing, prompt caching and better architecture. Batch APIs from Anthropic and OpenAI are half price for work that tolerates latency, such as nightly enrichment and backfills. Prompt caching can cut cached input cost to a tenth of the input price on Haiku 4.5, but only if the stable part of the prompt comes first, so avoid putting a timestamp at the top. Also pick the cheapest model tier that passes your eval, since tier spread matters more than vendor spread.

Lazar Milicevic

Lazar Milićević

Senior Technical Engineer. I build AI automation, GenAI/LLM systems and cloud architecture — autonomous systems that run while you sleep. Founder of BizFlowAI.

Building something hard with AI or automation? I am open to talk.

Get in touch

← All posts