Most businesses will never choose a chip.
- Silicon choice matters only when inference volume is high, workload predictable, and compute is a material cost.
- Custom accelerators are cheaper per token by shedding generality and optimizing tensor matrix multiplications.
- Claimed ASIC savings require sustained, predictable utilization; spot GPU pricing can beat ASICs for interruptible workloads.
- Custom silicon causes vendor lock-in through proprietary SDKs; estimate migration and engineering costs before committing.
- GPUs retain advantages in software ecosystem, portability, flexibility, and adapt better to changing model architectures.
You choose a cloud, a model provider, or an API, and the silicon decision is made for you several layers down. That is the correct outcome for the large majority of AI workloads, and any article urging you to evaluate tensor architectures is answering a question you do not have.
The useful questions are narrower: at what point does this abstraction leak, what changes when it does, and what does it cost to change your mind afterwards. That last one turns out to be the whole story.
When the Decision Becomes Yours
Silicon choice starts to matter when three conditions hold together:
- Inference volume is high and sustained, not bursty or experimental
- The workload is stable enough that you can predict utilisation months ahead
- Compute is a material line item, large enough that a 40% difference changes your unit economics
Below that threshold, the price difference between accelerator types is smaller than the engineering cost of caring about it. Above it, the difference compounds monthly and becomes one of the larger levers available.
Note that the trigger is inference, not training. Inference now represents roughly two-thirds of all AI compute, and it is where the economics of custom silicon actually bite.
Why Custom Chips Are Cheaper Per Token
The architectural difference is genuine rather than marketing.
GPUs use highly programmable cores under a SIMT execution model, dynamically distributing work. That flexibility supports diverse and evolving workloads, and it costs control overhead and energy per operation.
Custom AI accelerators are built around dedicated tensor pipelines and systolic-array-derived designs optimised for dense matrix multiplication — the operation that dominates neural network computation. Google’s TPU is, in the strict sense, a matrix multiplication engine rather than a general processor.
Strip out the generality and you gain efficiency on the specific operation. The cost is that anything outside that operation runs worse or not at all.
The 2026 Landscape
| Chip | Focus | Notable specification |
|---|---|---|
| NVIDIA B300 Blackwell Ultra | Training and inference | 288 GB HBM3e, ~15 PFLOPS dense FP4 |
| Google TPU v7 Ironwood | Training and inference | ~4,614 TFLOPS per chip; analyst consensus puts it near B200 for transformer workloads |
| AWS Trainium 3 | Training and inference | 2.52 PFLOPS FP8, 144 GB HBM3e |
| AWS Inferentia | Inference only | Optimised for high-throughput, low-latency serving |
| Microsoft Maia 200 | Inference | TSMC 3nm, 140B+ transistors, 216 GB HBM3e |
| Meta MTIA | Internal workloads | Four generations announced in 2026, RISC-V based |
The market split is stark. Custom ASIC shipments are projected to grow at around 44.6% against roughly 16.1% for standard GPUs, and some analysts project NVIDIA’s inference share falling substantially by 2028.
Worth noting that most of these are not built alone. Broadcom is Google’s long-standing TPU design partner under an agreement reported to run to 2031, and similar design-services relationships underpin several other programmes.
The Cost Claims, and Their Conditions
Published figures cluster in a wide band: AWS has claimed up to 50% savings against GPU inference, Trainium 3 materials cite up to 70% reduction, and independent analyses put the range at roughly 40 to 60%.
Treat all of these as vendor or vendor-adjacent, and read the condition attached to them: sustained, production-scale workloads where utilisation is predictable and continuous.
That qualifier does most of the work. Custom silicon wins on steady high-volume serving. It wins much less, or not at all, on bursty, experimental, or variable workloads — which describes most enterprise AI today.
One practical detail that cuts the other way: GPU capacity is available on spot markets at materially lower prices for batch, embedding, and asynchronous work. Trainium and TPU are generally on-demand only. For workloads that tolerate interruption, that pricing mechanism can outweigh the ASIC efficiency advantage entirely.
The Trade Nobody Puts in the Comparison Table
Here is the point a business reader should take away above all others.
The favourable price-per-FLOP is customer acquisition cost.
Cloud providers offer custom silicon at attractive rates because it captures your workload inside their ecosystem. Once your production inference stack is written against a proprietary SDK, switching costs rise to the point where short-term pricing differences stop mattering. The discount buys the lock-in.
That is not a criticism — it is a rational strategy and the economics can still favour you. But it should be priced into the decision, not discovered eighteen months later. The relevant question is not “what does this cost per million tokens today,” it is “what does it cost to leave in two years, and what happens to my rate when they know the answer.”
What GPUs Still Win On
Three moats have not been breached.
Software. CUDA represents decades of libraries, frameworks, and tooling. It remains the path of least resistance for almost every developer, and moving to a proprietary stack requires non-trivial code changes.
Portability. GPU workloads run across multiple clouds and on-premise. Custom accelerators generally run in exactly one place, with one vendor’s pricing.
Flexibility. Model architectures are still changing. A chip optimised for today’s dominant operation is a bet that the operation stays dominant, and GPUs absorb architectural change more gracefully.
There is also a supply dimension worth knowing: every major custom accelerator fabricates on the same leading-edge process, which has been running at full capacity with demand reported at several times supply. Availability, not just price, is a real constraint on both sides of this comparison.
Two Other Categories Worth Knowing
The debate is usually framed as GPUs versus hyperscaler ASICs, but two other accelerator types matter for specific situations.
FPGAs are reconfigurable, sitting between GPUs and fixed-function ASICs. They offer moderate efficiency gains while remaining changeable after deployment, which makes them useful for prototyping novel architectures and for workloads where the model will change but the deployment cannot easily be re-provisioned. They rarely win on raw throughput or on developer familiarity.
NPUs are the edge and on-device category — the accelerators appearing in laptops and phones. They prioritise low power for local inference rather than throughput at scale. If part of your product runs on the user’s device, this is the tier that determines what is feasible there, and it operates under completely different constraints from anything in a data centre.
Neither is a substitute for the main comparison. Both are relevant if your workload sits outside continuous cloud-based serving, and both are frequently omitted from the debate entirely.
It Is Not Either/Or
The framing that serves businesses best: AI has split computing into layers, and each layer wants different silicon.
A single model may be trained on GPUs, optimised in the cloud, served on inference accelerators, and run in reduced form on a device NPU. NVIDIA can lead in GPUs while Google succeeds with TPUs and Amazon sells both. These are not contradictions.
Even the largest AI developers run mixed estates, using custom silicon from more than one provider alongside GPU capacity. If organisations at that scale do not standardise on one architecture, a smaller organisation almost certainly should not either.
A Decision Framework
- Establish whether you are actually in scope. If compute is not a material line item, use whatever your provider defaults to and revisit annually.
- Separate training from inference. They have different economics and often warrant different answers.
- Measure utilisation predictability, since that is the single variable determining whether ASIC savings materialise.
- Price the exit before the entry. Estimate the engineering cost of migrating your inference stack, and treat that as the real switching cost.
- Check the pricing mechanisms, not just the rates. Spot availability can change the comparison entirely for interruptible work.
- Keep the abstraction layer where you can. Frameworks and serving layers that support multiple backends preserve optionality cheaply.
- Revisit each generation. Both sides are shipping annually, and a decision made two generations ago is stale.
Final Thoughts
Custom accelerators are winning inference on economics, and that shift is real rather than hype. But the comparison most businesses need is not architectural.
It is that GPUs cost more per token and keep your options open, while custom silicon costs less per token and closes them. Which trade is correct depends on whether your workload is stable enough for the savings to compound before the lock-in costs you more than it saved.
For most organisations today, the honest answer is that the decision is not yet yours to make — and the right move is keeping it that way for as long as the abstraction holds.
