New Horizon No. 233 / 2026-08-21 · Berlin
Interactive

What It Costs

Your hardware or someone else’s, priced on measured throughput and metered power.

Self-hosting a language model is usually sold as the cheap option. Whether it is depends entirely on volume. This calculator runs on throughput and wall power we measured on our own machine, against published list prices for hosted APIs, and returns whichever side is actually cheaper at the volume you set.

Both sides are counted in output tokens only. Hosted APIs bill input tokens as well, so the hosted figure here is a floor rather than a full invoice.


Where Your Volume Lands

250k output tokens per day
1k100k1M20M
Hosted API
0
per month, output tokens
Own hardware
0
1 card at 0% duty + electricity
Break-even:
Hosted API Own hardware cost per month · both axes logarithmic
Assumptions — change any of them

This calculator needs JavaScript. The short version: at 2026 list prices, one mid-range card beats a mid-tier hosted API from roughly 200,000 output tokens a day, beats a frontier model almost immediately, and never beats the cheapest hosted open-weight APIs on price at all.

Hosted prices are published list prices for representative models, in USD, converted at 0.866 €/$ on 2026-08-19. They move; the field is editable, so an out-of-date default can be corrected in place. Electricity defaults to the German SME band. Only the energy a generation actually draws is charged to it — the machine’s always-on baseline is a cost of the company, not of a token.


Numbers Off Our Own Wall Socket

The local side of this calculator is not a spec sheet. Throughput came from the inference server’s own token counters; the power figure is the difference between what the machine drew at rest and what it drew while generating, read off the UPS that powers it. Every number below was taken on 2026-08-19.

Throughput
38.7 t/s
  • 12B model, 3-bit quantised
  • One RTX 3060 Ti, 8 GB
  • Burst 39.5; 3.34M tokens per card-day
Power
+205.2 W
  • Rack at rest: 102.6 W
  • Rack generating: 307.8 W
  • Card alone accounts for 176.6 W
Energy per million
1.47 kWh
  • Per 1M output tokens
  • At the German SME rate: €0.40
  • This is the whole marginal cost

Buy Hardware for the Right Reason

Three regimes come out of the arithmetic. Against the cheapest hosted open-weight APIs, a card like ours does not pay for itself: the electricity to generate a million tokens costs more than buying those tokens outright. Against a mid-tier model it pays back somewhere around two hundred thousand output tokens a day, which is a few hundred documents. Against a frontier model it pays back almost immediately.

The frontier case needs a caveat. A 12B model on one mid-range card is not a frontier model, and a table that puts them in the same column compares two different products on one axis. The comparison that holds is narrower: for the high-volume, low-margin work a small model genuinely handles — classification, extraction, first-draft summaries, routing — you are otherwise paying frontier prices for a job that does not need them.

Which leaves the two reasons that survive the arithmetic entirely. The first is that nothing leaves the building — no prompt, no document, no customer record. The second is predictability: our own box held its throughput to under one percent across runs while a hosted frontier model swung by a factor of eleven on the same day. Neither of those appears in any price per million.

→ Work out which regime you are in → The speed side of the same box


The AI news that matters — in your inbox by 07:30 CET. Free, no spam.