Deciding whether to run a language model yourself comes down to three numbers: whether it fits, whether it is fast enough, and whether it is cheaper. Each of these tools answers one of them from measured figures rather than vendor arithmetic. All three run entirely in your browser — nothing is uploaded, and there is nothing to sign in to.
Paste a HuggingFace model and see what it actually needs — weights, KV cache and runtime overhead — against the machines people actually own. Reads the real file sizes rather than estimating from a table.
Every inference vendor quotes tokens per second and almost nobody knows what the number feels like. Move a slider and watch text arrive at exactly that rate, with reference marks measured on our own hardware.
A break-even calculator for self-hosting, built on throughput and wall power measured on our own machine against published hosted list prices. It returns whichever side is actually cheaper at your volume.
This company runs its own models on its own hardware, and each of these tools started as a question we had to answer for ourselves before committing to a machine or a model. Publishing them costs nothing and saves someone else the same afternoon.
The AI news that matters — in your inbox by 07:30 CET. Free, no spam.