O nas
About gputps.com
Spec sheets tell you what a GPU is. They do not tell you what it does with a language model. That is the gap we built this site to close.
What this is
What we do
For any GPU we estimate how many tokens per second it can generate running a given open-weight language model, whether that model actually fits in its memory, and at which quantisation. We publish these side by side with the specification tables, so a buying decision is grounded in an actual answer rather than raw VRAM and FLOPS numbers that do not, by themselves, say anything about inference speed.
Why an estimate, and not a single number
Generating a token means reading a model's weights out of memory once, so decode speed is bounded by memory bandwidth divided by model size — that is the calculation behind every figure on this site. But the same card and model vary 20–40% between inference engines, engine versions, and quantisation formats, and no spec sheet predicts that. Publishing one confident-looking number would be dishonest about how precise this kind of estimate can actually be. So every figure on this site ships as a range, with the method and inputs behind it, rather than a bare point value.
The vendor marketplace
Alongside the estimates, vendors can list GPUs for sale directly on the card's own page — next to the answer to "what does this actually run", which is where that decision gets made.
Start here
Both directions, one calculation
Start from a card you own or are considering, or from the model you want to run.
Skontaktuj się
Found something wrong?
Znalazłem błędną cyfrę, brakuje GPU lub masz sugestię? Chcielibyśmy to usłyszeć..