Over ons

About gputps.com

Spec sheets tell you what a GPU is. They do not tell you what it does with a language model. That is the gap we built this site to close.

What this is

What we do

For any GPU we estimate how many tokens per second it can generate running a given open-weight language model, whether that model actually fits in its memory, and at which quantisation. We publish these side by side with the specification tables, so a buying decision is grounded in an actual answer rather than raw VRAM and FLOPS numbers that do not, by themselves, say anything about inference speed.

Why an estimate, and not a single number

Generating a token means reading a model's weights out of memory once, so decode speed is bounded by memory bandwidth divided by model size — that is the calculation behind every figure on this site. But the same card and model vary 20–40% between inference engines, engine versions, and quantisation formats, and no spec sheet predicts that. Publishing one confident-looking number would be dishonest about how precise this kind of estimate can actually be. So every figure on this site ships as a range, with the method and inputs behind it, rather than a bare point value.

The vendor marketplace

Alongside the estimates, vendors can list GPUs for sale directly on the card's own page — next to the answer to "what does this actually run", which is where that decision gets made.

Learn about becoming a vendor.

Start here

Both directions, one calculation

Start from a card you own or are considering, or from the model you want to run.

Neem contact op

Found something wrong?

Iets past niet / zal niet passen.