The catalogue
GPU TPS calculator
Work out how many tokens per second any graphics card produces on a local AI model. Filter the catalogue below to the cards worth considering, then open one to see every language model it can run, how much memory each needs, and how fast it is likely to be.
818
cards on file
4
manufacturers
4 GB – 288 GB
memory range
8,190 GB/s
highest bandwidth
Every word has to match something, so more words narrow the list. A capacity such as 24gb matches on the card's memory even though no column stores it as text.
818 cards match
Updating| Memory type | Power | ||||||
|---|---|---|---|---|---|---|---|
| B300 NVIDIA · Blackwell Ultra | 288 GB | 8,000 GB/s | 2.03 GHz | 1.67 GHz | Sep 2025 | HBM3e | 1,400 W |
| Radeon Instinct MI350X AMD · CDNA 4.0 | 288 GB | 8,190 GB/s | 2.2 GHz | 1 GHz | Jan 2025 | HBM3e | 1,000 W |
| Radeon Instinct MI355X AMD · CDNA 4.0 | 288 GB | 8,190 GB/s | 2.4 GHz | 1 GHz | Jan 2025 | HBM3e | 1,400 W |
| Radeon Instinct MI325X AMD · CDNA 3.0 | 256 GB | 6,000 GB/s | 2.1 GHz | 1 GHz | Oct 2024 | HBM3e | 1,000 W |
| Radeon Instinct MI300X AMD · CDNA 3.0 | 192 GB | 5,325 GB/s | 2.1 GHz | 1 GHz | Dec 2023 | HBM3 | 750 W |
| Radeon Instinct MI308X AMD · CDNA 3.0 | 192 GB | 5,325 GB/s | 2.1 GHz | 1 GHz | Dec 2023 | HBM3 | 750 W |
| B200 NVIDIA · Blackwell | 180 GB | 8,000 GB/s | 1.97 GHz | 700 MHz | Jan 2024 | HBM3e | 1,000 W |
| H200 NVL NVIDIA · Hopper | 141 GB | 4,890 GB/s | 1.79 GHz | 1.37 GHz | Nov 2024 | HBM3e | 600 W |
| H200 SXM 141 GB NVIDIA · Hopper | 141 GB | 4,890 GB/s | 1.98 GHz | 1.5 GHz | Nov 2024 | HBM3e | 700 W |
| Data Center GPU Max 1550 Intel · Generation 12.5 | 128 GB | 3,280 GB/s | 1.6 GHz | 900 MHz | Jan 2023 | HBM2e | 600 W |
| Data Center GPU Max Subsystem Intel · Generation 12.5 | 128 GB | 3,210 GB/s | 1.6 GHz | 900 MHz | Jan 2023 | HBM2e | 2,400 W |
| GB10 NVIDIA · Blackwell 2.0 | 128 GB | 273 GB/s | 2.42 GHz | 1.67 GHz | Oct 2025 | LPDDR5X | 140 W |
| Jetson T5000 NVIDIA · Blackwell | 128 GB | 273 GB/s | 2.53 GHz | 1.67 GHz | Aug 2025 | LPDDR5X | 40 W |
| Radeon Instinct MI250 AMD · CDNA 2.0 | 128 GB | 3,280 GB/s | 1.7 GHz | 1 GHz | Nov 2021 | HBM2e | 500 W |
| Radeon Instinct MI250X AMD · CDNA 2.0 | 128 GB | 3,280 GB/s | 1.7 GHz | 1 GHz | Nov 2021 | HBM2e | 500 W |
| Radeon Instinct MI300 AMD · CDNA 3.0 | 128 GB | 6,550 GB/s | 1.7 GHz | 1 GHz | Jan 2023 | HBM3 | 600 W |
| Radeon Instinct MI300A AMD · CDNA 3.0 | 128 GB | 5,325 GB/s | 2.1 GHz | 1 GHz | Dec 2023 | HBM3 | 750 W |
| Data Center GPU Max 1350 Intel · Generation 12.5 | 96 GB | 2,460 GB/s | 1.55 GHz | 750 MHz | Jan 2023 | HBM2e | 450 W |
| H100 PCIe 96 GB NVIDIA · Hopper | 96 GB | 3,360 GB/s | 1.84 GHz | 1.67 GHz | Mar 2023 | HBM3 | 700 W |
| H100 SXM5 96 GB NVIDIA · Hopper | 96 GB | 3,360 GB/s | 1.98 GHz | 1.35 GHz | Mar 2023 | HBM3 | 700 W |
| RTX PRO 6000 Blackwell NVIDIA · Blackwell 2.0 | 96 GB | 1,790 GB/s | 2.62 GHz | 1.59 GHz | Mar 2025 | GDDR7 | 600 W |
| RTX PRO 6000 Blackwell Max-Q NVIDIA · Blackwell 2.0 | 96 GB | 1,790 GB/s | 2.28 GHz | 1.04 GHz | Mar 2025 | GDDR7 | 300 W |
| RTX PRO 6000 Blackwell Server NVIDIA · Blackwell 2.0 | 96 GB | 1,790 GB/s | 2.62 GHz | 1.59 GHz | Mar 2025 | GDDR7 | 600 W |
| RTX PRO 6000D Blackwell Max-Q NVIDIA · Blackwell 2.0 | 96 GB | 1,790 GB/s | 2.29 GHz | 1.59 GHz | Mar 2025 | GDDR7 | 300 W |
| H100 NVL 94 GB NVIDIA · Hopper | 94 GB | 3,940 GB/s | 1.79 GHz | 1.08 GHz | Mar 2023 | HBM3 | 400 W |
| H100 SXM5 94 GB NVIDIA · Hopper | 94 GB | 3,360 GB/s | 1.98 GHz | 1.35 GHz | Mar 2023 | HBM3 | 700 W |
| RTX 6000D NVIDIA · Blackwell 2.0 | 84 GB | 1,570 GB/s | 2.43 GHz | 1.59 GHz | Mar 2025 | GDDR7 | 600 W |
| A100 PCIe 80 GB NVIDIA · Ampere | 80 GB | 1,940 GB/s | 1.41 GHz | 1.07 GHz | Jun 2021 | HBM2e | 300 W |
| A100 SXM4 80 GB NVIDIA · Ampere | 80 GB | 2,040 GB/s | 1.41 GHz | 1.28 GHz | Nov 2020 | HBM2e | 400 W |
| A100X NVIDIA · Ampere | 80 GB | 2,040 GB/s | 1.44 GHz | 795 MHz | Jun 2021 | HBM2e | 300 W |
Step by step
How to calculate the TPS of your GPU
Nothing here needs working out by hand. The catalogue above already holds every card, and each card page already holds every model it can run with a tokens-per-second figure attached. Getting to your answer takes six steps.
-
01
Narrow the catalogue to cards worth considering
Use the filters above to set what you actually need — a minimum memory size, a manufacturer, a power ceiling your supply can feed, or a release year. Memory is the filter to start with, because it decides which models are possible at all.
-
02
Or search for a card you already own
Type the name into the search box. Every word has to match something, so more words narrow the list, and a capacity such as 24gb matches on the card's memory even though no column stores it as text.
-
03
Sort by the specification that decides your outcome
Sort by memory if the question is which models fit, or by bandwidth if the question is how fast they will run. Clock speeds and core counts are listed for completeness but barely affect text generation.
-
04
Open the card
Each card has its own page listing every language model it can run, with estimated tokens per second, the memory each model needs and the compression it runs at.
-
05
Set your context length and minimum quality
On the card page, set the conversation length you expect to work at and the lowest compression you are willing to accept. Both change the answer rather than filtering it — a longer conversation needs more memory, which can push a large model off the card entirely.
-
06
Read the range, not the single number
Every throughput figure is published with a confidence range around it. Take the range as the answer: the same card and model vary between inference engines and versions in ways no specification sheet predicts, and the headline number is only the middle of that spread.
What the numbers mean
What this catalogue is
This is a database of 818 graphics cards, filtered down to the ones with enough memory to hold a language model at all. Every card has its own page, and on that page every model we track has been assessed against it — whether it fits, at what compression, and how many tokens per second it is likely to produce.
The calculation runs when you open a card rather than being looked up from a stored table, which is what allows the context length and the minimum quality to be things you set rather than things we assumed on your behalf. Change either and every figure recomputes.
Why tokens per second is the number that matters
Running a language model on your own hardware became an ordinary thing to do in a short space of time. Open-weight models caught up far enough to be genuinely useful, and the tools to run them stopped requiring a research background. What did not arrive alongside that was a straight answer to the first question anyone asks: will it run on the card I already have, and will it be fast enough to be worth using.
Spec sheets do not answer it. They are written for graphics workloads and lead with the figures that decide frame rates, which are almost entirely the wrong ones. The two numbers that actually decide the outcome are memory capacity, which sets whether a model fits, and memory bandwidth, which sets how fast it generates once it does.
Tokens per second is the number that expresses the result in something you can feel. Reading speed is somewhere around ten tokens per second, so anything above that arrives faster than you can read it and anything well below feels like waiting. That is the figure this site calculates, for every combination of card and model it holds.
What the catalogue covers
Memory range
4 GB – 288 GB
Highest bandwidth
8,190 GB/s
Power range
6 – 2,400 W
Memory across the catalogue runs from 4 GB to 288 GB. That range is the difference between a card limited to the smallest usable models and one that holds models no desktop machine can touch — capacity is a hard gate rather than a gradual trade-off, so a model either fits or it does not run.
Memory bandwidth runs from 10 GB/s at the low end to 8,190 GB/s at the top, held by the Radeon Instinct MI355X. Because generating each token means reading the whole model out of memory once, that spread translates almost directly into a spread in generation speed: a card with ten times the bandwidth produces tokens roughly ten times faster on the same model.
Cards come from AMD, ATI, Intel and NVIDIA. The manufacturer matters more than it should: inference software has had far more attention paid to the CUDA path than to any other, so an AMD or Intel card of equivalent bandwidth generally reaches a smaller share of its theoretical ceiling. Our estimates apply a penalty for that rather than pretending the hardware is the only variable.
Power draw spans 6 W to 2,400 W. The upper end is datacentre hardware that a desktop supply cannot feed, and it is the constraint people tend to discover last — after the card fits the budget and the case.
Release dates run from 2009 to 2025, so the catalogue covers hardware people still own as well as hardware they might buy. An older card is not excluded for being old — if it has the memory, it can still run a model, and the page will say how fast.
Buying a card
Vendors can list cards for sale against any entry in this catalogue, so a card page can show both what the hardware will run and where it can be bought. Listings are attached to the specific board variant rather than to a model name, which matters here more than it usually would — the same name often covers several memory capacities, and capacity is the thing that decides which models run.
Leaderboards
Most memory
Capacity decides which models fit at all.
Highest bandwidth
Bandwidth decides how fast they run once they fit.
Most power-hungry
Check these against the supply you already own.
Reading the table
How to read this catalogue
- Memory
- How much the card can hold. A rough guide is a little over half a gigabyte per billion parameters at the compression most people use, plus room for the conversation. Not all of it is available — inference software reserves roughly a tenth for its own working space.
- Bandwidth
- How fast the card reads its own memory, in gigabytes per second. This is the single best predictor of generation speed anywhere on this site.
- Boost and base clock
- How fast the processor runs. Listed for completeness, and far less predictive here than on a gaming benchmark: generation waits on memory, not on arithmetic.
- Memory type and bus interface
- The memory technology and how the card connects to the machine. HBM parts are datacentre hardware with far more bandwidth than the GDDR used on desktop cards. The bus interface governs how fast a model loads, not how fast it runs.
- Power
- The board's rated draw in watts. Worth filtering on when the power supply or case is already decided, since the largest cards need considerably more than a typical desktop provides.
- Why a card appears more than once
- Each row is a board variant rather than a chip, so the same name can appear at several memory capacities. That distinction matters here more than anywhere else, because capacity is what decides whether a model runs.
Answers
Common questions
How do I calculate the tokens per second of my GPU?
You do not have to. Find your card in the table above — there are 818 in the catalogue — and open it. Its page lists every language model it can run with an estimated tokens-per-second figure for each, calculated from the card's memory bandwidth and the size of the model once compressed.
What is a good tokens per second for running AI locally?
Reading speed is roughly ten tokens per second, so anything above that produces text faster than you can read it and feels immediate. Between five and ten is usable but noticeably slow. Below five is uncomfortable for conversation, though still fine for work you leave running.
Which GPU specification matters most for AI?
Memory capacity first, because the entire model has to be held on the card before it can generate anything, and bandwidth second, because generating each token means reading that model out of memory once. Core counts and clock speeds decide graphics performance and barely register here.
How much VRAM do I need to run a language model?
As a rough guide, a little over half a gigabyte per billion parameters at the compression most people use, plus room for the conversation. That puts an eight-billion-parameter model comfortably on eight gigabytes and a thirty-billion one on twenty-four. Cards in this catalogue range from 4 GB to 288 GB.
How many graphics cards are in this database?
The catalogue holds 818 cards from AMD, ATI, Intel and NVIDIA. Integrated graphics are excluded because they have no memory of their own, and so are cards below four gigabytes, on which no model in our catalogue fits at any compression.
Which GPU has the highest memory bandwidth?
The Radeon Instinct MI355X, at 8,190 GB/s. Since generation speed follows bandwidth almost directly, it is also among the fastest cards here for producing text — though whether that matters depends first on whether your model fits.
Which GPU has the most memory?
The Radeon Instinct MI355X, with 288 GB. Capacity at that level is datacentre territory, and it is what allows the largest open-weight models to be held on a single card rather than split across several.
Is a gaming graphics card good enough for local AI?
For one person at a time, usually yes, and often better value than datacentre hardware. Consumer cards give up total memory rather than speed, so the limit is which models fit rather than how fast they run once they do.
Does the manufacturer matter for AI performance?
More than it should. Inference software has had far more work put into the CUDA path than into the alternatives, so an AMD or Intel card of equivalent bandwidth typically reaches a smaller share of its theoretical ceiling. Our estimates apply a penalty for that rather than assuming the hardware is the only variable.
Can I use two graphics cards together?
Yes, and it is often the point — two cards hold models neither could hold alone. It does not double generation speed the way it doubles memory, and every figure on this site describes a single card on its own.
How accurate are these tokens-per-second estimates?
They are calculated from specifications rather than measured, and every one is published with a range rather than as a single number. The same card and model vary by thirty to fifty per cent depending on which inference software you use and which version of it, which no calculation from spec sheets can predict. Treat the range as the honest answer.
I know which model I want. Can I search the other way round?
Yes. The AI models section lists every model on file, and each model's page names the graphics cards that can run it, smallest and fastest first. That is the same calculation approached from the opposite end.
The other direction
Looking at it from the other side?
This page starts from the hardware. If you already know which model you want to run and need to know what it takes, start there instead — lists every model on file, and each one names the cards that can run it.