Switch TPS-calculator
Elke kaart hieronder wordt beoordeeld op basis van dit model bij de contextlengte en minimumkwaliteit die je kiest. Snelheid is een schatting voor een enkele aanvraag, berekend op basis van de geheugendoorvoer van de kaart en de grootte van het model eenmaal gecomprimeerd..
Berekend voor dit model
Welke GPU's kunnen draaien Switch?
Stel de invoer in, lees het antwoord
Een langere conversatie heeft meer geheugen nodig, wat dit model van kleinere kaarten kan duwen..
Verbergt kaarten die alleen in het model zouden passen door deze onder dit punt te comprimeren..
0 kaarten passen
Berekenen| Behoeften | Kwantisatie | Pas | |||||
|---|---|---|---|---|---|---|---|
|
Geen kaart in onze catalogus kan dit model met deze instellingen draaien.. |
|||||||
Snelheden zijn schattingen voor een enkele aanvraag - één gesprek tegelijk - berekend uit geheugenbandbreedte, modelgrootte en kwantisatie. De werkelijke doorvoer varieert met de inferentietijd en de versie. Cijfers gepubliceerd door hardwareleveranciers meten veel gelijktijdige aanvragen en zijn veel hoger..
Opname
Volledige specificatie
Alles op record voor dit model. Het meeste beschrijft hoe het is getraind in plaats van hoe het draait — nuttige context om te beoordelen hoeveel werk erin is gestoken en hoe het zich verhoudt tot modellen die op een andere schaal zijn gebouwd..
Origine
Wie heeft dit model gebouwd, waar, en wanneer het is gepubliceerd.
- Organisatie
- Organisatietype
- Industry
- Land
- United States of America
- Gepubliceerd
- 11 January 2021
- Vertalers
- William Fedus, Barret Zoph, Noam Shazeer
Wat het doet
De probleemgebieden waarvoor het model is gebouwd. Een model kan er verschillende van elk dragen.
- Domein
- Language
- Model/modelen kunnen niet passen in de GPU VRAM. Tight = bijna uitvoerbaar, comfortabel = draait eenvoudig met ruimte. Wat AI-model kan de GPU draaien?
- Text autocompletion
- Benadering
- Self-supervised learning
- BF16
Grootte
Hoe groot het model is en hoeveel data erop is getraind. Parameters zijn de cijfers die bepalen of het past op een bepaalde grafische kaart.
- Parameters
- 1.6T
- Training gegevens
- 86,400,000,000 tokens
"Combining expert, model and data parallelism, we design two large Switch Transformer models, one with 395 billion and 1.6 trillion parameters" Table 9 gives more precise count of 1571B parameters
"In our protocol we pre-train with 2^20 (1,048,576) tokens per batch for 550k steps amounting to 576B total tokens." 1 token ~ 0.75 words
Trainingsberekeningen
De rekensom die is uitgevoerd om het model te trainen, gemeten in drijvende-komma-bewerkingen. Het is een maat voor de kosten van de trainingsrun, niet voor hoe snel het voltooide model je antwoorden geeft.
- Trainingsberekeningen
- 8.2 × 10²² FLOP
- Hoe het is vastgesteld
- Third-party estimation
Table 4 https://arxiv.org/ftp/arxiv/papers/2104/2104.10350.pdf
De trainingsrun
Wat het fysiek kostte om te trainen: welke chips, hoeveel, hoe lang, en wat dat van het net trok.
- Trainingshardware
- Google TPU v3
- Gebruikte chips
- 1,024
- Chip-uren
- 663,552
- Muur-klok tijd
- 648 hours (27 days)
- Hardware-utilisatie
- HFU 28.0%
- Vermogen verbruik
- 935.4 kW
- Compute cost
- $145,101
see table 4 in https://arxiv.org/ftp/arxiv/papers/2104/2104.10350.pdf
Table 4 in https://arxiv.org/pdf/2104.10350 gives measured performance of 34.4 TFLOP/s, vs. peak achievable FLOP/s of 123 TFLOP/s on the TPUv3 being used. HFU = 34.4/123 = 0.27967
Beschikbaarheid
Of u het model kunt verkrijgen en het op uw eigen hardware kunt draaien, wat bepaalt of een van de specificaties van de GPU op deze pagina van toepassing is.
- Gewichten
- Open — downloadable
- Model toegang
- Open weights (unrestricted)
- Training code
- Unreleased
Apache 2 for weights: https://huggingface.co/google/switch-c-2048 paper links to this repo but not clear that the training hyperparams for Switch are here: https://github.com/google-research/t5x
Hoe het is geclassificeerd
Labels die op de bron-dataset van toepassing zijn bij het volgen van opvallende modellen, en hoe zeker het is van de invoer.
- Frontier-model
- Yes
- Waarom het wordt gevolgd
- Highly cited,SOTA improvement
- Model/modellen past niet / past niet.
- Confident
- Model/modellen passen niet / zal niet passen. Te krap / comfortabel. GPU kan het AI-model draaien.
- 3,888
" On ANLI (Nie et al., 2019), Switch XXL improves over the prior state-of-the-art to get a 65.7 accuracy versus the prior best of 49.4 (Yang et al., 2020)... Finally, we also conduct an early examination of the model’s knowledge with three closed-book knowledge-based tasks: Natural Questions, WebQuestions and TriviaQA, without additional pre-training using Salient Span Masking (Guu et al., 2020). In all three cases, we observe improvements over the prior stateof-the-art T5-XXL model (without SSM…
Bronnen
Waar deze registratie vandaan komt en wanneer deze voor het laatst is gecontroleerd.
- Referentie
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Laatste update
- 25 May 2026
Wat de getallen betekenen
Wat je nodig hebt om het te draaien
At 1.6T parameters, Switch is beyond what any single graphics card holds. Running it means either splitting it across several cards or renting hardware built for the job — 0 of the cards we track can hold it on their own, and all of them are datacentre parts.
Achtergrond
Switch was published by Google, in United States of America, in January 2021. The organisation is categorised as industry.
It works in Language, and is recorded as doing text autocompletion.
The weights are published, so it can be downloaded and run on your own hardware indefinitely, offline, with no account attached.
Training en afkomst
The training run consumed about 8.2 × 10²² FLOP, on Google TPU v3. That figure describes the cost of creating it and has no bearing on how quickly it generates text.
Around 86,400,000,000 tokens went into training it.
It is tracked in the underlying dataset for one reason in particular: highly cited,SOTA improvement.
Stap voor stap
Hoe een GPU te kiezen voor Switch
De bovenstaande tabel heeft elke kaart die we hebben beoordeeld met specificaties tegen dit model. Naar je antwoord komen kost zes stappen..
-
01
Lees eerst de geheugenfiguur
Look at what Switch actually needs. No amount of processing power compensates for a card that cannot hold it.
-
02
Bepaal hoe lang uw gesprekken duren
The conversation occupies memory too, and grows as it goes. Set the slider to the length you expect: at long context Switch can slip off a card that handles short questions easily.
-
03
Bepaal hoeveel compressie je wilt accepteren.
Compression is what makes Switch fit smaller cards, at some cost in accuracy. A minimum quality removes the ones that go too far.
-
04
Sorteer op snelheid
Sort by speed to see how cards rank for Switch. It will not match a gaming ordering — generation is bound by memory bandwidth.
-
05
Lees de pas column als laatste
The fit column separates cards that just manage Switch from those with room to spare. Buy for the second if the context might grow.
-
06
Bekijk wat die kaart nog meer kan draaien
Following a card through to its own page shows every other model it can hold, which is the question that follows once Switch is settled.
Antwoorden
Switch — Voor veelgebruikte vragen
What is Switch used for?
Switch works in Language, and is recorded as handling text autocompletion. These are the areas it was designed around; they describe intent rather than a hard boundary.
Where can I download Switch?
The weights for Switch are published, though we do not hold a repository link for it. This site calculates hardware requirements rather than hosting model files.
How much compute was used to train Switch?
Around 8.2 × 10²² FLOP, on Google TPU v3. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.
Can I run Switch if it does not fit in my GPU?
It can be split between the card and system memory, but Switch generates painfully slowly that way — the nearest miss we calculate is short by 691.9 GB. Nothing on this page assumes offloading.
Would two GPUs run Switch faster?
Two cards buy memory rather than speed. That matters for Switch only if one card cannot hold it — 0 can, so a second adds little.
Why does the quantisation differ between cards for Switch?
A larger card holds a more accurate copy. Across the cards that run Switch, 1 compression levels are used; the floor control above pins it to one.
How accurate are these Switch speed estimates?
These are estimates with real error bars. The fastest result here, the range beneath each figure, could reasonably land anywhere in its published range depending on which runtime you use.
Is Switch open source?
Its weights are published, so Switch can be downloaded and run on your own hardware. Note that open weights is not the same as open source in the full sense — it says nothing about the training data, the training code, or the commercial terms attached.
How many parameters does Switch have?
Switch has 1.6T parameters. "Combining expert, model and data parallelism, we design two large Switch Transformer models, one with 395 billion and 1.6 trillion parameters" Table 9 gives more precise count of 1571B parameters. That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.
Who created Switch?
Switch was published by Google, based in United States of America, categorised as industry.
When was Switch released?
Switch was published in January 2021. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.
Bron
Model/modelen kunnen niet passen in de GPU VRAM. Tight = nauwelijks uitvoerbaar, comfortabel = draait gemakkelijk met ruimte. Welke AI model kan de GPU draaien? 25 May 2026
De andere richting
Kijk je er van de andere kant naar?
Deze pagina begint bij het model. Als je al een kaart hebt en alles wilt weten wat het kan draaien., model/modellen draait/draaien past niet/past niet goed strak = nauwelijks uitvoerbaar, comfortabel = draait gemakkelijk met speling.