Switch kalkulator TPS

Otwórz wagi Google 1.6T Parametry January 2021

Każda karta poniżej jest oceniana w stosunku do tego modelu przy wybranej długości kontekstu i minimalnej jakości. Prędkość to szacunkowa wartość dla pojedynczego żądania, obliczona na podstawie przepustowości pamięci karty i rozmiaru modelu po kompresji..

Obliczone dla tego modelu

0 818 karty, które mogą to uruchomić

Które GPUn mogą uruchomić Switch?

Ustaw wejścia, odczytaj odpowiedź

Dłuższa rozmowa potrzebuje więcej pamięci, co może wykluczyć ten model z mniejszych kart..

Ukrywa karty, które mogłyby pasować do modelu, kompresując je poniżej tego punktu.

0 karty pasują

Obliczanie
Wymagana Kwantyzacja Dopasowane

Żaden model w naszym katalogu nie może uruchomić tego modelu z tymi ustawieniami..

Prędkości są szacunkowe dla jednego zapytania — jedna rozmowa na raz — obliczone z szerokości pasma pamięci, rozmiaru modelu i kwantyzacji. Rzeczywista przepustowość różni się w zależności od czasu działania inferencji i jej wersji. Liczby publikowane przez dostawców sprzętu mierzą wiele jednoczesnych zapytań i są znacznie wyższe..

Na rekord

Pełna specyfikacja

Wszystko w rejestrze dla tego modelu. Większość z tego opisuje, jak został wytrenowany, a nie jak działa — przydatny kontekst do oceny, ile pracy włożono w jego budowę i jak wypada w porównaniu z modelami stworzonymi w innym skali..

Origin

Kto zbudował ten model, gdzie i kiedy został opublikowany.

Organizacja
Google
Typ organizacji
Industry
Kraj
United States of America
Opublikowano
11 January 2021
Autorzy
William Fedus, Barret Zoph, Noam Shazeer

Co to robi

Obszary problemowe, dla których model został stworzony. Model może obsługiwać kilka z każdego.

Domena
Language
Zadanie
Text autocompletion
Podejście
Self-supervised learning
NumerFormat
BF16

Rozmiar

Jak duży jest model i ile danych został wytrenowany. Parametry to wartość, która decyduje, czy mieści się na danej karcie graficznej.

Parametry
1.6T

"Combining expert, model and data parallelism, we design two large Switch Transformer models, one with 395 billion and 1.6 trillion parameters" Table 9 gives more precise count of 1571B parameters

Trening danych
86,400,000,000 tokens

"In our protocol we pre-train with 2^20 (1,048,576) tokens per batch for 550k steps amounting to 576B total tokens." 1 token ~ 0.75 words

Obliczenia treningowe

Arytmetyka wykonywana w celu wytrenowania modelu, mierzone w operacjach na liczbach zmiennoprzecinkowych. Jest to miara kosztu treningu, a nie tego, jak szybko gotowy model odpowiada.

Obliczenia treningowe
8.2 × 10²² FLOP

Table 4 https://arxiv.org/ftp/arxiv/papers/2104/2104.10350.pdf

Jak to zostało ustalone
Third-party estimation

Trening uruchomienia

Co fizycznie było potrzebne do treningu: które chipy, ile, przez jak długo i ile to pobrało z sieci.

Sprzęt treningowy
Google TPU v3
Użyte chipy
1,024
Chip-godziny
663,552
Czas rzeczywisty
648 hours (27 days)

see table 4 in https://arxiv.org/ftp/arxiv/papers/2104/2104.10350.pdf

Wykorzystanie sprzętu
HFU 28.0%

Table 4 in https://arxiv.org/pdf/2104.10350 gives measured performance of 34.4 TFLOP/s, vs. peak achievable FLOP/s of 123 TFLOP/s on the TPUv3 being used. HFU = 34.4/123 = 0.27967

Pobór mocy
935.4 kW
Compute cost
$145,101

Dostępność

Czy możesz uzyskać model i uruchomić go na swoim sprzęcie, co decyduje o tym, czy jakiekolwiek z danych dotyczących karty graficznej na tej stronie mają zastosowanie.

Wagi
Open — downloadable
Dostęp do modelu
Open weights (unrestricted)
Kod treningowy
Unreleased

Apache 2 for weights: https://huggingface.co/google/switch-c-2048 paper links to this repo but not clear that the training hyperparams for Switch are here: https://github.com/google-research/t5x

Jak jest klasyfikowany

Etykiety zastosowane w źródłowym zbiorze danych przy śledzeniu istotnych modeli oraz jak pewny jest wpisu.

Model Frontier
Yes
Dlaczego to jest śledzone
Highly cited,SOTA improvement

" On ANLI (Nie et al., 2019), Switch XXL improves over the prior state-of-the-art to get a 65.7 accuracy versus the prior best of 49.4 (Yang et al., 2020)... Finally, we also conduct an early examination of the model’s knowledge with three closed-book knowledge-based tasks: Natural Questions, WebQuestions and TriviaQA, without additional pre-training using Salient Span Masking (Guu et al., 2020). In all three cases, we observe improvements over the prior stateof-the-art T5-XXL model (without SSM…

Model/modele nie pasuje do VRAM GPU.
Confident
Model/modeli nie pasuje do VRAM GPU. Zbyt ciasne. Wygodne.
3,888

Źródła

Gdzie ten rekord pochodził i kiedy był ostatnio sprawdzany.

Wersja modelu nie pasuje do VRAM GPU.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Ostatnia aktualizacja
25 May 2026

Co oznaczają liczby

Co potrzebujesz, aby to uruchomić

At 1.6T parameters, Switch is beyond what any single graphics card holds. Running it means either splitting it across several cards or renting hardware built for the job — 0 of the cards we track can hold it on their own, and all of them are datacentre parts.

Kontekst

Switch was published by Google, in United States of America, in January 2021. The organisation is categorised as industry.

It works in Language, and is recorded as doing text autocompletion.

Wagi są opublikowane, więc można je pobrać i uruchomić na własnym sprzęcie przez czas nieokreślony, offline, bez powiązanego konta.

Szkolenie i pochodzenie

The training run consumed about 8.2 × 10²² FLOP, on Google TPU v3. That figure describes the cost of creating it and has no bearing on how quickly it generates text.

Around 86,400,000,000 tokens went into training it.

It is tracked in the underlying dataset for one reason in particular: highly cited,SOTA improvement.

Krok po kroku

Jak wybrać GPU dla Switch

Tabela powyżej już oceniła każdą kartę, dla której posiadamy specyfikacje, w porównaniu do tego modelu. Dotarcie do twojej odpowiedzi zajmuje sześć kroków..

  1. 01

    Przeczytaj najpierw wartość pamięci

    Look at what Switch actually needs. No amount of processing power compensates for a card that cannot hold it.

  2. 02

    Zdecyduj, jak długo będą trwały twoje rozmowy.

    The conversation occupies memory too, and grows as it goes. Set the slider to the length you expect: at long context Switch can slip off a card that handles short questions easily.

  3. 03

    Zdecyduj, ile kompresji zaakceptujesz.

    Compression is what makes Switch fit smaller cards, at some cost in accuracy. A minimum quality removes the ones that go too far.

  4. 04

    Sortuj według prędkości

    Sort by speed to see how cards rank for Switch. It will not match a gaming ordering — generation is bound by memory bandwidth.

  5. 05

    Przeczytaj ostatnią kolumnę dopasowania.

    The fit column separates cards that just manage Switch from those with room to spare. Buy for the second if the context might grow.

  6. 06

    Zobacz, co ta karta może uruchomić.

    Following a card through to its own page shows every other model it can hold, which is the question that follows once Switch is settled.

Odpowiedzi

Switch — Powszechne pytania

01

What is Switch used for?

Switch works in Language, and is recorded as handling text autocompletion. These are the areas it was designed around; they describe intent rather than a hard boundary.

02

Where can I download Switch?

The weights for Switch are published, though we do not hold a repository link for it. This site calculates hardware requirements rather than hosting model files.

03

How much compute was used to train Switch?

Around 8.2 × 10²² FLOP, on Google TPU v3. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.

04

Can I run Switch if it does not fit in my GPU?

It can be split between the card and system memory, but Switch generates painfully slowly that way — the nearest miss we calculate is short by 691.9 GB. Nothing on this page assumes offloading.

05

Would two GPUs run Switch faster?

Two cards buy memory rather than speed. That matters for Switch only if one card cannot hold it — 0 can, so a second adds little.

06

Why does the quantisation differ between cards for Switch?

A larger card holds a more accurate copy. Across the cards that run Switch, 1 compression levels are used; the floor control above pins it to one.

07

How accurate are these Switch speed estimates?

These are estimates with real error bars. The fastest result here, the range beneath each figure, could reasonably land anywhere in its published range depending on which runtime you use.

08

Is Switch open source?

Its weights are published, so Switch can be downloaded and run on your own hardware. Note that open weights is not the same as open source in the full sense — it says nothing about the training data, the training code, or the commercial terms attached.

09

How many parameters does Switch have?

Switch has 1.6T parameters. "Combining expert, model and data parallelism, we design two large Switch Transformer models, one with 395 billion and 1.6 trillion parameters" Table 9 gives more precise count of 1571B parameters. That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.

10

Who created Switch?

Switch was published by Google, based in United States of America, categorised as industry.

11

When was Switch released?

Switch was published in January 2021. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.

Źródło

Oryginalna publikacja

Model/modeli nie pasuje do VRAM GPU. Zbyt ciasne. Model działa wygodnie z zapasem. Jak modele AI może GPU uruchomić? 25 May 2026

Inna strona

Patrząc na to z drugiej strony?

Ta strona zaczyna się od modelu. Jeśli już posiadasz kartę i chcesz wiedzieć, co wszystko uruchomi., Model/modele nie mieści się w VRAM GPU..