Switch TPS-Rechner

Offene Gewichte Google 1.6T Parameter January 2021

Jede Karte unten wird gegen dieses Modell in der gewählten Kontextlänge und minimalen Qualität bewertet. Die Geschwindigkeit ist eine Schätzung für eine einzelne Anfrage, berechnet aus der Speicherbandbreite der Karte und der Größe des Modells, nachdem es komprimiert wurde..

Für dieses Modell berechnet

0 der 818 Karten, die es ausführen können

Welche GPUs können laufen Switch?

Setze die Eingaben, lies die Antwort

Ein längerer Dialog benötigt mehr Speicher, was dieses Modell von kleineren Karten drängen kann..

Versteckt Karten, die nur für das Modell passen würden, indem es unter diesen Punkt komprimiert wird..

0 Karten stimmen überein

Berechnen
Benötigt Quantisierung Fit

Kein Modell in unserem Katalog kann dieses Modell mit diesen Einstellungen ausführen..

Die Geschwindigkeiten sind Schätzungen für eine einzelne Anfrage – ein Gespräch zur Zeit – berechnet aus der Speicherbandbreite, der Modellgröße und der Quantisierung. Der tatsächliche Durchsatz variiert mit der Inferenzlaufzeit und deren Version. Die von Hardwareanbietern veröffentlichten Zahlen messen viele gleichzeitige Anfragen und sind wesentlich höher..

Aufzeichnen

Vollständige Spezifikation

Alles, was für dieses Modell aufgezeichnet wurde. Der Großteil beschreibt, wie es trainiert wurde, anstatt wie es läuft – nützlicher Kontext, um zu beurteilen, wie viel Arbeit darin steckt und wie es sich mit Modellen vergleicht, die in unterschiedlichem Maßstab erstellt wurden..

Ursprung

Wer hat dieses Modell erstellt, wo und wann es veröffentlicht wurde.

Organisation
Google
Typ der Organisation
Industry
Land
United States of America
Veröffentlicht
11 January 2021
Autoren
William Fedus, Barret Zoph, Noam Shazeer

Was es tut

Die Problembereiche, für die das Modell entwickelt wurde. Ein Modell kann mehrere von jedem tragen.

Domäne
Language
Aufgabe
Text autocompletion
Ansatz
Self-supervised learning
Numerisches Format
BF16

Größe

Wie groß das Modell ist und wie viele Daten es trainiert wurde. Parameter sind die Werte, die entscheiden, ob es auf eine bestimmte Grafikkarte passt.

Parameter
1.6T

"Combining expert, model and data parallelism, we design two large Switch Transformer models, one with 395 billion and 1.6 trillion parameters" Table 9 gives more precise count of 1571B parameters

Modell/Modelle können auf der GPU nicht passen. Eng = kaum auszuführen, komfortabel = läuft problemlos mit Spielraum.
86,400,000,000 tokens

"In our protocol we pre-train with 2^20 (1,048,576) tokens per batch for 550k steps amounting to 576B total tokens." 1 token ~ 0.75 words

Training-Berechnung

Die Berechnungen, die zum Trainieren des Modells durchgeführt werden, gemessen in Gleitkommaoperationen. Es ist ein Maß dafür, was der Trainingslauf gekostet hat, nicht dafür, wie schnell das fertige Modell Ihnen antwortet.

Training-Berechnung
8.2 × 10²² FLOP

Table 4 https://arxiv.org/ftp/arxiv/papers/2104/2104.10350.pdf

Wie es etabliert wurde
Third-party estimation

Der Trainingslauf

Was es physisch benötigte, um zu trainieren: welche Chips, wie viele, wie lange und was das aus der Steckdose zog.

Trainingshardware
Google TPU v3
Verwendete Chips
1,024
Chip-hours
663,552
Wall-clock time
648 hours (27 days)

see table 4 in https://arxiv.org/ftp/arxiv/papers/2104/2104.10350.pdf

Hardware utilisation
HFU 28.0%

Table 4 in https://arxiv.org/pdf/2104.10350 gives measured performance of 34.4 TFLOP/s, vs. peak achievable FLOP/s of 123 TFLOP/s on the TPUv3 being used. HFU = 34.4/123 = 0.27967

Leistungsaufnahme
935.4 kW
Compute cost
$145,101

Verfügbarkeit

Ob Sie das Modell erhalten und es auf Ihrer eigenen Hardware ausführen können, was entscheidet, ob eine der Grafikkartenfiguren auf dieser Seite zutrifft.

Gewichte
Open — downloadable
Modellzugriff
Open weights (unrestricted)
Trainingscode
Unreleased

Apache 2 for weights: https://huggingface.co/google/switch-c-2048 paper links to this repo but not clear that the training hyperparams for Switch are here: https://github.com/google-research/t5x

Wie es klassifiziert ist

Labels, die auf das Quell-Dataset angewendet werden, wenn bemerkenswerte Modelle verfolgt werden, und wie zuversichtlich es in den Eintrag ist.

Frontier model
Yes
Warum es verfolgt wird
Highly cited,SOTA improvement

" On ANLI (Nie et al., 2019), Switch XXL improves over the prior state-of-the-art to get a 65.7 accuracy versus the prior best of 49.4 (Yang et al., 2020)... Finally, we also conduct an early examination of the model’s knowledge with three closed-book knowledge-based tasks: Natural Questions, WebQuestions and TriviaQA, without additional pre-training using Salient Span Masking (Guu et al., 2020). In all three cases, we observe improvements over the prior stateof-the-art T5-XXL model (without SSM…

Modell/Modelle können auf der GPU ausgeführt werden. Es passt nicht / wird nicht passen. Eng = kaum ausführbar, komfortabel = läuft problemlos mit Spielraum.
Confident
3,888

Model/Modelle laufen nicht / wird nicht passen

Woher dieser Datensatz stammt und wann er zuletzt überprüft wurde.

Referenz
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Modell/Modelle können von der GPU nicht ausgeführt werden. Es passt nicht. Zu eng. Lauffähig. Bequem.
25 May 2026

Was die Zahlen bedeuten

Was Sie benötigen, um es auszuführen

At 1.6T parameters, Switch is beyond what any single graphics card holds. Running it means either splitting it across several cards or renting hardware built for the job — 0 of the cards we track can hold it on their own, and all of them are datacentre parts.

Hintergrund

Switch was published by Google, in United States of America, in January 2021. The organisation is categorised as industry.

It works in Language, and is recorded as doing text autocompletion.

The weights are published, so it can be downloaded and run on your own hardware indefinitely, offline, with no account attached.

Training und Herkunft

The training run consumed about 8.2 × 10²² FLOP, on Google TPU v3. That figure describes the cost of creating it and has no bearing on how quickly it generates text.

Around 86,400,000,000 tokens went into training it.

It is tracked in the underlying dataset for one reason in particular: highly cited,SOTA improvement.

Schritt für Schritt

Wie man eine GPU für Switch

Die obige Tabelle hat bereits jede Karte, für die wir Spezifikationen haben, mit diesem Modell bewertet. Um zu Ihrer Antwort zu gelangen, sind sechs Schritte erforderlich..

  1. 01

    Lies zuerst die Speicherzahl

    Look at what Switch actually needs. No amount of processing power compensates for a card that cannot hold it.

  2. 02

    Entscheide, wie lange deine Gespräche laufen

    The conversation occupies memory too, and grows as it goes. Set the slider to the length you expect: at long context Switch can slip off a card that handles short questions easily.

  3. 03

    Entscheiden Sie, wie viel Kompression Sie akzeptieren.

    Compression is what makes Switch fit smaller cards, at some cost in accuracy. A minimum quality removes the ones that go too far.

  4. 04

    Nach Geschwindigkeit sortieren

    Sort by speed to see how cards rank for Switch. It will not match a gaming ordering — generation is bound by memory bandwidth.

  5. 05

    Modelle passen nicht / werden nicht passen.

    The fit column separates cards that just manage Switch from those with room to spare. Buy for the second if the context might grow.

  6. 06

    Sieh dir an, was diese Karte noch laufen kann.

    Following a card through to its own page shows every other model it can hold, which is the question that follows once Switch is settled.

Antworten

Switch — Häufige Fragen

01

What is Switch used for?

Switch works in Language, and is recorded as handling text autocompletion. These are the areas it was designed around; they describe intent rather than a hard boundary.

02

Where can I download Switch?

The weights for Switch are published, though we do not hold a repository link for it. This site calculates hardware requirements rather than hosting model files.

03

How much compute was used to train Switch?

Around 8.2 × 10²² FLOP, on Google TPU v3. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.

04

Can I run Switch if it does not fit in my GPU?

It can be split between the card and system memory, but Switch generates painfully slowly that way — the nearest miss we calculate is short by 691.9 GB. Nothing on this page assumes offloading.

05

Would two GPUs run Switch faster?

Two cards buy memory rather than speed. That matters for Switch only if one card cannot hold it — 0 can, so a second adds little.

06

Why does the quantisation differ between cards for Switch?

A larger card holds a more accurate copy. Across the cards that run Switch, 1 compression levels are used; the floor control above pins it to one.

07

How accurate are these Switch speed estimates?

These are estimates with real error bars. The fastest result here, the range beneath each figure, could reasonably land anywhere in its published range depending on which runtime you use.

08

Is Switch open source?

Its weights are published, so Switch can be downloaded and run on your own hardware. Note that open weights is not the same as open source in the full sense — it says nothing about the training data, the training code, or the commercial terms attached.

09

How many parameters does Switch have?

Switch has 1.6T parameters. "Combining expert, model and data parallelism, we design two large Switch Transformer models, one with 395 billion and 1.6 trillion parameters" Table 9 gives more precise count of 1571B parameters. That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.

10

Who created Switch?

Switch was published by Google, based in United States of America, categorised as industry.

11

When was Switch released?

Switch was published in January 2021. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.

Quelle

Originalveröffentlichung

Modell/Modelle läuft nicht / wird nicht passen 25 May 2026

Die andere Richtung

Von der anderen Seite betrachten?

Diese Seite beginnt vom Modell. Wenn Sie bereits eine Karte besitzen und alles erfahren möchten, was sie ausführen kann., Modell/Modelle laufen GPU AI-Modell. passt nicht/wird nicht passen. eng = kaum lauffähig, komfortabel = läuft problemlos mit Spielraum..