GLaM
Keine Schätzung
Keine Hardwareanforderungen für dieses Modell
Die Gewichte für dieses Modell wurden nicht veröffentlicht, sodass es weder heruntergeladen noch auf Ihrer eigenen Hardware in irgendeiner Größe ausgeführt werden kann. Es ist nur über seinen Anbieter zugänglich, und keine Grafikkarte ändert daran etwas.
Aufzeichnung
Vollständige Spezifikation
Alles aufzeichnungsfähig für dieses Modell. Die meisten Informationen beschreiben, wie es trainiert wurde, anstatt wie es läuft — ein nützlicher Kontext, um zu beurteilen, wie viel Arbeit investiert wurde und wie es im Vergleich zu auf einer anderen Skala entwickelten Modellen abschneidet..
Ursprung
Wer hat dieses Modell gebaut, wo und wann es veröffentlicht wurde.
- Organisation
- Organisationstyp
- Industry
- Land
- United States of America
- Veröffentlicht
- 13 December 2021
- Autoren
- Nan Du, Yanping Huang, Andrew M. Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, Barret Zoph, Liam Fedus, Maarten Bosma, Zongwei Zhou, Tao Wang, Yu Emma Wang, Kellie Webster, Marie Pellat, Kevin Robinson, Kathleen Meier-Hellstern, Toju Duke, Lucas Dixon, Kun Zhang, Quoc V Le, Yonghui Wu, Zhifeng Chen, Claire Cui
Was es tut
Die Problemfelder, für die das Modell entwickelt wurde. Ein Modell kann mehrere von jedem tragen.
- Domäne
- Language
- Aufgabe
- Language modeling/generation, Question answering
- Numerical format
- BF16
Größe
Wie groß das Modell ist und wie viele Daten es trainiert wurde. Die Parameter sind die Größe, die entscheidet, ob es auf eine bestimmte Grafikkarte passt.
- Parameter
- 1.2T
- Trainingsdaten
- 600,000,000,000 tokens
- Batchgröße
- 1,000,000
1.2 trillion parameters
The dataset is made of 1.6 trillion tokens, but later in the paper they say they only train the largest model for 600b tokens. 600b / 0.75 words/token = 800b words. "The complete GLaM training using 600B tokens consumes only 456 MWh and emits 40.2 net tCO2e."
"We use a maximum sequence length of 1024 tokens, and pack each input example to have up to 1 million tokens per batch."
Training-Computing
Die zur Ausbildung des Modells durchgeführte Arithmetik, gemessen in Gleitkommaoperationen. Es ist ein Maß dafür, was der Trainingslauf gekostet hat, nicht wie schnell das fertige Modell Ihnen antwortet.
- Training-Computing
- 3.6 × 10²³ FLOP
- Wie es etabliert wurde
- Operation counting,Hardware
The network activates 96.6 billion parameters per token and trained for 600B tokens. 6 * 600B * 96.6B = 3.478e23 Digitizing figure 4 (d) indicates 139.67 TPU-years of training. 2.75e14 * 139.67 * 365.25 * 24 * 3600 * 0.3 = 3.636e23 Since these are close, we will use the 6NC estimate and derive hardware utilization from the training time information. Later they say they measured 326W power usage per chip, which could maybe be used to estimate utilization.
Der Trainingslauf
Was physisch notwendig war, um zu trainieren: welche Chips, wie viele, wie lange und was das aus der Steckdose zog.
- Trainingshardware
- Google TPU v4
- Verwendete Chips
- 1,024
- Chip-Stunden
- 1,398,784
- Wand-Uhrzeit
- 1,366 hours (56.9 days)
- Hardware utilisation
- MFU 28.7%
- Stromverbrauch
- 701.5 kW
- Rechenkosten
- $541,437
Note that they give several energy estimates. Use the complete training figures for 600B tokens, not the GPT-3 comparison values with 280B tokens. "326W measured system power per TPU-v4 chip" "The complete GLaM training using 600B tokens consumes only 456 MWh" 1024 TPU v4 chips (456 MWh) / (326W/chip * 1024 chips) = 1366 hours
6ND: 6 * 600B * 96.6B = 3.478e23 Hardware: Digitizing figure 4 (d) indicates 139.67 TPU-years of training. 2.75e14 * 139.67 * 365.25 * 24 * 3600 = 1.212e24 at full utilization Implies utilization of 0.287 MFU = 0.2870
Verfügbarkeit
Ob Sie das Modell erhalten und es auf Ihrer eigenen Hardware ausführen können, entscheidet darüber, ob eine der Grafikkartenfiguren auf dieser Seite zutrifft.
- Gewichte
- Closed — provider access only
- Modellzugang
- Unreleased
- Trainingscode
- Unreleased
Wie es klassifiziert ist
Labels der Quelldatenbank, die beim Verfolgen bemerkenswerter Modelle angewendet werden, und wie zuversichtlich sie in den Eintrag ist.
- Foundation model
- Yes
- Wahrscheinlich über 10²³ FLOP
- Yes
- Warum es nachverfolgt wird
- SOTA improvement
- Aufzeichnungszuversicht
- Confident
- Zitationen
- 1,198
"As shown in Table 5, GLaM (64B/64E) is better than the dense model and outperforms the previous finetuned state-of-the-art (SOTA) on this dataset in the open-domain setting"
Quellen
Woher diese Aufzeichnung stammt und wann sie zuletzt überprüft wurde.
- Referenz
- GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
- Zuletzt aktualisiert
- 25 May 2026
Was die Zahlen bedeuten
Hintergrund
GLaM was published by Google, in United States of America, in December 2021. It comes out of industry.
Es funktioniert in Sprache und wird als Sprachmodellierung/Generierung, Beantwortung von Fragen aufgezeichnet.
This is a closed model: the trained values stayed with whoever produced them, and there is no local version to run.
Wie es trainiert wurde
Producing it required around 3.6 × 10²³ FLOP of arithmetic, on Google TPU v4, which is a statement about the training budget rather than about inference.
It was trained on about 600,000,000,000 tokens of text.
Es wird im zugrunde liegenden Datensatz aus einem bestimmten Grund verfolgt: sOTA-Verbesserung.
Antworten
GLaM — Häufig gestellte Fragen
How much compute was used to train GLaM?
Around 3.6 × 10²³ FLOP, on Google TPU v4. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.
What GPU do I need to run GLaM?
None. GLaM is a closed model — its weights were never published, so it cannot be downloaded or run on your own hardware at any price. It is reachable only through its provider.
Is GLaM open source?
No. GLaM has not had its weights published, so it exists only as a service controlled by its owner.
How many parameters does GLaM have?
GLaM has 1.2T parameters. 1.2 trillion parameters. That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.
Who created GLaM?
GLaM was published by Google, based in United States of America, categorised as industry.
When was GLaM released?
GLaM was published in December 2021. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.
What is GLaM used for?
GLaM works in Language, and is recorded as handling language modeling/generation, Question answering. Models frequently carry more than one of each, and the tags describe purpose rather than capability limits.
Die andere Richtung
Wenn Sie es von der anderen Seite betrachten?
Diese Seite beginnt beim Modell. Wenn Sie bereits eine Karte besitzen und alles erfahren möchten, was es ausführen kann, Beginnen Sie stattdessen mit der Hardware.