ST-MoE

Zárt súlyok Google,Google Brain,Google Research 269B Paraméterek February 2022

Nincs becslés

Ennekhez a modellhez nincs hardverkövetelmény

A modellhez tartozó súlyokat nem tették közzé, így azt semmilyen méretben nem lehet letölteni vagy saját hardveren futtatni. Csak az adattovábbítóján keresztül érhető el, és semmilyen grafikus kártya nem változtat ezen..

Felvételen

Teljes specifikáció

Minden adatok ezen a modellen. A legtöbb leírás arról szól, hogyan volt kiképzve, nem pedig arról, hogyan fut — hasznos háttérinformáció arra vonatkozóan, hogy mennyi munka került bele, és hogyan viszonyul más méretű modellekhez..

Eredet

Ki építette ezt a modellt, hol, és mikor publikálták.

Szervezet
Google,Google Brain,Google Research
Szervezet típusa
Industry,Industry,Industry
Ország
United States of America
Kiadva
17 February 2022
Szerzők
Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, William Fedus

Mit csinál

A modell által épített probléma területek. Egy modell több különbözőt is tartalmazhat.

Irányítószám
Language
Feladat
Language modeling/generation
Numerical format
BF16

Méret

Mekkora a modell és mennyi adatot használtak a betanításához. A paraméterek azok a számok, amelyek meghatározzák, hogy illeszkedik-e egy adott grafikus kártyára.

Paraméterek
269B

269B. it's called ST-MoE-32B because it's equivalent to a 32B dense model.

Képzési adatok
1,500,000,000,000 tokens

"We pre-train for 1.5T tokens on a mixture of English-only C4 dataset (Raffel et al., 2019) and the dataset from GLaM (Du et al., 2021) summarized in Appendix E"

Tételméret
1,000,000

" We use 1M tokens per batch"

Képzés számítás

A modell betanításához végzett aritmetikai műveletek, amelyet lebegőpontos műveletekben mérnek. Ez a tréning futásának költségét méri, nem azt, hogy mennyire gyorsan válaszol a kész modell.

Képzés számítás
2.9 × 10²³ FLOP

The paper claims "scaling a sparse model to 269B parameters, with a computational cost comparable to a 32B dense encoder-decoder". If this is true for training cost, then 6*32e9*1.5e12 = 2.9e23

Hogyan jött létre
Operation counting

Elérhetőség

Hogy megszerezheted-e a modellt és futtathatod a saját hardvereden, ami eldönti, hogy a ezen az oldalon szereplő grafikus kártyák adatai alkalmazhatóak-e.

Súlyok
Closed — provider access only
Modell hozzáférés
Unreleased
Képzés kód
Open source

Apache License 2.0 Code for our models is available at https://github.com/tensorflow/mesh/blob/master/mesh_tensorflow/transformer/moe.py

Hogyan van besorolva

Címkék, amelyeket a forráshalmaz alkalmaz, amikor figyelemmel kíséri a figyelemre méltó modelleket, és mennyire biztos a bejegyzésben.

Valószínűleg 10²³ FLOP felett
Yes
Miért követik nyomon
SOTA improvement

"ST-MoE-32B improves the current state-of-the-art on the test server submissions for both ARC Easy (92.7 → 94.8) and ARC Challenge (81.4 → 86.5)."

Rögzítési bizalom
Likely
Citations
375

Források

Honnan származik ez a rekord és mikor ellenőrizték utoljára.

Referencia
ST-MoE: Designing Stable and Transferable Sparse Expert Models
Utoljára frissítve
25 May 2026

Mit jelentek a számok

Background

ST-MoE was published by Google,Google Brain,Google Research, in United States of America, in February 2022. industry,Industry,Industry is the category the publisher falls under.

It works in Language, and is recorded as doing language modeling/generation.

Because the weights are not available, none of the hardware figures elsewhere on this site apply to it.

Hogyan lett betanítva

Producing it required around 2.9 × 10²³ FLOP of arithmetic, which is a statement about the training budget rather than about inference.

The training set ran to roughly 1,500,000,000,000 tokens.

Its inclusion criterion is sOTA improvement.

Válaszok

ST-MoE — Gyakran ismételt kérdések

01

When was ST-MoE released?

ST-MoE was published in February 2022. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.

02

What is ST-MoE used for?

ST-MoE works in Language, and is recorded as handling language modeling/generation. Models frequently carry more than one of each, and the tags describe purpose rather than capability limits.

03

How much compute was used to train ST-MoE?

Around 2.9 × 10²³ FLOP. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.

04

What GPU do I need to run ST-MoE?

None. ST-MoE is a closed model — its weights were never published, so it cannot be downloaded or run on your own hardware at any price. It is reachable only through its provider.

05

Is ST-MoE open source?

No. ST-MoE has not had its weights published, so it exists only as a service controlled by its owner.

06

How many parameters does ST-MoE have?

ST-MoE has 269B parameters. 269B. it's called ST-MoE-32B because it's equivalent to a 32B dense model. That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.

07

Who created ST-MoE?

ST-MoE was published by Google,Google Brain,Google Research, based in United States of America, categorised as industry,Industry,Industry.

Forrás

Eredeti közzététel

Utolsó frissítés időpontja 25 May 2026

A másik irány

Másik oldalról nézve?

Ez az oldal a modellel kezdődik. Ha már rendelkezik egy kártyával és szeretné tudni, hogy minden mit fog futtatni, kérjük, folytassa., kezdje a hardverrel inkább.