ST-MoE
Nincs becslés
Ennekhez a modellhez nincs hardverkövetelmény
A modellhez tartozó súlyokat nem tették közzé, így azt semmilyen méretben nem lehet letölteni vagy saját hardveren futtatni. Csak az adattovábbítóján keresztül érhető el, és semmilyen grafikus kártya nem változtat ezen..
Felvételen
Teljes specifikáció
Minden adatok ezen a modellen. A legtöbb leírás arról szól, hogyan volt kiképzve, nem pedig arról, hogyan fut — hasznos háttérinformáció arra vonatkozóan, hogy mennyi munka került bele, és hogyan viszonyul más méretű modellekhez..
Eredet
Ki építette ezt a modellt, hol, és mikor publikálták.
- Szervezet
- Google,Google Brain,Google Research
- Szervezet típusa
- Industry,Industry,Industry
- Ország
- United States of America
- Kiadva
- 17 February 2022
- Szerzők
- Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, William Fedus
Mit csinál
A modell által épített probléma területek. Egy modell több különbözőt is tartalmazhat.
- Irányítószám
- Language
- Feladat
- Language modeling/generation
- Numerical format
- BF16
Méret
Mekkora a modell és mennyi adatot használtak a betanításához. A paraméterek azok a számok, amelyek meghatározzák, hogy illeszkedik-e egy adott grafikus kártyára.
- Paraméterek
- 269B
- Képzési adatok
- 1,500,000,000,000 tokens
- Tételméret
- 1,000,000
269B. it's called ST-MoE-32B because it's equivalent to a 32B dense model.
"We pre-train for 1.5T tokens on a mixture of English-only C4 dataset (Raffel et al., 2019) and the dataset from GLaM (Du et al., 2021) summarized in Appendix E"
" We use 1M tokens per batch"
Képzés számítás
A modell betanításához végzett aritmetikai műveletek, amelyet lebegőpontos műveletekben mérnek. Ez a tréning futásának költségét méri, nem azt, hogy mennyire gyorsan válaszol a kész modell.
- Képzés számítás
- 2.9 × 10²³ FLOP
- Hogyan jött létre
- Operation counting
The paper claims "scaling a sparse model to 269B parameters, with a computational cost comparable to a 32B dense encoder-decoder". If this is true for training cost, then 6*32e9*1.5e12 = 2.9e23
Elérhetőség
Hogy megszerezheted-e a modellt és futtathatod a saját hardvereden, ami eldönti, hogy a ezen az oldalon szereplő grafikus kártyák adatai alkalmazhatóak-e.
- Súlyok
- Closed — provider access only
- Modell hozzáférés
- Unreleased
- Képzés kód
- Open source
Apache License 2.0 Code for our models is available at https://github.com/tensorflow/mesh/blob/master/mesh_tensorflow/transformer/moe.py
Hogyan van besorolva
Címkék, amelyeket a forráshalmaz alkalmaz, amikor figyelemmel kíséri a figyelemre méltó modelleket, és mennyire biztos a bejegyzésben.
- Valószínűleg 10²³ FLOP felett
- Yes
- Miért követik nyomon
- SOTA improvement
- Rögzítési bizalom
- Likely
- Citations
- 375
"ST-MoE-32B improves the current state-of-the-art on the test server submissions for both ARC Easy (92.7 → 94.8) and ARC Challenge (81.4 → 86.5)."
Források
Honnan származik ez a rekord és mikor ellenőrizték utoljára.
- Referencia
- ST-MoE: Designing Stable and Transferable Sparse Expert Models
- Utoljára frissítve
- 25 May 2026
Mit jelentek a számok
Background
ST-MoE was published by Google,Google Brain,Google Research, in United States of America, in February 2022. industry,Industry,Industry is the category the publisher falls under.
It works in Language, and is recorded as doing language modeling/generation.
Because the weights are not available, none of the hardware figures elsewhere on this site apply to it.
Hogyan lett betanítva
Producing it required around 2.9 × 10²³ FLOP of arithmetic, which is a statement about the training budget rather than about inference.
The training set ran to roughly 1,500,000,000,000 tokens.
Its inclusion criterion is sOTA improvement.
Válaszok
ST-MoE — Gyakran ismételt kérdések
When was ST-MoE released?
ST-MoE was published in February 2022. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.
What is ST-MoE used for?
ST-MoE works in Language, and is recorded as handling language modeling/generation. Models frequently carry more than one of each, and the tags describe purpose rather than capability limits.
How much compute was used to train ST-MoE?
Around 2.9 × 10²³ FLOP. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.
What GPU do I need to run ST-MoE?
None. ST-MoE is a closed model — its weights were never published, so it cannot be downloaded or run on your own hardware at any price. It is reachable only through its provider.
Is ST-MoE open source?
No. ST-MoE has not had its weights published, so it exists only as a service controlled by its owner.
How many parameters does ST-MoE have?
ST-MoE has 269B parameters. 269B. it's called ST-MoE-32B because it's equivalent to a 32B dense model. That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.
Who created ST-MoE?
ST-MoE was published by Google,Google Brain,Google Research, based in United States of America, categorised as industry,Industry,Industry.
A másik irány
Másik oldalról nézve?
Ez az oldal a modellel kezdődik. Ha már rendelkezik egy kártyával és szeretné tudni, hogy minden mit fog futtatni, kérjük, folytassa., kezdje a hardverrel inkább.