Seq2Seq LSTM
Nincs becslés
Ennekhez a modellhez nincs hardverkövetelmény
A modellhez tartozó súlyokat nem tették közzé, így azt semmilyen méretben nem lehet letölteni vagy saját hardveren futtatni. Csak az adattovábbítóján keresztül érhető el, és semmilyen grafikus kártya nem változtat ezen..
Felvételen
Teljes specifikáció
Minden adatok ezen a modellen. A legtöbb leírás arról szól, hogyan volt kiképzve, nem pedig arról, hogyan fut — hasznos háttérinformáció arra vonatkozóan, hogy mennyi munka került bele, és hogyan viszonyul más méretű modellekhez..
Eredet
Ki építette ezt a modellt, hol, és mikor publikálták.
- Szervezet
- Szervezet típusa
- Industry
- Ország
- United States of America
- Kiadva
- 10 September 2014
- Szerzők
- Ilya Sutskever, Oriol Vinyals, Quoc V. Le
Mit csinál
A modell által épített probléma területek. Egy modell több különbözőt is tartalmazhat.
- Irányítószám
- Language
- Feladat
- Translation
Méret
Mekkora a modell és mennyi adatot használtak a betanításához. A paraméterek azok a számok, amelyek meghatározzák, hogy illeszkedik-e egy adott grafikus kártyára.
- Paraméterek
- 1.9B
- Képzési adatok
- 870,000,000 tokens
- Epókák
- 7.5
The resulting LSTM has 384M parameters of which 64M are pure recurrent connections (32M for the “encoder” LSTM and 32M for the “decoder” LSTM). The paper uses an ensemble of 5 LSTMs.
[WORDS] "We used the WMT’14 English to French dataset. We trained our models on a subset of 12M sentences consisting of 348M French words and 304M English words, which is a clean “selected” subset from [29]."
Képzés számítás
A modell betanításához végzett aritmetikai műveletek, amelyet lebegőpontos műveletekben mérnek. Ez a tréning futásának költségét méri, nem azt, hogy mennyire gyorsan válaszol a kész modell.
- Képzés számítás
- 5.6 × 10¹⁹ FLOP
- Hogyan jött létre
- Operation counting,Hardware,Third-party estimation
384E+6 parameters * 2 FLOP/parameter * (348E+6 + 304E+6 points per epoch) * 7.5 epochs * 3 FLOP/point ~= 1.126656e+19 FLOP Times 5 independent models in ensemble => 5.6E+19 FLOP If we assume NVIDIA K40 (in use at the time): 10 days * 24 * 60 * 60 seconds/day * 8 GPUs * 33% * 5e12 FLOP/s * 5 models in ensemble ~= 5.7E+19 FLOP Authors of "AI and Memory Wall" estimated model's training compute as 11,000 PFLOPS = 1.1*10^19 FLOPS (https://github.com/amirgholami/ai_and_memory_wall)
A képzési futam
A kiképzéshez fizikailag szükséges volt: milyen chipek, hány, meddig és mennyi áramot vett le a hálózatról.
- Wall-clock time
- 240 hours (10 days)
Training took about 10 days
Hogyan van besorolva
Címkék, amelyeket a forráshalmaz alkalmaz, amikor figyelemmel kíséri a figyelemre méltó modelleket, és mennyire biztos a bejegyzésben.
- Frontier model
- Yes
- Miért követik nyomon
- Highly cited
- Rögzítési bizalom
- Confident
- Citations
- 22,025
Források
Honnan származik ez a rekord és mikor ellenőrizték utoljára.
- Referencia
- Sequence to Sequence Learning with Neural Networks
- Utoljára frissítve
- 25 May 2026
Mit jelentek a számok
Where it came from
Seq2Seq LSTM was published by Google, in United States of America, in September 2014. industry is the category the publisher falls under.
It works in Language, and is recorded as doing translation.
Súlyait soha nem publikálták, így csak a szolgáltatóján keresztül érhető el. Nincs grafikus kártya, amely ezt megváltoztatja.
Mi került a megépítésébe
Producing it required around 5.6 × 10¹⁹ FLOP of arithmetic, which is a statement about the training budget rather than about inference.
The training set ran to roughly 870,000,000 tokens.
Its inclusion criterion is highly cited.
Válaszok
Seq2Seq LSTM — Gyakran ismételt kérdések
Is Seq2Seq LSTM open source?
The licensing for Seq2Seq LSTM was never recorded in our source data. We treat unstated licensing as closed, because an unrecorded licence is not one to rely on.
How many parameters does Seq2Seq LSTM have?
Seq2Seq LSTM has 1.9B parameters. The resulting LSTM has 384M parameters of which 64M are pure recurrent connections (32M for the “encoder” LSTM and 32M for the “decoder” LSTM). The paper uses an ensemble of 5 LSTMs. That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.
Who created Seq2Seq LSTM?
Seq2Seq LSTM was published by Google, based in United States of America, categorised as industry.
When was Seq2Seq LSTM released?
Seq2Seq LSTM was published in September 2014. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.
What is Seq2Seq LSTM used for?
Seq2Seq LSTM works in Language, and is recorded as handling translation. Models frequently carry more than one of each, and the tags describe purpose rather than capability limits.
How much compute was used to train Seq2Seq LSTM?
Around 5.6 × 10¹⁹ FLOP. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.
What GPU do I need to run Seq2Seq LSTM?
None. Seq2Seq LSTM is a closed model — its weights were never published, so it cannot be downloaded or run on your own hardware at any price. It is reachable only through its provider.
A másik irány
Másik oldalról nézve?
Ez az oldal a modellel kezdődik. Ha már rendelkezik egy kártyával és szeretné tudni, hogy minden mit fog futtatni, kérjük, folytassa., kezdje a hardverrel inkább.