Seq2Seq LSTM
Brak szacunków
Brak wymagań sprzętowych dla tego modelu
Wagi tego modelu nie zostały opublikowane, dlatego nie można ich pobrać ani uruchomić na własnym sprzęcie w żadnym rozmiarze. Jest dostępny wyłącznie poprzez jego dostawcę, a żadna karta graficzna tego nie zmieni.
Na nagraniu
Pełna specyfikacja
Wszystko, co jest dostępne na temat tego modelu. Większość z tego opisuje, jak został wytrenowany, a nie jak działa — przydatny kontekst do oceny, ile pracy w to włożono i jak wypada na tle modeli zbudowanych na innym poziomie skali..
Pochodzenie
Kto zbudował ten model, gdzie i kiedy został opublikowany.
- Organizacja
- Typ organizacji
- Industry
- Kraj
- United States of America
- Opublikowane
- 10 September 2014
- Autorzy
- Ilya Sutskever, Oriol Vinyals, Quoc V. Le
Co to robi
Obszary problemowe, dla których stworzono model. Model może zawierać kilka z każdej.
- Domena
- Language
- Zadanie
- Translation
Rozmiar
Jakiej wielkości jest model i na ile danych był szkolony. Parametry to liczba, która decyduje, czy mieści się na danej karcie graficznej.
- Parametry
- 1.9B
- Dane treningowe
- 870,000,000 tokens
- Epoki
- 7.5
The resulting LSTM has 384M parameters of which 64M are pure recurrent connections (32M for the “encoder” LSTM and 32M for the “decoder” LSTM). The paper uses an ensemble of 5 LSTMs.
[WORDS] "We used the WMT’14 English to French dataset. We trained our models on a subset of 12M sentences consisting of 348M French words and 304M English words, which is a clean “selected” subset from [29]."
Obliczenia szkoleniowe
Aritmetyka wykonywana w celu wytrenowania modelu, mierzone w operacjach zmiennoprzecinkowych. Jest to miara kosztu przeprowadzonego treningu, a nie tego, jak szybko gotowy model odpowiada na twoje zapytania.
- Obliczenia szkoleniowe
- 5.6 × 10¹⁹ FLOP
- Jak to zostało ustalone
- Operation counting,Hardware,Third-party estimation
384E+6 parameters * 2 FLOP/parameter * (348E+6 + 304E+6 points per epoch) * 7.5 epochs * 3 FLOP/point ~= 1.126656e+19 FLOP Times 5 independent models in ensemble => 5.6E+19 FLOP If we assume NVIDIA K40 (in use at the time): 10 days * 24 * 60 * 60 seconds/day * 8 GPUs * 33% * 5e12 FLOP/s * 5 models in ensemble ~= 5.7E+19 FLOP Authors of "AI and Memory Wall" estimated model's training compute as 11,000 PFLOPS = 1.1*10^19 FLOPS (https://github.com/amirgholami/ai_and_memory_wall)
Trening
Co fizycznie zajęło szkolenie: które chipy, ile ich, na jak długo i ile to pobrało z sieci.
- Wall-clock time
- 240 hours (10 days)
Training took about 10 days
Jak jest klasyfikowane
Etykiety stosowane przez źródłowy zbiór danych podczas śledzenia znaczących modeli oraz jak pewny jest w danym wpisie.
- Frontier model
- Yes
- Dlaczego jest śledzone
- Highly cited
- Zaufanie do nagrania
- Confident
- Cytacje
- 22,025
Źródła
Skąd pochodzi ten rekord i kiedy był ostatnio sprawdzany.
- Referencja
- Sequence to Sequence Learning with Neural Networks
- Ostatnia aktualizacja
- 25 May 2026
Co oznaczają liczby
Skąd to przyszło
Seq2Seq LSTM was published by Google, in United States of America, in September 2014. industry is the category the publisher falls under.
It works in Language, and is recorded as doing translation.
Jego wagi nigdy nie zostały opublikowane, więc można go osiągnąć tylko przez jego dostawcę. Żadna karta graficzna tego nie zmieni.
Co poszło w budowę tego
Producing it required around 5.6 × 10¹⁹ FLOP of arithmetic, which is a statement about the training budget rather than about inference.
The training set ran to roughly 870,000,000 tokens.
Its inclusion criterion is highly cited.
Odpowiedzi
Seq2Seq LSTM — Często zadawane pytania
Is Seq2Seq LSTM open source?
The licensing for Seq2Seq LSTM was never recorded in our source data. We treat unstated licensing as closed, because an unrecorded licence is not one to rely on.
How many parameters does Seq2Seq LSTM have?
Seq2Seq LSTM has 1.9B parameters. The resulting LSTM has 384M parameters of which 64M are pure recurrent connections (32M for the “encoder” LSTM and 32M for the “decoder” LSTM). The paper uses an ensemble of 5 LSTMs. That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.
Who created Seq2Seq LSTM?
Seq2Seq LSTM was published by Google, based in United States of America, categorised as industry.
When was Seq2Seq LSTM released?
Seq2Seq LSTM was published in September 2014. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.
What is Seq2Seq LSTM used for?
Seq2Seq LSTM works in Language, and is recorded as handling translation. Models frequently carry more than one of each, and the tags describe purpose rather than capability limits.
How much compute was used to train Seq2Seq LSTM?
Around 5.6 × 10¹⁹ FLOP. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.
What GPU do I need to run Seq2Seq LSTM?
None. Seq2Seq LSTM is a closed model — its weights were never published, so it cannot be downloaded or run on your own hardware at any price. It is reachable only through its provider.
W innym kierunku
Patrząc na to z drugiej strony?
Ta strona zaczyna się od modelu. Jeśli już posiadasz kartę i chcesz wiedzieć, wszystko co będzie ona obsługiwać, Rozpocznij od sprzętu..