Seq2Seq LSTM

Zamknięte wagi Google 1.9B Parametry September 2014

Brak szacunków

Brak wymagań sprzętowych dla tego modelu

Wagi tego modelu nie zostały opublikowane, dlatego nie można ich pobrać ani uruchomić na własnym sprzęcie w żadnym rozmiarze. Jest dostępny wyłącznie poprzez jego dostawcę, a żadna karta graficzna tego nie zmieni.

Na nagraniu

Pełna specyfikacja

Wszystko, co jest dostępne na temat tego modelu. Większość z tego opisuje, jak został wytrenowany, a nie jak działa — przydatny kontekst do oceny, ile pracy w to włożono i jak wypada na tle modeli zbudowanych na innym poziomie skali..

Pochodzenie

Kto zbudował ten model, gdzie i kiedy został opublikowany.

Organizacja
Google
Typ organizacji
Industry
Kraj
United States of America
Opublikowane
10 September 2014
Autorzy
Ilya Sutskever, Oriol Vinyals, Quoc V. Le

Co to robi

Obszary problemowe, dla których stworzono model. Model może zawierać kilka z każdej.

Domena
Language
Zadanie
Translation

Rozmiar

Jakiej wielkości jest model i na ile danych był szkolony. Parametry to liczba, która decyduje, czy mieści się na danej karcie graficznej.

Parametry
1.9B

The resulting LSTM has 384M parameters of which 64M are pure recurrent connections (32M for the “encoder” LSTM and 32M for the “decoder” LSTM). The paper uses an ensemble of 5 LSTMs.

Dane treningowe
870,000,000 tokens

[WORDS] "We used the WMT’14 English to French dataset. We trained our models on a subset of 12M sentences consisting of 348M French words and 304M English words, which is a clean “selected” subset from [29]."

Epoki
7.5

Obliczenia szkoleniowe

Aritmetyka wykonywana w celu wytrenowania modelu, mierzone w operacjach zmiennoprzecinkowych. Jest to miara kosztu przeprowadzonego treningu, a nie tego, jak szybko gotowy model odpowiada na twoje zapytania.

Obliczenia szkoleniowe
5.6 × 10¹⁹ FLOP

384E+6 parameters * 2 FLOP/parameter * (348E+6 + 304E+6 points per epoch) * 7.5 epochs * 3 FLOP/point ~= 1.126656e+19 FLOP Times 5 independent models in ensemble => 5.6E+19 FLOP If we assume NVIDIA K40 (in use at the time): 10 days * 24 * 60 * 60 seconds/day * 8 GPUs * 33% * 5e12 FLOP/s * 5 models in ensemble ~= 5.7E+19 FLOP Authors of "AI and Memory Wall" estimated model's training compute as 11,000 PFLOPS = 1.1*10^19 FLOPS (https://github.com/amirgholami/ai_and_memory_wall)

Jak to zostało ustalone
Operation counting,Hardware,Third-party estimation

Trening

Co fizycznie zajęło szkolenie: które chipy, ile ich, na jak długo i ile to pobrało z sieci.

Wall-clock time
240 hours (10 days)

Training took about 10 days

Jak jest klasyfikowane

Etykiety stosowane przez źródłowy zbiór danych podczas śledzenia znaczących modeli oraz jak pewny jest w danym wpisie.

Frontier model
Yes
Dlaczego jest śledzone
Highly cited
Zaufanie do nagrania
Confident
Cytacje
22,025

Źródła

Skąd pochodzi ten rekord i kiedy był ostatnio sprawdzany.

Referencja
Sequence to Sequence Learning with Neural Networks
Ostatnia aktualizacja
25 May 2026

Co oznaczają liczby

Skąd to przyszło

Seq2Seq LSTM was published by Google, in United States of America, in September 2014. industry is the category the publisher falls under.

It works in Language, and is recorded as doing translation.

Jego wagi nigdy nie zostały opublikowane, więc można go osiągnąć tylko przez jego dostawcę. Żadna karta graficzna tego nie zmieni.

Co poszło w budowę tego

Producing it required around 5.6 × 10¹⁹ FLOP of arithmetic, which is a statement about the training budget rather than about inference.

The training set ran to roughly 870,000,000 tokens.

Its inclusion criterion is highly cited.

Odpowiedzi

Seq2Seq LSTM — Często zadawane pytania

01

Is Seq2Seq LSTM open source?

The licensing for Seq2Seq LSTM was never recorded in our source data. We treat unstated licensing as closed, because an unrecorded licence is not one to rely on.

02

How many parameters does Seq2Seq LSTM have?

Seq2Seq LSTM has 1.9B parameters. The resulting LSTM has 384M parameters of which 64M are pure recurrent connections (32M for the “encoder” LSTM and 32M for the “decoder” LSTM). The paper uses an ensemble of 5 LSTMs. That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.

03

Who created Seq2Seq LSTM?

Seq2Seq LSTM was published by Google, based in United States of America, categorised as industry.

04

When was Seq2Seq LSTM released?

Seq2Seq LSTM was published in September 2014. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.

05

What is Seq2Seq LSTM used for?

Seq2Seq LSTM works in Language, and is recorded as handling translation. Models frequently carry more than one of each, and the tags describe purpose rather than capability limits.

06

How much compute was used to train Seq2Seq LSTM?

Around 5.6 × 10¹⁹ FLOP. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.

07

What GPU do I need to run Seq2Seq LSTM?

None. Seq2Seq LSTM is a closed model — its weights were never published, so it cannot be downloaded or run on your own hardware at any price. It is reachable only through its provider.

Źródło

Publikacja oryginalna

Ostatnia aktualizacja rekordu 25 May 2026

W innym kierunku

Patrząc na to z drugiej strony?

Ta strona zaczyna się od modelu. Jeśli już posiadasz kartę i chcesz wiedzieć, wszystko co będzie ona obsługiwać, Rozpocznij od sprzętu..