Seq2Seq LSTM

Chiuso pesi Google 1.9B Parametri September 2014

Nessuna stima

Nessun requisito hardware per questo modello

I pesi di questo modello non sono stati pubblicati, quindi non può essere scaricato o eseguito sul vostro hardware a qualsiasi dimensione. È accessibile solo attraverso il suo fornitore, e nessuna scheda grafica può cambiare questo.

In registrazione

Specifiche complete

Tutto ciò che riguarda questo modello. La maggior parte delle informazioni descrive come è stato addestrato piuttosto che come viene eseguito — un contesto utile per valutare la quantità di lavoro impiegata e come si confronta con modelli realizzati su scala diversa.

Origine

Chi ha costruito questo modello, dove e quando è stato pubblicato.

Organizzazione
Google
Tipo di organizzazione
Industry
Paese
United States of America
Pubblicato
10 September 2014
Autori
Ilya Sutskever, Oriol Vinyals, Quoc V. Le

Cosa fa

Le aree problematiche per cui è stato costruito il modello. Un modello può contenere diversi di ciascuno.

Dominio
Language
Compito
Translation

Dimensione

Quanto è grande il modello e su quanti dati è stato addestrato. I parametri sono il valore che decide se si adatta a una determinata scheda grafica.

Parametri
1.9B

The resulting LSTM has 384M parameters of which 64M are pure recurrent connections (32M for the “encoder” LSTM and 32M for the “decoder” LSTM). The paper uses an ensemble of 5 LSTMs.

Dati di addestramento
870,000,000 tokens

[WORDS] "We used the WMT’14 English to French dataset. We trained our models on a subset of 12M sentences consisting of 348M French words and 304M English words, which is a clean “selected” subset from [29]."

Epoche
7.5

Addestramento computazionale

L'aritmetica eseguita per addestrare il modello, misurata in operazioni in virgola mobile. È una misura di quanto costato l'addestramento, non di quanto velocemente il modello finito ti risponde.

Addestramento computazionale
5.6 × 10¹⁹ FLOP

384E+6 parameters * 2 FLOP/parameter * (348E+6 + 304E+6 points per epoch) * 7.5 epochs * 3 FLOP/point ~= 1.126656e+19 FLOP Times 5 independent models in ensemble => 5.6E+19 FLOP If we assume NVIDIA K40 (in use at the time): 10 days * 24 * 60 * 60 seconds/day * 8 GPUs * 33% * 5e12 FLOP/s * 5 models in ensemble ~= 5.7E+19 FLOP Authors of "AI and Memory Wall" estimated model's training compute as 11,000 PFLOPS = 1.1*10^19 FLOPS (https://github.com/amirgholami/ai_and_memory_wall)

Come è stato stabilito
Operation counting,Hardware,Third-party estimation

Il training run

Cosa è stato fisicamente necessario per addestrare: quali chip, quanti, per quanto tempo e cosa ha richiesto dalla rete elettrica.

Tempo di orologio muro
240 hours (10 days)

Training took about 10 days

Come è classificato

Etichette che il dataset sorgente applica quando si monitorano modelli notevoli e quanto è fiducioso nell'entry.

Frontier model
Yes
Perché viene tracciato
Highly cited
Registrare fiducia
Confident
Citazioni
22,025

Fonti

Da dove proviene questo record e quando è stato controllato l'ultima volta.

Riferimento
Sequence to Sequence Learning with Neural Networks
Ultimo aggiornamento
25 May 2026

Cosa significano i numeri

Where it came from

Seq2Seq LSTM was published by Google, in United States of America, in September 2014. industry is the category the publisher falls under.

It works in Language, and is recorded as doing translation.

I suoi pesi non sono mai stati pubblicati, quindi può essere raggiunto solo attraverso il suo fornitore. Nessuna scheda grafica cambia ciò.

Cosa ci è voluto per costruirlo

Producing it required around 5.6 × 10¹⁹ FLOP of arithmetic, which is a statement about the training budget rather than about inference.

The training set ran to roughly 870,000,000 tokens.

Its inclusion criterion is highly cited.

Risposte

Seq2Seq LSTM — Domande frequenti

01

Is Seq2Seq LSTM open source?

The licensing for Seq2Seq LSTM was never recorded in our source data. We treat unstated licensing as closed, because an unrecorded licence is not one to rely on.

02

How many parameters does Seq2Seq LSTM have?

Seq2Seq LSTM has 1.9B parameters. The resulting LSTM has 384M parameters of which 64M are pure recurrent connections (32M for the “encoder” LSTM and 32M for the “decoder” LSTM). The paper uses an ensemble of 5 LSTMs. That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.

03

Who created Seq2Seq LSTM?

Seq2Seq LSTM was published by Google, based in United States of America, categorised as industry.

04

When was Seq2Seq LSTM released?

Seq2Seq LSTM was published in September 2014. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.

05

What is Seq2Seq LSTM used for?

Seq2Seq LSTM works in Language, and is recorded as handling translation. Models frequently carry more than one of each, and the tags describe purpose rather than capability limits.

06

Quanto calcolo è stato utilizzato per addestrare Seq2Seq LSTM?

Around 5.6 × 10¹⁹ FLOP. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.

07

What GPU do I need to run Seq2Seq LSTM?

None. Seq2Seq LSTM is a closed model — its weights were never published, so it cannot be downloaded or run on your own hardware at any price. It is reachable only through its provider.

Fonte

Pubblicazione originale

Ultimo aggiornamento del record 25 May 2026

L'altra direzione

Guardando la questione dall'altra parte?

Questa pagina inizia dal modello. Se possedete già una scheda e desiderate conoscere tutto ciò che può eseguire, inizia dall'hardware invece.