Chinchilla

Poids fermés DeepMind 70B Paramètres March 2022

Aucune estimation

Aucun besoin matériel pour ce modèle

Les poids de ce modèle n'ont pas été publiés, il ne peut donc pas être téléchargé ou exécuté sur votre propre matériel, quelle que soit la taille. Il n'est accessible que via son fournisseur, et aucune carte graphique ne change cela.

Enregistrement terminé

Spécification complète

Tout ce qui concerne ce modèle. La plupart décrit comment il a été entraîné plutôt que comment il fonctionne — contexte utile pour évaluer le travail fourni et sa comparaison avec des modèles construits à une échelle différente.

Origine

Qui a construit ce modèle, où et quand il a été publié.

Organisation
DeepMind
Type d'organisation
Industry
Pays
United Kingdom of Great Britain and Northern Ireland
Publié
29 March 2022
Auteurs
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, Laurent Sifre

Ce qu'il fait

Les domaines problématiques pour lesquels le modèle a été conçu. Un modèle peut en porter plusieurs de chacun.

Domaine
Language
Tâche
Language modeling
Approach
Self-supervised learning
Format numérique
BF16

Taille

Quelle est la taille du modèle et combien de données il a été formé. Les paramètres sont le chiffre qui détermine s'il tient sur une carte graphique donnée.

Paramètres
70B

"We test this hypothesis by training a predicted compute-optimal model, \chinchilla, that uses the same compute budget as \gopher but with 70B parameters and 4× more more data. \chinchilla uniformly and significantly outperforms \Gopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks."

Données d'entraînement
1,400,000,000,000 tokens

Table 1 shows Chinchilla was training on 1.4 trillion tokens 1 token ~ 0.75 words

Époques
1
Batch size
3,000,000

Table 1. "1.5M → 3M"

Calcul de formation

L'arithmétique effectuée pour entraîner le modèle, mesurée en opérations à virgule flottante. C'est une mesure du coût de l'exécution de l'entraînement, et non de la rapidité avec laquelle le modèle fini vous répond.

Calcul de formation
5.8 × 10²³ FLOP

"Both Chinchilla and Gopher have been trained for the same number of FLOPs but differ in the size of the model and the number of training tokens." We see the number of flops in table 3

Comment cela a été établi
Reported

L'entraînement

Ce qu'il a physiquement fallu pour entraîner : quels puces, combien, pendant combien de temps, et ce que cela a tiré du mur.

Matériel d'entraînement
Google TPU v4,Google TPU v3

Disponibilité

Que vous puissiez obtenir le modèle et l'exécuter sur votre propre matériel, ce qui décide si l'une des chiffres de carte graphique sur cette page s'applique.

Poids
Closed — provider access only
Accès au modèle
Unreleased
Code d'entraînement
Unreleased

Comment il est classé

Étiquettes que le jeu de données source applique lors du suivi des modèles notables, et à quel point il est confiant dans l'entrée.

Foundation model
Yes
Probablement au-dessus de 10²³ FLOP
Yes
Pourquoi c'est suivi
SOTA improvement,Historical significance

Proposes new scaling law, with good empirical results

Confiance d'enregistrement
Confident
Citations
3,188
Benchmark data
Chinchilla

Sources

D'où provient cet enregistrement et quand a-t-il été vérifié pour la dernière fois.

Référence
Training Compute-Optimal Large Language Models
Dernière mise à jour
25 May 2026

Ce que signifient les chiffres

À propos de ce modèle

Chinchilla was published by DeepMind, in United Kingdom of Great Britain and Northern Ireland, in March 2022. The organisation is categorised as industry.

It works in Language, and is recorded as doing language modeling.

Ses poids n'ont jamais été publiés, donc on ne peut y accéder que par son fournisseur. Aucun changement de carte graphique ne change cela.

Ce qui a contribué à sa construction

Producing it required around 5.8 × 10²³ FLOP of arithmetic, on Google TPU v4,Google TPU v3, which is a statement about the training budget rather than about inference.

The training set ran to roughly 1,400,000,000,000 tokens.

Its inclusion criterion is sOTA improvement,Historical significance.

Réponses

Chinchilla — Questions fréquentes

01

How much compute was used to train Chinchilla?

Around 5.8 × 10²³ FLOP, on Google TPU v4,Google TPU v3. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.

02

What GPU do I need to run Chinchilla?

None. Chinchilla is a closed model — its weights were never published, so it cannot be downloaded or run on your own hardware at any price. It is reachable only through its provider.

03

Is Chinchilla open source?

No. Chinchilla has not had its weights published, so it exists only as a service controlled by its owner.

04

How many parameters does Chinchilla have?

Chinchilla has 70B parameters. "We test this hypothesis by training a predicted compute-optimal model, \chinchilla, that uses the same compute budget as \gopher but with 70B parameters and 4× more more data. \chinchilla uniformly and significantly outperforms \Gopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks.". That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.

05

Who created Chinchilla?

Chinchilla was published by DeepMind, based in United Kingdom of Great Britain and Northern Ireland, categorised as industry.

06

When was Chinchilla released?

Chinchilla was published in March 2022. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.

07

What is Chinchilla used for?

Chinchilla works in Language, and is recorded as handling language modeling. These are the areas it was designed around; they describe intent rather than a hard boundary.

Source

Publication originale

Dernière mise à jour de l'enregistrement 25 May 2026

Dans l'autre sens

Regardez-la de l'autre côté ?

Cette page commence par le modèle. Si vous possédez déjà une carte et souhaitez connaître tout ce qu'elle pourra exécuter, commencez plutôt par le matériel.