Chinchilla
Aucune estimation
Aucun besoin matériel pour ce modèle
Les poids de ce modèle n'ont pas été publiés, il ne peut donc pas être téléchargé ou exécuté sur votre propre matériel, quelle que soit la taille. Il n'est accessible que via son fournisseur, et aucune carte graphique ne change cela.
Enregistrement terminé
Spécification complète
Tout ce qui concerne ce modèle. La plupart décrit comment il a été entraîné plutôt que comment il fonctionne — contexte utile pour évaluer le travail fourni et sa comparaison avec des modèles construits à une échelle différente.
Origine
Qui a construit ce modèle, où et quand il a été publié.
- Organisation
- DeepMind
- Type d'organisation
- Industry
- Pays
- United Kingdom of Great Britain and Northern Ireland
- Publié
- 29 March 2022
- Auteurs
- Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, Laurent Sifre
Ce qu'il fait
Les domaines problématiques pour lesquels le modèle a été conçu. Un modèle peut en porter plusieurs de chacun.
- Domaine
- Language
- Tâche
- Language modeling
- Approach
- Self-supervised learning
- Format numérique
- BF16
Taille
Quelle est la taille du modèle et combien de données il a été formé. Les paramètres sont le chiffre qui détermine s'il tient sur une carte graphique donnée.
- Paramètres
- 70B
- Données d'entraînement
- 1,400,000,000,000 tokens
- Époques
- 1
- Batch size
- 3,000,000
"We test this hypothesis by training a predicted compute-optimal model, \chinchilla, that uses the same compute budget as \gopher but with 70B parameters and 4× more more data. \chinchilla uniformly and significantly outperforms \Gopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks."
Table 1 shows Chinchilla was training on 1.4 trillion tokens 1 token ~ 0.75 words
Table 1. "1.5M → 3M"
Calcul de formation
L'arithmétique effectuée pour entraîner le modèle, mesurée en opérations à virgule flottante. C'est une mesure du coût de l'exécution de l'entraînement, et non de la rapidité avec laquelle le modèle fini vous répond.
- Calcul de formation
- 5.8 × 10²³ FLOP
- Comment cela a été établi
- Reported
"Both Chinchilla and Gopher have been trained for the same number of FLOPs but differ in the size of the model and the number of training tokens." We see the number of flops in table 3
L'entraînement
Ce qu'il a physiquement fallu pour entraîner : quels puces, combien, pendant combien de temps, et ce que cela a tiré du mur.
- Matériel d'entraînement
- Google TPU v4,Google TPU v3
Disponibilité
Que vous puissiez obtenir le modèle et l'exécuter sur votre propre matériel, ce qui décide si l'une des chiffres de carte graphique sur cette page s'applique.
- Poids
- Closed — provider access only
- Accès au modèle
- Unreleased
- Code d'entraînement
- Unreleased
Comment il est classé
Étiquettes que le jeu de données source applique lors du suivi des modèles notables, et à quel point il est confiant dans l'entrée.
- Foundation model
- Yes
- Probablement au-dessus de 10²³ FLOP
- Yes
- Pourquoi c'est suivi
- SOTA improvement,Historical significance
- Confiance d'enregistrement
- Confident
- Citations
- 3,188
- Benchmark data
- Chinchilla
Proposes new scaling law, with good empirical results
Sources
D'où provient cet enregistrement et quand a-t-il été vérifié pour la dernière fois.
- Référence
- Training Compute-Optimal Large Language Models
- Dernière mise à jour
- 25 May 2026
Ce que signifient les chiffres
À propos de ce modèle
Chinchilla was published by DeepMind, in United Kingdom of Great Britain and Northern Ireland, in March 2022. The organisation is categorised as industry.
It works in Language, and is recorded as doing language modeling.
Ses poids n'ont jamais été publiés, donc on ne peut y accéder que par son fournisseur. Aucun changement de carte graphique ne change cela.
Ce qui a contribué à sa construction
Producing it required around 5.8 × 10²³ FLOP of arithmetic, on Google TPU v4,Google TPU v3, which is a statement about the training budget rather than about inference.
The training set ran to roughly 1,400,000,000,000 tokens.
Its inclusion criterion is sOTA improvement,Historical significance.
Réponses
Chinchilla — Questions fréquentes
How much compute was used to train Chinchilla?
Around 5.8 × 10²³ FLOP, on Google TPU v4,Google TPU v3. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.
What GPU do I need to run Chinchilla?
None. Chinchilla is a closed model — its weights were never published, so it cannot be downloaded or run on your own hardware at any price. It is reachable only through its provider.
Is Chinchilla open source?
No. Chinchilla has not had its weights published, so it exists only as a service controlled by its owner.
How many parameters does Chinchilla have?
Chinchilla has 70B parameters. "We test this hypothesis by training a predicted compute-optimal model, \chinchilla, that uses the same compute budget as \gopher but with 70B parameters and 4× more more data. \chinchilla uniformly and significantly outperforms \Gopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks.". That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.
Who created Chinchilla?
Chinchilla was published by DeepMind, based in United Kingdom of Great Britain and Northern Ireland, categorised as industry.
When was Chinchilla released?
Chinchilla was published in March 2022. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.
What is Chinchilla used for?
Chinchilla works in Language, and is recorded as handling language modeling. These are the areas it was designed around; they describe intent rather than a hard boundary.
Dans l'autre sens
Regardez-la de l'autre côté ?
Cette page commence par le modèle. Si vous possédez déjà une carte et souhaitez connaître tout ce qu'elle pourra exécuter, commencez plutôt par le matériel.