Weight Decay
Aucune estimation
Aucun besoin matériel pour ce modèle
Les poids de ce modèle n'ont pas été publiés, il ne peut donc pas être téléchargé ou exécuté sur votre propre matériel, quelle que soit la taille. Il n'est accessible que via son fournisseur, et aucune carte graphique ne change cela.
Enregistrement terminé
Spécification complète
Tout ce qui concerne ce modèle. La plupart décrit comment il a été entraîné plutôt que comment il fonctionne — contexte utile pour évaluer le travail fourni et sa comparaison avec des modèles construits à une échelle différente.
Origine
Qui a construit ce modèle, où et quand il a été publié.
- Publié
- 2 December 1991
- Auteurs
- A. Krogh, J. Hertz
Ce qu'il fait
Les domaines problématiques pour lesquels le modèle a été conçu. Un modèle peut en porter plusieurs de chacun.
- Domaine
- Speech
- Tâche
- Speech synthesis
Taille
Quelle est la taille du modèle et combien de données il a été formé. Les paramètres sont le chiffre qui détermine s'il tient sur une carte graphique donnée.
- Paramètres
- 8.4K
- Données d'entraînement
- 25,000 tokens
- Époques
- 300
7*26*40+40+40*26+26=8386 "The network had 7 x 26 input units, 40 hidden units and 26 output units"
"It was trained on 400 to 5000 random words from the data base of around 20.000 words,"
Calcul de formation
L'arithmétique effectuée pour entraîner le modèle, mesurée en opérations à virgule flottante. C'est une mesure du coût de l'exécution de l'entraînement, et non de la rapidité avec laquelle le modèle fini vous répond.
- Calcul de formation
- 7.5 × 10¹⁰ FLOP
- Comment cela a été établi
- Operation counting
2*8386*3*1500000=75474000000=7.55e10 "It was trained on 400 to 5000 random words from the data base of around 20.000 words," "The top full line corresponds to the generalization error after 300 epochs"
Comment il est classé
Étiquettes que le jeu de données source applique lors du suivi des modèles notables, et à quel point il est confiant dans l'entrée.
- Frontier model
- Yes
- Pourquoi c'est suivi
- Highly cited,Historical significance
- Confiance d'enregistrement
- Confident
Sources
D'où provient cet enregistrement et quand a-t-il été vérifié pour la dernière fois.
- Référence
- A Simple Weight Decay Can Improve Generalization
- Dernière mise à jour
- 28 November 2025
Ce que signifient les chiffres
Background
Weight Decay was published by its authors, in December 1991.
It works in Speech, and is recorded as doing speech synthesis.
This is a closed model: the trained values stayed with whoever produced them, and there is no local version to run.
Ce qui a contribué à sa construction
Producing it required around 7.5 × 10¹⁰ FLOP of arithmetic, which is a statement about the training budget rather than about inference.
It was trained on about 25,000 tokens of text.
The reason it appears in this catalogue at all is highly cited,Historical significance.
Réponses
Weight Decay — Questions fréquentes
When was Weight Decay released?
Weight Decay was published in December 1991. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.
What is Weight Decay used for?
Weight Decay works in Speech, and is recorded as handling speech synthesis. These are the areas it was designed around; they describe intent rather than a hard boundary.
How much compute was used to train Weight Decay?
Around 7.5 × 10¹⁰ FLOP. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.
What GPU do I need to run Weight Decay?
None. Weight Decay is a closed model — its weights were never published, so it cannot be downloaded or run on your own hardware at any price. It is reachable only through its provider.
Is Weight Decay open source?
The licensing for Weight Decay was never recorded in our source data. We treat unstated licensing as closed, because an unrecorded licence is not one to rely on.
How many parameters does Weight Decay have?
Weight Decay has 8.4K parameters. 7*26*40+40+40*26+26=8386 "The network had 7 x 26 input units, 40 hidden units and 26 output units". That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.
Dans l'autre sens
Regardez-la de l'autre côté ?
Cette page commence par le modèle. Si vous possédez déjà une carte et souhaitez connaître tout ce qu'elle pourra exécuter, commencez plutôt par le matériel.