GPT-2 (1.5B, Curriculum Learning 45K)
Nincs becslés
Ennekhez a modellhez nincs hardverkövetelmény
A modellhez tartozó súlyokat nem tették közzé, így azt semmilyen méretben nem lehet letölteni vagy saját hardveren futtatni. Csak az adattovábbítóján keresztül érhető el, és semmilyen grafikus kártya nem változtat ezen..
Felvételen
Teljes specifikáció
Minden adatok ezen a modellen. A legtöbb leírás arról szól, hogyan volt kiképzve, nem pedig arról, hogyan fut — hasznos háttérinformáció arra vonatkozóan, hogy mennyi munka került bele, és hogyan viszonyul más méretű modellekhez..
Eredet
Ki építette ezt a modellt, hol, és mikor publikálták.
- Szervezet
- Microsoft
- Szervezet típusa
- Industry
- Ország
- United States of America
- Kiadva
- 13 August 2021
- Szerzők
- Conglong Li, Minjia Zhang, Yuxiong He
Mit csinál
A modell által épített probléma területek. Egy modell több különbözőt is tartalmazhat.
- Irányítószám
- Language
- Feladat
- Language modeling/generation
Méret
Mekkora a modell és mennyi adatot használtak a betanításához. A paraméterek azok a számok, amelyek meghatározzák, hogy illeszkedik-e egy adott grafikus kártyára.
- Paraméterek
- 1.5B
- Képzési adatok
- 157,000,000,000 tokens
- Epókák
- 1.6
- Tételméret
- 4,000,000
Table 2 SLW 45K bsz4K-seqlen1K 58.8K 157B (1x) training time: 155Hr (2.2x) batch size 4000 sequence length 1000 58.8K steps 157B tokens 4*10^6*58800/157*10^9 = 1.6 epochs
Képzés számítás
A modell betanításához végzett aritmetikai műveletek, amelyet lebegőpontos műveletekben mérnek. Ez a tréning futásának költségét méri, nem azt, hogy mennyire gyorsan válaszol a kész modell.
- Képzés számítás
- 2.4 × 10²¹ FLOP
- Hogyan jött létre
- Comparison with other models,Hardware,Operation counting
"Experiments replicating GPT-2 models (117M and 1.5B) show that <..> our method reduces the required number of training tokens and wall clock time by up to 2.2x and 3.7x, respectively." GPT-2 estimated training compute: 1.9200000000009998e+21 FLOP (Speculative confidence) -> 1.9200000000009998e+21 / 3.7 = 5.1891892e+20 FLOP " All of the experiments are performed on 128 NVIDIA V100 GPUs (32GB memory). There are 16 nodes and 8 GPUs per node" 155 hours (Table 2) 125000000000000 FLOP/sec * 128 …
A képzési futam
A kiképzéshez fizikailag szükséges volt: milyen chipek, hány, meddig és mennyi áramot vett le a hálózatról.
- Képzési hardver
- NVIDIA V100
- Felhasznált chipek
- 128
- Wall-clock time
- 155 hours
- Energiafogyasztás
- 77.6 kW
Elérhetőség
Hogy megszerezheted-e a modellt és futtathatod a saját hardvereden, ami eldönti, hogy a ezen az oldalon szereplő grafikus kártyák adatai alkalmazhatóak-e.
- Súlyok
- Closed — provider access only
- Modell hozzáférés
- Unreleased
- Képzés kód
- Unreleased
there's a repo for the technique but I don't see training code for this model: https://github.com/microsoft/DeepSpeed
Hogyan van besorolva
Címkék, amelyeket a forráshalmaz alkalmaz, amikor figyelemmel kíséri a figyelemre méltó modelleket, és mennyire biztos a bejegyzésben.
- Rögzítési bizalom
- Likely
- Citations
- 55
- Benchmark data
- GPT-2 (1.5B, Curriculum Learning 45K)
Források
Honnan származik ez a rekord és mikor ellenőrizték utoljára.
- Referencia
- Curriculum Learning: A Regularization Method for Efficient and Stable Billion-Scale GPT Model Pre-Training
- Utoljára frissítve
- 25 May 2026
Mit jelentek a számok
Mit jelent ez a modell?
GPT-2 (1.5B, Curriculum Learning 45K) was published by Microsoft, in United States of America, in August 2021. industry is the category the publisher falls under.
It works in Language, and is recorded as doing language modeling/generation.
Because the weights are not available, none of the hardware figures elsewhere on this site apply to it.
Hogyan lett betanítva
Training it took roughly 2.4 × 10²¹ FLOP of computation, on NVIDIA V100 — a measure of what producing the model cost, not of how fast it answers.
The training set ran to roughly 157,000,000,000 tokens.
Válaszok
GPT-2 (1.5B, Curriculum Learning 45K) — Gyakran ismételt kérdések
When was GPT-2 (1.5B, Curriculum Learning 45K) released?
GPT-2 (1.5B, Curriculum Learning 45K) was published in August 2021. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.
What is GPT-2 (1.5B, Curriculum Learning 45K) used for?
A GPT-2 (1.5B, Curriculum Learning 45K) nyelven működik, és nyelvi modellezés/generálás kezelésére van feljegyezve. Egy modell többből is hordozhat, így ezek azok a területek, amelyekre létrehozták, nem pedig a korlát arra, amit meg fog kísérelni.
How much compute was used to train GPT-2 (1.5B, Curriculum Learning 45K)?
Around 2.4 × 10²¹ FLOP, on NVIDIA V100. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.
What GPU do I need to run GPT-2 (1.5B, Curriculum Learning 45K)?
None. GPT-2 (1.5B, Curriculum Learning 45K) is a closed model — its weights were never published, so it cannot be downloaded or run on your own hardware at any price. It is reachable only through its provider.
Is GPT-2 (1.5B, Curriculum Learning 45K) open source?
No. GPT-2 (1.5B, Curriculum Learning 45K) has not had its weights published, so it exists only as a service controlled by its owner.
How many parameters does GPT-2 (1.5B, Curriculum Learning 45K) have?
GPT-2 (1.5B, Curriculum Learning 45K) has 1.5B parameters. That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.
Who created GPT-2 (1.5B, Curriculum Learning 45K)?
GPT-2 (1.5B, Curriculum Learning 45K) was published by Microsoft, based in United States of America, categorised as industry.
A másik irány
Másik oldalról nézve?
Ez az oldal a modellel kezdődik. Ha már rendelkezik egy kártyával és szeretné tudni, hogy minden mit fog futtatni, kérjük, folytassa., kezdje a hardverrel inkább.