GPT-2 (1.5B, Curriculum Learning 45K)

Zárt súlyok Microsoft 1.5B Paraméterek August 2021

Nincs becslés

Ennekhez a modellhez nincs hardverkövetelmény

A modellhez tartozó súlyokat nem tették közzé, így azt semmilyen méretben nem lehet letölteni vagy saját hardveren futtatni. Csak az adattovábbítóján keresztül érhető el, és semmilyen grafikus kártya nem változtat ezen..

Felvételen

Teljes specifikáció

Minden adatok ezen a modellen. A legtöbb leírás arról szól, hogyan volt kiképzve, nem pedig arról, hogyan fut — hasznos háttérinformáció arra vonatkozóan, hogy mennyi munka került bele, és hogyan viszonyul más méretű modellekhez..

Eredet

Ki építette ezt a modellt, hol, és mikor publikálták.

Szervezet
Microsoft
Szervezet típusa
Industry
Ország
United States of America
Kiadva
13 August 2021
Szerzők
Conglong Li, Minjia Zhang, Yuxiong He

Mit csinál

A modell által épített probléma területek. Egy modell több különbözőt is tartalmazhat.

Irányítószám
Language
Feladat
Language modeling/generation

Méret

Mekkora a modell és mennyi adatot használtak a betanításához. A paraméterek azok a számok, amelyek meghatározzák, hogy illeszkedik-e egy adott grafikus kártyára.

Paraméterek
1.5B
Képzési adatok
157,000,000,000 tokens

Table 2 SLW 45K bsz4K-seqlen1K 58.8K 157B (1x) training time: 155Hr (2.2x) batch size 4000 sequence length 1000 58.8K steps 157B tokens 4*10^6*58800/157*10^9 = 1.6 epochs

Epókák
1.6
Tételméret
4,000,000

Képzés számítás

A modell betanításához végzett aritmetikai műveletek, amelyet lebegőpontos műveletekben mérnek. Ez a tréning futásának költségét méri, nem azt, hogy mennyire gyorsan válaszol a kész modell.

Képzés számítás
2.4 × 10²¹ FLOP

"Experiments replicating GPT-2 models (117M and 1.5B) show that <..> our method reduces the required number of training tokens and wall clock time by up to 2.2x and 3.7x, respectively." GPT-2 estimated training compute: 1.9200000000009998e+21 FLOP (Speculative confidence) -> 1.9200000000009998e+21 / 3.7 = 5.1891892e+20 FLOP " All of the experiments are performed on 128 NVIDIA V100 GPUs (32GB memory). There are 16 nodes and 8 GPUs per node" 155 hours (Table 2) 125000000000000 FLOP/sec * 128 …

Hogyan jött létre
Comparison with other models,Hardware,Operation counting

A képzési futam

A kiképzéshez fizikailag szükséges volt: milyen chipek, hány, meddig és mennyi áramot vett le a hálózatról.

Képzési hardver
NVIDIA V100
Felhasznált chipek
128
Wall-clock time
155 hours
Energiafogyasztás
77.6 kW

Elérhetőség

Hogy megszerezheted-e a modellt és futtathatod a saját hardvereden, ami eldönti, hogy a ezen az oldalon szereplő grafikus kártyák adatai alkalmazhatóak-e.

Súlyok
Closed — provider access only
Modell hozzáférés
Unreleased
Képzés kód
Unreleased

there's a repo for the technique but I don't see training code for this model: https://github.com/microsoft/DeepSpeed

Hogyan van besorolva

Címkék, amelyeket a forráshalmaz alkalmaz, amikor figyelemmel kíséri a figyelemre méltó modelleket, és mennyire biztos a bejegyzésben.

Rögzítési bizalom
Likely
Citations
55
Benchmark data
GPT-2 (1.5B, Curriculum Learning 45K)

Források

Honnan származik ez a rekord és mikor ellenőrizték utoljára.

Referencia
Curriculum Learning: A Regularization Method for Efficient and Stable Billion-Scale GPT Model Pre-Training
Utoljára frissítve
25 May 2026

Mit jelentek a számok

Mit jelent ez a modell?

GPT-2 (1.5B, Curriculum Learning 45K) was published by Microsoft, in United States of America, in August 2021. industry is the category the publisher falls under.

It works in Language, and is recorded as doing language modeling/generation.

Because the weights are not available, none of the hardware figures elsewhere on this site apply to it.

Hogyan lett betanítva

Training it took roughly 2.4 × 10²¹ FLOP of computation, on NVIDIA V100 — a measure of what producing the model cost, not of how fast it answers.

The training set ran to roughly 157,000,000,000 tokens.

Válaszok

GPT-2 (1.5B, Curriculum Learning 45K) — Gyakran ismételt kérdések

01

When was GPT-2 (1.5B, Curriculum Learning 45K) released?

GPT-2 (1.5B, Curriculum Learning 45K) was published in August 2021. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.

02

What is GPT-2 (1.5B, Curriculum Learning 45K) used for?

A GPT-2 (1.5B, Curriculum Learning 45K) nyelven működik, és nyelvi modellezés/generálás kezelésére van feljegyezve. Egy modell többből is hordozhat, így ezek azok a területek, amelyekre létrehozták, nem pedig a korlát arra, amit meg fog kísérelni.

03

How much compute was used to train GPT-2 (1.5B, Curriculum Learning 45K)?

Around 2.4 × 10²¹ FLOP, on NVIDIA V100. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.

04

What GPU do I need to run GPT-2 (1.5B, Curriculum Learning 45K)?

None. GPT-2 (1.5B, Curriculum Learning 45K) is a closed model — its weights were never published, so it cannot be downloaded or run on your own hardware at any price. It is reachable only through its provider.

05

Is GPT-2 (1.5B, Curriculum Learning 45K) open source?

No. GPT-2 (1.5B, Curriculum Learning 45K) has not had its weights published, so it exists only as a service controlled by its owner.

06

How many parameters does GPT-2 (1.5B, Curriculum Learning 45K) have?

GPT-2 (1.5B, Curriculum Learning 45K) has 1.5B parameters. That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.

07

Who created GPT-2 (1.5B, Curriculum Learning 45K)?

GPT-2 (1.5B, Curriculum Learning 45K) was published by Microsoft, based in United States of America, categorised as industry.

Forrás

Eredeti közzététel

Utolsó frissítés időpontja 25 May 2026

A másik irány

Másik oldalról nézve?

Ez az oldal a modellel kezdődik. Ha már rendelkezik egy kártyával és szeretné tudni, hogy minden mit fog futtatni, kérjük, folytassa., kezdje a hardverrel inkább.