Transformer + GFM

Κλειστά βάρη Nanjing University 185.2M Παραμέτροι December 2022

Χωρίς εκτίμηση

Δεν υπάρχουν απαιτήσεις υλικού για αυτό το μοντέλο

Τα βάρη αυτού του μοντέλου δεν έχουν δημοσιευθεί, επομένως δεν μπορεί να κατεβεί ή να τρέξει σε δικό σας υλικό οποιουδήποτε μεγέθους. Είναι προσβάσιμο μόνο μέσω του παρόχου του και καμία κάρτα γραφικών δεν το αλλάζει αυτό.

Σε εγγραφή

Πλήρης προδιαγραφή

Όλα τα δεδομένα σχετικά με αυτό το μοντέλο. Τα περισσότερα περιγράφουν πώς εκπαιδεύτηκε παρά πώς τρέχει — χρήσιμο πλαίσιο για την αξιολόγηση του πόσο δουλειά απαιτήθηκε και πώς συγκρίνεται με μοντέλα που έχουν κατασκευαστεί σε διαφορετική κλίμακα.

Προέλευση

Ποιος δημιούργησε αυτό το μοντέλο, πού και πότε δημοσιεύθηκε;

Οργάνωση
Nanjing University
Τύπος οργάνωσης
Academia
Χώρα
China
Δημοσιευμένο
1 December 2022
Συγγραφείς
Hao Yu, Jianxin Wu

Τι κάνει

Οι προβληματικές περιοχές για τις οποίες κατασκευάστηκε το μοντέλο. Ένα μοντέλο μπορεί να περιέχει αρκετές από κάθε μία.

Τομέας
Language
Εντολή
Language modeling
Βασικό μοντέλο
FAIRSEQ Adaptive Inputs

Μέγεθος

Πόσο μεγάλο είναι το μοντέλο και πόσα δεδομένα εκπαιδεύτηκε. Οι παράμετροι είναι το νούμερο που αποφασίζει αν χωράει σε μια δεδομένη κάρτα γραφικών.

Παραμέτροι
185.2M

185.2M (Table 4) "We implemented our methods based on fairseq (Ott et al. 2019). The original transformer model follows the architectural choice described in Baevski and Auli (2018), which includes 16 decoder blocks and sinusoidal position embeddings in the input layer. Each MHSA module has 8 heads and adaptive input representations have three bands of size 20K, 40K and 200K. The embedding layer and FFN’s hidden-state have dimensions of 1024 and 4096, respectively. We sampled 4K sentences to fo…

Δεδομένα εκπαίδευσης
103,000,000 tokens

Υπολογιστική εκπαίδευση

Η αριθμητική που εκτελείται για την εκπαίδευση του μοντέλου, μετρημένη σε αριθμητικές πράξεις κινητής υποδιαστολής. Είναι ένα μέτρο του κόστους της εκπαίδευσης, όχι πόσο γρήγορα απαντά το ολοκληρωμένο μοντέλο.

Υπολογιστική εκπαίδευση
7.7 × 10¹⁸ FLOP

7.30e18 FLOP [base transformer] + 4.3694162e+17 FLOP [GFM] = 7.7369416e+18 FLOP _______________ Estimations from the Algorithmic progress paper (upd - a100 gpu were assumed while the paper reports 3090 gpus): SOURCE: Compression of Baevski et al. transformer, impute During the GFM process, we removed layer dropout and trained on 8 GPUs. We limited the number of tokens per GPU to a maximum threshold 1536, which means each GPU processes 1536 tokens using the same model parameters. We accumulat…

Πώς ιδρύθηκε
Operation counting,Hardware
Καλή ρύθμιση υπολογισμού
4.4 × 10¹⁷ FLOP

6 FLOP/token/parameter * 185000000 parameters * 103000000 tokens = 1.1433e+17 FLOP or 0.6 hours - reported training time for another model in the paper, I assumed, this training was similar in time length 0.6 hours * 3600 sec / hour * 8 GPUs * 35580000000000 FLOP/s [assumed precision fp16] * 0.3 [assumed utilization] = 1.8444672e+17 FLOP or "During the GFM process, we removed layer dropout and trained on 8 GPUs. We limited the number of tokens per GPU to a maximum threshold 1536, which me…

Η εκπαίδευση

Τι χρειάστηκε φυσικά για να εκπαιδεύσουμε: ποια τσιπ, πόσα, για πόσο καιρό, και τι αυτό τραβούσε από τον τοίχο.

Εκπαίδευση υλικού
NVIDIA GeForce RTX 3090
Τσιπ που χρησιμοποιούνται
8
Τσιπ-ώρες
1
Κατανάλωση ενέργειας
5.6 kW

Διαθεσιμότητα

Εάν μπορείτε να αποκτήσετε το μοντέλο και να το εκτελέσετε στον δικό σας εξοπλισμό, κάτι που αποφασίζει αν κάποια από τα στοιχεία της κάρτας γραφικών σε αυτή τη σελίδα εφαρμόζονται.

Βάρη
Closed — provider access only
Πρόσβαση μοντέλου
Unreleased
Κώδικας εκπαίδευσης
Unreleased

Πώς ταξινομείται

Ετικέτες που εφαρμόζει το πηγαίο σύνολο δεδομένων κατά την παρακολούθηση σημαντικών μοντέλων, και πόσο σίγουρο είναι για την είσοδο.

Καταγραφή εμπιστοσύνης
Confident
Benchmark data
Transformer + GFM

Πηγές

Από πού προήλθε αυτή η καταγραφή και πότε ελέγχθηκε τελευταία.

Αναφορά
"Compressing Transformers: Features Are Low-Rank, but Weights Are Not"
Τελευταία ενημέρωση
28 November 2025

Τι σημαίνουν οι αριθμοί

Από πού προήλθε

Transformer + GFM was published by Nanjing University, in China, in December 2022. academia is the category the publisher falls under.

It works in Language, and is recorded as doing language modeling.

It is derived from FAIRSEQ Adaptive Inputs rather than trained from scratch, which is the usual way a specialised model is produced.

Because the weights are not available, none of the hardware figures elsewhere on this site apply to it.

Εκπαίδευση και προέλευση

Producing it required around 7.7 × 10¹⁸ FLOP of arithmetic, on NVIDIA GeForce RTX 3090, which is a statement about the training budget rather than about inference.

It was trained on about 103,000,000 tokens of text.

Απαντήσεις

Transformer + GFM — Συχνές ερωτήσεις

01

What is Transformer + GFM used for?

Transformer + GFM works in Language, and is recorded as handling language modeling. Models frequently carry more than one of each, and the tags describe purpose rather than capability limits.

02

How much compute was used to train Transformer + GFM?

Around 7.7 × 10¹⁸ FLOP, on NVIDIA GeForce RTX 3090. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.

03

What GPU do I need to run Transformer + GFM?

None. Transformer + GFM is a closed model — its weights were never published, so it cannot be downloaded or run on your own hardware at any price. It is reachable only through its provider.

04

Is Transformer + GFM open source?

No. Transformer + GFM has not had its weights published, so it exists only as a service controlled by its owner.

05

How many parameters does Transformer + GFM have?

Transformer + GFM has 185.2M parameters. 185.2M (Table 4) "We implemented our methods based on fairseq (Ott et al. 2019). The original transformer model follows the architectural choice described in Baevski and Auli (2018), which includes 16 decoder blocks and sinusoidal position embeddings in the input layer. Each MHSA module has 8 heads and adaptive input representations have three bands of size 20K, 40K and 200K. The embedding layer and FFN’s hidden-state have dimensions of 1024 and 4096, respectively. We sampled 4K sentences to form the proxy dataset D. Then, we reduced the parameters by 15%, 20% and 25%.". That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.

06

Who created Transformer + GFM?

Transformer + GFM was published by Nanjing University, based in China, categorised as academia.

07

When was Transformer + GFM released?

Transformer + GFM was published in December 2022. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.

Πηγή

Αρχική δημοσίευση

Τελευταία ενημέρωση αρχείου 28 November 2025

Η άλλη κατεύθυνση

Βλέποντάς το από την άλλη πλευρά;

Αυτή η σελίδα ξεκινάει από το μοντέλο. Αν ήδη διαθέτετε μια κάρτα και θέλετε να γνωρίζετε όλα όσα θα τρέξει, Ξεκινήστε από το υλικό αντί..