Switch υπολογιστής TPS
Κάθε κάρτα παρακάτω αξιολογείται σύμφωνα με αυτό το μοντέλο στο μήκος πλαισίου και την ελάχιστη ποιότητα που επιλέγετε. Η ταχύτητα είναι μια εκτίμηση για ένα μόνο αίτημα, υπολογισμένη από το εύρος μνήμης της κάρτας και το μέγεθος του μοντέλου μόλις συμπιεστεί..
Υπολογίστηκε για αυτό το μοντέλο
Ποιες GPU μπορούν να τρέξουν Switch?
Ορίστε τα εισερχόμενα, διάβασε την απάντηση
Μια μακρύτερη συνομιλία χρειάζεται περισσότερη μνήμη, γεγονός που μπορεί να ωθήσει αυτό το μοντέλο εκτός από μικρότερα κάρτες..
Κρύβει κάρτες που θα ταίριαζαν μόνο στο μοντέλο συμπιέζοντάς το κάτω από αυτό το σημείο.
0 οι μοντέλα μπορούν να τρέξουν στο GPU. αυτό το μοντέλο δεν ταιριάζει. τρέχει άνετα.
Υπολογισμός| Χρειάζεται | Ποσοτικοποίηση | Κατάλληλο | |||||
|---|---|---|---|---|---|---|---|
|
No card in our catalogue can run this model with these settings. |
|||||||
Οι ταχύτητες είναι εκτιμήσεις για μια μόνο αίτηση — μια συνομιλία τη φορά — υπολογισμένες από την εύρος ζώνης μνήμης, το μέγεθος του μοντέλου και την ποσοστοποίηση. Η πραγματική διαμεταγωγή ποικίλλει με τον χρόνο εκτέλεσης της συμπερασματολογίας και την έκδοσή της. Οι αριθμοί που δημοσιεύονται από τις εταιρείες υλικού μετρούν πολλές ταυτόχρονες αιτήσεις και είναι πολύ υψηλότεροι..
Σε καταγραφή
πλήρης προδιαγραφή
Τα πάντα είναι καταγεγραμμένα για αυτό το μοντέλο. Το μεγαλύτερο μέρος περιγράφει πώς έχει εκπαιδευτεί παρά πώς εκτελείται — χρήσιμος πλαίσιο για να κρίνουμε πόση δουλειά έχει γίνει σε αυτό, και πώς συγκρίνεται με μοντέλα που έχουν κατασκευαστεί σε διαφορετική κλίμακα..
Προέλευση
Ποιος δημιούργησε αυτό το μοντέλο, πού και πότε δημοσιεύθηκε.
- Οργάνωση
- Τύπος οργάνωσης
- Industry
- Χώρα
- United States of America
- Δημοσιευμένο
- 11 January 2021
- Authors
- William Fedus, Barret Zoph, Noam Shazeer
Τι κάνει
Οι προβληματικές περιοχές για τις οποίες χτίστηκε το μοντέλο. Ένα μοντέλο μπορεί να φέρει αρκετές από κάθε μία.
- Δικτύου
- Language
- Text autocompletion
- Προσέγγιση
- Self-supervised learning
- ΔΕΝ ΤΑΙΡΙΑΖΕΙ / ΔΕΝ ΘΑ ΤΑΙΡΙΑΞΕΙ
- BF16
Μέγεθος
Πόσο μεγάλο είναι το μοντέλο και πόσα δεδομένα έχει εκπαιδευτεί. Οι παράμετροι είναι ο αριθμός που αποφασίζει αν θα χωρέσει σε μια δεδομένη κάρτα γραφικών.
- παραμέτρους
- 1.6T
- Εκπαίδευση δεδομένων
- 86,400,000,000 tokens
"Combining expert, model and data parallelism, we design two large Switch Transformer models, one with 395 billion and 1.6 trillion parameters" Table 9 gives more precise count of 1571B parameters
"In our protocol we pre-train with 2^20 (1,048,576) tokens per batch for 550k steps amounting to 576B total tokens." 1 token ~ 0.75 words
Υπολογιστική εκπαίδευση
Η αριθμητική που εκτελείται για την εκπαίδευση του μοντέλου, μετρημένη σε floating-point operations. Είναι ένα μέτρο του κόστους της εκπαίδευσης, όχι της ταχύτητας που απαντά το ολοκληρωμένο μοντέλο.
- Υπολογιστική εκπαίδευση
- 8.2 × 10²² FLOP
- Πώς ιδρύθηκε
- Third-party estimation
Table 4 https://arxiv.org/ftp/arxiv/papers/2104/2104.10350.pdf
Η εκπαίδευση τρέχει
Τι χρειάστηκε φυσικά για την εκπαίδευση: ποια τσιπ, πόσα, για πόσο καιρό, και τι ρεύμα αντλήθηκε από το δίκτυο.
- Υλικό εκπαίδευσης
- Google TPU v3
- Χρησιμοποιούμενοι επεξεργαστές
- 1,024
- Χρόνια-τσιπ
- 663,552
- Χρόνος ρολογιού
- 648 hours (27 days)
- Χρήση υλικού
- HFU 28.0%
- Κατανάλωση ισχύος
- 935.4 kW
- Compute cost
- $145,101
see table 4 in https://arxiv.org/ftp/arxiv/papers/2104/2104.10350.pdf
Table 4 in https://arxiv.org/pdf/2104.10350 gives measured performance of 34.4 TFLOP/s, vs. peak achievable FLOP/s of 123 TFLOP/s on the TPUv3 being used. HFU = 34.4/123 = 0.27967
Διαθεσιμότητα
Εάν μπορείτε να αποκτήσετε το μοντέλο και να το τρέξετε στον δικό σας υλικό, το οποίο είναι αυτό που αποφασίζει αν κάποια από τα στοιχεία της κάρτας γραφικών σε αυτή τη σελίδα ισχύουν.
- Βάρη
- Open — downloadable
- Πρόσβαση μοντέλου
- Open weights (unrestricted)
- Κωδικός εκπαίδευσης
- Unreleased
Apache 2 for weights: https://huggingface.co/google/switch-c-2048 paper links to this repo but not clear that the training hyperparams for Switch are here: https://github.com/google-research/t5x
Πώς ταξινομείται
Ετικέτες που εφαρμόζονται στο σύνολο δεδομένων προέλευσης κατά την παρακολούθηση σημαντικών μοντέλων και πόσο σίγουρο είναι για την καταχώρηση.
- Μοντέλο Frontier
- Yes
- Γιατί παρακολουθείται
- Highly cited,SOTA improvement
- Το μοντέλο δεν ταιριάζει στο VRAM του GPU.
- Confident
- 3,888
" On ANLI (Nie et al., 2019), Switch XXL improves over the prior state-of-the-art to get a 65.7 accuracy versus the prior best of 49.4 (Yang et al., 2020)... Finally, we also conduct an early examination of the model’s knowledge with three closed-book knowledge-based tasks: Natural Questions, WebQuestions and TriviaQA, without additional pre-training using Salient Span Masking (Guu et al., 2020). In all three cases, we observe improvements over the prior stateof-the-art T5-XXL model (without SSM…
Ποιό μοντέλο μπορεί να τρέξει η GPU; Δεν ταιριάζει. Τρέχει άνετα.
Όταν αυτό το αρχείο δημιουργήθηκε και πότε ελέγχθηκε τελευταία φορά.
- Αναφορά
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Τι μοντέλο μπορεί να τρέξει η GPU; Δεν χωράει. Πολύ σφιχτό. Άνετο.
- 25 May 2026
Τι σημαίνουν οι αριθμοί
What you need to run it
At 1.6T parameters, Switch is beyond what any single graphics card holds. Running it means either splitting it across several cards or renting hardware built for the job — 0 of the cards we track can hold it on their own, and all of them are datacentre parts.
Ιστορικό
Switch was published by Google, in United States of America, in January 2021. The organisation is categorised as industry.
It works in Language, and is recorded as doing text autocompletion.
The weights are published, so it can be downloaded and run on your own hardware indefinitely, offline, with no account attached.
Εκπαίδευση και προέλευση
The training run consumed about 8.2 × 10²² FLOP, on Google TPU v3. That figure describes the cost of creating it and has no bearing on how quickly it generates text.
Around 86,400,000,000 tokens went into training it.
It is tracked in the underlying dataset for one reason in particular: highly cited,SOTA improvement.
Βήμα προς βήμα
Πώς να επιλέξετε μία GPU για Switch
Ο πίνακας παραπάνω έχει ήδη αξιολογήσει κάθε κάρτα για την οποία έχουμε προδιαγραφές ενάντια σε αυτό το μοντέλο. Η απάντησή σας απαιτεί έξι βήματα..
-
01
Διαβάστε πρώτα το νούμερο μνήμης.
Look at what Switch actually needs. No amount of processing power compensates for a card that cannot hold it.
-
02
Αποφασίστε πόσο καιρό θα τρέχουν οι συνομιλίες σας
The conversation occupies memory too, and grows as it goes. Set the slider to the length you expect: at long context Switch can slip off a card that handles short questions easily.
-
03
Αποφασίστε πόση συμπίεση θα αποδεχθείτε
Compression is what makes Switch fit smaller cards, at some cost in accuracy. A minimum quality removes the ones that go too far.
-
04
Ταξινόμηση κατά ταχύτητα
Sort by speed to see how cards rank for Switch. It will not match a gaming ordering — generation is bound by memory bandwidth.
-
05
Read the fit column last
The fit column separates cards that just manage Switch from those with room to spare. Buy for the second if the context might grow.
-
06
Δες τι άλλο τρέχει αυτή η κάρτα
Following a card through to its own page shows every other model it can hold, which is the question that follows once Switch is settled.
Απαντήσεις
Switch — Αυτό το μοντέλο δεν χωράει στη VRAM της GPU.
What is Switch used for?
Switch works in Language, and is recorded as handling text autocompletion. These are the areas it was designed around; they describe intent rather than a hard boundary.
Where can I download Switch?
The weights for Switch are published, though we do not hold a repository link for it. This site calculates hardware requirements rather than hosting model files.
How much compute was used to train Switch?
Around 8.2 × 10²² FLOP, on Google TPU v3. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.
Can I run Switch if it does not fit in my GPU?
It can be split between the card and system memory, but Switch generates painfully slowly that way — the nearest miss we calculate is short by 691.9 GB. Nothing on this page assumes offloading.
Would two GPUs run Switch faster?
Two cards buy memory rather than speed. That matters for Switch only if one card cannot hold it — 0 can, so a second adds little.
Why does the quantisation differ between cards for Switch?
A larger card holds a more accurate copy. Across the cards that run Switch, 1 compression levels are used; the floor control above pins it to one.
How accurate are these Switch speed estimates?
These are estimates with real error bars. The fastest result here, the range beneath each figure, could reasonably land anywhere in its published range depending on which runtime you use.
Is Switch open source?
Its weights are published, so Switch can be downloaded and run on your own hardware. Note that open weights is not the same as open source in the full sense — it says nothing about the training data, the training code, or the commercial terms attached.
How many parameters does Switch have?
Switch has 1.6T parameters. "Combining expert, model and data parallelism, we design two large Switch Transformer models, one with 395 billion and 1.6 trillion parameters" Table 9 gives more precise count of 1571B parameters. That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.
Who created Switch?
Switch was published by Google, based in United States of America, categorised as industry.
When was Switch released?
Switch was published in January 2021. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.
Η άλλη κατεύθυνση
Κοιτάζοντας το από την άλλη πλευρά;
Αυτή η σελίδα ξεκινά από το μοντέλο. Εάν έχετε ήδη μια κάρτα και θέλετε να ξέρετε τα πάντα που μπορεί να εκτελέσει., δεν χωράει.