GLM-5.2 calcolatore TPS

Pesi aperti Z.ai (Zhipu AI) 744B parametri June 2026

Ogni scheda qui sotto è valutata rispetto a questo modello alla lunghezza di contesto e alla qualità minima che scegli. La velocità è una stima per una singola richiesta, calcolata dalla larghezza di banda della memoria della scheda e dalla dimensione del modello una volta compresso..

Calcolato per questo modello

0 di 818 schede che possono eseguirlo

Quali GPU possono eseguire GLM-5.2?

Imposta gli input, leggi la risposta

Una conversazione più lunga richiede più memoria, il che può spingere questo modello fuori da schede più piccole..

Nasconde le schede che si adatterebbero solo al modello comprimendole al di sotto di questo punto..

0 le modelli non si adattano al VRAM della GPU

Calcolo
Modello/models non si adatta / non si adatterà al VRAM della GPU. Tight = quasi eseguibile, comodo = esegue facilmente con margine. Quantizzazione Modello/modelli non adatti per la VRAM della GPU. Esecuzione non riuscita, richiede meno di TPS. Modello funziona comodamente, ampio margine.

Nessuna scheda nel nostro catalogo può eseguire questo modello con queste impostazioni.

Le velocità sono stime per una singola richiesta — una conversazione alla volta — calcolate dalla larghezza di banda della memoria, dalle dimensioni del modello e dalla quantizzazione. Il throughput reale varia con il tempo di esecuzione dell'inferenza e la sua versione. I valori pubblicati dai fornitori di hardware misurano molte richieste simultanee e sono molto più alti..

In registrazione

Specifiche complete

Tutto registrato per questo modello. La maggior parte descrive come è stato addestrato piuttosto che come funziona — contesto utile per valutare quanto lavoro ci sia stato dietro e come si confronta con modelli costruiti su scala differente..

Origine

Chi ha costruito questo modello, dove e quando è stato pubblicato.

Organizzazione
Z.ai (Zhipu AI)
Tipo di organizzazione
Industry
Paese
China
Pubblicato
16 June 2026

Cosa fa

Le aree problematiche per cui il modello è stato creato. Un modello può gestire diversi di ciascuno.

Domini
Language
Spazio
Language modeling/generation

Dimensione

Quanto è grande il modello e su quanti dati è stato addestrato. I parametri sono la cifra che decide se si adatta a una data scheda grafica.

parametri
744B
Dati di addestramento
tokens

Disponibilità

Se puoi ottenere il modello e farlo girare sul tuo hardware, il che decide se alcune delle figure della scheda grafica su questa pagina si applicano.

Pesi
Open — downloadable
Accesso al modello
Open weights (unrestricted)
Hugging Face
zai-org

Come è classificato

Etichette sui dataset di origine applicate per il tracciamento dei modelli notevoli e quanto è fiduciosa nell'entry.

Modello/modelli non si adatta / non si adatterà GPU VRAM. Tight = barely runnable, comodo = runs easily with headroom.
Confident

Modello/modelli GPU/AI che può eseguire. Non si adatta / non starà. Stretto = appena eseguibile, comodo = esegue facilmente con margine.

Dove è stato registrato questo record e quando è stato controllato l'ultima volta.

Modello/modelli GPU può eseguire? Non si adatta / non si adatterà. Limitato = difficilmente eseguibile, comodo = esegue facilmente con margine.
GLM-5.2: Built for Long-Horizon Tasks
Ultimo aggiornamento
10 July 2026

Cosa significano i numeri

Cosa serve per eseguire questo modello

At 744B parameters, GLM-5.2 is beyond what any single graphics card holds. Running it means either splitting it across several cards or renting hardware built for the job — 0 of the cards we track can hold it on their own, and all of them are datacentre parts.

Riguardo a questo modello

GLM-5.2 was published by Z.ai (Zhipu AI), in China, in June 2026. industry is the category the publisher falls under.

Funziona in lingua ed è registrato come modello/generazione di lingua.

The weights are published, so it can be downloaded and run on your own hardware indefinitely, offline, with no account attached. It is published under the zai-org organisation on Hugging Face.

Passo dopo passo

Come scegliere una GPU per GLM-5.2

La tabella sopra ha già valutato ogni scheda di cui possediamo specifiche rispetto a questo modello. Arrivare alla tua risposta richiede sei passaggi..

  1. 01

    Leggi prima il valore di memoria

    Look at what GLM-5.2 actually needs. No amount of processing power compensates for a card that cannot hold it.

  2. 02

    Modello/modelli GPU compatibili: GPU può eseguire modelli: Non va bene / non si adatta Esecuzione leggera: Esecuzione confortevole: Esecuzione intensa: Esecuzione difficile: In seguito: Utilizzo VRAM: Prestazioni TPS: OK Calcoli CUDA: Inferenza Tensor: Quantizzazione: LoRA: GGUF: FP16: INT4: INT8:

    Set the context to what you will actually use. The cache grows with the conversation, and it is the usual reason GLM-5.2 stops fitting a card that seemed fine.

  3. 03

    Decidi quanto compressione accetterai

    The quantisation column varies by card, because a bigger card holds a more accurate copy of GLM-5.2. Set a floor to hold the comparison at one level.

  4. 04

    Confronta TPS, non specifiche

    The speed ordering for GLM-5.2 is effectively an ordering by memory bandwidth.

  5. 05

    Controlla il verdetto di compatibilità prima di acquistare

    The fit column separates cards that just manage GLM-5.2 from those with room to spare. Buy for the second if the context might grow.

  6. 06

    Apri la scheda su cui hai deciso

    Every card name links to its own page, which runs the same calculation across the whole model catalogue. Worth a look before buying for GLM-5.2 alone — a card is usually bought for more than one model.

Modeli/modelli sempre significa modelli AI. Esegui/Esplora significa eseguire un modello AI. Non si adatta / non si adatterà significa che il modello AI non si adatta alla VRAM della GPU. Stretto = appena eseguibile, comodo = funziona facilmente con margine. Termini tecnici come GPU, AI, TPS, VRAM, CUDA, Tensor, Inference, Quantization, LoRA, GGUF, FP16, INT4, INT8 rimangono in inglese. GPU, AI, TPS, VRAM, CUDA devono sempre essere in maiuscolo. Se viene chiesto di tradurre il nome di una lingua (ad es. spagnolo, inglese, lingua), tradurre la parola correttamente. SOLO output la traduzione pura in italiano. Nessun commento, nessuna domanda, nient'altro. Se vuoto rispondi con un singolo spazio: Risposte.

GLM-5.2 — modello/modelli può GPU eseguire. non si adatta / non si adatterà. stretto = appena eseguibile, confortevole = esegue facilmente con margine.

01

Can I run GLM-5.2 if it does not fit in my GPU?

Only by offloading, which is usually a false economy: the part in system memory drags the whole thing down — the nearest miss we calculate is short by 191.6 GB. Our figures for GLM-5.2 assume it is fully resident.

02

Would two GPUs run GLM-5.2 faster?

Two cards buy memory rather than speed. That matters for GLM-5.2 only if one card cannot hold it — 0 can, so a second adds little.

03

Why does the quantisation differ between cards for GLM-5.2?

A larger card holds a more accurate copy. Across the cards that run GLM-5.2, 1 compression levels are used; the floor control above pins it to one.

04

How accurate are these GLM-5.2 speed estimates?

Ogni cifra è derivata dalla larghezza di banda della memoria e dalla dimensione del modello, non da benchmark. Ecco perché ciascuna è pubblicata come un intervallo, come l'intervallo sotto ogni cifra piuttosto che un singolo numero.

05

Is GLM-5.2 open source?

Its weights are published, so GLM-5.2 can be downloaded and run on your own hardware. Note that open weights is not the same as open source in the full sense — it says nothing about the training data, the training code, or the commercial terms attached.

06

How many parameters does GLM-5.2 have?

GLM-5.2 has 744B parameters. That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.

07

Who created GLM-5.2?

GLM-5.2 was published by Z.ai (Zhipu AI), based in China, categorised as industry.

08

When was GLM-5.2 released?

GLM-5.2 was published in June 2026.

09

What is GLM-5.2 used for?

GLM-5.2 works in Language, and is recorded as handling language modeling/generation. These are the areas it was designed around; they describe intent rather than a hard boundary.

10

Where can I download GLM-5.2?

Its weights are published under the zai-org organisation on Hugging Face. We do not host model files — this site calculates what hardware is needed to run them.

Fonte

Pubblicazione originale

Record last updated 10 July 2026

L'altra direzione

Guardandolo dall'altro lato?

Questa pagina inizia dal modello. Se possiedi già una scheda e vuoi sapere tutto ciò che eseguirà, inizia dall'hardware invece.