Demist-2

Zárt súlyok Darktrace 95M Paraméterek April 2025

Nincs becslés

Ennekhez a modellhez nincs hardverkövetelmény

A modellhez tartozó súlyokat nem tették közzé, így azt semmilyen méretben nem lehet letölteni vagy saját hardveren futtatni. Csak az adattovábbítóján keresztül érhető el, és semmilyen grafikus kártya nem változtat ezen..

Felvételen

Teljes specifikáció

Minden adatok ezen a modellen. A legtöbb leírás arról szól, hogyan volt kiképzve, nem pedig arról, hogyan fut — hasznos háttérinformáció arra vonatkozóan, hogy mennyi munka került bele, és hogyan viszonyul más méretű modellekhez..

Eredet

Ki építette ezt a modellt, hol, és mikor publikálták.

Szervezet
Darktrace
Szervezet típusa
Industry
Ország
United Kingdom of Great Britain and Northern Ireland
Kiadva
17 April 2025
Szerzők
Miguel Neves Fonseca, Philip Sellars, Tim Bazalgette

Mit csinál

A modell által épített probléma területek. Egy modell több különbözőt is tartalmazhat.

Irányítószám
Language
Feladat
Language modeling/generation, Text classification, Question answering, Entity embedding

Méret

Mekkora a modell és mennyi adatot használtak a betanításához. A paraméterek azok a számok, amelyek meghatározzák, hogy illeszkedik-e egy adott grafikus kártyára.

Paraméterek
95M

95M

Képzési adatok
351,000,000,000 tokens

C4: We extracted a 305 billion token subset URIs: We extracted the available URIs from several Common Crawl web scrapes. This allowed us to create a large corpus of 10 billion tokens. Linked Hostnames: Using the Common Crawl dataset, we extracted the observed links between web pages. <..> We further processed this data into rows of hostnames that shared linking properties which totaled 11 billion tokens. HTTP Data: 25 billion tokens. 305 + 10 + 11 + 25 = 351 b tokens "1.2 million optimization …

Epókák
1.8

Képzés számítás

A modell betanításához végzett aritmetikai műveletek, amelyet lebegőpontos műveletekben mérnek. Ez a tréning futásának költségét méri, nem azt, hogy mennyire gyorsan válaszol a kész modell.

Képzés számítás
4.6 × 10²⁰ FLOP

6 FLOP / parameter / token * 95 * 10^6 parameters * 351 * 10^9 tokens [see dataset size notes] *1.8 epochs = 3.60126e+20 FLOP 312000000000000 FLOP / GPU / sec [A100 reported] * 8 GPUs * 216 hours [9 days reported] * 3600 sec / hour * 0.3 [assumed utilization] = 5.8226688e+20 FLOP sqrt(3.60126e+20 * 5.8226688e+20) = 4.579186e+20 FLOP

Hogyan jött létre
Operation counting,Hardware

A képzési futam

A kiképzéshez fizikailag szükséges volt: milyen chipek, hány, meddig és mennyi áramot vett le a hálózatról.

Képzési hardver
NVIDIA A100
Felhasznált chipek
8
Wall-clock time
216 hours (9 days)

"DEMIST-2 was trained on 8 A100 GPUs over a nine-day period using the AWS Sagemaker platform." 9 days = 216 hours

Energiafogyasztás
6.3 kW
Cloud vendor
AWS,Amazon Web Services

Elérhetőség

Hogy megszerezheted-e a modellt és futtathatod a saját hardvereden, ami eldönti, hogy a ezen az oldalon szereplő grafikus kártyák adatai alkalmazhatóak-e.

Súlyok
Closed — provider access only
Modell hozzáférés
Hosted access (no API)
Képzés kód
Unreleased

Hogyan van besorolva

Címkék, amelyeket a forráshalmaz alkalmaz, amikor figyelemmel kíséri a figyelemre méltó modelleket, és mennyire biztos a bejegyzésben.

Rögzítési bizalom
Confident

Források

Honnan származik ez a rekord és mikor ellenőrizték utoljára.

Referencia
DEMIST-2: Darktrace Embedding Model for Investigation of Security Threats
Utoljára frissítve
28 November 2025

Mit jelentek a számok

Where it came from

Demist-2 was published by Darktrace, in United Kingdom of Great Britain and Northern Ireland, in April 2025. The organisation is categorised as industry.

It works in Language, and is recorded as doing language modeling/generation, Text classification, Question answering, Entity embedding.

Because the weights are not available, none of the hardware figures elsewhere on this site apply to it.

Training and provenance

Training it took roughly 4.6 × 10²⁰ FLOP of computation, on NVIDIA A100 — a measure of what producing the model cost, not of how fast it answers.

Around 351,000,000,000 tokens went into training it.

Válaszok

Demist-2 — Gyakran ismételt kérdések

01

Who created Demist-2?

Demist-2 was published by Darktrace, based in United Kingdom of Great Britain and Northern Ireland, categorised as industry.

02

When was Demist-2 released?

Demist-2 was published in April 2025.

03

What is Demist-2 used for?

Demist-2 works in Language, and is recorded as handling language modeling/generation, Text classification, Question answering, Entity embedding. These are the areas it was designed around; they describe intent rather than a hard boundary.

04

How much compute was used to train Demist-2?

Around 4.6 × 10²⁰ FLOP, on NVIDIA A100. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.

05

What GPU do I need to run Demist-2?

None. Demist-2 is a closed model — its weights were never published, so it cannot be downloaded or run on your own hardware at any price. It is reachable only through its provider.

06

Is Demist-2 open source?

No. Demist-2 has not had its weights published, so it exists only as a service controlled by its owner.

07

How many parameters does Demist-2 have?

Demist-2 has 95M parameters. 95M. That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.

Forrás

Eredeti közzététel

Utolsó frissítés időpontja 28 November 2025

A másik irány

Másik oldalról nézve?

Ez az oldal a modellel kezdődik. Ha már rendelkezik egy kártyával és szeretné tudni, hogy minden mit fog futtatni, kérjük, folytassa., kezdje a hardverrel inkább.