LIMA

Geschlossene Gewichte Meta AI,Carnegie Mellon University (CMU),University of Southern California,Tel Aviv University 65B Parameter May 2023

Keine Schätzung

Keine Hardwareanforderungen für dieses Modell

Die Gewichte für dieses Modell wurden nicht veröffentlicht, sodass es weder heruntergeladen noch auf Ihrer eigenen Hardware in irgendeiner Größe ausgeführt werden kann. Es ist nur über seinen Anbieter zugänglich, und keine Grafikkarte ändert daran etwas.

Aufzeichnung

Vollständige Spezifikation

Alles aufzeichnungsfähig für dieses Modell. Die meisten Informationen beschreiben, wie es trainiert wurde, anstatt wie es läuft — ein nützlicher Kontext, um zu beurteilen, wie viel Arbeit investiert wurde und wie es im Vergleich zu auf einer anderen Skala entwickelten Modellen abschneidet..

Ursprung

Wer hat dieses Modell gebaut, wo und wann es veröffentlicht wurde.

Organisation
Meta AI,Carnegie Mellon University (CMU),University of Southern California,Tel Aviv University
Organisationstyp
Industry,Academia,Academia,Academia
Land
United States of America, Israel
Veröffentlicht
18 May 2023
Autoren
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, Omer Levy

Was es tut

Die Problemfelder, für die das Modell entwickelt wurde. Ein Modell kann mehrere von jedem tragen.

Domäne
Language
Aufgabe
Chat, Language modeling/generation
Basis-Modell
LLaMA-65B

Größe

Wie groß das Modell ist und wie viele Daten es trainiert wurde. Die Parameter sind die Größe, die entscheidet, ob es auf eine bestimmte Grafikkarte passt.

Parameter
65B

"We train LIMA (Less Is More for Alignment) using the following protocol. Starting from LLaMa 65B [Touvron et al., 2023], we fine-tune on our 1,000-example alignment training set," according to page 4 of https://arxiv.org/pdf/2305.11206.

Trainingsdaten
685,300 tokens

The total amount of training data is roughly 750,000 tokens, split over exactly 1,000 sequences.

Epochen
15

Training-Computing

Die zur Ausbildung des Modells durchgeführte Arithmetik, gemessen in Gleitkommaoperationen. Es ist ein Maß dafür, was der Trainingslauf gekostet hat, nicht wie schnell das fertige Modell Ihnen antwortet.

Training-Computing
5.5 × 10²³ FLOP

Finetune: 4.39e18 FLOP Base model: 5.5e+23 FLOP Total: 4.39e18+5.5e+23=5.5e+23 FLOP

Wie es etabliert wurde
Operation counting
Feinabstimmung Rechenleistung
4.4 × 10²⁰ FLOP

“The total amount of training data is roughly 750,000 tokens, split over exactly 1,000 sequences,” according to page 2 of https://arxiv.org/pdf/2305.11206. “We train LIMA (Less Is More for Alignment) using the following protocol. Starting from LLaMa 65B [Touvron et al., 2023], we fine-tune on our 1,000-example alignment training set. To differentiate between each speaker (user and assistant), we introduce a special end-of-turn token (EOT) at the end of each utterance; this token plays the same …

Verfügbarkeit

Ob Sie das Modell erhalten und es auf Ihrer eigenen Hardware ausführen können, entscheidet darüber, ob eine der Grafikkartenfiguren auf dieser Seite zutrifft.

Gewichte
Closed — provider access only
Modellzugang
Unreleased
Trainingscode
Unreleased

Wie es klassifiziert ist

Labels der Quelldatenbank, die beim Verfolgen bemerkenswerter Modelle angewendet werden, und wie zuversichtlich sie in den Eintrag ist.

Wahrscheinlich über 10²³ FLOP
Yes
Aufzeichnungszuversicht
Confident
Zitationen
1,274

Quellen

Woher diese Aufzeichnung stammt und wann sie zuletzt überprüft wurde.

Referenz
LIMA: Less Is More for Alignment
Zuletzt aktualisiert
25 May 2026

Was die Zahlen bedeuten

Was dieses Modell ist

LIMA was published by Meta AI,Carnegie Mellon University (CMU),University of Southern California,Tel Aviv University, in United States of America, in May 2023. industry,Academia,Academia,Academia is the category the publisher falls under.

It works in Language, and is recorded as doing chat, Language modeling/generation.

Its starting point was LLaMA-65B — most models at this scale are adapted from an existing base rather than built from nothing.

Because the weights are not available, none of the hardware figures elsewhere on this site apply to it.

Was in den Aufbau geflossen ist

Producing it required around 5.5 × 10²³ FLOP of arithmetic, which is a statement about the training budget rather than about inference.

The training set ran to roughly 685,300 tokens.

Antworten

LIMA — Häufig gestellte Fragen

01

What is LIMA used for?

LIMA works in Language, and is recorded as handling chat, Language modeling/generation. These are the areas it was designed around; they describe intent rather than a hard boundary.

02

How much compute was used to train LIMA?

Around 5.5 × 10²³ FLOP. That measures what producing the model cost and says nothing about how quickly it answers once trained — inference speed comes from memory bandwidth, not from the training budget.

03

What GPU do I need to run LIMA?

None. LIMA is a closed model — its weights were never published, so it cannot be downloaded or run on your own hardware at any price. It is reachable only through its provider.

04

Is LIMA open source?

No. LIMA has not had its weights published, so it exists only as a service controlled by its owner.

05

How many parameters does LIMA have?

LIMA has 65B parameters. "We train LIMA (Less Is More for Alignment) using the following protocol. Starting from LLaMa 65B [Touvron et al., 2023], we fine-tune on our 1,000-example alignment training set," according to page 4 of https://arxiv.org/pdf/2305.11206. That figure is the total, and it is what decides how much memory the model needs — roughly half a gigabyte per billion at the compression most people use.

06

Who created LIMA?

LIMA was published by Meta AI,Carnegie Mellon University (CMU),University of Southern California,Tel Aviv University, based in United States of America, categorised as industry,Academia,Academia,Academia.

07

When was LIMA released?

LIMA was published in May 2023. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.

Quelle

Originalveröffentlichung

Letzte Aktualisierung des Datensatzes 25 May 2026

Die andere Richtung

Wenn Sie es von der anderen Seite betrachten?

Diese Seite beginnt beim Modell. Wenn Sie bereits eine Karte besitzen und alles erfahren möchten, was es ausführen kann, Beginnen Sie stattdessen mit der Hardware.