MetaLM
No estimate
No hardware requirements for this model
The weights for this model have not been published, so it cannot be downloaded or run on your own hardware at any size. It is reachable only through its provider, and no graphics card changes that.
On record
Full specification
Everything on record for this model. Most of it describes how it was trained rather than how it runs — useful context for judging how much work went into it, and how it compares with models built at a different scale.
Origin
Who built this model, where, and when it was published.
- Organisation
- Microsoft Research
- Organisation type
- Industry
- Country
- United States of America
- Published
- 13 June 2022
- Authors
- Yaru Hao, Haoyu Song, Li Dong, Shaohan Huang, Zewen Chi, Wenhui Wang, Shuming Ma, Furu Wei
What it does
The problem areas the model was built for. A model can carry several of each.
- Domain
- Multimodal, Language, Vision
- Task
- Language modeling, Visual question answering, Language modeling/generation, Image captioning
- Approach
- Self-supervised learning
Size
How large the model is and how much data it was trained on. Parameters are the figure that decides whether it fits on a given graphics card.
- Training data
- 646,707,200,000 tokens
- Batch size
- 2,097,152
1024 * 2048 - "The maximum input lengths for non-causal and semi-causal models are 512 and 2048, respectively." (Sec. 3.2, p. 9) and "We pretrain METALM for 300k steps with a batch size of 1024"
Availability
Whether you can obtain the model and run it on your own hardware, which is what decides if any of the graphics-card figures on this page apply.
- Weights
- Closed — provider access only
- Model access
- Unreleased
- Training code
- Unreleased
I don't see neither code nor weights here https://github.com/microsoft/unilm/tree/master/metalm
How it is classified
Labels the source dataset applies when tracking notable models, and how confident it is in the entry.
- Likely above 10²³ FLOP
- Yes
- Why it is tracked
- SOTA improvement
- Record confidence
- Speculative
- Citations
- 110
Abstract: "Experimental results across various language-only and vision-language benchmarks show that our model outperforms or is competitive with specialized models on finetuning, zero-shot generalization, and few-shot learning." "Table 7 and Table 8 show the zero-shot captioning results on COCO Karpathy test split, NoCaps validation set, and Flickr30k test set. METALM outperforms recent strong methods on three image captioning datasets." Table 12
Sources
Where this record came from and when it was last checked.
- Reference
- Language Models are General-Purpose Interfaces
- Last updated
- 25 May 2026
What the numbers mean
Background
MetaLM was published by Microsoft Research, in United States of America, in June 2022. It comes out of industry.
It works in Multimodal, Language, Vision, and is recorded as doing language modeling, Visual question answering, Language modeling/generation, Image captioning.
Its weights were never published, so it can only be reached through its provider. No graphics card changes that.
Training and provenance
It was trained on about 646,707,200,000 tokens of text.
The reason it appears in this catalogue at all is sOTA improvement.
Answers
MetaLM — common questions
Is MetaLM open source?
No. MetaLM has not had its weights published, so it exists only as a service controlled by its owner.
How many parameters does MetaLM have?
No parameter count has been published for MetaLM, which is why no memory or speed figure appears on this page.
Who created MetaLM?
MetaLM was published by Microsoft Research, based in United States of America, categorised as industry.
When was MetaLM released?
MetaLM was published in June 2022. Capability per parameter has improved considerably since, so a newer model of the same size is often the better use of the same hardware.
What is MetaLM used for?
MetaLM works in Multimodal, Language, Vision, and is recorded as handling language modeling, Visual question answering, Language modeling/generation, Image captioning. These are the areas it was designed around; they describe intent rather than a hard boundary.
What GPU do I need to run MetaLM?
None. MetaLM is a closed model — its weights were never published, so it cannot be downloaded or run on your own hardware at any price. It is reachable only through its provider.
The other direction
Looking at it from the other side?
This page starts from the model. If you already own a card and want to know everything it will run, start from the hardware instead.