Architectures, access, and reported results from original sources.
A separate reference from the daily newsletter. Open source indicates an openly licensed release; open weights indicates downloadable weights with custom restrictions. Proprietary models do not release weights. Consult each exact license for terms.
Benchmarks are reported results, not a ranking. Scores from different datasets, versions, prompting, training, or evaluation setups are not directly comparable. “Not reported” means the cited sources do not state the detail; “Not available” means no linked artifact is supplied.
Benchmark sorting uses the displayed metric only. Evaluation setups differ; this order does not establish comparable model performance. Models without a reported score appear last.
No models match these filters. Try a different selection or clear the filters.
Downloadable weights; custom license and acceptable-use restrictions; gated access on Hugging Face.
Autoregressive Transformer with grouped-query attention; SFT and RLHF instruction tuning.
MMLU: 69.4% Instruction-tuned 8B; 5-shot; macro_avg/acc; Meta internal evaluation library. Dataset revision and library version: Not reported. Meta model card
HumanEval: 72.6% Instruction-tuned 8B; 0-shot; pass@1; Meta internal evaluation library. Dataset revision and sampling settings: Not reported. Meta model card
No released weights. At launch: text access through ChatGPT and an API waitlist; image access limited. Historical release, not a current availability guarantee.
Transformer trained for next-token prediction and aligned with RLHF; detailed architecture: Not reported.
MMLU: 86.4% Original 2023 technical report, Table 2; 5-shot accuracy over 57 subjects; text evaluation. Dataset revision: Not reported. Report notes minor differences from standard evaluation setups. Technical report, Table 2
Image classifier trained on ImageNet without external training data. This entry is the baseline model; the paper also introduces attention-based distillation variants.
ImageNet-1K: 81.8% top-1; 95.6% top-5 Official baseline DeiT-base checkpoint; ImageNet 2012 training only, no external data; 224px, single-crop validation. Reference inference uses timm 0.3.2. Distilled and 384px variants have different scores. Official model zoo and evaluation instructions
Downloadable weights; custom license and acceptable-use restrictions; gated access on Hugging Face.
LLM (text → text)
Autoregressive Transformer with grouped-query attention; SFT and RLHF instruction tuning.
8 billion (reported model size)
MMLU: 69.4% Instruction-tuned 8B; 5-shot; macro_avg/acc; Meta internal evaluation library. Dataset revision and library version: Not reported. Meta model card
HumanEval: 72.6% Instruction-tuned 8B; 0-shot; pass@1; Meta internal evaluation library. Dataset revision and sampling settings: Not reported. Meta model card
No released weights. At launch: text access through ChatGPT and an API waitlist; image access limited. Historical release, not a current availability guarantee.
Multimodal LLM / VLM (image + text → text)
Transformer trained for next-token prediction and aligned with RLHF; detailed architecture: Not reported.
Not reported
MMLU: 86.4% Original 2023 technical report, Table 2; 5-shot accuracy over 57 subjects; text evaluation. Dataset revision: Not reported. Report notes minor differences from standard evaluation setups. Technical report, Table 2
ImageNet-1K: 81.8% top-1; 95.6% top-5 Official baseline DeiT-base checkpoint; ImageNet 2012 training only, no external data; 224px, single-crop validation. Reference inference uses timm 0.3.2. Distilled and 384px variants have different scores. Official model zoo and evaluation instructions