| Llama 3.1 8B Instruct | 2024-07-23 | M Meta | Downloadable weights; custom license and acceptable-use restrictions; gated access on Hugging Face. | LLM (text → text) | Autoregressive Transformer with grouped-query attention; SFT and RLHF instruction tuning. | 8 billion (reported model size) | MMLU: 69.4% Instruction-tuned 8B; 5-shot; macro_avg/acc; Meta internal evaluation library. Dataset revision and library version: Not reported. Meta model card HumanEval: 72.6% Instruction-tuned 8B; 0-shot; pass@1; Meta internal evaluation library. Dataset revision and sampling settings: Not reported. Meta model card | 📄 Paper: The Llama 3 Herd of Models 💻 Code: Meta reference implementation | Meta model card |
|---|
| GPT-4 (original, 2023) | 2023-03-14 | O OpenAI | No released weights. At launch: text access through ChatGPT and an API waitlist; image access limited. Historical release, not a current availability guarantee. | Multimodal LLM / VLM (image + text → text) | Transformer trained for next-token prediction and aligned with RLHF; detailed architecture: Not reported. | Not reported | MMLU: 86.4% Original 2023 technical report, Table 2; 5-shot accuracy over 57 subjects; text evaluation. Dataset revision: Not reported. Report notes minor differences from standard evaluation setups. Technical report, Table 2 | 📄 Paper: Technical report, Table 2 💻 Code: Not available | Launch announcement Technical report, Table 2 |
|---|
| CLIP ViT-B/32 | 2021-01-05 | O OpenAI | Public code and pretrained weights; training dataset is not released. | VLM (image–text embedding model) | Dual encoder: ViT-B/32 image encoder and Transformer text encoder; contrastive image–text training. | Not reported | ImageNet-1K: 63.2% top-1 CLIP paper Table 17; ViT-B/32, 224px input; zero-shot classification with prompt ensembling, no ImageNet classifier training. Evaluation uses ImageNet validation; dataset revision: Not reported. CLIP paper, Table 17 | 📄 Paper: Learning Transferable Visual Models From Natural Language Supervision 💻 Code: OpenAI CLIP | Release announcement CLIP paper, Table 17 Official model card |
|---|
| DeiT-Base (224, non-distilled) | Not reported (paper first submitted 2020-12-23) | M Meta FAIR | Public code and baseline pretrained checkpoints linked in the official repository. | Vision model (image classification) | Vision Transformer, base size; 16×16 image patches; 224×224 input; non-distilled baseline. | 86 million (official model zoo) | ImageNet-1K: 81.8% top-1; 95.6% top-5 Official baseline DeiT-base checkpoint; ImageNet 2012 training only, no external data; 224px, single-crop validation. Reference inference uses timm 0.3.2. Distilled and 384px variants have different scores. Official model zoo and evaluation instructions | 📄 Paper: Training data-efficient image transformers & distillation through attention 💻 Code: Official DeiT repository | Official model zoo and evaluation instructions Paper and author affiliations |
|---|
| Jev | 2026-09-15 | T TypeSafe AI | Hosted API in early access; weights are not publicly released. | Decision model (structured decisions) | System One decision model trained with Reinforcement Learning for Calibrated Decisions (RLCD); implementation details are not reported. | Not reported | Not reported | 📄 Paper: Not available 💻 Code: Not available | Official launch post Official product page |
|---|
| DINO ViT-S/16 | 2021-04-29 | M Meta FAIR | Public training code and pretrained model weights; the official repository is archived but remains available. | Vision foundation model | Vision Transformer Small with 16×16 patches, trained by self-distillation with no labels. | 21 million | ImageNet-1K linear: 77.0% top-1 ViT-S/16 frozen backbone with a supervised linear classifier; ImageNet validation top-1 result from the official repository. Dataset revision: Not reported. Official DINO repository | 📄 Paper: Emerging Properties in Self-Supervised Vision Transformers 💻 Code: Official DINO repository | Official DINO repository DINO paper |
|---|
| DINOv2 ViT-L/14 | 2023-04-14 | M Meta FAIR | Public training code and pretrained model weights. | Vision foundation model | Distilled Vision Transformer with 14×14 patches; embedding dimension 1024; trained without labels on LVD-142M. | 300 million | ImageNet-1K linear: 86.3% top-1 ViT-L/14 without registers; linear-classifier evaluation reported in the official model zoo. Dataset revision and training details for the head are documented in the repository. Official DINOv2 model zoo | 📄 Paper: DINOv2: Learning Robust Visual Features without Supervision 💻 Code: Official DINOv2 repository | Official model card Official model zoo |
|---|
| DINOv3 ViT-7B | 2025-08-14 | M Meta FAIR | Pretrained backbones and training code are downloadable under the DINOv3 license. | Vision foundation model | Vision Transformer trained with self-supervised learning on 1.7 billion images. | 7 billion | Not reported | 📄 Paper: DINOv3 💻 Code: Official DINOv3 repository | Official release Official research page |
|---|
| Perception Encoder Core G/14 448 | 2025-04-17 | M Meta FAIR | Public code and downloadable PE-Core-G14-448 checkpoint on Hugging Face. | VLM (image–text embedding model) | Contrastively trained dual encoder with a 50-layer G/14 Vision Transformer, attention pooling, and a 24-layer text Transformer; 448px image input. | 1.88B vision + 0.47B text | ImageNet-1K: 85.4% zero-shot top-1 PE-Core-G14-448 zero-shot image classification at 448px, as reported in the official model table. Prompt and dataset revision details: Not reported in the table. Official Perception Encoder model table | 📄 Paper: Perception Encoder: The Best Visual Embeddings Are Not at the Output of the Network 💻 Code: Official Perception Models repository | Meta research page Official implementation and model table |
|---|
| I-JEPA ViT-H/14 | 2023-01-19 | M Meta FAIR | Public training code and pretrained ViT-H/14 checkpoint; noncommercial license restrictions apply. | Vision foundation model | Vision Transformer Huge with 14×14 patches, trained with a non-generative joint-embedding prediction objective on ImageNet-1K. | Not reported | Not reported | 📄 Paper: Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture 💻 Code: Official I-JEPA repository | Official I-JEPA repository and model zoo I-JEPA paper |
|---|
| V-JEPA ViT-H/16 384 | 2024-02-15 | M Meta FAIR | Public training code, configuration, and pretrained ViT-H/16 checkpoint; noncommercial license restrictions apply. | Video foundation model | Video Vision Transformer Huge with 2×16×16 spatiotemporal patches at 384px, trained by predicting masked features in latent space on VideoMix2M. | Not reported | Kinetics-400: 81.9% top-1 Frozen V-JEPA ViT-H/16 384px backbone with an attentive probe; 16×8×3 evaluation views, as reported in the official model zoo. Official V-JEPA model zoo | 📄 Paper: Revisiting Feature Prediction for Learning Visual Representations from Video 💻 Code: Official V-JEPA repository | Meta V-JEPA release Official V-JEPA model zoo |
|---|
| GPT-5.6 Sol | 2026-07-09 | O OpenAI | Available as a hosted OpenAI model; weights and detailed architecture are not released. | Multimodal LLM / VLM (image + text → text) | Not reported. | Not reported | Agents’ Last Exam: 53.6 OpenAI evaluation of long-running professional workflows across 55 fields; GPT-5.6 Sol result. Exact public harness version and statistical uncertainty: Not reported. Official GPT-5.6 release | 📄 Paper: Not available 💻 Code: Not available | Official release |
|---|
| GPT-6 Astra | 2026-09-03 | O OpenAI | Available through OpenAI-hosted products and services; weights are not released. | Multimodal LLM / VLM (image + text → text) | Not reported. | Not reported | Not reported | 📄 Paper: Not available 💻 Code: Not available | Official safety overview |
|---|
| Claude Opus 5.5 | 2026-09-22 | A Anthropic | Available through Claude and the Claude Platform; weights are not released. | Multimodal LLM / VLM (image + text → text) | Not reported. | Not reported | Terminal-Bench 4.0: 66.4% Claude Opus 5.5 at xhigh effort; production safeguards enabled; standard error ±2.6 points; Anthropic-reported evaluation. Official Claude Opus 5.5 release Humanity's Last Exam: 67.7% with tools Adaptive thinking at max effort with tools; production safeguards enabled; Anthropic-reported evaluation. Official Claude Opus 5.5 release | 📄 Paper: Claude Opus 5.5 system card 💻 Code: Not available | Official release |
|---|
| Claude Fable 5.1 | 2026-09-01 | A Anthropic | Available through Anthropic-hosted products and APIs; weights are not released. | Multimodal LLM / VLM (image + text → text) | Not reported. | Not reported | Not reported | 📄 Paper: Not available 💻 Code: Not available | Anthropic newsroom release listing |
|---|
| Pi 3.0 | Not reported | I Inflection AI | Hosted in Pi and available through the Inflection API as inflection_3_pi; weights are not released. | LLM (text → text) | Custom fine-tuned Inflection foundation model; detailed architecture is not reported. | Not reported | Not reported | 📄 Paper: Not available 💻 Code: Not available | Official Inflection API documentation Official Pi product page |
|---|
| Gemini 3.8 Flash | 2026-09-02 | G Google DeepMind | Available through Google AI Studio and Google-hosted services; weights are not released. | Multimodal LLM / VLM (image + text → text) | Not reported. | Not reported | HLE-Verified: 54.9% Google-reported result for Gemini 3.8 Flash. The release describes a multi-step reasoning evaluation; exact harness and effort configuration: Not reported. Official Gemini 3.8 Flash release | 📄 Paper: Not available 💻 Code: Not available | Official release |
|---|
| Gemma 4 12B Unified | 2026-06-03 | G Google DeepMind | Downloadable weights under the Gemma terms; local and hosted deployment options vary. | Multimodal LLM / VLM (image + text → text) | Not reported in the cited release log. | 12 billion | Not reported | 📄 Paper: Not available 💻 Code: Not available | Official Gemma release log |
|---|
| Llama 3.3 70B Instruct | 2024-12-06 | M Meta | Downloadable weights under the Llama 3.3 license; available locally through Ollama as llama3.3. | LLM (text → text) | Autoregressive Transformer with grouped-query attention; instruction tuning. | 70 billion | Not reported | 📄 Paper: Llama 3 model paper 💻 Code: Meta Llama models repository | Official Ollama listing Official Meta repository |
|---|
| Mistral Small 3.1 24B | 2025-03-17 | M Mistral AI | Public base and instruction checkpoints; available locally through Ollama as mistral-small3.1. | Multimodal LLM / VLM (image + text → text) | Dense Transformer with a vision encoder and up to 128k context. | 24 billion | Not reported | 📄 Paper: Not available 💻 Code: Official model repository | Official release Official Ollama listing |
|---|
| Phi-4 14B | 2024-12-12 | M Microsoft | Downloadable weights under MIT; available locally through Ollama as phi4. | LLM (text → text) | Dense decoder-only Transformer. | 14 billion | Not reported | 📄 Paper: Phi-4 technical report 💻 Code: Official Phi model repository | Microsoft Phi cookbook Official Ollama listing |
|---|
| Qwen3 32B | 2025-04-29 | Q Alibaba Qwen | Downloadable weights; supported by local runtimes including Ollama as qwen3:32b. | LLM (text → text) | Dense decoder-only Transformer; 64 layers, 64 query heads, 8 key/value heads, 128k context. | 32 billion | Not reported | 📄 Paper: Qwen3 technical report 💻 Code: Official Qwen3 repository | Official Qwen3 release Official Ollama listing |
|---|
| DeepSeek-V3.2 | 2025-12-02 | D DeepSeek | Downloadable model weights and inference code; very large hardware requirements apply. | LLM (text → text) | Mixture-of-experts Transformer with DeepSeek Sparse Attention. | 685 billion total parameters | GPQA Diamond: 82.4 Score listed in the official Hugging Face model repository evaluation results; exact prompting, sampling, and harness version: Not reported. Official DeepSeek-V3.2 model repository | 📄 Paper: DeepSeek-V3.2 technical report 💻 Code: Official model repository | Official model repository |
|---|
| Qwen3-VL 235B-A22B Instruct | 2025-09-23 | Q Alibaba Qwen | Downloadable model weights and public inference code; the 235B-total-parameter checkpoint requires substantial accelerator memory. | Multimodal LLM / VLM (image + video + text → text) | Mixture-of-experts Qwen3 language backbone with a SigLIP 2 vision encoder, MLP vision-language merger, Interleaved-MRoPE, DeepStack feature fusion, and explicit video timestamps; 256k native context. | 235 billion total; 22 billion active per token | Not reported | 📄 Paper: Qwen3-VL Technical Report 💻 Code: Official Qwen3-VL repository | Official Qwen3-VL repository Official model repository |
|---|
| Llama 4 Maverick 17B-128E Instruct | 2025-04-05 | M Meta | Downloadable weights under the Llama 4 Community License and acceptable-use policy; multimodal license rights include a European Union restriction described in the use policy. | Multimodal LLM / VLM (image + text → text) | Autoregressive early-fusion mixture-of-experts Transformer with 128 routed experts, one shared expert, and a MetaCLIP-derived vision encoder; 1M-token context. | 400 billion total; 17 billion active | MMLU: 85.5% Meta model card result for the pretrained Maverick model; 5-shot macro-average character accuracy; bf16 evaluation. Exact harness version and dataset revision: Not reported. Official Llama 4 model card | 📄 Paper: Not available 💻 Code: Meta Llama models repository | Official Llama 4 model card Meta release announcement |
|---|
| Molmo 2 8B | 2025-12-11 | A Ai2 | Downloadable weights and inference examples. Ai2 reports that training code, evaluations, and intermediate checkpoints will be released later; users should review the licenses of underlying third-party training datasets. | Multimodal LLM / VLM (image + video + text → text + grounding) | Qwen3-8B language backbone paired with a SigLIP 2 vision encoder and trained for image, multi-image, and video understanding and grounding. | 8 billion class | 15 academic video benchmarks: 63.1 average Ai2-reported average across 15 academic benchmarks for Molmo2-8B; individual tasks and evaluation details are documented in the technical report. Official Molmo 2 model card | 📄 Paper: Molmo2 technical report 💻 Code: Official model repository and examples | Ai2 release announcement Official model repository |
|---|
| Mistral Small 4 119B-A6B | 2026-03-16 | M Mistral AI | Downloadable weights under Apache 2.0, plus hosted access through Mistral AI Studio and the Mistral API. | Multimodal LLM / VLM (image + text → text) | Mixture-of-experts Transformer with 128 experts and four active experts per token, native image input, configurable reasoning effort, and a 256k context window. | 119 billion total; 6 billion active per token (8 billion including embeddings and output layers) | Not reported | 📄 Paper: Not available 💻 Code: Official model repository | Official release announcement Official model repository |
|---|
| V-JEPA 2 ViT-g/16 384 | 2025-06-11 | M Meta FAIR | Public pretrained checkpoints, evaluation probes, training and evaluation configurations, PyTorch code, and Hugging Face integration. | Video foundation model | ViT-g/16 video encoder trained by predicting latent representations of masked video regions; this entry uses the 384-pixel, 64-frame checkpoint. | 1 billion | SSv2: 77.3% Official attentive-probe result using frozen V-JEPA 2 ViT-g/16 384 features; training and inference configurations are linked from the repository. Official V-JEPA 2 evaluation table | 📄 Paper: V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning 💻 Code: Official V-JEPA 2 repository | Official V-JEPA 2 repository Official model repository |
|---|
| Qwen3-Omni 30B-A3B Instruct | 2025-09-22 | Q Alibaba Qwen | Downloadable weights, public inference code, Docker image, cookbooks, and hosted API access; local use has substantial GPU-memory requirements. | Omnimodal model (text + image + audio + video → text + speech) | End-to-end mixture-of-experts Thinker–Talker architecture with audio-text pretraining and a multi-codebook speech generator for real-time streaming. | 30 billion total; 3 billion active | Video-MME: 70.5 Official repository result for Qwen3-Omni-30B-A3B-Instruct in the vision-to-text video-understanding table; exact harness version: Not reported. Official Qwen3-Omni evaluation table | 📄 Paper: Qwen3-Omni Technical Report 💻 Code: Official Qwen3-Omni repository | Official Qwen3-Omni repository Official model repository |
|---|