A reference for AI models

Model Directory.

Architectures, access, and reported results from original sources.

A separate reference from the daily newsletter. Open source indicates an openly licensed release; open weights indicates downloadable weights with custom restrictions. Proprietary models do not release weights. Consult each exact license for terms.

Benchmarks are reported results, not a ranking. Scores from different datasets, versions, prompting, training, or evaluation setups are not directly comparable. “Not reported” means the cited sources do not state the detail; “Not available” means no linked artifact is supplied.

LLM (text → text)

Llama 3.1 8B Instruct

  • Release date: 2024-07-23
  • Organization: Meta
  • Parameters: 8 billion (reported model size)
License and availability:

Text assistant supporting eight languages, with a 128k context window and instruction tuning for dialogue.

Reported scores; setups differ. Expand for evaluation details.

Show moreShow less about Llama 3.1 8B Instruct
Downloadable weights; custom license and acceptable-use restrictions; gated access on Hugging Face.
Autoregressive Transformer with grouped-query attention; SFT and RLHF instruction tuning.

MMLU: 69.4%
Instruction-tuned 8B; 5-shot; macro_avg/acc; Meta internal evaluation library. Dataset revision and library version: Not reported.
Meta model card

HumanEval: 72.6%
Instruction-tuned 8B; 0-shot; pass@1; Meta internal evaluation library. Dataset revision and sampling settings: Not reported.
Meta model card

Multimodal LLM / VLM (image + text → text)

GPT-4 (original, 2023)

  • Release date: 2023-03-14
  • Organization: OpenAI
  • Parameters: Parameters: Not reported

Multimodal model for image and text understanding and text generation. This entry describes the original report, rather than later GPT-4 variants.

Reported scores; setups differ. Expand for evaluation details.

Show moreShow less about GPT-4 (original, 2023)
No released weights. At launch: text access through ChatGPT and an API waitlist; image access limited. Historical release, not a current availability guarantee.
Transformer trained for next-token prediction and aligned with RLHF; detailed architecture: Not reported.

MMLU: 86.4%
Original 2023 technical report, Table 2; 5-shot accuracy over 57 subjects; text evaluation. Dataset revision: Not reported. Report notes minor differences from standard evaluation setups.
Technical report, Table 2

VLM (image–text embedding model)

CLIP ViT-B/32

  • Release date: 2021-01-05
  • Organization: OpenAI
  • Parameters: Parameters: Not reported
License and availability:

Contrastively trained image and text encoders for retrieval and zero-shot classification using natural-language class descriptions.

Reported scores; setups differ. Expand for evaluation details.

Show moreShow less about CLIP ViT-B/32
Public code and pretrained weights; training dataset is not released.
Dual encoder: ViT-B/32 image encoder and Transformer text encoder; contrastive image–text training.

ImageNet-1K: 63.2% top-1
CLIP paper Table 17; ViT-B/32, 224px input; zero-shot classification with prompt ensembling, no ImageNet classifier training. Evaluation uses ImageNet validation; dataset revision: Not reported.
CLIP paper, Table 17

Vision model (image classification)

DeiT-Base (224, non-distilled)

  • Release date: Not reported (paper first submitted 2020-12-23)
  • Organization: Meta FAIR
  • Parameters: 86 million (official model zoo)
License and availability:

Image classifier trained on ImageNet without external training data. This entry is the baseline model; the paper also introduces attention-based distillation variants.

Reported scores; setups differ. Expand for evaluation details.

Show moreShow less about DeiT-Base (224, non-distilled)
Public code and baseline pretrained checkpoints linked in the official repository.
Vision Transformer, base size; 16×16 image patches; 224×224 input; non-distilled baseline.

ImageNet-1K: 81.8% top-1; 95.6% top-5
Official baseline DeiT-base checkpoint; ImageNet 2012 training only, no external data; 224px, single-crop validation. Reference inference uses timm 0.3.2. Distilled and 384px variants have different scores.
Official model zoo and evaluation instructions

Decision model (structured decisions)

Jev

  • Release date: 2026-09-15
  • Organization: TypeSafe AI
  • Parameters: Parameters: Not reported

TypeSafe's System One model for fast, bounded, typed decisions such as choices, scores, and yes/no probabilities rather than free-form generation.

  • Benchmarks: Not reported
Show moreShow less about Jev
Hosted API in early access; weights are not publicly released.
System One decision model trained with Reinforcement Learning for Calibrated Decisions (RLCD); implementation details are not reported.
Benchmarks: Not reported

Vision foundation model

DINO ViT-S/16

  • Release date: 2021-04-29
  • Organization: Meta FAIR
  • Parameters: 21 million
License and availability:

The original DINO self-supervised vision model, using self-distillation without labels to learn reusable image features and emergent object segmentation cues.

Reported scores; setups differ. Expand for evaluation details.

Show moreShow less about DINO ViT-S/16
Public training code and pretrained model weights; the official repository is archived but remains available.
Vision Transformer Small with 16×16 patches, trained by self-distillation with no labels.

ImageNet-1K linear: 77.0% top-1
ViT-S/16 frozen backbone with a supervised linear classifier; ImageNet validation top-1 result from the official repository. Dataset revision: Not reported.
Official DINO repository

Vision foundation model

DINOv2 ViT-L/14

  • Release date: 2023-04-14
  • Organization: Meta FAIR
  • Parameters: 300 million
License and availability:

Self-supervised vision backbone designed to provide reusable image features for classification, retrieval, depth estimation, and segmentation.

Reported scores; setups differ. Expand for evaluation details.

Show moreShow less about DINOv2 ViT-L/14
Public training code and pretrained model weights.
Distilled Vision Transformer with 14×14 patches; embedding dimension 1024; trained without labels on LVD-142M.

ImageNet-1K linear: 86.3% top-1
ViT-L/14 without registers; linear-classifier evaluation reported in the official model zoo. Dataset revision and training details for the head are documented in the repository.
Official DINOv2 model zoo

Vision foundation model

DINOv3 ViT-7B

  • Release date: 2025-08-14
  • Organization: Meta FAIR
  • Parameters: 7 billion
License and availability:

Large self-supervised vision backbone producing high-resolution features for classification and dense tasks such as detection, segmentation, tracking, and depth.

  • Benchmarks: Not reported
Show moreShow less about DINOv3 ViT-7B
Pretrained backbones and training code are downloadable under the DINOv3 license.
Vision Transformer trained with self-supervised learning on 1.7 billion images.
Benchmarks: Not reported

VLM (image–text embedding model)

Perception Encoder Core G/14 448

  • Release date: 2025-04-17
  • Organization: Meta FAIR
  • Parameters: 1.88B vision + 0.47B text
License and availability:

Meta's largest Perception Encoder core checkpoint for general image and video embeddings, zero-shot classification and retrieval, and downstream spatial or language alignment.

Reported scores; setups differ. Expand for evaluation details.

Show moreShow less about Perception Encoder Core G/14 448
Public code and downloadable PE-Core-G14-448 checkpoint on Hugging Face.
Contrastively trained dual encoder with a 50-layer G/14 Vision Transformer, attention pooling, and a 24-layer text Transformer; 448px image input.

ImageNet-1K: 85.4% zero-shot top-1
PE-Core-G14-448 zero-shot image classification at 448px, as reported in the official model table. Prompt and dataset revision details: Not reported in the table.
Official Perception Encoder model table

Vision foundation model

I-JEPA ViT-H/14

  • Release date: 2023-01-19
  • Organization: Meta FAIR
  • Parameters: Parameters: Not reported
License and availability:

Image-based Joint-Embedding Predictive Architecture that learns semantic image representations by predicting masked regions in representation space rather than reconstructing pixels.

  • Benchmarks: Not reported
Show moreShow less about I-JEPA ViT-H/14
Public training code and pretrained ViT-H/14 checkpoint; noncommercial license restrictions apply.
Vision Transformer Huge with 14×14 patches, trained with a non-generative joint-embedding prediction objective on ImageNet-1K.
Benchmarks: Not reported

Video foundation model

V-JEPA ViT-H/16 384

  • Release date: 2024-02-15
  • Organization: Meta FAIR
  • Parameters: Parameters: Not reported
License and availability:

Video Joint-Embedding Predictive Architecture trained without labels to predict masked video representations and learn motion- and appearance-aware features.

Reported scores; setups differ. Expand for evaluation details.

Show moreShow less about V-JEPA ViT-H/16 384
Public training code, configuration, and pretrained ViT-H/16 checkpoint; noncommercial license restrictions apply.
Video Vision Transformer Huge with 2×16×16 spatiotemporal patches at 384px, trained by predicting masked features in latent space on VideoMix2M.

Kinetics-400: 81.9% top-1
Frozen V-JEPA ViT-H/16 384px backbone with an attentive probe; 16×8×3 evaluation views, as reported in the official model zoo.
Official V-JEPA model zoo

Multimodal LLM / VLM (image + text → text)

GPT-5.6 Sol

  • Release date: 2026-07-09
  • Organization: OpenAI
  • Parameters: Parameters: Not reported

OpenAI flagship in the GPT-5.6 family for coding, knowledge work, cybersecurity, science, and long-running agent workflows.

Reported scores; setups differ. Expand for evaluation details.

Show moreShow less about GPT-5.6 Sol
Available as a hosted OpenAI model; weights and detailed architecture are not released.
Not reported.

Agents’ Last Exam: 53.6
OpenAI evaluation of long-running professional workflows across 55 fields; GPT-5.6 Sol result. Exact public harness version and statistical uncertainty: Not reported.
Official GPT-5.6 release

Multimodal LLM / VLM (image + text → text)

GPT-6 Astra

  • Release date: 2026-09-03
  • Organization: OpenAI
  • Parameters: Parameters: Not reported

OpenAI frontier model broadly deployed with stronger coding, scientific, and cybersecurity capabilities and strengthened safeguards.

  • Benchmarks: Not reported
Show moreShow less about GPT-6 Astra
Available through OpenAI-hosted products and services; weights are not released.
Not reported.
Benchmarks: Not reported

Multimodal LLM / VLM (image + text → text)

Claude Opus 5.5

  • Release date: 2026-09-22
  • Organization: Anthropic
  • Parameters: Parameters: Not reported

Anthropic flagship for long-running coding, computer use, knowledge work, research, and complex agentic tasks.

Reported scores; setups differ. Expand for evaluation details.

Show moreShow less about Claude Opus 5.5
Available through Claude and the Claude Platform; weights are not released.
Not reported.

Terminal-Bench 4.0: 66.4%
Claude Opus 5.5 at xhigh effort; production safeguards enabled; standard error ±2.6 points; Anthropic-reported evaluation.
Official Claude Opus 5.5 release

Humanity's Last Exam: 67.7% with tools
Adaptive thinking at max effort with tools; production safeguards enabled; Anthropic-reported evaluation.
Official Claude Opus 5.5 release

Multimodal LLM / VLM (image + text → text)

Claude Fable 5.1

  • Release date: 2026-09-01
  • Organization: Anthropic
  • Parameters: Parameters: Not reported

Anthropic frontier model for coding, knowledge work, vision, scientific research, and extended agentic tasks.

  • Benchmarks: Not reported
Show moreShow less about Claude Fable 5.1
Available through Anthropic-hosted products and APIs; weights are not released.
Not reported.
Benchmarks: Not reported

LLM (text → text)

Pi 3.0

  • Release date: Not reported
  • Organization: Inflection AI
  • Parameters: Parameters: Not reported
License and availability:

Inflection AI personal assistant model tuned for emotionally intelligent, conversational support, customer service, productivity, and safety.

  • Benchmarks: Not reported
Show moreShow less about Pi 3.0
Hosted in Pi and available through the Inflection API as inflection_3_pi; weights are not released.
Custom fine-tuned Inflection foundation model; detailed architecture is not reported.
Benchmarks: Not reported

Multimodal LLM / VLM (image + text → text)

Gemini 3.8 Flash

  • Release date: 2026-09-02
  • Organization: Google DeepMind
  • Parameters: Parameters: Not reported

Google workhorse model for fast reasoning, coding, and long-horizon autonomous agent workflows.

Reported scores; setups differ. Expand for evaluation details.

Show moreShow less about Gemini 3.8 Flash
Available through Google AI Studio and Google-hosted services; weights are not released.
Not reported.

HLE-Verified: 54.9%
Google-reported result for Gemini 3.8 Flash. The release describes a multi-step reasoning evaluation; exact harness and effort configuration: Not reported.
Official Gemini 3.8 Flash release

Multimodal LLM / VLM (image + text → text)

Gemma 4 12B Unified

  • Release date: 2026-06-03
  • Organization: Google DeepMind
  • Parameters: 12 billion
License and availability:
Open weights
Gemma Terms of Use

Google open-weight Gemma-family model intended for local and customizable generative AI applications.

  • Benchmarks: Not reported
Show moreShow less about Gemma 4 12B Unified
Downloadable weights under the Gemma terms; local and hosted deployment options vary.
Not reported in the cited release log.
Benchmarks: Not reported

LLM (text → text)

Llama 3.3 70B Instruct

  • Release date: 2024-12-06
  • Organization: Meta
  • Parameters: 70 billion
License and availability:

Multilingual instruction-tuned text model optimized for dialogue, tool use, and local or hosted deployment.

  • Benchmarks: Not reported
Show moreShow less about Llama 3.3 70B Instruct
Downloadable weights under the Llama 3.3 license; available locally through Ollama as llama3.3.
Autoregressive Transformer with grouped-query attention; instruction tuning.
Benchmarks: Not reported

Multimodal LLM / VLM (image + text → text)

Mistral Small 3.1 24B

  • Release date: 2025-03-17
  • Organization: Mistral AI
  • Parameters: 24 billion
License and availability:

Multimodal small model for instruction following, conversation, image understanding, function calling, and local deployment.

  • Benchmarks: Not reported
Show moreShow less about Mistral Small 3.1 24B
Public base and instruction checkpoints; available locally through Ollama as mistral-small3.1.
Dense Transformer with a vision encoder and up to 128k context.
Benchmarks: Not reported

LLM (text → text)

Phi-4 14B

  • Release date: 2024-12-12
  • Organization: Microsoft
  • Parameters: 14 billion
License and availability:
Open weights
MIT License

Microsoft small language model focused on strong mathematical, reasoning, and coding capabilities at a locally deployable size.

  • Benchmarks: Not reported
Show moreShow less about Phi-4 14B
Downloadable weights under MIT; available locally through Ollama as phi4.
Dense decoder-only Transformer.
Benchmarks: Not reported

LLM (text → text)

Qwen3 32B

  • Release date: 2025-04-29
  • Organization: Alibaba Qwen
  • Parameters: 32 billion
License and availability:

Dense bilingual and multilingual reasoning model with switchable thinking and non-thinking modes for dialogue, coding, math, and agents.

  • Benchmarks: Not reported
Show moreShow less about Qwen3 32B
Downloadable weights; supported by local runtimes including Ollama as qwen3:32b.
Dense decoder-only Transformer; 64 layers, 64 query heads, 8 key/value heads, 128k context.
Benchmarks: Not reported

LLM (text → text)

DeepSeek-V3.2

  • Release date: 2025-12-02
  • Organization: DeepSeek
  • Parameters: 685 billion total parameters
License and availability:
Open source
MIT License

Open-weight reasoning and agent model combining sparse attention with scalable reinforcement learning for efficient long-context use.

Reported scores; setups differ. Expand for evaluation details.

Show moreShow less about DeepSeek-V3.2
Downloadable model weights and inference code; very large hardware requirements apply.
Mixture-of-experts Transformer with DeepSeek Sparse Attention.

GPQA Diamond: 82.4
Score listed in the official Hugging Face model repository evaluation results; exact prompting, sampling, and harness version: Not reported.
Official DeepSeek-V3.2 model repository

Multimodal LLM / VLM (image + video + text → text)

Qwen3-VL 235B-A22B Instruct

  • Release date: 2025-09-23
  • Organization: Alibaba Qwen
  • Parameters: 235 billion total; 22 billion active per token
License and availability:

Flagship open vision-language model for image, multi-image, document, video, spatial, and GUI-agent understanding with a native 256k context window.

  • Benchmarks: Not reported
Show moreShow less about Qwen3-VL 235B-A22B Instruct
Downloadable model weights and public inference code; the 235B-total-parameter checkpoint requires substantial accelerator memory.
Mixture-of-experts Qwen3 language backbone with a SigLIP 2 vision encoder, MLP vision-language merger, Interleaved-MRoPE, DeepStack feature fusion, and explicit video timestamps; 256k native context.
Benchmarks: Not reported

Multimodal LLM / VLM (image + text → text)

Llama 4 Maverick 17B-128E Instruct

  • Release date: 2025-04-05
  • Organization: Meta
  • Parameters: 400 billion total; 17 billion active
License and availability:

Native multimodal Llama model for multilingual text and image understanding, visual reasoning, assistant tasks, and code generation with a 1M-token context window.

Reported scores; setups differ. Expand for evaluation details.

Show moreShow less about Llama 4 Maverick 17B-128E Instruct
Downloadable weights under the Llama 4 Community License and acceptable-use policy; multimodal license rights include a European Union restriction described in the use policy.
Autoregressive early-fusion mixture-of-experts Transformer with 128 routed experts, one shared expert, and a MetaCLIP-derived vision encoder; 1M-token context.

MMLU: 85.5%
Meta model card result for the pretrained Maverick model; 5-shot macro-average character accuracy; bf16 evaluation. Exact harness version and dataset revision: Not reported.
Official Llama 4 model card

Multimodal LLM / VLM (image + video + text → text + grounding)

Molmo 2 8B

  • Release date: 2025-12-11
  • Organization: Ai2
  • Parameters: 8 billion class

Open vision-language model for image, multi-image, and video understanding and grounding, including pointing, counting, tracking, captioning, and question answering.

Reported scores; setups differ. Expand for evaluation details.

Show moreShow less about Molmo 2 8B
Downloadable weights and inference examples. Ai2 reports that training code, evaluations, and intermediate checkpoints will be released later; users should review the licenses of underlying third-party training datasets.
Qwen3-8B language backbone paired with a SigLIP 2 vision encoder and trained for image, multi-image, and video understanding and grounding.

15 academic video benchmarks: 63.1 average
Ai2-reported average across 15 academic benchmarks for Molmo2-8B; individual tasks and evaluation details are documented in the technical report.
Official Molmo 2 model card

Multimodal LLM / VLM (image + text → text)

Mistral Small 4 119B-A6B

  • Release date: 2026-03-16
  • Organization: Mistral AI
  • Parameters: 119 billion total; 6 billion active per token (8 billion including embeddings and output layers)
License and availability:

Open multimodal model combining general chat, configurable reasoning, coding agents, function calling, and image understanding in one checkpoint.

  • Benchmarks: Not reported
Show moreShow less about Mistral Small 4 119B-A6B
Downloadable weights under Apache 2.0, plus hosted access through Mistral AI Studio and the Mistral API.
Mixture-of-experts Transformer with 128 experts and four active experts per token, native image input, configurable reasoning effort, and a 256k context window.
Benchmarks: Not reported

Video foundation model

V-JEPA 2 ViT-g/16 384

  • Release date: 2025-06-11
  • Organization: Meta FAIR
  • Parameters: 1 billion
License and availability:

Self-supervised video foundation model that learns predictive world representations from video for motion understanding, anticipation, video question answering, and robot planning.

Reported scores; setups differ. Expand for evaluation details.

Show moreShow less about V-JEPA 2 ViT-g/16 384
Public pretrained checkpoints, evaluation probes, training and evaluation configurations, PyTorch code, and Hugging Face integration.
ViT-g/16 video encoder trained by predicting latent representations of masked video regions; this entry uses the 384-pixel, 64-frame checkpoint.

SSv2: 77.3%
Official attentive-probe result using frozen V-JEPA 2 ViT-g/16 384 features; training and inference configurations are linked from the repository.
Official V-JEPA 2 evaluation table

Omnimodal model (text + image + audio + video → text + speech)

Qwen3-Omni 30B-A3B Instruct

  • Release date: 2025-09-22
  • Organization: Alibaba Qwen
  • Parameters: 30 billion total; 3 billion active
License and availability:

End-to-end omni-modal model that understands text, images, audio, and video and streams responses as text or natural speech.

Reported scores; setups differ. Expand for evaluation details.

Show moreShow less about Qwen3-Omni 30B-A3B Instruct
Downloadable weights, public inference code, Docker image, cookbooks, and hosted API access; local use has substantial GPU-memory requirements.
End-to-end mixture-of-experts Thinker–Talker architecture with audio-text pretraining and a multi-codebook speech generator for real-time streaming.

Video-MME: 70.5
Official repository result for Qwen3-Omni-30B-A3B-Instruct in the vision-to-text video-understanding table; exact harness version: Not reported.
Official Qwen3-Omni evaluation table

On small screens, scroll horizontally to read every field.

Model directory — scores use the evaluation setups shown
ModelRelease dateOrganizationLicense & availabilityModel typeArchitectureParametersBenchmarksPaper & codeSources
Llama 3.1 8B Instruct2024-07-23MetaDownloadable weights; custom license and acceptable-use restrictions; gated access on Hugging Face.LLM (text → text)Autoregressive Transformer with grouped-query attention; SFT and RLHF instruction tuning.8 billion (reported model size)

MMLU: 69.4%
Instruction-tuned 8B; 5-shot; macro_avg/acc; Meta internal evaluation library. Dataset revision and library version: Not reported.
Meta model card

HumanEval: 72.6%
Instruction-tuned 8B; 0-shot; pass@1; Meta internal evaluation library. Dataset revision and sampling settings: Not reported.
Meta model card

📄 Paper: The Llama 3 Herd of Models
💻 Code: Meta reference implementation
Meta model card
GPT-4 (original, 2023)2023-03-14OpenAINo released weights. At launch: text access through ChatGPT and an API waitlist; image access limited. Historical release, not a current availability guarantee.Multimodal LLM / VLM (image + text → text)Transformer trained for next-token prediction and aligned with RLHF; detailed architecture: Not reported.Not reported

MMLU: 86.4%
Original 2023 technical report, Table 2; 5-shot accuracy over 57 subjects; text evaluation. Dataset revision: Not reported. Report notes minor differences from standard evaluation setups.
Technical report, Table 2

📄 Paper: Technical report, Table 2
💻 Code: Not available
Launch announcement
Technical report, Table 2
CLIP ViT-B/322021-01-05OpenAIPublic code and pretrained weights; training dataset is not released.VLM (image–text embedding model)Dual encoder: ViT-B/32 image encoder and Transformer text encoder; contrastive image–text training.Not reported

ImageNet-1K: 63.2% top-1
CLIP paper Table 17; ViT-B/32, 224px input; zero-shot classification with prompt ensembling, no ImageNet classifier training. Evaluation uses ImageNet validation; dataset revision: Not reported.
CLIP paper, Table 17

📄 Paper: Learning Transferable Visual Models From Natural Language Supervision
💻 Code: OpenAI CLIP
Release announcement
CLIP paper, Table 17
Official model card
DeiT-Base (224, non-distilled)Not reported (paper first submitted 2020-12-23)Meta FAIRPublic code and baseline pretrained checkpoints linked in the official repository.Vision model (image classification)Vision Transformer, base size; 16×16 image patches; 224×224 input; non-distilled baseline.86 million (official model zoo)

ImageNet-1K: 81.8% top-1; 95.6% top-5
Official baseline DeiT-base checkpoint; ImageNet 2012 training only, no external data; 224px, single-crop validation. Reference inference uses timm 0.3.2. Distilled and 384px variants have different scores.
Official model zoo and evaluation instructions

📄 Paper: Training data-efficient image transformers & distillation through attention
💻 Code: Official DeiT repository
Official model zoo and evaluation instructions
Paper and author affiliations
Jev2026-09-15TypeSafe AIHosted API in early access; weights are not publicly released.Decision model (structured decisions)System One decision model trained with Reinforcement Learning for Calibrated Decisions (RLCD); implementation details are not reported.Not reportedNot reported📄 Paper: Not available
💻 Code: Not available
Official launch post
Official product page
DINO ViT-S/162021-04-29Meta FAIRPublic training code and pretrained model weights; the official repository is archived but remains available.Vision foundation modelVision Transformer Small with 16×16 patches, trained by self-distillation with no labels.21 million

ImageNet-1K linear: 77.0% top-1
ViT-S/16 frozen backbone with a supervised linear classifier; ImageNet validation top-1 result from the official repository. Dataset revision: Not reported.
Official DINO repository

📄 Paper: Emerging Properties in Self-Supervised Vision Transformers
💻 Code: Official DINO repository
Official DINO repository
DINO paper
DINOv2 ViT-L/142023-04-14Meta FAIRPublic training code and pretrained model weights.Vision foundation modelDistilled Vision Transformer with 14×14 patches; embedding dimension 1024; trained without labels on LVD-142M.300 million

ImageNet-1K linear: 86.3% top-1
ViT-L/14 without registers; linear-classifier evaluation reported in the official model zoo. Dataset revision and training details for the head are documented in the repository.
Official DINOv2 model zoo

📄 Paper: DINOv2: Learning Robust Visual Features without Supervision
💻 Code: Official DINOv2 repository
Official model card
Official model zoo
DINOv3 ViT-7B2025-08-14Meta FAIRPretrained backbones and training code are downloadable under the DINOv3 license.Vision foundation modelVision Transformer trained with self-supervised learning on 1.7 billion images.7 billionNot reported📄 Paper: DINOv3
💻 Code: Official DINOv3 repository
Official release
Official research page
Perception Encoder Core G/14 4482025-04-17Meta FAIRPublic code and downloadable PE-Core-G14-448 checkpoint on Hugging Face.VLM (image–text embedding model)Contrastively trained dual encoder with a 50-layer G/14 Vision Transformer, attention pooling, and a 24-layer text Transformer; 448px image input.1.88B vision + 0.47B text

ImageNet-1K: 85.4% zero-shot top-1
PE-Core-G14-448 zero-shot image classification at 448px, as reported in the official model table. Prompt and dataset revision details: Not reported in the table.
Official Perception Encoder model table

📄 Paper: Perception Encoder: The Best Visual Embeddings Are Not at the Output of the Network
💻 Code: Official Perception Models repository
Meta research page
Official implementation and model table
I-JEPA ViT-H/142023-01-19Meta FAIRPublic training code and pretrained ViT-H/14 checkpoint; noncommercial license restrictions apply.Vision foundation modelVision Transformer Huge with 14×14 patches, trained with a non-generative joint-embedding prediction objective on ImageNet-1K.Not reportedNot reported📄 Paper: Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
💻 Code: Official I-JEPA repository
Official I-JEPA repository and model zoo
I-JEPA paper
V-JEPA ViT-H/16 3842024-02-15Meta FAIRPublic training code, configuration, and pretrained ViT-H/16 checkpoint; noncommercial license restrictions apply.Video foundation modelVideo Vision Transformer Huge with 2×16×16 spatiotemporal patches at 384px, trained by predicting masked features in latent space on VideoMix2M.Not reported

Kinetics-400: 81.9% top-1
Frozen V-JEPA ViT-H/16 384px backbone with an attentive probe; 16×8×3 evaluation views, as reported in the official model zoo.
Official V-JEPA model zoo

📄 Paper: Revisiting Feature Prediction for Learning Visual Representations from Video
💻 Code: Official V-JEPA repository
Meta V-JEPA release
Official V-JEPA model zoo
GPT-5.6 Sol2026-07-09OpenAIAvailable as a hosted OpenAI model; weights and detailed architecture are not released.Multimodal LLM / VLM (image + text → text)Not reported.Not reported

Agents’ Last Exam: 53.6
OpenAI evaluation of long-running professional workflows across 55 fields; GPT-5.6 Sol result. Exact public harness version and statistical uncertainty: Not reported.
Official GPT-5.6 release

📄 Paper: Not available
💻 Code: Not available
Official release
GPT-6 Astra2026-09-03OpenAIAvailable through OpenAI-hosted products and services; weights are not released.Multimodal LLM / VLM (image + text → text)Not reported.Not reportedNot reported📄 Paper: Not available
💻 Code: Not available
Official safety overview
Claude Opus 5.52026-09-22AnthropicAvailable through Claude and the Claude Platform; weights are not released.Multimodal LLM / VLM (image + text → text)Not reported.Not reported

Terminal-Bench 4.0: 66.4%
Claude Opus 5.5 at xhigh effort; production safeguards enabled; standard error ±2.6 points; Anthropic-reported evaluation.
Official Claude Opus 5.5 release

Humanity's Last Exam: 67.7% with tools
Adaptive thinking at max effort with tools; production safeguards enabled; Anthropic-reported evaluation.
Official Claude Opus 5.5 release

📄 Paper: Claude Opus 5.5 system card
💻 Code: Not available
Official release
Claude Fable 5.12026-09-01AnthropicAvailable through Anthropic-hosted products and APIs; weights are not released.Multimodal LLM / VLM (image + text → text)Not reported.Not reportedNot reported📄 Paper: Not available
💻 Code: Not available
Anthropic newsroom release listing
Pi 3.0Not reportedInflection AIHosted in Pi and available through the Inflection API as inflection_3_pi; weights are not released.LLM (text → text)Custom fine-tuned Inflection foundation model; detailed architecture is not reported.Not reportedNot reported📄 Paper: Not available
💻 Code: Not available
Official Inflection API documentation
Official Pi product page
Gemini 3.8 Flash2026-09-02Google DeepMindAvailable through Google AI Studio and Google-hosted services; weights are not released.Multimodal LLM / VLM (image + text → text)Not reported.Not reported

HLE-Verified: 54.9%
Google-reported result for Gemini 3.8 Flash. The release describes a multi-step reasoning evaluation; exact harness and effort configuration: Not reported.
Official Gemini 3.8 Flash release

📄 Paper: Not available
💻 Code: Not available
Official release
Gemma 4 12B Unified2026-06-03Google DeepMind
Open weights
Gemma Terms of Use
Downloadable weights under the Gemma terms; local and hosted deployment options vary.
Multimodal LLM / VLM (image + text → text)Not reported in the cited release log.12 billionNot reported📄 Paper: Not available
💻 Code: Not available
Official Gemma release log
Llama 3.3 70B Instruct2024-12-06MetaDownloadable weights under the Llama 3.3 license; available locally through Ollama as llama3.3.LLM (text → text)Autoregressive Transformer with grouped-query attention; instruction tuning.70 billionNot reported📄 Paper: Llama 3 model paper
💻 Code: Meta Llama models repository
Official Ollama listing
Official Meta repository
Mistral Small 3.1 24B2025-03-17Mistral AIPublic base and instruction checkpoints; available locally through Ollama as mistral-small3.1.Multimodal LLM / VLM (image + text → text)Dense Transformer with a vision encoder and up to 128k context.24 billionNot reported📄 Paper: Not available
💻 Code: Official model repository
Official release
Official Ollama listing
Phi-4 14B2024-12-12Microsoft
Open weights
MIT License
Downloadable weights under MIT; available locally through Ollama as phi4.
LLM (text → text)Dense decoder-only Transformer.14 billionNot reported📄 Paper: Phi-4 technical report
💻 Code: Official Phi model repository
Microsoft Phi cookbook
Official Ollama listing
Qwen3 32B2025-04-29Alibaba QwenDownloadable weights; supported by local runtimes including Ollama as qwen3:32b.LLM (text → text)Dense decoder-only Transformer; 64 layers, 64 query heads, 8 key/value heads, 128k context.32 billionNot reported📄 Paper: Qwen3 technical report
💻 Code: Official Qwen3 repository
Official Qwen3 release
Official Ollama listing
DeepSeek-V3.22025-12-02DeepSeek
Open source
MIT License
Downloadable model weights and inference code; very large hardware requirements apply.
LLM (text → text)Mixture-of-experts Transformer with DeepSeek Sparse Attention.685 billion total parameters

GPQA Diamond: 82.4
Score listed in the official Hugging Face model repository evaluation results; exact prompting, sampling, and harness version: Not reported.
Official DeepSeek-V3.2 model repository

📄 Paper: DeepSeek-V3.2 technical report
💻 Code: Official model repository
Official model repository
Qwen3-VL 235B-A22B Instruct2025-09-23Alibaba QwenDownloadable model weights and public inference code; the 235B-total-parameter checkpoint requires substantial accelerator memory.Multimodal LLM / VLM (image + video + text → text)Mixture-of-experts Qwen3 language backbone with a SigLIP 2 vision encoder, MLP vision-language merger, Interleaved-MRoPE, DeepStack feature fusion, and explicit video timestamps; 256k native context.235 billion total; 22 billion active per tokenNot reported📄 Paper: Qwen3-VL Technical Report
💻 Code: Official Qwen3-VL repository
Official Qwen3-VL repository
Official model repository
Llama 4 Maverick 17B-128E Instruct2025-04-05MetaDownloadable weights under the Llama 4 Community License and acceptable-use policy; multimodal license rights include a European Union restriction described in the use policy.Multimodal LLM / VLM (image + text → text)Autoregressive early-fusion mixture-of-experts Transformer with 128 routed experts, one shared expert, and a MetaCLIP-derived vision encoder; 1M-token context.400 billion total; 17 billion active

MMLU: 85.5%
Meta model card result for the pretrained Maverick model; 5-shot macro-average character accuracy; bf16 evaluation. Exact harness version and dataset revision: Not reported.
Official Llama 4 model card

📄 Paper: Not available
💻 Code: Meta Llama models repository
Official Llama 4 model card
Meta release announcement
Molmo 2 8B2025-12-11Ai2Downloadable weights and inference examples. Ai2 reports that training code, evaluations, and intermediate checkpoints will be released later; users should review the licenses of underlying third-party training datasets.Multimodal LLM / VLM (image + video + text → text + grounding)Qwen3-8B language backbone paired with a SigLIP 2 vision encoder and trained for image, multi-image, and video understanding and grounding.8 billion class

15 academic video benchmarks: 63.1 average
Ai2-reported average across 15 academic benchmarks for Molmo2-8B; individual tasks and evaluation details are documented in the technical report.
Official Molmo 2 model card

📄 Paper: Molmo2 technical report
💻 Code: Official model repository and examples
Ai2 release announcement
Official model repository
Mistral Small 4 119B-A6B2026-03-16Mistral AIDownloadable weights under Apache 2.0, plus hosted access through Mistral AI Studio and the Mistral API.Multimodal LLM / VLM (image + text → text)Mixture-of-experts Transformer with 128 experts and four active experts per token, native image input, configurable reasoning effort, and a 256k context window.119 billion total; 6 billion active per token (8 billion including embeddings and output layers)Not reported📄 Paper: Not available
💻 Code: Official model repository
Official release announcement
Official model repository
V-JEPA 2 ViT-g/16 3842025-06-11Meta FAIRPublic pretrained checkpoints, evaluation probes, training and evaluation configurations, PyTorch code, and Hugging Face integration.Video foundation modelViT-g/16 video encoder trained by predicting latent representations of masked video regions; this entry uses the 384-pixel, 64-frame checkpoint.1 billion

SSv2: 77.3%
Official attentive-probe result using frozen V-JEPA 2 ViT-g/16 384 features; training and inference configurations are linked from the repository.
Official V-JEPA 2 evaluation table

📄 Paper: V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
💻 Code: Official V-JEPA 2 repository
Official V-JEPA 2 repository
Official model repository
Qwen3-Omni 30B-A3B Instruct2025-09-22Alibaba QwenDownloadable weights, public inference code, Docker image, cookbooks, and hosted API access; local use has substantial GPU-memory requirements.Omnimodal model (text + image + audio + video → text + speech)End-to-end mixture-of-experts Thinker–Talker architecture with audio-text pretraining and a multi-codebook speech generator for real-time streaming.30 billion total; 3 billion active

Video-MME: 70.5
Official repository result for Qwen3-Omni-30B-A3B-Instruct in the vision-to-text video-understanding table; exact harness version: Not reported.
Official Qwen3-Omni evaluation table

📄 Paper: Qwen3-Omni Technical Report
💻 Code: Official Qwen3-Omni repository
Official Qwen3-Omni repository
Official model repository