Overview
Llama 4 Scout is the efficiency-focused member of the Llama 4 herd: 109B total parameters across 16 experts with 17B active, small enough to run on a single NVIDIA H100 yet natively multimodal. Its headline feature is an industry-first 10M-token context window — roughly enough to ingest entire codebases or vast document sets in one pass.
Scout is built for speed, efficiency and ultra-long-context tasks, beating comparably sized models like Gemma 3 and Gemini 2.0 Flash-Lite across many benchmarks at launch. For developers who want a capable open multimodal model that fits a single GPU and handles enormous inputs, Scout is a remarkable package.
Key capabilities
- Industry-first 10M-token context window
- Runs on a single H100 GPU
- Natively multimodal
- Strong for its size and efficiency
At a glance
Context
10M tokens
Params
109B total / 17B active
Hardware
Single H100
Pricing: Very low — efficient, single-GPU deployment.
Pros & cons
What we like
- Unmatched context length
- Fits on one GPU
- Fast and efficient
Trade-offs
- Lower ceiling than Maverick on hard tasks
- Long-context quality varies in practice
The verdict
Llama 4 Scout is a clever engineering win — a single-GPU, multimodal open model with a 10M-token context that no rival matched at launch. For ultra-long-context and cost-sensitive deployments, it's a standout choice.
Best for: Ultra-long-context tasks on a single GPU.
Frequently asked questions
How long is Llama 4 Scout's context window?
Up to 10 million tokens — an industry first at launch — making it ideal for ingesting entire codebases or large document collections at once.