Model

Llama 4 Scout: EU self-hosting spotlight

EU deployment outline for Llama 4 Scout: licence limits, self-hosting posture, benchmark trade-offs and provider fallback routes.

Llama 4 Scout is Meta's smaller Llama 4 open-weight model for long-context, multimodal inference: 17B activated parameters, 109B total parameters, image and text input, and a 10M-token context window. LLM Radar's read as of 2026-06-09: Scout is technically attractive for EU teams that want local control over inference, but the compliance verdict is Conditional, not EU-ready by default.

The tension is simple. Downloadable weights and single-H100 int4 positioning make Scout look like a self-hosting answer, but EU readiness turns on licence terms, data governance, AI Act documentation and operational evidence, not only on where the GPU sits.

What it is

Scout is the smaller released member of Meta's Llama 4 family, positioned below Llama 4 Maverick. The official model card identifies Llama 4 Scout Instruct as a mixture-of-experts model with 17B activated parameters and 109B total parameters, multilingual text plus image input, multilingual text and code output, roughly 40T training tokens, a 10M-token context window, and an August 2024 knowledge cutoff (as of 2026-06-09).

Meta released Scout on April 5, 2025. This article concerns the Llama 4 Scout Instruct checkpoint for assistant-style chat and visual reasoning. The base checkpoint matters mainly for teams adapting, evaluating or fine-tuning under the Llama 4 licence.

The core deployment case is not frontier reasoning. Scout is better read as a long-context, multimodal, open-weight candidate for internal assistants, document-heavy workflows and controlled inference environments where local operation is part of the compliance strategy.

Origin and licence

Scout is not EU-origin. Meta is a US-headquartered vendor, although the Llama 4 Community License Agreement says Meta Platforms Ireland Limited is the contracting Meta entity for users located in the EEA or Switzerland. That matters for procurement, but it does not turn Scout into an EU-origin model.

The licence is a custom commercial community licence. It is not Apache-2.0, MIT or another OSI-style permissive licence. It grants a worldwide, royalty-free limited licence to use, reproduce, distribute, modify and create derivative works, including commercial use, but it also imposes Llama-specific attribution, acceptable-use compliance, trademark limits, indemnity, a 700M monthly-active-user threshold requiring a separate Meta licence, and California governing law.

LLM Radar's licence verdict is Conditional. Most EU companies below the scale threshold can likely use Scout commercially if procurement accepts the terms, but the licence is not neutral for regulated procurement. It is also meaningfully different from permissive open-weight alternatives and from research-only releases such as Mistral Large 2, where open weights sit under the Mistral Research License and commercial deployment requires a separate paid licence.

The EU caveat is also narrower than some commentary suggests. LLM Radar does not read the current official licence text reviewed on 2026-06-09 as a blanket prohibition on EU-domiciled use. The better compliance issue is contractual and operational: teams distributing a product, service, derivative model or model package that includes Llama Materials need to carry the agreement, display "Built with Llama", retain required notices, and use "Llama" at the beginning of derivative model names where Llama outputs or materials are used to improve a distributed model.

Under the AI Act, Meta is the upstream general-purpose AI model provider. The EU deployer remains responsible for downstream deployment controls. For self-hosted regulated use, the defensible posture requires keeping the model card, licence, benchmark notes, intended-use mapping, risk assessment, logging policy and post-market monitoring evidence in the technical file.

Strengths

The case for Scout is pragmatic: long context, multimodal input, fast MoE inference and commercially available weights. Meta reports pre-trained scores of MMLU 79.6, MMLU-Pro 58.2, MATH 50.3 and MBPP 67.8, with instruction-tuned scores of MMMU 69.4, MathVista 70.7, ChartQA 88.8, DocVQA 94.4, LiveCodeBench 32.8, MMLU-Pro 74.3, GPQA Diamond 57.2 and MGSM 90.6 (as of 2026-06-09).

Artificial Analysis gives Scout an Intelligence Index of 14, median output speed around 118 tokens/s, median time-to-first-token around 0.84s, and median API pricing of $0.17 per 1M input tokens and $0.66 per 1M output tokens (as of 2026-06-09). Those figures make Scout interesting for cost-sensitive chat, high-throughput summarisation and workloads where latency matters more than maximum reasoning quality.

Best-fit use cases include long-document summarisation, retrieval over very large internal corpora, visual document triage, multilingual assistant workloads in supported languages, and non-reasoning support flows bounded by retrieval and policy checks. Scout compares favourably with older 70B-class open-weight deployments such as Llama 3.3 70B on multimodality and context length, but it should not be treated as the same quality proposition as Maverick.

The 10M context claim is the headline feature. It is also the feature that needs local verification. Teams relying on long context for production legal, claims or medical analysis should test usable context under their actual serving stack, quantization approach and attention implementation before making it part of the control design.

Limitations

Scout is open-weight, not compliance-ready by default. The model card says training data includes publicly available data, licensed data, and information from Meta products and services, including publicly shared Instagram and Facebook posts and interactions with Meta AI. That is not a fully auditable training-data bill of materials.

Self-hosting in an EU-controlled environment can reduce or avoid processor-transfer issues for inference data. It does not answer training-data provenance, output-risk, bias, explainability or data-subject-rights questions by itself. For GDPR, the inference architecture is only one part of the posture.

Safety is also delegated heavily to deployers. Meta states that safety testing cannot cover all scenarios and that deployers should perform application-specific safety testing and tuning. For banking, healthcare, employment, education and public-sector use, that means use-case classification under the AI Act, documented human oversight, logging, fallback paths and model-output validation.

Benchmark-wise, Scout trails Maverick on Meta-reported headline metrics and is classified as a non-reasoning model by Artificial Analysis. Meta's benchmark table is useful, but it is vendor-supplied. The defensible read is to pair it with independent measurements and domain evaluation before relying on Scout for production decisions.

Operationally, the single-H100 int4 positioning helps pilots. Production self-hosting still needs quantization validation, throughput testing, monitoring, red-team prompts, rollback options and evidence that model updates or serving changes did not silently alter risk.

When to use it

The verdict here is Conditional: ship it for non-sensitive and moderately sensitive EU workloads where inference is self-hosted inside an EU-controlled environment, licence obligations are accepted by procurement, and the deployment team can document application-specific safety controls.

Scout is not EU-ready for regulated deployment merely because the weights can run locally. For regulated personal data, the defensible read is to self-host in EU infrastructure or use a provider with EU region pinning, a signed DPA, subprocessor list, logging controls and no unmanaged onward transfer.

Use Scout for internal knowledge assistants, large-context summarisation, visual document routing and multilingual support flows where outputs are reviewed or bounded by retrieval and policy checks. Avoid it as the primary model for high-stakes autonomous decisions, unsupported languages, production code generation without tests, or legal and medical conclusions without expert review.

For EU self-hosting, the deployment note should cover GPU location, operator jurisdiction, access controls, prompt-retention policy, incident process, licence acceptance, attribution obligations and AI Act role mapping. API fallback via Groq, AWS Bedrock, Azure or Google Vertex can be useful for latency and capacity, but each hosted route changes the GDPR analysis. Artificial Analysis tracks Groq as the fastest measured Scout provider, with Amazon, Azure and Google Vertex among additional hosted routes (as of 2026-06-09). LLM Radar would only mark a hosted deployment EU-ready after checking region pinning and contractual posture.

Comparable models

ModelOrigin/vendorLicence classContextModalitySelf-hosting practicalityEU compliance read
Llama 4 ScoutMeta, USCustom commercial community licence10MText and image inputStrong for controlled self-hosting pilotsConditional
Llama 4 MaverickMeta, USSame Llama 4 licence family1MText and image inputHeavier than ScoutConditional
Llama 3.3 70BMeta, USLlama community licence familyLower than ScoutTextMature 70B-class pathConditional
Mistral Large 2Mistral AI, FranceMistral Research License; paid commercial licence requiredNot the Scout long-context propositionTextStrong model, weaker open-weight commercial clarityConditional until paid commercial terms are in place

Comparable open-weight alternatives include Maverick, Llama 3.3 70B, Mistral Large 2 and smaller permissive models where sovereignty and licence clarity matter more than top benchmark scores. Maverick is the obvious quality upgrade inside the same Llama 4 licence family, with stronger Meta-reported headline results but a 1M rather than 10M context window. Llama 3.3 70B is simpler and widely deployed, but lacks native multimodality.

Mistral Large 2 is the EU-vendor counter-argument: stronger sovereignty optics for Europe, but open weights under a research licence and commercial deployment only with a separate paid licence. LLM Radar's read on Scout is therefore narrow: Conditional for EU self-hosting, defensible for controlled non-sensitive production and internal workflows, but not EU-ready for regulated personal-data deployment without documented licence review, local hosting controls and use-case-specific AI Act/GDPR evidence.

Sources