Model

Qwen 3.5 — LLM Radar spotlight

Alibaba's Apache-2.0 multimodal MoE flagship is strong on reasoning and vision, but China-origin governance keeps EU use Conditional.

Qwen 3.5 is Alibaba's open-weight multimodal MoE flagship for long-context reasoning, coding, vision-language and tool workflows. LLM Radar's read on Qwen 3.5 is Conditional: the Apache 2.0 licence is clean, but China-origin governance, undisclosed training data and hosted API jurisdiction prevent a default EU-ready verdict for regulated European deployment.

The tracking row "Qwen 3.5" maps to the public model name Qwen3.5-397B-A17B. Architecture, benchmark and licence claims below are treated as current as of 2026-06-09.

What it is

Qwen 3.5 is not a small local model and not a cosmetic Qwen 2.5 refresh. The official Hugging Face card lists Qwen3.5-397B-A17B as a causal language model with a vision encoder, 397B total parameters and 17B active parameters per token, using 60 layers, 512 experts and a hybrid Gated DeltaNet plus gated-attention MoE layout (as of 2026-06-09).

The model's native context is 262,144 tokens, extendable to roughly 1,010,000 tokens with YaRN. Inputs cover text, images and video, while output remains text (as of 2026-06-09). Public reporting places the first Qwen 3.5 397B-A17B release in February 2026.

Technically, this is a credible candidate for long-context retrieval, document-heavy agent systems, multimodal assistants and code workflows. Compliance-wise, the weights being portable does not settle the question. For European teams, the decisive evidence is where inference runs, what contracts govern the data path, and whether provider-side documentation is enough for the intended risk class.

Origin and licence

The model is published by Qwen, Alibaba's model group, and LLM Radar tracks origin as China. Hugging Face marks the repository licence as apache-2.0 (as of 2026-06-09). That matters: Apache 2.0 is a permissive licence allowing commercial use, modification, distribution and derivative work, subject to preserving notices, licence terms and the licence's patent and warranty provisions.

Licence clarity is green. Compared with commercial-restricted model licences, Qwen 3.5 gives EU teams a cleaner self-hosting and fine-tuning route. There is no field-of-use restriction in the model licence itself that blocks ordinary commercial deployment.

The counterweight is governance, not licence. Training data remains undisclosed, Alibaba is China-linked, and the official hosted Alibaba route is not a simple answer for European personal data unless the customer can evidence EU data residency, DPA coverage, subprocessors and lawful transfer mechanisms. Self-hosting in the EU can be defensible; routing personal data through Alibaba-hosted services is not defensible without reviewed contractual and residency evidence.

AI Act posture is similarly incomplete from public sources. Deployers building GPAI-based downstream systems retain their own obligations, and public provider-side technical documentation and training-data transparency do not appear sufficient for a low-friction regulated procurement file.

The verdict here is Conditional: ship it for non-sensitive workloads, but document origin, hosting and training-data uncertainty before regulated use.

Strengths

The official model card reports strong vendor-reported figures: MMLU-Pro 87.8, MMLU-Redux 94.9, GPQA 88.4, AIME26 91.3, LiveCodeBench v6 83.6, BFCL-V4 72.9, TAU2-Bench 86.7, SWE-bench Verified 76.4, MMMLU 88.5, MMLU-ProX 84.7, MathVision 88.6, MMMU 85.0, OmniDocBench1.5 90.8, OCRBench 93.1 and VideoMME with subtitles 87.5 (as of 2026-06-09).

Those figures point to plausible production utility in multilingual enterprise assistants, code review agents, document analysis, OCR-heavy pipelines and visual workflow automation. The architecture also makes the efficiency story credible: 397B total parameters but 17B active per token gives Qwen 3.5 high headline scale without dense-model inference cost.

Artificial Analysis adds a more useful independent comparison layer. It scored Qwen3.5-397B-A17B at 45 on its Intelligence Index, ranking third among open-weight models in that snapshot behind GLM-5 reasoning at 50 and Kimi K2.5 reasoning at 47, while using fewer active parameters than several peers (as of 2026-06-09).

The defensible read is that Qwen 3.5 is attractive where latency, long context and multimodal breadth matter together: public-content assistants, synthetic-data workflows, internal engineering copilots on EU-controlled infrastructure, long-context research synthesis and agent prototypes.

Limitations

Benchmark strength does not cure GDPR transfer risk. The largest limitation for regulated EU use is not whether Qwen 3.5 can answer hard questions; it is whether the deployment can prove where prompts, logs, files, images and video frames are processed and retained.

Training data remains undisclosed in LLM Radar's tracking data, and public materials do not provide the dataset provenance a regulated EU buyer would normally want. China-origin model governance can also trigger internal risk review even where the weights are Apache-licensed and self-hosted.

Independent evaluation also flags hallucination risk. Artificial Analysis reported Qwen3.5's AA-Omniscience Index at -32, improved from Qwen3 235B but still behind Kimi K2.5 and GLM-5, and attributed the improvement more to accuracy than to a lower hallucination rate (as of 2026-06-09). Refusal behavior, over-answering and source discipline need local evaluation before regulated deployment.

The deployment footprint is non-trivial. A 262K native context window is useful, but cost, latency and memory pressure rise quickly. YaRN-to-1M should be treated as a specialist configuration, not default production posture. Current serving normally implies serious GPU capacity and modern stacks such as SGLang, vLLM or managed inference.

For high-risk domains, teams should require red-team evaluation, source-grounded RAG, logging controls, output review and clear human escalation. Model-card safety claims are not enough.

When to use it

LLM Radar's read on Qwen 3.5 is Conditional.

The strongest route is EU self-hosting for document processing, code agents and multilingual assistants where prompts, logs, fine-tunes and uploaded media remain under EU control. Apache 2.0 removes much of the licence friction for commercial deployment and adaptation.

Managed US providers such as Together AI may be acceptable for public, synthetic or non-personal data, but personal data requires transfer impact analysis, DPA coverage, retention review and subprocessor documentation. NVIDIA NIM is best understood as an infrastructure route: the decisive question is where the NIM endpoint actually runs, not the brand.

Alibaba's own hosted route should be treated as higher-risk for EU personal data unless a European deployment contract, data-residency posture and transfer mechanism are documented.

Procurement review should cover physical hosting region, DPA, subprocessors, retention defaults, logging opt-out, abuse-monitoring path, model update policy and whether multimodal uploads are stored. Blocked scenarios include European personal data sent to an API with China-only processing, unclear subprocessors or no DPA and transfer mechanism.

Comparable models

ModelLicenceOriginReported strengthEU deployment read
Qwen 3.5Apache 2.0ChinaMultimodal, long context, active-parameter efficiencyConditional; strongest via EU self-hosting
DeepSeek V3.2Open-weight posture varies by releaseChinaReasoning, coding and agent workloadsConditional; similar origin and governance questions
Kimi K2Open-weight peerChinaStrong independent benchmark positioningConditional; diligence depends on hosting and contracts
GLM-5.1Open-weight peerChinaIntelligence-index and hallucination comparison pointConditional; deployment route remains decisive

Comparable open-weight alternatives include DeepSeek V3.2, Kimi K2 and GLM-5.1. Qwen 3.5 should not be framed as beating every alternative. The stronger claim is narrower: it is competitive on multimodal breadth, long context and MoE efficiency.

The counter-argument is real. Apache-2.0 self-hosting reduces vendor data-access risk substantially, and for some teams that will be enough. LLM Radar's response is also straightforward: origin, training-data opacity and safety calibration still matter for regulated deployment, even when the licence is permissive.

We do not require readers to agree with the verdict. We require that the evidence be visible enough for engineering, procurement and compliance teams to document the decision.

Sources