Per the published Apache 2.0 LICENSE, the Step-3.7-Flash weights ship without commercial restriction — a 198B / 11B-active vision-language MoE with a 256k-token context aimed at tool-heavy and agentic workflows. The remaining EU-readiness gaps are the entirely undisclosed training corpus and StepFun's Shanghai-based vendor jurisdiction; deploy on self-managed EU infrastructure and document the Art. 50 transparency story for any synthetic or biometric output.
Sovereignty
Licence: Apache 2.0Commercial: UnrestrictedTraining data: UndisclosedOrigin: China (Shanghai)
Licence facts
Parameters
198B total / ~11B active (sparse MoE)
Architecture
Vision-language MoE: 196B language backbone + 1.8B vision encoder
Context window
256,000 tokens
Modality
Image + text in, text out
Languages
Chinese, English, plus additional languages (unspecified)
Local deployment
~120-128GB unified memory minimum for GGUF quantisations
Training corpus is entirely undisclosed; with a vision-language model the rights-clearance and biometric-image provenance questions are sharper, and the AI Act Art. 53 transparency file is the deployer's responsibility.
Visual modality raises AI Act Art. 50 transparency obligations for synthetic image-derived outputs and any biometric inference — bake watermarking and end-user disclosure into the application layer.
StepFun is headquartered in Shanghai: vendor-hosted endpoints would trigger Chapter V GDPR scrutiny, so EU rollouts should run on EU infrastructure under a self-managed DPIA.