Ant Group's Ling-3.1-flash: 560B MoE, 25B Active, 1M Context, Open-Source After Trial — Which Agent LLM to Pick?
Ant Group's InclusionAI released Ling-3.1-flash on September 30: a 560B total / 25B active MoE model with a 1M-token design context window, targeting agent tasks, search, office, and specialist applications. A two-week free trial is live (capped at 256K context), with full context and open-source weights planned post-trial. Eight vendor-reported benchmarks cover coding, agentic, and knowledge categories, but no independent verification exists yet.
Public-source compilation
Synthesized from public posts/docs. Prefer the original source for primary claims.
TL;DR
Ant Group's InclusionAI (Ant Ling / 蚂蚁百灵) officially released Ling-3.1-flash on September 30, 2026. The model uses a Mixture of Experts (MoE) architecture with approximately 560 billion total parameters and roughly 25 billion activated parameters per token — an activation rate under 5%. The design context window reaches 1 million tokens, though the current two-week free trial limits context to 256,000 tokens. The model targets agent tasks, search, office software, and specialist domains including healthcare, finance, and materials science. After the trial period, InclusionAI plans to enable the full 1M context and release model weights as open source. The model is currently accessible via Vercel AI Gateway (via Novita) and OpenRouter with the API model ID inclusionai/ling-3.1-flash.
Key Specifications
| Dimension | Value |
|---|---|
| Release date | September 30, 2026 |
| Developer | InclusionAI (Ant Group, 蚂蚁百灵) |
| Architecture | Mixture of Experts (MoE) |
| Total parameters | ~560 billion (560B) |
| Active parameters per token | ~25 billion (25B) |
| Activation rate | < 5% |
| Design context window | 1,000,000 tokens (1M) |
| Trial context window | 262,144 tokens (256K) |
| Maximum output | 32,768 tokens |
| Pricing | Free during trial (through ~Oct 13); paid pricing not yet announced |
| Open-source plan | Weights to be released after trial (license not yet confirmed) |
| Available via | Vercel AI Gateway (Novita), OpenRouter |
| API model ID | inclusionai/ling-3.1-flash |
| Predecessor | Ling-3.0-flash (124B total / 5.1B active) |
Architecture and Positioning
Ling-3.1-flash continues the flash series' MoE approach but with a massive parameter scale-up. The previous generation Ling-3.0-flash had 124B total parameters with 5.1B active; the new generation scales total parameters by ~4.5x and active parameters by ~5x. The core design philosophy is clear: use a larger parameter reserve for stronger model capability, while controlling per-inference compute cost through an extremely low activation rate.
The million-token context window is another headline feature. Ling-3.1-flash's design target of 1M tokens enables it to process very long documents, complete codebases, and extended agent workflows with deep task histories. The trial period's 256K token limit is sufficient for most long-document and agent pilot scenarios, but the full 1M window is the key differentiator for production use cases.
InclusionAI positions Ling-3.1-flash across four directions: agent tasks, search, office software, and specialist applications. This positioning aligns with Ant Group's recent product strategy. The previously released Ling-3.0-flash-Fin (MIT-licensed, 124B total / 5.5B active) targeted investment research workflows, while Ling-3.0-flash-VL was the first native multimodal model in the series. Ling-3.1 serves as a general-purpose foundation model, upgrading both parameter scale and context capability to provide stronger infrastructure for long-chain tasks.
Benchmark Data: 8 Self-Reported Scores, No Independent Verification
According to BenchLM tracking data, Ling-3.1-flash has published 8 benchmark scores, all sourced from Ant Ling's launch screenshots on X/Twitter (@AntLingAGI). No third-party independent evaluation exists yet. The scores and comparison with currently verified leaders are as follows:
Coding
| Benchmark | Ling-3.1-flash | Verified Leader | Gap |
|---|---|---|---|
| TerminalBench | 40.4% | Ling-3.1-flash itself (highest verified) | — |
| SWE-Atlas Codebase QnA | 55.9% | Muse Spark 1.3 (59.4%) | -3.5 |
Agentic
| Benchmark | Ling-3.1-flash | Verified Leader | Gap |
|---|---|---|---|
| AutomationBench | 52.5% | DeepSeek V4.1 Flash (54.8%) | -2.3 |
| SkillsBench | 68.7% | Qwen3.8 Max (70.2%) | -1.5 |
| CyberGym | 87.9% | MiMo-V2.6-Flash (95.1%) | -7.2 |
| Finance Agent v2 | 57.9% | Gemini 4 Argon (65.4%) | -7.5 |
| DRACO | 85.5% | Claude Opus 5 (88.6%) | -3.1 |
Knowledge
| Benchmark | Ling-3.1-flash | Verified Leader | Gap |
|---|---|---|---|
| HealthBench Professional | 65.3% | Claude Sonnet 5.5 (69.2%) | -3.9 |
Important caveats: The evaluation harnesses and grading settings for these benchmarks are not documented. HealthBench Professional was evaluated in the "AQ environment," and announced healthcare capabilities are currently available only in AQ. These are provider-reported results, not independent measurements or clinical certifications.
Open-Source Strategy and Timeline
Ling-3.1-flash adopts a "try first, open-source later" release rhythm. The two-week free trial is the most practical element of this launch: developers can test the model without payment, and the 256K token limit during the trial is sufficient for most long-document and agent pilot scenarios.
After the trial period, InclusionAI plans to enable the full 1M token context window and release the model weights as open source. Based on OrcaRouter's records, Ling-3.0-Flash had approximately a two-week gap between announcement and open-source weight release, suggesting Ling-3.1-flash's open-source release could come around mid-October. However, the specific license and weight release timeline have not been confirmed — these two factors will directly determine whether Ling-3.1-flash can truly enter the practical ranks of China's open-source first tier.
Ant Group's InclusionAI has previously open-sourced Ling series tiny and flash foundation models (e.g., Ling-3.0-flash-Fin under MIT license). If Ling-3.1-flash follows with MIT or Apache 2.0 licensing, a 560B-parameter-class open-source model would provide Chinese developers with a new option alongside DeepSeek V4 and Qwen3.8.
China's Open-Source LLM Landscape
Ling-3.1-flash's release further enriches the Chinese LLM matrix. The current open-source first tier includes:
- DeepSeek V4: 1M context, OpenAI/Anthropic API compatible, zero CUDA dependency
- Qwen3.8 series: Multiple sizes from 27B to 2.4T, Apache 2.0 license, leading Hugging Face downloads
- Kimi-K3: Moonshot's long-context model
- Ling series: Ant Group's InclusionAI, with specialized variants for finance and multimodal
Ling-3.1-flash's differentiation lies in its 560B total / 25B active MoE configuration, which leads the peer group in total parameter count while keeping active parameters at the 25B level — theoretically deployable on reasonable hardware. The million-token context window is a key capability benchmarked against Gemini 4 Argon (1M output tokens) and DeepSeek V4 (1M context).
Selection Guidance
For developers and enterprises evaluating agent LLMs, Ling-3.1-flash is worth validating in the following scenarios:
- Long-context agent scenarios: Agent workflows requiring processing of dozens of documents and multi-round tool calls, where a million-token window is a core requirement
- Finance/healthcare vertical applications: Ant Group's prior accumulation in finance (Ling-3.0-flash-Fin) may carry forward in this model
- Cost-sensitive large-scale deployment: A 25B-active MoE has inherent advantages in inference cost; once open-sourced, private deployment hardware requirements can be assessed
When validating, focus on two dimensions: effective memory length under long context (don't just look at the nominal window — test how information retrieval accuracy changes as context grows), and stability and tool-calling accuracy in multi-step agent tasks.
Also watch for two key signals after open-source release: the license (MIT/Apache 2.0 is commercially friendly; custom licenses require careful review) and the weight release timeline. These signals determine whether Ling-3.1-flash is "yet another parameter race participant" or "a genuinely production-deployable option."
Sources
- TechNode report (2026-09-30): https://technode.com/2026/09/30/ant-group-launches-ling-3-1-flash-with-560-billion-parameters/
- IT Home report (2026-09-30): https://www.ithome.com/1/008/907.htm
- InclusionAI official X/Twitter: https://x.com/AntLingAGI
- Vercel AI Gateway model page: https://vercel.com/ai-gateway/models/ling-3.1-flash
- BenchLM benchmark tracking: https://benchlm.ai/models/ling-3-1-flash
- LLM Reference model page: https://www.llmreference.com/model/ling-3.1-flash
- AI Weekly report (2026-09-30): https://aiweekly.co/alerts/ants-inclusionai-ships-ling-31-flash-with-560b-parameters
Topic hub
AI Agent Tutorials & Workflow Guides
Evergreen how-tos for coding agents, content pipelines, and n8n automation—linked to news context and real earn cases.
Explore AI Agent Tutorials & Workflow Guides →Monetization angle
How can you make money from this trend?
WayToClawEarn focuses on verified earn playbooks—not just news. Start from these cases.
DeepSeek + Claude Code Micro SaaS
Run multiple small products on cheap inference
Claude Code bug bounty
Productize agent skills into security services
Related tutorials
Related news
- AWS Strands Decider 2B Open-Sourced: 2B-Param Decision Model for Agent Routing, Tool Gating at 115ms
- Cloudflare Open-Sources Clef Decision Models: Returns Probabilities Not Text, How to Cut AI Agent Costs
- DeepSeek Harness v0.2 Desktop Launch: Zero-Config AI Coding Agent, How to Choose?
- ElevenLabs Doubles to $22B Valuation with Eleven v4 Turbo: How to Choose Voice AI Agents?