WayToClawEarn
Medium impactTechNode

Ant Group's Ling-3.1-flash: 560B MoE, 25B Active, 1M Context, Open-Source After Trial — Which Agent LLM to Pick?

Ant Group's InclusionAI released Ling-3.1-flash on September 30: a 560B total / 25B active MoE model with a 1M-token design context window, targeting agent tasks, search, office, and specialist applications. A two-week free trial is live (capped at 256K context), with full context and open-source weights planned post-trial. Eight vendor-reported benchmarks cover coding, agentic, and knowledge categories, but no independent verification exists yet.

Edisen Lu · WayToClawEarnVia TechNodePublished Oct 3, 2026

Reviewed from public sources · AI-assisted drafting under editorial oversight. How we work · Original source

Public-source compilation

Synthesized from public posts/docs. Prefer the original source for primary claims.

How we review content · Primary source · TechNode

TL;DR

Ant Group's InclusionAI (Ant Ling / 蚂蚁百灵) officially released Ling-3.1-flash on September 30, 2026. The model uses a Mixture of Experts (MoE) architecture with approximately 560 billion total parameters and roughly 25 billion activated parameters per token — an activation rate under 5%. The design context window reaches 1 million tokens, though the current two-week free trial limits context to 256,000 tokens. The model targets agent tasks, search, office software, and specialist domains including healthcare, finance, and materials science. After the trial period, InclusionAI plans to enable the full 1M context and release model weights as open source. The model is currently accessible via Vercel AI Gateway (via Novita) and OpenRouter with the API model ID inclusionai/ling-3.1-flash.

Key Specifications

DimensionValue
Release dateSeptember 30, 2026
DeveloperInclusionAI (Ant Group, 蚂蚁百灵)
ArchitectureMixture of Experts (MoE)
Total parameters~560 billion (560B)
Active parameters per token~25 billion (25B)
Activation rate< 5%
Design context window1,000,000 tokens (1M)
Trial context window262,144 tokens (256K)
Maximum output32,768 tokens
PricingFree during trial (through ~Oct 13); paid pricing not yet announced
Open-source planWeights to be released after trial (license not yet confirmed)
Available viaVercel AI Gateway (Novita), OpenRouter
API model IDinclusionai/ling-3.1-flash
PredecessorLing-3.0-flash (124B total / 5.1B active)

Architecture and Positioning

Ling-3.1-flash continues the flash series' MoE approach but with a massive parameter scale-up. The previous generation Ling-3.0-flash had 124B total parameters with 5.1B active; the new generation scales total parameters by ~4.5x and active parameters by ~5x. The core design philosophy is clear: use a larger parameter reserve for stronger model capability, while controlling per-inference compute cost through an extremely low activation rate.

The million-token context window is another headline feature. Ling-3.1-flash's design target of 1M tokens enables it to process very long documents, complete codebases, and extended agent workflows with deep task histories. The trial period's 256K token limit is sufficient for most long-document and agent pilot scenarios, but the full 1M window is the key differentiator for production use cases.

InclusionAI positions Ling-3.1-flash across four directions: agent tasks, search, office software, and specialist applications. This positioning aligns with Ant Group's recent product strategy. The previously released Ling-3.0-flash-Fin (MIT-licensed, 124B total / 5.5B active) targeted investment research workflows, while Ling-3.0-flash-VL was the first native multimodal model in the series. Ling-3.1 serves as a general-purpose foundation model, upgrading both parameter scale and context capability to provide stronger infrastructure for long-chain tasks.

Benchmark Data: 8 Self-Reported Scores, No Independent Verification

According to BenchLM tracking data, Ling-3.1-flash has published 8 benchmark scores, all sourced from Ant Ling's launch screenshots on X/Twitter (@AntLingAGI). No third-party independent evaluation exists yet. The scores and comparison with currently verified leaders are as follows:

Coding

BenchmarkLing-3.1-flashVerified LeaderGap
TerminalBench40.4%Ling-3.1-flash itself (highest verified)—
SWE-Atlas Codebase QnA55.9%Muse Spark 1.3 (59.4%)-3.5

Agentic

BenchmarkLing-3.1-flashVerified LeaderGap
AutomationBench52.5%DeepSeek V4.1 Flash (54.8%)-2.3
SkillsBench68.7%Qwen3.8 Max (70.2%)-1.5
CyberGym87.9%MiMo-V2.6-Flash (95.1%)-7.2
Finance Agent v257.9%Gemini 4 Argon (65.4%)-7.5
DRACO85.5%Claude Opus 5 (88.6%)-3.1

Knowledge

BenchmarkLing-3.1-flashVerified LeaderGap
HealthBench Professional65.3%Claude Sonnet 5.5 (69.2%)-3.9

Important caveats: The evaluation harnesses and grading settings for these benchmarks are not documented. HealthBench Professional was evaluated in the "AQ environment," and announced healthcare capabilities are currently available only in AQ. These are provider-reported results, not independent measurements or clinical certifications.

Open-Source Strategy and Timeline

Ling-3.1-flash adopts a "try first, open-source later" release rhythm. The two-week free trial is the most practical element of this launch: developers can test the model without payment, and the 256K token limit during the trial is sufficient for most long-document and agent pilot scenarios.

After the trial period, InclusionAI plans to enable the full 1M token context window and release the model weights as open source. Based on OrcaRouter's records, Ling-3.0-Flash had approximately a two-week gap between announcement and open-source weight release, suggesting Ling-3.1-flash's open-source release could come around mid-October. However, the specific license and weight release timeline have not been confirmed — these two factors will directly determine whether Ling-3.1-flash can truly enter the practical ranks of China's open-source first tier.

Ant Group's InclusionAI has previously open-sourced Ling series tiny and flash foundation models (e.g., Ling-3.0-flash-Fin under MIT license). If Ling-3.1-flash follows with MIT or Apache 2.0 licensing, a 560B-parameter-class open-source model would provide Chinese developers with a new option alongside DeepSeek V4 and Qwen3.8.

China's Open-Source LLM Landscape

Ling-3.1-flash's release further enriches the Chinese LLM matrix. The current open-source first tier includes:

  • DeepSeek V4: 1M context, OpenAI/Anthropic API compatible, zero CUDA dependency
  • Qwen3.8 series: Multiple sizes from 27B to 2.4T, Apache 2.0 license, leading Hugging Face downloads
  • Kimi-K3: Moonshot's long-context model
  • Ling series: Ant Group's InclusionAI, with specialized variants for finance and multimodal

Ling-3.1-flash's differentiation lies in its 560B total / 25B active MoE configuration, which leads the peer group in total parameter count while keeping active parameters at the 25B level — theoretically deployable on reasonable hardware. The million-token context window is a key capability benchmarked against Gemini 4 Argon (1M output tokens) and DeepSeek V4 (1M context).

Selection Guidance

For developers and enterprises evaluating agent LLMs, Ling-3.1-flash is worth validating in the following scenarios:

  1. Long-context agent scenarios: Agent workflows requiring processing of dozens of documents and multi-round tool calls, where a million-token window is a core requirement
  2. Finance/healthcare vertical applications: Ant Group's prior accumulation in finance (Ling-3.0-flash-Fin) may carry forward in this model
  3. Cost-sensitive large-scale deployment: A 25B-active MoE has inherent advantages in inference cost; once open-sourced, private deployment hardware requirements can be assessed

When validating, focus on two dimensions: effective memory length under long context (don't just look at the nominal window — test how information retrieval accuracy changes as context grows), and stability and tool-calling accuracy in multi-step agent tasks.

Also watch for two key signals after open-source release: the license (MIT/Apache 2.0 is commercially friendly; custom licenses require careful review) and the weight release timeline. These signals determine whether Ling-3.1-flash is "yet another parameter race participant" or "a genuinely production-deployable option."

Sources

Ant GroupInclusionAILing 3.1 FlashMoEopen source LLMAI agentChinese AI modelslarge language model

View source →

Educational reference only: cases summarize public sources and may use AI-assisted drafting under editorial review. Not financial advice; outcomes are not guaranteed.