Norway trains sovereign LLM with 2PB Huawei all-flash: the real challenge of petabyte-scale data pipelines
The National Library of Norway disclosed at Huawei ID Forum 2026 that it is using a 2PB Huawei OceanStor Dorado all-flash array to train a Norwegian language sovereign-level large model. Project reveals an underestimated engineering challenge - moving petabytes of data from a 60PB archiving system to an AI training pipeline. Three key takeaways: non-English speaking countries need autonomous LLM to defend cultural sovereignty, Huawei storage is penetrating European government-level infrastructure, and 'data pipeline throughput, not computing power, is the real bottleneck.
Core conclusion
The National Library of Norway is using a 2PB Huawei OceanStor Dorado all-flash array to train a sovereign-grade Norwegian large language model (LLM). This project revealed three key trends: any non-English-speaking country needs autonomous large models to protect language and cultural heritage; the movement of petabytes of data from archiving systems to training pipelines is a real bottleneck that "no one talks about"; and Huawei storage is playing an increasingly important role in mainstream IT infrastructure in Europe.
Key Points
- Event time: May 22, 2026 (disclosed at the Huawei ID Forum 2026 Paris meeting)
- Systems involved: 2PB Huawei OceanStor Dorado all-flash array + Nvidia DGX H200 + Norwegian national supercomputer Sigma2 Olivia
- Data scale: The library has 20PB of unique digital data (3-2-1 storage, totaling about 60PB)
- Core challenge: data movement across storage tiers from archiving system to training pipeline, not lack of computing power
Project background: Why does Norway need its own LLM?
Marius Husnes, Head of IT Platform at the National Library of Norway, revealed at Huawei ID Forum 2026 that no commercial LLM provider is developing a Norwegian large model. He bluntly stated: Any country with its own language without a sovereign-level LLM trained in that language will be at a disadvantage - a globally trained English LLM will not understand the country's history, news and culture as described in the local language.
The Norwegian Ministry of Culture therefore commissioned the National Library to build sovereign-grade AI. The reason is simple: the library has the largest collection of digitized Norwegian-language books in the country – digitization efforts began in 2005 and have accumulated 20 petabytes of unique data, including books, newspapers, web pages, audio, video and more.
A key advantage comes from the copyright side: the library has an agreement with a Norwegian newspaper to allow LLM training on copyrighted content. "No private company has this," Husnes stressed in his speech.
Technical architecture: Complete data pipeline from archiving to training
The core of the entire system is not the limitation of computing power - Husnes clearly stated that the bottleneck is data quality and pipeline throughput, not calculation.
| Hierarchy | System | Capacity | Role |
|---|---|---|---|
| Digital archiving (long-term preservation) | Disk + tape hybrid system | 20PB of unique data (60PB with copies) | Low speed, high durability, low cost |
| AI preprocessing environment | Huawei OceanStor Dorado all-flash array | 2PB flash memory | High-speed data cleaning, deduplication, and format standardization |
| Local Computing | Nvidia DGX H200 + 384-Core CPU Cluster | — | Data Pipeline Processing and Training Preparation |
| Training supercomputing | HPE Cray Supercomputing EX (Sigma2 Olivia) | 5.3PB Cray ClusterStor E1000 | Actual model training (448 GPU + 64,512 CPU cores) |
Data pipeline process
Data starts from the archive system and goes through a 6-stage pipeline:
- Data Ingest – Read raw data from 60PB archive system
- Clean – Remove noise and low-quality text
- Deduplication——Eliminate redundant content
- Format Standardization - unified into a training compatible format
- Validation – Check data integrity and quality
- Preparation——Package and send to Sigma2 supercomputer for training
Husnes specifically points out: No one is talking about the actual problem of moving data from petabyte-scale archiving systems to AI training pipelines. The 60PB archive system is optimized for durability and cost (high read latency, low frequency access), while AI training requires high throughput, low latency parallel data IO. His team had to figure out how to flow data between these two completely different storage systems on its own.
Still working out: Assessment, Governance and Orchestration
LLM training is still ongoing, and Husnes summarizes three areas of continued learning:
1. Assessment Difficulties
There are no standard tools for evaluating a sovereign Norwegian LLM. Norwegian has two written forms (Bokmål and Nynorsk), multiple dialects and historical variations. Teams are left building their own assessment tools on the fly.
2. Governance issues
Who has access to a sovereign LLM? Who decides where it can be used? This is an institutional and political question with no easy technical answer.
3. Three-system arrangement
Getting three independent systems—60PB archive library + 2PB flash AI preprocessing environment + Sigma2 national supercomputer—to work together is an ongoing engineering challenge.
Implications for the AI industry
This project provides several important insights:
- Sovereign Data Revaluation: Institutions (libraries, archives, media groups) with unique local data sets are becoming key asset holders for AI training
- Storage Architecture Cracks: The IO performance gap between archive systems and AI systems is a real and underestimated engineering problem, and petabyte-level data movement is not a simple "cp -r"
- Huawei Storage’s European Penetration: The OceanStor Dorado series is playing a central role in European government-level projects, which is of special significance in the context of U.S. technology restrictions on China
- Non-English market blue ocean: Almost all commercial LLMs focus on English, and minority language sovereign LLMs have a large amount of untapped demand and policy-driven opportunities.
Husnes's conclusion is worth pondering: "Norway is a small country, but we are solving a problem that every non-English speaking country will face - how do you build AI that reflects your language, culture and history? AI needs custodians, not just builders."
Tool entry
Relevant tools appearing in the text: n8n (data pipeline orchestration), ChatGPT (general LLM comparison), Hugging Face (model evaluation ecology), DeepSeek (local LLM deployment reference path).
Related reading
Internal link guidance
- Want to learn how to run AI tasks with low-cost local models? See: DeepSeek Reasonix in action: Build an AI programming agent at zero cost (30-minute tutorial)
- Real case: A security researcher uses Claude Code to mine vulnerabilities and earns $10,000 | Security researcher uses Claude Code for vulnerability mining: a real case of monthly income $10,000 per month
- Want to learn how to build a complete AI data automation pipeline? Watch: AI Agent-Driven Content Automation: n8n MCP Building Guide from Scratch
Monetization angle
How can you make money from this trend?
WayToClawEarn focuses on verified earn playbooks—not just news. Start from these cases.
n8n + OpenAI affiliate site
Automate content and affiliate monetization
Claude + n8n automation agency
Charge monthly for agent workflow builds