Karpathy Proposes 4-Rung LLM Output Ladder: STE100, Diagrams, HTML, Video — How Humans Understand Autonomous Agent Work?
Andrej Karpathy published on October 2 a four-rung ladder for LLM output formats: ASD-STE100 controlled writing → diagrams → HTML interactive pages → bespoke explainer videos. Core thesis: as LLMs work more autonomously, the bottleneck shifts from model capability to human understanding — output format is the key lever determining how fast humans can review and verify agent work. STE100 caps sentences at 20 words with 900 approved words.
Public-source compilation
Synthesized from public posts/docs. Prefer the original source for primary claims.
How we review content · Primary source · ExplainX / Karpathy Thread
TL;DR
On October 2, 2026, Andrej Karpathy published a four-rung ladder framework for LLM output formats: ASD-STE100 controlled writing → diagrams → HTML interactive pages → bespoke explainer videos. The thread garnered approximately 469K views within hours.
Karpathy's core thesis: as LLMs work more autonomously, the bottleneck shifts from model capability to human understanding. Output format is the key lever determining how fast humans can comprehend and verify agent work — "as models do more of the work, a lot more of our work will rise up the abstractions into oversight and understanding."
He also noted: as "intelligence and code are increasingly abundant," you can ask LLMs for "large, custom, discardable software artifacts (e.g. web apps, video explainers)" — artifacts that exist solely to transfer understanding into your head, not to be maintained.
The Four Rungs
Each rung is connected to the next with "but even better", signaling an ascending ladder of comprehension speed.
Rung 1: ASD-STE100 Writing
What is STE100: ASD-STE100 is a controlled language specification originally built for aircraft maintenance documentation. Started in 1979 with AECMA, became ASD-STE100 in 2005.
Spec details:
- Procedural sentences capped at 20 words
- Descriptive sentences capped at 25 words
- Paragraphs limited to 6 sentences
- Noun clusters limited to 3 words
- Dictionary of ~900 approved words, each with one meaning and one part of speech
Value for AI output: Short sentences, one meaning per word, active voice — forces clarity and removes ambiguity.
Key rules:
- Use active voice in procedures
- Do not leave out words like "the," "a," and "this"
- Use one word for one meaning, every time
- Write one topic in each paragraph
- Prefer "use" over "utilize," "start" over "commence"
Karpathy's practical tip: Since full STE is stringent, he sometimes asks for "80% of the way to ASD-STE100" — a good default since full-strength STE can read like a maintenance card.
Important caveat: A style instruction applied mid-task can shape the thinking itself, not just the wording. Fix: let the agent work in normal dense mode, then run a second pass that rewrites the result in STE for the reader.
Rung 2: Diagrams
Value: Makes structure visible rather than buried in paragraphs. Better when the answer has components and flows, states and transitions, a timeline, or a hierarchy. Worse when the answer is a nuanced argument — "a box-and-arrow picture of 'it depends' tells you nothing."
Practical tips:
- Ask for SVG diagrams with numbered arrows showing data flow, labeled with real component names
- Ask for Mermaid sequence diagrams, then list the three places the diagram is least certain — a diagram looks authoritative even when a label is wrong, so asking the model to flag its own weak spots is key
Rung 3: HTML Pages
Value: Turns a wall of Markdown into tabs, table of contents, collapsible sections, sliders, animations, and color-coded status boards. A single .html file opens in any browser and is shareable.
Why it's the best return for effort: One prompt change yields the biggest readability jump for the least setup. The blog recommends this as the starting rung for most users.
Practical examples:
- "Read this 40-page paper and create a single self-contained HTML file that explains it. Add: a one-paragraph summary at the top, a section per major claim, one interactive figure I can drag, and a 'what would change my mind' box at the end. No external dependencies."
- "Turn this PR diff into an HTML review page: group changes by intent, color risky hunks, and add a checklist of things I must test by hand."
Caveat: HTML is one more file to open and harder to diff than Markdown. For content that lives in git, gets diffed in pull requests, or is read by other agents, Markdown still wins.
Rung 4: Bespoke Explainer Videos
Value: Narrated, paced, visual, and tailored to a specific topic — the most powerful format for teaching a concept from zero.
Karpathy's quote: "fully custom / bespoke explainer videos generated on any arbitrary topic," adding that it "is actually starting to work!"
Three required components:
| Component | What does the job |
|---|---|
| Animation as code | HTML/canvas or Manim-style scenes the agent writes and renders |
| HTML to MP4 | Deterministic render of web animation to video |
| Narration | Paid TTS (e.g., ElevenLabs) or a local model |
Key technical shift: The video is "code, not diffusion" — the agent writes a program that draws each frame, so equations, graphs, and labels come out exact and are editable by changing a line.
Pre-trust checklist:
- Spot-check math or claims against a primary source (pick the three most surprising statements)
- Ask for the script separately — mistakes are easier to see in text
- Keep it discardable — if you plan to publish it, it needs the review a published piece needs
Core Thesis: Human Understanding Is the Scarce Resource
Karpathy's framework isn't just about prettier answers. It addresses a deeper problem:
"We'll be spending a lot more time trying to understand the outputs of language models."
"As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding."
Core risk: Better formats make a wrong answer easier to believe — persuasion without accuracy. This means:
- When agents autonomously generate HTML reports or video explainers, users are more easily persuaded by smooth presentation
- Format quality can mask factual errors
- Therefore, Karpathy's framework is fundamentally an audit methodology for the agent age, not just output beautification advice
Format Selection Guide
| If the answer is mostly... | Use |
|---|---|
| A procedure or explanation you will re-read | STE-style text |
| Structure, flows, or relationships | Diagram |
| Exploration, comparison, or many parameters | HTML page |
| A concept you need to learn from zero | Explainer video |
Developer Impact
For AI Agent Developers
Karpathy's framework directly impacts agent system design:
- Output format should be a configurable parameter: Different task types should default to different formats — HTML for code review, diagrams for architecture, STE for operational manuals
- Multi-format output combinations: Complex tasks can first generate an STE summary for quick scanning, then HTML details for deep study, and finally a video for zero-to-one learning
- Auto-uncertainty labeling: The practice of asking the model to flag its own weak spots in diagrams should extend to all output formats — this is key to reducing the "persuasion trap"
For Enterprise AI Adopters
If you're deploying AI agents in enterprise environments:
- Audit processes need upgrading: Traditional text review processes don't work for HTML reports and video explainers. Multi-format review standards are needed.
- Format as cost control: STE-style output reduces misreading risk through ambiguity reduction — worth making the default in compliance-sensitive contexts.
- "Discardable" principle: Karpathy emphasizes that generated artifacts are "custom, discardable" — don't try to maintain these artifacts. Their sole purpose is transferring understanding.
Experiment Suggestion
Karpathy implies a 15-minute experiment:
- Take something you don't fully understand (an inherited repo, a saved paper)
- Run all four rungs on it and time yourself:
- STE-style summary and read it
- Structure diagram
- HTML explorer with one interactive element
- Two-minute narrated video (if you have a TTS source)
- Note which one made you say "oh, now I get it" first
- For most readers, it will be rung 3 or 4 — that tells you which format to default to
Industry Context
Why It Matters Now
In H2 2026, AI agents are rapidly moving from single-turn chat to multi-step autonomous execution. OpenAI's Dots, Anthropic's Claude Code, and Google's Gemini Agents all push toward "give a goal, the agent completes all steps autonomously."
Against this backdrop, human work shifts from "executing each step" to "reviewing agent output at the endpoint." Karpathy's framework directly addresses this shift — when humans no longer execute step-by-step but review at the endpoint, output format determines review efficiency.
Echoes With Other Thinkers
Karpathy's perspective resonates with other industry discussions:
- Anthropic's "Teaching Claude Why" research focuses on agent behavior alignment
- Major labs' interpretability research focuses on model internal mechanisms
- Karpathy's framework focuses on the output layer — how humans understand agent work
These three cover input alignment, internal understanding, and output review respectively, together forming a complete human-agent collaboration framework for the AI agent age.
Conclusion
Karpathy's four-rung ladder framework is simple and practical. It doesn't ask you to choose one format — it provides a progressive understanding-acceleration ladder: from controlled writing to diagrams to interactive pages to video, each step increases human comprehension speed but also increases generation complexity and "persuasion trap" risk.
For AI practitioners, the framework's core insight is: in the agent age, the return on investing in output format optimization may be higher than investing in prompt engineering — because the former directly determines whether you can quickly trust and verify agent work.
Topic hub
YouTube AI Content Policy Hub
Answer-style evergreen hub for AI labels, auto detection, and disclosure—not just breaking news.
Explore YouTube AI Content Policy Hub →Monetization angle
How can you make money from this trend?
WayToClawEarn focuses on verified earn playbooks—not just news. Start from these cases.
ChatGPT ads + content distribution
Sell compliance checklists and automated distribution
OpenClaw Agent short-video growth
Lean into hybrid workflows as labels get stricter
Related tutorials
- How to build an AI content automated distribution system with n8n + ChatGPT: a complete 30-minute tutorial
- Claude Code automated writing practice: build an AI content production pipeline in 30 minutes
- AI Agent drives automated website operations: Build a fully automatic content pipeline in 30 minutes
- AI Agent-Driven Content Automation: n8n MCP Building Guide from Scratch
Related news
- Tavus Griffin Passes Video Turing Test: 48% Mistook AI for Human, How to Choose Real-Time Video AI
- Hikvision Guanlan AI Applications: From Vision Recognition to Industry Agents
- Overmind Open-Sources Specialized SLM Training Platform: Ex-Intelligence Officers Launch, How to Beat Frontier Models With Your Own Data?
- Comfy Org Launches Comfy Agent: Autonomous Workflow Building on the Canvas, How to Automate ComfyUI Pipelines?