Sources
The essay claims to separate what is established from what is merely believed. This page is the bookkeeping behind that claim: every source, grouped by the section it supports, with the weak ones marked ⚠ and one corrected error left on the record.
References used while writing this essay, grouped by the section they support.
How to read this list, and its limits
Two different kinds of claim appear in this essay, and they were sourced differently.
Recent claims — 2026 model names, prices, parameter counts, chip specifications, energy figures, export-control status — were gathered through web searches run in August 2026, because they change too fast to state from memory. The results below are what those searches returned.
Foundational claims — the history from 1972 to roughly 2023, and how backpropagation, transformers and training actually work — come from the established literature. The canonical papers are cited in the last section so a reader can go to the primary text.
Three honest caveats:
- Not every URL below was opened and read in full. Search results were used partly as summaries. Anything load-bearing should be re-verified against the primary source before being quoted elsewhere.
- Source quality is uneven, and is marked below. Official announcements, institutional research (IEA, CSIS, Gartner, Goldman Sachs) and vendor technical documentation are solid. Weaker results — SEO-style listicle sites — are marked ⚠ and their numbers deserve independent confirmation. The Chinese open-weight section was originally sourced this way and has since been re-verified against the labs' own papers and model cards; doing so corrected a genuine error about Qwen's context length and model sizes, which is a fair indication of how much the aggregators should be trusted elsewhere.
- This will age fast. Figures were current in August 2026 and some were already provisional then.
01 · History, 1972 → 2026
Historical material is drawn from the primary literature (see Foundational papers below). The 2024–2026 portion of the timeline draws on the sources in sections 05–07 of this file.
02 · Foundations · 03 · Transformers
Entirely from the foundational papers listed at the end. The gradient-descent, neural-network and polynomial-fitting demos implement textbook algorithms directly, with no external source beyond those papers.
04 · How models learn — training and interpretability
- Transformer Circuits Thread — Anthropic's interpretability publication venue; the primary source for feature extraction, circuit tracing and attribution graphs. Primary.
- Anthropic Fellows Program / alignment research — open-sourced attribution-graph tooling. Primary.
- Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models — peer-reviewed interpretability benchmarking. Primary.
- Making Sense of the Unsensible: Reflection, Survey, and Challenges for XAI in LLMs — survey of explainability limits. Primary.
- Secondary commentary on Anthropic's interpretability results, used for orientation only: IntuitionLabs overview ⚠
Claims specifically checked against these: that circuit tracing yielded satisfying insight on only about a quarter of prompts examined; the rhyme-planning and multi-step-reasoning findings; the unfaithful chain-of-thought result; and the 2025 introspective-awareness experiments. The "171 emotion concept vectors" figure surfaced in secondary reporting and is treated in the essay as functional representation rather than evidence of felt experience — deliberately hedged.
05 · Vocabulary of 2026
- Agentic AI in 2026: What Every Developer Needs to Know
- All Agent Harnesses: The Live Comparison
- AI Agents and Agentic AI — Navigating a Plethora of Concepts — Primary, academic treatment of the agent/goal vocabulary.
- Harness-engineering commentary: Lyzr ⚠
The "goals" definition (a persistent, mechanically verifiable objective pursued across long autonomous loops) reflects how the term was being used for Codex's /goal feature and equivalents in mid-2026.
06 · The players
Anthropic — Primary: Introducing Claude Fable 5 and Claude Mythos 5, Redeploying Claude Fable 5. Journalism: TechCrunch, 9 June 2026, CNBC, 9 June 2026, CNBC on export controls being lifted, 30 June 2026.
OpenAI — Primary: GPT-5.6, Previewing GPT-5.6 Sol, Introducing GPT-5.5, Model release notes, GPT-5.6 Preview System Card. Journalism: CNBC, 8 July 2026.
Google DeepMind — Primary: DeepMind models, Gemini, Gemini API changelog, Google AI updates, July 2026. Journalism: TechCrunch, 21 July 2026 — source for the Gemini 3.6 Flash release and the confirmation that Gemini 4 is in pre-training.
Meta — Primary: The Llama 4 herd, The future of AI: Built with Llama, Llama 4 developer page, meta-llama on Hugging Face.
xAI / SpaceXAI — added August 2026, after a reader noticed the lab was named in section 07 without ever being introduced in section 06.
The merger is the load-bearing claim and is confirmed by two independent outlets: CNBC, 2 February 2026 and CNBC on the $1.25T combined valuation, plus Bloomberg. xAI became a wholly owned SpaceX subsidiary in an all-stock deal valuing it at $250B; the group now trades as SpaceXAI.
Models — Primary: xAI release notes, which give Grok 4.6 as the current flagship (500k context) and Grok 4.5 as the July 2026 release. Parameter counts circulating for these models come from aggregators only and are not stated by xAI, so the essay does not quote any. ⚠ Secondary roundups claiming 1.5T for Grok 4.5 or 6–10T for an unreleased Grok 5 should be treated with suspicion.
Colossus — Wikipedia) and SemiAnalysis on Colossus 2. The figures used — ~100,000 GPUs in about four months for the first cluster, ~555,000 GPUs and roughly a gigawatt by 2026 — are consistent across these and industry trackers ⚠, but no primary xAI disclosure confirms them; the essay hedges accordingly ("quelque", "de l'ordre du gigawatt").
Grok-1's 314B weights under Apache 2.0 (March 2024) are independently checkable on the model's public repository.
Mistral — TechTimes on the July 2026 open-weight early access. Remaining Mistral details (Large 3, Small 4) came from aggregator round-ups ⚠ and would benefit from confirmation against Mistral's own release notes.
DeepSeek and the Chinese ecosystem — this section was originally sourced from comparison/listicle sites and has since been re-verified against primary sources. All figures below now come from the labs' own papers, model cards and release posts.
DeepSeek — Primary: DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence (technical report), DeepSeek-V4-Pro model card, DeepSeek-V4-Flash model card, DeepSeek API release notes. V4 is a two-model family, not one: V4-Pro at 1.6T total / 49B activated and V4-Flash at 284B total / 13B activated, both at 1M-token context, using a hybrid of Compressed Sparse Attention and Heavily Compressed Attention.
Moonshot AI — Primary: Kimi-K3 model card, MoonshotAI/Kimi-K3 on GitHub. 2.8T parameters, released 16 July 2026, weights under the Kimi K3 License. Built on Kimi Delta Attention with native vision and a 1M-token context. Notably it uses quantization-aware training from the SFT stage in MXFP4, so the released weights occupy ~1.4TB instead of the ~5.6TB FP16 would require. The earlier "largest open-source model ever released" phrasing has been replaced with the more defensible "first open-weight model in the three-trillion-parameter class".
Alibaba / Qwen — Primary: Qwen3.6-35B-A3B model card, Qwen3.6-27B model card, QwenLM/Qwen3.6 on GitHub, Qwen release post. This corrected a real error. Aggregators described a "Qwen 3.6 Plus" with a native 1M context; the actual open-weight releases are Qwen3.6-27B (dense) and Qwen3.6-35B-A3B (35B total, 3B activated MoE), with a 262,144-token native context extensible to ~1,010,000, under Apache 2.0.
Zhipu / Z.ai — Primary: GLM-5.2 model card, GLM-5.2: Built for Long-Horizon Tasks, zai-org/GLM-5 on GitHub, GLM-5 technical report. MoE with 256 routed + 1 shared expert (8 per token), 78 layers, 1M-token context, released under MIT.
The January 2025 DeepSeek R1 market reaction (~$600B of Nvidia market value in a day) is widely reported and independently checkable.
07 · Infrastructure and geopolitics
Lithography physics:
- *Physics of laser-driven tin plasma sources of EUV radiation for nanolithography*, Plasma Sources Science and Technology — Primary. Source for the tin-droplet laser-produced plasma and its 13.5 nm emission. The paper gives the plasma temperature in electronvolts — "several tens of eV" — which is a few hundred thousand degrees; popular accounts quote anywhere from 220,000 °C to 500,000 °C, so the essay gives the order of magnitude rather than pick one.
Semiconductors and export controls:
- CSIS — Understanding the Biden Administration's Updated Export Controls — Primary/institutional.
- International Center for Law & Economics — US Export Controls on AI and Semiconductors — Primary/institutional.
- Silicon Analysts — China AI semiconductor localization — industry analysis; source for the CXMT HBM trajectory.
- Semiconductors Insight — H200 export policy shift, 2026 ⚠
- Foreign Affairs Forum — TSMC's $265B investment and the AI power map
GPUs and memory:
- NVIDIA — Inside the Vera Rubin Platform — Primary; source for 336B transistors, 288GB HBM4, ~22 TB/s, 50 PFLOPS NVFP4.
- Arc Compute — Preparing data centers for Rubin and HBM4
- Introl — Rubin enters full production, CES 2026
- VRLA Tech — NVIDIA GPU roadmap 2026–2030 ⚠
Energy:
- IEA — Energy demand from AI — Primary/institutional, the most authoritative source in this section.
- Gartner — data center electricity consumption to grow 26% in 2026 — Primary; source for the ~132 GW / 565 TWh 2026 figures.
- Goldman Sachs — US data center power demand to double by 2027 — Primary.
- Brookings — Global energy demands within the AI regulatory landscape — Primary/institutional.
The comparison of ~130 GW to Germany's average national draw is an order-of-magnitude illustration, not a precise equivalence.
Foundational papers
Cited from the established literature rather than from the August 2026 searches. These are the primary texts behind sections 01–04.
| Year | Work | Relevance | |---|---|---| | 1973 | Lighthill, Artificial Intelligence: A General Survey | Triggered the first AI winter | | 1986 | Rumelhart, Hinton & Williams, "Learning representations by back-propagating errors", Nature 323 | Backpropagation enters the mainstream | | 1989 | LeCun et al., "Backpropagation Applied to Handwritten Zip Code Recognition" | First convolutional network in production use | | 1997 | Hochreiter & Schmidhuber, "Long Short-Term Memory", Neural Computation | Memory across time | | 2009 | Deng, Dong, Socher, Li, Li & Fei-Fei, "ImageNet: A Large-Scale Hierarchical Image Database" | The dataset bet | | 2012 | Krizhevsky, Sutskever & Hinton, "ImageNet Classification with Deep Convolutional Neural Networks" | AlexNet; 15.3% vs 26.2% top-5 error | | 2014 | Goodfellow et al., "Generative Adversarial Networks" (arXiv:1406.2661) | Networks that generate | | 2016 | Silver et al., "Mastering the game of Go with deep neural networks and tree search", Nature 529 | AlphaGo | | 2017 | Vaswani et al., "Attention Is All You Need" (arXiv:1706.03762) | The transformer | | 2018 | Devlin et al., "BERT" (arXiv:1810.04805) | Pre-training economics | | 2020 | Kaplan et al., "Scaling Laws for Neural Language Models" (arXiv:2001.08361) | Intelligence acquires a price curve | | 2020 | Brown et al., "Language Models are Few-Shot Learners" (arXiv:2005.14165) | GPT-3 | | 2022 | Power et al., "Grokking: Generalization Beyond Overfitting" (arXiv:2201.02177) | The delayed jump to a general algorithm | | 2022 | Bai et al., "Constitutional AI" (arXiv:2212.08073) | Training against written principles | | 2024 | Templeton et al., "Scaling Monosemanticity" (Transformer Circuits) | Feature extraction; Golden Gate Claude | | 2025 | Anthropic, circuit-tracing and attribution-graph work (Transformer Circuits) | Rhyme planning, multi-step reasoning, unfaithful reasoning |
On the interactive components
The two computation-heavy demos implement their algorithms directly and are not driven by any external data: NetworkPlayground trains a multilayer perceptron with backpropagation and SGD in the browser, and FitPlayground solves least squares via normal equations with light ridge regularization.
AttentionPlayground and TokenPredictor use hand-authored illustrative values, not measured model outputs. The attention weights are shaped to resemble published visualisations of real attention heads, simplified to a single head for legibility; the token logits are plausible rather than sampled from a live model, though the softmax-with-temperature computed over them is exact. Both are labelled as illustrative in the interface itself.