General Operating Principles
As of this version, six genuine cross-topic patterns have emerged clearly enough across multiple, independent source entries to record here:
- Treat numeric claims and benchmark scores as reported, not verified. Specific figures (benchmark scores, cost comparisons, usage statistics) circulating in AI-focused YouTube and newsletter content are frequently self-reported by the company being discussed, or repeated secondhand by the presenter from another source. This Handbook records what a source claims, not confirmed fact — particularly for speculative or rumor-driven content, where the original creator may themselves flag the number as unconfirmed (DK-5, DK-6, DK-8) (DK-5, DK-6, DK-8). NEW (v11) — v11 adds a failure mode beyond self-reporting and secondhand repetition: a figure with no human source at all. A produced explainer on AI energy demand states on air that two of its central numbers were obtained by asking Grok, then presents both as findings (DK-89). Both are excluded from this Handbook entirely rather than hedged, because hedging implies a source that can be checked and there is none. The test to apply is not only "is this verified" but "could this have been verified by anyone, including the person saying it".
- Content-moderation posture is becoming a standard comparison axis. Multiple sources now evaluate AI models not just on output quality but on how restrictive or permissive their safety/censorship policies are — this is treated as a first-class comparison dimension alongside benchmark performance, not an afterthought (DK-6, DK-8).
- Cross-checking a claim across more than one independent source — including running the same question through a second AI model — is a recurring, recommended verification technique, not a one-off tip. It shows up both in AI-literacy content aimed at spotting manipulated media and in general AI-usage advice aimed at catching hallucinations, suggesting it's a durable practice worth applying to this Handbook's own sourcing as well (DK-9, DK-15). A concrete example: two independent creators covering the same WWDC 2026 Siri AI reveal on the same day corroborated the same core feature set (DK-38, DK-39), and a third creator's hands-on review weeks later reinforced the same feature set again, including matching performance figures (DK-41).
- Long-range speculative or predictive content should be explicitly tiered by confidence — separating what's already happening, what's plausible but unconfirmed, and what's highly speculative — rather than presented as a single flat forecast. This discipline, modeled by more than one source's own methodology, is applied whenever this Handbook incorporates futurist or predictive material rather than treating a prediction with the same weight as a reported fact (DK-6, DK-24).
- NEW (v10) — Frontier AI releases are increasingly gated by pre-launch government review, not just post-hoc regulation, and this now cuts across model launches, funding structures, and safety testing alike. The pattern first appeared with Fable 5/Mythos 5's reported export-control restriction (DK-6), and v10's newsletter backfill shows it recurring on a near-monthly cadence: the Commerce Department lifting restrictions on Anthropic's models (DK-70), the Trump administration granting a delayed green light before GPT-5.6 could launch publicly (DK-72), a newly completed White House voluntary framework letting the government get up to 30 days' early access to "covered frontier models" for cybersecurity testing (DK-79), and OpenAI's proposed 5% government ownership stake — explicitly read by commentary as trading equity for political goodwill given the timing (DK-71). Treat this as a standing feature of how frontier releases now work in the US, not a one-off Anthropic story.
- NEW (v12) — A benchmark score measures what its index chose to weigh, and the index changes. This earns a place here rather than inside AI Model Landscape because it now appears across three topics with independent sources, and because it governs how every other numeric claim in this Handbook should be read. Two reputable suites ranked the same two models in opposite orders in the same week, because one leant towards mathematics and factual recall and the other towards broad economic work (DK-91). When GPT-6 Astra scored poorly on the first, Artificial Analysis published a revised index within days that weighted agentic tasks more heavily, after which the same model led everything except Fable 5.1 (DK-99) — the model did not change, the ruler did. In image generation the same problem appears as a change in what "good" means at all: a reasoning model scores correctly on instruction-following where a diffusion model produces a more attractive picture that ignores the instruction (DK-95). The practical rules this Handbook now applies: a claim that a model is "the best" is incomplete unless it names the benchmark; a lab's choice of which benchmarks to publish is itself evidence of what it believes it is selling; and an index revised shortly after an embarrassing result should be recorded as revised, not quietly adopted. This extends rather than replaces the standing principle that numeric claims are recorded as reported and not verified. NEW (v13) — v13 adds the mechanism, and it is worse than a changing ruler: a public benchmark can be beaten without ever being trained on. SemiAnalysis's argument, applied to Gemini 3.8 Flash and Muse Spark 1.3, is that a lab need not touch the tasks directly because it can buy data from reinforcement-learning environment startups built to mimic them as closely as possible — the net effect is the same, and the tell is failure to generalise, with both models comparable to the frontier on terminal-bench 2.1 and markedly worse on 4.0 (DK-109). Meta's own AI chief conceded the substance in public. A second mechanism arrived the same week: the referee can be on the payroll, with Epoch AI reported to have disclosed that OpenAI funded Frontier Math and holds exclusive access to part of it, and OpenAI reported to have noted its Claude comparison runs used modified evaluation settings (DK-107). And the divergence now reproduces from one consistent hands-on tester across two weeks and two different models: the model topping the coding leaderboard produced the weakest practical output of four tested (DK-101), and a week later the same complaint recurred on a new model — "I don't quite understand how these models are scoring better and better on this benchmark when what I'm seeing doesn't really compare" (DK-104). The rule this Handbook now applies: a public benchmark score is weak evidence by construction, and the only scores carrying much signal are those from a benchmark too new to have been farmed, or from a task you ran yourself and can judge with your own eyes.
This section will continue to fill in only as further evidence reinforces a pattern across independent sources — not from a single entry alone.
- NEW (v11) — Electrical power, not compute, is now named as the binding constraint by independent sources across four separate topics in this Handbook, which is what earns this a place here rather than inside any one of them. It appears as capital expenditure and grid capacity in industry analysis, with a single site put at 1.2 gigawatts — more than a mid-sized city — condensed into a footprint orders of magnitude smaller (DK-86); as the explicit thesis of an energy arms race in which every major player needs its own Colossus-scale facility and the US grid cannot absorb that on a timescale of years (DK-89); as the reason a leading researcher expects to need "more power and more resources" and cannot say whether they will be available (DK-85); as the design driver behind Tesla's AI5 chip being optimised for roughly 250W to extend Optimus's operating time (DK-53); and as the entire premise of the orbital-compute thesis, where the stated advantages are constant solar power and free cooling (DK-66). The useful form of the principle is this: capability announcements are increasingly downstream of energy procurement, so an energy constraint is now a legitimate reason to discount a capability roadmap. Note that this Handbook holds no verified figure for any specific facility's draw — see the numeric-claims principle above.
Version 13, compiled 15 September 2026. 112 source entries (DK-2 – DK-113). Claims are recorded as their sources framed them; figures described as claimed or reported are not independently verified.
