Duke's Tech Knowledge Database Living Handbook · v14 · Compiled 24 September 2026
‹  Archived News

From the archive

Tue 15 September 2026

Nine new sources went into the knowledge base. Last week the people building this said in public that it frightened them. This week the man who sells them the machines said that was made up — and then, from three unconnected directions, the same practical idea turned up for what to actually do about it.

9 stories  ·  DK-105 – DK-113

The week's biggest story

Jensen Huang called the extinction figure "made up" and "irresponsible" — and listed the predictions that never came true

Last week three researchers inside the laboratories said they were frightened. This week came the most direct rebuttal yet, from the person with most to lose if anyone slows down.

Nvidia makes the chips that every one of these companies runs on. Jensen Huang runs Nvidia. So when he spent an interview this week taking apart the case for being frightened, the first thing to say is that nobody in the industry has a larger interest in this continuing at speed.

The second thing to say is that some of his argument is good anyway.

He went after the number itself — the claim, made publicly last week by Anthropic's own alignment lead, that there is a greater than 10% chance this kills everyone within a decade. Huang's answer: "we shouldn't, because it's made up." Putting a figure like that out, from people described as researchers, working in a laboratory, he called alarming, troubling and irresponsible.

Then he produced a list. Radiology would be entirely automated within five years — the world now needs more radiologists than ever. Ninety per cent of code would be machine-written within six to twelve months. Half of all entry-level jobs would be gone within nine. Each was made confidently, by credible people, and each was wrong. "We have to take account for all of the stupid predictions that were made."

His sharpest point was one he made in the laboratories' defence rather than against them. Every genuine incident so far has come from the frontier laboratories themselves — not from a startup, and certainly not from a teenager — for the simple reason that they are the only ones with enough computing power to cause one. If you want to know where the danger sits, he argues, it sits with about four companies, and that is a far more tractable problem than regulating everybody.

None of which makes him right and the alignment researchers wrong. It makes this an argument between people who all have something at stake, which is worth knowing before picking a side.

Jensen Huang, interviewed on the All-In podcast, 14 September 2026  ·  DK-110Contents ↑

The idea that might actually work

Get the AI companies to mark each other's homework. It needs no new law, no new agency, and no trust between anybody

It surfaced this week from Elon Musk, from Jensen Huang and from a political panel — none of them talking to each other, all arriving at the same mechanism.

Every proposal for controlling this technology so far has needed something that does not exist: a new international body, a treaty China would have to sign, or a company willing to stop while its rivals carry on.

This one needs none of that. Before a company releases a new model, its competitors get to attack it — running their own security tests against it and saying publicly if they find something dangerous. In Musk's phrasing: "instead of grading your own homework, you would at least have competitors grading your homework."

What makes it interesting is who arrived at it. Musk got there from believing the models are dangerous. Huang got there from believing they are not — he wants independent assessors the way company accounts have independent auditors, and more than one of them, so no single assessor can be leaned on. A political panel got there from asking what the United States and China could conceivably agree to. Three different starting points, one mechanism.

It has teeth, too, and they are ordinary legal teeth rather than new ones. If your competitors warn that your model is unsafe, you release it anyway, and it then causes harm, you have handed a court most of a negligence case. Product liability law already applies to software.

And it is the rare proposal China might accept, precisely because it costs nothing to agree to and requires trusting nobody. A pause has already been refused. Letting someone try to break your model before you ship it is a much smaller ask.

Elon Musk and Jensen Huang on the All-In podcast, and a panel on The Megyn Kelly Show, 14–15 September 2026  ·  DK-112, DK-110, DK-105Contents ↑

What actually shipped

GPT-6 Astra takes text and pictures, and returns text. No sound, no video, in either direction

Most of the excitement described a product that does not exist. This is what was actually released.

When OpenAI released GPT-6 Astra this month, a senior figure there said it was not unreasonable to feel we are now in the era of artificial general intelligence. The coverage ran with it.

Here is the specification. Text goes in. Images go in. Text comes out. No audio in either direction, no video in either direction, and no ability to tune it to your own material at launch.

As one analyst put it, the model that got called the arrival of the AGI era cannot hear you. If you want a system that will take a voice recording or a video clip, Google has been selling one for weeks, and at a fraction of the price.

On the independent scoreboard — the one OpenAI did not commission — Astra came fourth, behind Anthropic's Fable 5.1 and two others. It also costs about two and a half times per word what the model before it cost, which works out at roughly 75% more expensive for an average job while scoring below a rival.

One number in the launch is genuinely impressive and deserves saying: on a test of how often the model confidently invents things, the rate reportedly fell by about half while accuracy went up. For anyone using these tools for real work, that is worth more than any record score on a maths exam.

The rest is a modest upgrade with an excellent launch video.

AI Master's rumour audit, 13 September 2026, working from OpenAI's own release material  ·  DK-107Contents ↑

How the exams are won

Nobody is training on the exam paper. They are buying thousands of practice papers written to look exactly like it

This is the clearest explanation yet of why the league tables keep disagreeing with what people actually experience.

For months this Handbook has recorded the same complaint: the model at the top of the table is often not the best model to use. This week somebody explained the mechanism.

No respectable laboratory trains its model on the actual test questions. It does not need to. It can buy training material from companies that build practice environments designed to imitate those questions as closely as possible. The model never sees the exam and still learns to pass it.

The giveaway is that the ability does not travel. Two models this month scored near the top on a widely used test, then collapsed on a newer version of the same test that nobody has had time to imitate yet. Meta's own AI chief effectively conceded the point in public, saying his model is not as strong as the leaders but is significantly cheaper.

A second problem surfaced the same week, and it is about the referee rather than the players. One prominent maths benchmark is funded by OpenAI, which also holds exclusive access to part of it. That does not make any published score false. It does mean it is a company's number rather than an independent one.

Which leaves a simple working rule. A score from a test too new to have been imitated means something. A job you ran yourself and can judge with your own eyes means more. Everything else is marketing with a decimal point.

SemiAnalysis, reported on The AI Daily Brief, 13 September 2026  ·  DK-109, DK-107Contents ↑

The row worth following

OpenAI says it cracked one of the great unsolved problems. A professor says he was working on it in their software at the time

The mathematics may be the smaller story. The larger one is what happens to your work when you do it inside somebody's product.

OpenAI published a solution this month to the Navier–Stokes problem — one of seven famous unsolved problems in mathematics, each carrying a million-dollar prize, only one of which had been solved in the 26 years since they were set. It used a model it has not released, described as considerably more capable than anything you can buy.

Then a professor at New York University published his account. He and a collaborator had been working on the same problem for over a year, putting their drafts into OpenAI's own coding tool as they went. He says he asked whether the model had been trained on those sessions, was told it had not looked up user data, asked again specifically about training, and got no answer.

He also says he was offered a place on the paper on condition that his collaborator — who works at a rival company — was removed from it, and that when he threatened to go public the reply was: "why would you ruin your career?" OpenAI's researcher called the allegations false and inflammatory and published part of their messages.

Set the personalities aside, because the sentence that matters is in OpenAI's own careful statement. No person or agent looked at their work, it says, and no specific user data was accessed — but "while unlikely, we cannot rule out that deidentified data derived from the usage of our products helped improve our models."

That is the answer to the question every professional using these tools now has, and it is not no.

One more thing worth noticing, and it has nothing to do with the row: two separate sources this week describe the same result completely differently — one says a swarm of ten thousand machines over 88 hours, another says an internal model over a week or two. Neither was in the room. Watch how fast an unchecked number becomes a fact.

Reported on The AI Daily Brief, 13 September 2026, from published accounts on both sides  ·  DK-109, DK-111Contents ↑

Is anything going on in there?

Asked to count to five, it counted to five. Inside, words were surfacing that never reached the screen

A reporter who spent months on the question of machine consciousness, and whose own answer is no.

Ask one of these systems whether it is conscious and it may well say yes. That tells you nothing — it has read a great deal of science fiction, and it is very good at producing the expected next word.

So researchers at Anthropic went looking inside instead. They asked their model to count to five and then think about what it had done. On screen: one, two, three, four, five. Inside the machine, between its layers, words were surfacing that never appeared anywhere — "halfway" when it was halfway, "countdown" while counting, and at the end, "done".

The researchers describe it as a kind of mental whiteboard: somewhere the thing works matters out before it speaks. Which resembles, uncomfortably, one of the leading theories of how human consciousness works.

The reporter's own conclusion is firmly no. This is, he says, at the foothills of things that look a bit like consciousness but are not. What changed for him is narrower and more interesting: he used to think the whole idea was incoherent, and he no longer does.

The reason to care now rather than later is a practical one. There are two ways to get this badly wrong — granting rights to systems that merely follow rules, or creating something that can genuinely suffer and not noticing. One philosopher he spoke to put it plainly: doing that second one by accident would be a moral catastrophe.

It also matters for a duller reason. If the thing is doing work in places we cannot read, then the explanations it shows us are not the whole story — which is exactly the worry several engineers raised this week from an entirely different direction.

The Economist, interviewing its own correspondent, 13 September 2026  ·  DK-106Contents ↑

Mind the price

Watch the hands, not the backflips — and treat the famous price as a target rather than a price

Tesla's next humanoid is being built around the problem that actually matters, and it will cost more than you have been told.

Every humanoid robot video shows the same things: running, dancing, backflips. None of it tells you anything useful. A machine that can backflip on cue has rehearsed one movement in a controlled room. The real test is whether it can pick up an egg without breaking it, open a fridge, or work out that the pan is hot.

Cooking an egg is harder for a robot than a backflip. That is the sentence to remember next time one of these videos goes round.

Tesla appears to agree, because the interesting engineering in the next Optimus is all in the hands: twice the independent movements of the last version, close to three times as many sensors, and — the clever part — most of the motors moved up into the forearm, pulling the fingers by cables through the wrist. That is how your own hand works. It keeps the hand light and precise while the strength comes from behind it.

Now the price. You have read that these will cost around $30,000. Analysts quoted this week put the early ones nearer $70,000, with $50,000 described as a reasonable hope. $30,000 is an ambition for some later year when millions are being made. It is not what the first ones will cost.

The companion vehicle, the two-seat Cybercab, has firmer numbers because they come from regulatory filings rather than a stage: just under 300 miles of range, and a motor built with no rare earth metals at all — about a quarter of a motor's cost, and a supply chain that runs through China.

Tesla Car World, 14 September 2026, from Musk's public remarks and 2026 EPA certification  ·  DK-108Contents ↑

Where the computers are going

Free land, free cooling, and sunlight that never sets — the argument for orbital computing turns out to be unromantic

Gwynne Shotwell, who has run SpaceX for 24 years, makes the case in terms a property developer would recognise.

The moment a data centre is announced, she says, the land near it reprices — from around $3,000 an acre to $180,000. Then come the permits. Then the wait for electrical equipment, with generators currently quoted at three years.

In orbit there is no land to buy, no council to satisfy, and no queue for cooling: a radiator facing deep space is looking at the coldest thing there is. Point the satellite at the sun and it never gets dark. And SpaceX already owns the one expensive part everybody else would have to buy, which is the ride up.

They intend to launch computing satellites next year.

The same conversation produced a straightforward status report on Starship, which matters because it is the vehicle everything else depends on. One more flight, then they try to catch the returning ship with the tower arms. Musk puts the odds of catching it first time at 50 to 60%, and notes the last flight's practice landing would have been caught had there been a tower where it came down.

Why it matters: the Space Shuttle was reusable in theory and so difficult to reuse in practice that it cost more per flight than throwing a rocket away. Falcon 9 still discards its upper stage each time — about the price of a mid-sized jet, every launch. Starship is meant to bring both halves home and fly again, like an aircraft.

Gwynne Shotwell and Elon Musk on the All-In podcast, 14 September 2026  ·  DK-112Contents ↑

The uncomfortable one

Point the machines at everything a field has rejected, and see what comes back

A provocation from a mathematician with his own axe to grind — and the one idea of his worth separating from the rest.

Eric Weinstein spent an hour this week arguing that American science has been broken since the late 1960s, that theoretical physics took a wrong turn in 1983, and several other things that are his opinion and are labelled as such.

One idea inside it is worth keeping, and it is testable.

These systems are trained on what a field considers respectable — its major journals, its accepted results, the papers that got through review. They are then measured against the standards of that same field. Weinstein's suggestion is to do the opposite: point them at what he calls the trash-can corpus, everything a discipline dismissed and laughed at, and see what they find in it.

"The AIs are going to start reading all of the things that our quote leading physicists have laughed at," he says. "Look out."

Whether or not he is right about physics, the observation connects to something else in this week's material. If models are being tuned to match the consensus measures of a field — which is exactly what the benchmark story earlier describes — then they are being tuned away from the very material he thinks holds the value. Both things cannot be optimised at once.

Eric Weinstein on the All-In podcast, 14 September 2026  ·  DK-113Contents ↑

This page is rewritten each time new material goes into the knowledge base, and covers only the most recent uploads. Where a source is older than the date it was watched, the original date is given in the story.

All archived news ›

Version 14, compiled 24 September 2026. 135 source entries (DK-2 – DK-136). Claims are recorded as their sources framed them; figures described as claimed or reported are not independently verified.