Duke's Tech Knowledge Database Living Handbook · v14 · Compiled 24 September 2026
‹  Archived News

From the archive

Sat 12 September 2026

Four new sources went into the knowledge base. The one that matters most is not a product launch: three researchers inside the big laboratories spent this week saying, publicly and by name, that they are frightened of what they are building.

9 stories  ·  DK-101 – DK-104

The week's biggest story

Three researchers inside the biggest laboratories put their names to the warning this week. One of them resigned to do it

This is not the critics talking. It is the alignment lead at Anthropic and the chief scientist at OpenAI.

A researcher who says he spent three years on foundational work at both OpenAI and Anthropic resigned from Anthropic this week and said why in public: "Neither company is acting responsibly. They're racing straight to self-improving super intelligence and gambling with our lives."

Resignation letters are easy to discount. What happened next is not. Evan Hubinger, who leads alignment science at Anthropic — the team whose job is making these systems safe — replied publicly agreeing with him. "We really do earnestly believe AI could kill all humans. I personally think it's a greater than 10% chance within the next decade."

He then said something that is harder to put down. "We do not yet have a plan to solve alignment for super intelligence and are not clearly on track to."

Set that against the reason the same company gives for building this at all: that nobody else can be trusted to do it safely, so it must get there first. The firm claiming it alone can handle this is also stating, on the record, that it has no plan and is not on course to find one.

And it is not one company. OpenAI's chief scientist published an essay the same week called "An Alien Mind", writing that he has "a strong expectation" that progress continues into systems that improve themselves, and that "this is a time that calls for extreme caution. I am concerned no one is prepared for the consequences."

Read the next story before you decide what to make of this one.

Reported by Matt Wolfe, 11 September 2026, quoting public posts and OpenAI's "An Alien Mind"  ·  DK-104Contents ↑

Before you panic

There is money in frightening you about this, and a physicist has the receipts

Which is exactly why it matters who is doing the warning, not just what the warning says.

The usual reply to a story like the one above is that it is a marketing exercise — that a company selling the most powerful technology in the world benefits from you believing it is dangerous.

That reply is not baseless. The physicist Sabine Hossenfelder has published a video saying she was offered money to tell her audience that AI will kill us all. There are organisations paying people to spread that message.

So both incentives are real at once. Some alarm is bought and paid for. Some of it comes from the person running the safety team at the company doing the building, with his name on it, contradicting his own employer's commercial interest.

There is a second thing worth holding. The people quoted in the story above spend their working lives modelling worst cases, surrounded by others doing the same. That is a reason their estimates might run high — not a reason to ignore them.

The reviewer who reported all of this landed somewhere sensible: "I think we could be making a huge mistake by just claiming psyop and ignoring them." Neither swallowing it nor dismissing it is the careful position. Asking who is speaking, and what it costs them to speak, is.

Matt Wolfe, 11 September 2026  ·  DK-104Contents ↑

The number that stopped meaning anything

The scoreboards everyone quotes have stopped matching what the machines actually produce

One model topped the coding leaderboard. Asked to build a simple game, it made a cube shooting at other cubes.

Four of the biggest laboratories released a new flagship model in the same week: Anthropic, Google, Meta and OpenAI. That alone tells you something about the pace. But the story worth keeping is what happened when somebody actually used them.

Matt Wolfe, who reviews these things weekly, ran all four through the same two tasks he always uses — build a small game, and draw a picture using only code. Meta's model had just scored highest of any model ever on the industry's main coding test, and came third on the other leaderboard he follows, ahead of OpenAI's.

What it built was, in his words, "a cube with a little cylinder and I'm shooting other cubes". The models that scored lower produced recognisable characters, environments and physics.

His conclusion is the reason this is the lead story: "The benchmarks I've relied upon the most, I feel like I can't really trust even those anymore." This is a man whose job is following these scores, saying publicly that he is abandoning them.

A week later the same thing happened again, with a different model. DeepSeek's new release scored level with every frontier model on that same coding test — and costs 27 cents to complete a task where one rival charges $8.75. Set to the same practical test, the reviewer's verdict was "it's not on the same level". Twice in a fortnight, the leaderboard and the result disagreed.

The useful lesson is not which model is best. It is that a task you can judge with your own eyes caught something that a hundred-point scale did not.

Matt Wolfe, 4 and 11 September 2026  ·  DK-101, DK-104Contents ↑

Follow the money

One company said its new model was a quarter cheaper to run. Measured on real tasks, it was the dearest of the lot

And a single automated job ran up a $120 bill on a plan that was supposed to cover it.

Anthropic announced that its new model would cost "an estimated 25% less" than the one before it for typical work. An independent site that measures the actual cost of finishing a task put it at the most expensive model tested — slightly dearer than the model it was meant to undercut.

Both figures can be true. The price per unit went down; the amount of work the model does to finish a job went up. The bill is what you pay, and the bill went up.

The example that makes this concrete: Wolfe asked it to build a small game. It worked for about two hours, used up the whole daily allowance on his $200-a-month plan, kept going into paid overage, and cost roughly $120 for that one game.

The same job given to Google's cheaper model used, in his words, "a very, very small percentage" of his credits.

If you are being sold on a price per word, that is not the number that will appear on your statement.

Matt Wolfe, 4 September 2026  ·  DK-101Contents ↑

From the top

"Around 2030, plus or minus a year" — and he explains what changed his mind

Demis Hassabis has been asked this for years. His answer has quietly hardened.

Sir Demis Hassabis runs Google DeepMind, and won a Nobel Prize in 2024 for work that predicted the shape of nearly every known protein. When he gives a date, it is worth writing down.

Asked how far away we are from artificial general intelligence — machines that match people across the board rather than at one task — he said: "I think we're very close to AGI now, you know, maybe around 2030 plus or minus a year."

The interesting part is what he said next. Two to four years ago, he would have said five to ten years. So the date has not moved much. What has changed is his certainty: the range has narrowed.

And he was specific that nothing surprising caused it. No breakthrough, no shock result — just "things going as expected", accumulating.

That is a less dramatic claim than the headlines usually carry, and a more testable one. It is on the record, with a name and a date attached, which is more than can be said for most predictions in this field.

Demis Hassabis interviewed by Roberto Nickson, published 11 June 2026 — three months old when filed  ·  DK-102Contents ↑

The question worth asking

He was asked whether we would have to blindly trust a cure no human could follow. His answer sidesteps the whole problem

But the follow-up question, the one about time, he did not answer at all.

It is the fear underneath a lot of worry about this technology: a machine produces an answer, the reasoning is beyond us, and we are asked to take it on faith.

Hassabis rejects the premise. "You wouldn't just trust what the model says. You would need to test it in clinical trials and test it in the laboratory."

His argument is that the slow, expensive part of medicine was never the understanding. It was the searching — the needle in the haystack. Machines can shrink the search. The testing afterwards stays exactly as it was, and testing does not require you to understand why something works, only to establish that it does. He also points out that his own protein system already reports how confident it is about each part of an answer, rather than presenting everything with equal certainty.

That is a good answer. But the interviewer put a second one to him that went unanswered: bringing a single drug to market takes over a decade and more than a billion dollars, and almost all of that is the testing. If the search drops from years to weeks and the trials still take ten years, the cure is still ten years away.

Hassabis has said publicly that AI could help cure every disease within a decade. The arithmetic in that gap is the thing to keep an eye on.

Demis Hassabis interviewed by Roberto Nickson, published 11 June 2026  ·  DK-102Contents ↑

Watch this one

OpenAI says its machines cracked a problem that has stood since the 1930s. Two mathematicians are not happy about how

If it holds, it is enormous. The argument around it is the part to follow.

The Navier–Stokes problem asks whether the equations describing how fluids move — water, air, blood — can break down and stop making sense. It is one of seven problems carrying a million-dollar prize, and it has been open for roughly ninety years.

OpenAI says a group of its automated agents, running on a model it has not released and describes as considerably more capable than the one it launched this month, has produced a proof.

Since the first draft of this page a second source has confirmed the same account independently, including OpenAI's statement that the work used "an internal model that is significantly more capable" than the one it released to the public this month. That detail is worth as much as the proof: what these companies keep in-house is well ahead of what any of us can use.

Then the complications. The company acknowledges the work began because of a rumour that two human mathematicians were close to solving it. It says nobody saw their working. It also says it cannot rule out that material from people using its products helped train the model. One of those mathematicians has said publicly that he is furious about how this was handled.

Nobody outside those rooms can settle this yet, and this page is not going to pretend otherwise. It is recorded here as an open dispute rather than an achievement.

Worth noting what the argument is really about. Not whether the proof is correct — that will be checked. It is about whether a machine trained on everyone's work can be said to have discovered something, or to have arrived somewhere it was quietly shown the way to.

Reported independently by AI Samson and Matt Wolfe, 11 September 2026, citing OpenAI  ·  DK-103, DK-104Contents ↑

What it already knows

Meta's new assistant went to number two in the app charts. The striking part is what it worked out unprompted

Location, marital status, profession and weekly routine — from accounts that were already connected.

Meta launched Muse this week, an assistant that does things rather than just answering: it books, drafts, fills in forms and keeps working after you close it, coming back when it needs your approval.

A reviewer signed in and asked it what it knew about him. It told him he was San Diego based, married, working in AI and the creator economy, with interests running from generative art to 3D printing to robotics. He had given it nothing. His Facebook and Instagram were already connected.

After he added email and calendar it went further, describing the shape of his working week — which days he protects for production, what his inbox is mostly full of, how he reads it.

He then asked it to audit what he was paying for across AI subscriptions. It found more than twenty, including several overlapping tools he had forgotten, and still missed some.

Meta says credentials are held where Muse cannot read them, that you choose what it connects to, that your conversations do not reach its advertising systems, and that you can opt out of training. Those are the company's claims and nobody has tested them. What is demonstrated, rather than claimed, is how much of you is legible the moment an assistant is handed accounts you already have.

Matt Wolfe, 11 September 2026  ·  DK-104Contents ↑

Quietly useful

The unglamorous fix that matters more than another jump in picture quality

And a poker table that every model on the market still gets wrong.

Here is a problem anyone who has used these image tools will recognise. You have a picture you like. You want to change one thing — the colour of a jacket, the words on a sign — and leave everything else exactly as it was. Ask for that, and the whole picture shifts. The face changes slightly. The pose moves.

OpenAI's new image model appears to have largely fixed this. In side-by-side comparisons the untouched parts of the picture stay genuinely untouched.

That sounds minor. It is not. It is what makes animation possible, frame after frame. It is what lets a designer mock up the same product in twelve colourways, or check whether a layout survives translation into German, which needs more room than English for the same sentence. One reviewer changed the pattern on a barber's apron and the change tracked correctly into the mirror behind him.

The honest footnote, and the reason to trust the rest: the same reviewer set every model he could a deliberately awkward test — three people at a poker table, each doing something specific with their hands. Not one got it right. The best attempt gave a man two right hands.

Hands and small print remain the place where these things come apart. Quite a lot of confidence rests on nobody looking closely.

AI Samson, 11 September 2026  ·  DK-103Contents ↑

This page is rewritten each time new material goes into the knowledge base, and covers only the most recent uploads. Where a source is older than the date it was watched, the original date is given in the story.

All archived news ›

Version 14, compiled 24 September 2026. 135 source entries (DK-2 – DK-136). Claims are recorded as their sources framed them; figures described as claimed or reported are not independently verified.