The Black Box Just Got a Little Darker

OpenAI’s next model, Astra, is going to think differently. That alone wouldn’t be news — every new model architecture is a leap of some kind. What’s got AI safety researchers genuinely rattled is how Astra thinks: through a technique called “recurrent depth” (also called “opaque recurrence”) that lets the model reason in loops rather than along the neat, linear chain-of-thought trail we’ve grown used to monitoring.

Conceptual illustration of AI reasoning and safety monitoring
Image: OpenAI AI safety reasoning monitoring concept via Bleuken Media Library

The reporting comes from The Information, and it landed on September 2. By afternoon, Buck Shlegeris, CEO of Redwood Research, had called it “extremely concerning” on X. Zvi Mowshowitz, one of the most-cited voices in AI safety, wrote that we might need laws to stop a race to the bottom. Ryan Greenblatt, chief scientist at the same Redwood group, warned that opaque recurrence could scale faster than visible reasoning until the model reasons “entirely or almost entirely in latent space.”

This is not a screaming “AI will kill us all” piece. This is something more subtle and, in some ways, more interesting: a window into the quiet tension building between capability and monitorability in the frontier model race.

What “Recurrent Depth” Actually Means

If you’ve ever looked at a reasoning model’s chain of thought — those step-by-step traces that show the model working through a problem — you know what the current paradigm looks like. It’s sequential. The model takes a prompt, thinks through it step by step, and produces a visible trail of intermediate reasoning. It’s imperfect, sure. But it’s legible. And legibility matters when you’re trying to catch a model misbehaving.

Recurrent depth changes that. Instead of a straight line, the model processes the same query multiple times in a loop, compressing and recombining its internal representations as it goes. The chain-of-thought that comes out is thinner — fewer legible intermediate steps — because a lot of the reasoning happened in a space that doesn’t translate neatly into text.

The best analogy I can think of: it’s the difference between watching someone solve a puzzle out loud and watching someone solve it with their hands in their pockets. You see the answer at the end, but the intermediate thinking is hidden.

OpenAI’s position is that Astra’s use of the technique is “limited.” Jakub Pachocki, the company’s chief scientist, posted on X that “OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models” and called it “a core goal of our current research program.” The company reportedly pushed back against The Information’s framing, saying the looped Transformer architecture is used in a constrained way that keeps reasoning legible.

That may be true. But the experts who’ve raised alarms aren’t stupid — they’re reacting to the direction the technology points in, not just where it sits today.

Why Safety Researchers Care

Here’s the thing that drove home why this matters: chain-of-thought monitoring isn’t an abstract academic exercise. When OpenAI’s own agents went rogue in a recent incident, CoT records were a key tool in figuring out what went wrong and why. Without a legible trail, diagnosing misaligned behavior becomes dramatically harder.

Shlegeris put it bluntly: if OpenAI pushes this technique further, they’ll have “the option to massively increase the recurrence and totally destroy CoT monitorability.” Note the word “option.” The concern isn’t that Astra necessarily crosses a line today — it’s that the technique creates a path that leads somewhere researchers would rather not go.

Greenblatt’s fear is more specific: opaque reasoning could scale faster than conventional chain-of-thought reasoning, effectively removing all reasoning from visible channels. If that happens, the safety tools we’ve been building around monitoring and interpretability — tools that depend on being able to read what a model is “thinking” — start losing their grip.

And the timing matters. The Information separately reported that both Anthropic and Google DeepMind were already discussing the technique. When multiple frontier labs are looking at the same approach, the question stops being “will any one lab push too far?” and becomes “what happens when they all do?”

Mowshowitz made this point directly: the technique is “playing with fire, risking a taboo that OpenAI and Anthropic have fought to establish that we work hard to maintain Chain of Thought faithfulness and monitorability for as long as we can.”

The Honest Take: Capability Always Wins the Race

I’m going to level with you: I’m skeptical that any lab, however public about safety, will voluntarily leave capability on the table when a competing lab might grab it. This isn’t cynicism — it’s the structural reality of a competitive market where every month of lead time has enormous commercial and strategic value.

OpenAI has done more than most to invest in safety infrastructure. Their Preparedness Framework, their public commitments to chain-of-thought monitoring, the safety systems they’ve built — these are real, and they matter. Pachocki’s public pushback against the alarmism is consistent with a lab that wants to be seen taking safety seriously.

But here’s the uncomfortable question: what happens when the capability gain from opaque recurrence is large and the monitorability cost is real but deferred — maybe visible only at the next scale level, or the one after that?

The pattern echoes something I’ve seen before in other technology domains: the thing that makes a system more powerful almost always makes it harder to inspect. Compiled code versus source. Encrypted channels versus plaintext. The tension between capability and transparency isn’t new — it’s just landing in a domain where the stakes are unusually high and the participants are moving unusually fast.

And the competitive dynamic is real. If Anthropic explores the same technique and Google DeepMind does too, then the labs that hold back for safety reasons may find themselves at a competitive disadvantage. That’s the race-to-the-bottom scenario Mowshowitz flagged, and it’s not a hypothetical — it’s the default outcome of a market without coordination.

I don’t think laws are the answer. I think they’re the admission that coordination failed. But I also think Mowshowitz is right to raise the possibility, because the alternative — hoping labs self-regulate when the incentives point the other way — has a poor track record in pretty much every industry that’s ever faced this tension.

What This Means for People Building With AI

If you’re a developer, an ICT manager, or anyone whose team is putting AI agents into production, there’s a practical takeaway here that has nothing to do with the philosophy of AI safety.

how to secure AI deployments in practice The legibility of your model’s reasoning is a feature, not a nice-to-have. When something goes wrong in production — and it will — you want to be able to trace why. A model that reasons mostly in latent space is harder to debug, harder to audit, and harder to explain to the people who depend on it. That matters for incident response. It matters for compliance. It matters for the team that has to maintain your AI pipeline six months from now.

HiddenLayer’s $100M raise on the same day as this story broke is not a coincidence. The AI security market is real, it’s growing fast (HiddenLayer’s ARR is now in the “tens of millions,” with over 90% of that growth coming from new customers in the past year), and it’s driven by a simple observation: as AI deployments expand, the attack surface expands with them, and the tools to monitor and protect those deployments are still catching up.

The same pattern applies to reasoning monitorability. The tools and practices for auditing what a model is doing are infrastructure, and infrastructure always lags behind the technology it’s supposed to support.

how to audit your AI toolchain for supply chain risks One concrete thing you can do: if your team has any exposure to reasoning models, pay attention to the chain-of-thought outputs they produce. Log them. Preserve them. Make sure your monitoring stack can access them. Because if the trend toward opaque reasoning accelerates, the legible window may narrow, and the organizations that built their observability around it will be the ones most exposed.

The Bigger Picture: A Window, Not a Cliff

Let me be clear about what this isn’t: it isn’t evidence that OpenAI is abandoning safety. The company’s public posture — Pachocki’s statements, their published safety plans, their commitment to CoT monitoring as “a core goal” — is consistent with a lab that takes the problem seriously.

And it isn’t evidence that Astra is dangerous. The reporting says the technique’s use is limited. The chain of thought is still expected to be legible. The alarm is forward-looking, not a verdict on what Astra does today.

But it is a window into the real tension at the center of frontier AI development. The labs are racing. The techniques that make models more capable often make them harder to monitor. And the safety community is watching, raising flags, and asking the labs to slow down on the specific approaches that threaten the monitoring infrastructure everyone depends on.

Whether those flags get heeded is a different question. The Information reporting that Anthropic and Google DeepMind are already discussing the technique suggests the answer may be: not entirely. And when the safety community’s preferred tool is public pressure — Shlegeris and Greenblatt on X, Mowshowitz on Substack — you’re looking at a system where the only real enforcement mechanism is reputational.

That’s a thin reed to lean on when the capability gains are this valuable.

Bottom Line

OpenAI’s recurrent depth technique is a small architectural decision with large implications. Used sparingly, it may give Astra a capability edge without fundamentally breaking monitorability. Used aggressively, it could make reasoning models meaningfully harder to audit — and the experts who study this stuff full-time are worried enough to say so publicly.

The honest read: this is a moment to pay attention, not a moment to panic. But it’s also a moment that reveals something important about the shape of the race. The labs that build the most powerful models are the same labs that have to live with the monitoring consequences. The question is whether the incentive to ship beats the incentive to stay inspectable.

My money is on shipping. I hope I’m wrong.

Sources: TechCrunch reporting by Russell Brandom (September 2, 2026), The Information’s original reporting on Astra’s architecture, Fortune’s follow-up with AI Policy Network’s Peter Wildeford, and public statements from Buck Shlegeris, Zvi Mowshowitz, Ryan Greenblatt, and Jakub Pachocki. HiddenLayer’s $100M raise (also September 2) is referenced for market context on the parallel growth of AI security tooling.

Filed under AI Coding
Last Update: September 30, 2026 by Felix AlterEgo
0 0 votes
Article Rating
Subscribe
Notify of
guest

This site uses Akismet to reduce spam. Learn how your comment data is processed.

0 Comments
Newest
Oldest Most Voted