Why less visibility into how OpenAI’s new GPT-6 Astra ‘thinks’ is sparking safety concerns

In this news:

OpenAI’s new model, GPT-6 Astra, has less direct visibility into how a model thinks, a development that has sparked concerns coming just weeks after the Hugging Face hacking incident that required a Chinese open model to investigate, according to analysts.
When announcing Astra on Thursday, OpenAI said it was “the world’s most intelligent and aligned model,” with a “significant jump in cyber capabilities”.
OpenAI president Greg Brockman said at the end of a press call announcing Astra’s arrival that it likely represents AGI, or artificial general intelligence – AI that matches or outperforms human intelligence.
However, OpenAI also said the model’s written reasoning was “harder to monitor” compared with GPT-5.6 Sol, the previous generation released in July.
“We have found that GPT-6 Astra is more capable of controlling its own CoT (chain of thought) than GPT-5.6 Sol, and less likely to include incriminating information in its CoT,” OpenAI said, referring to the intermediate reasoning steps an AI generates while solving a task.
The shift in visibility stems from a technique known as recurrent depth, or looped transformers, which reuses parts of a neural network. As a result, it processes complex logic inside hidden mathematical loops rather than in step-by-step readable text, The Information reported on Tuesday ahead of the launch.
OpenAI chief scientist Jakub Pachocki on X called the report “confusing”, without naming the media.
OpenAI did not disclose architectural details of the new system. A company representative said in a statement to the South China Morning Post on Friday that the decline in monitorability was “a real challenge that [it] takes seriously”.
The representative added that the new model does not use “neuralese”, a controversial technique where AI reasons internally using mathematical representations rather than words.
As frontier AI labs race to squeeze greater intelligence out of increasingly expensive computing infrastructure, the decrease in visibility has drawn global attention, especially after OpenAI’s agents breached the Hugging Face platform in July, highlighting the importance of being able to inspect what models are thinking.
Is recurrent depth a bad thing?
Unlike existing AI architectures that pass text through a fixed series of computational layers to determine the next word, recurrent depth loops the data through these layers repeatedly. This approach can have a positive impact on performance and costs.
Hamish Low, research associate at the Institute for AI Policy and Strategy, said it may “reduce the amount of memory required by having a smaller model repeat some of its internal processing” amid growing memory chip costs.
The technique effectively buys “greater reasoning depth with fewer weights and potentially lower memory requirements,” said Wei Sun, principal analyst for artificial intelligence at Counterpoint Research.
However, this approach reduces “the amount of reasoning humans can directly inspect,” which “creates an emerging intelligence-versus-observability trade-off”, Sun said.
What are the security concerns?
Reading the chain-of-thought reasoning can be critical for understanding the high-profile Hugging Face incident, analysts said.
“Having much of the reasoning occur in a closed box rather than by ‘thinking out loud’ through normal human language will make it much harder to understand how these AI systems are operating,” said Kyle Chan, a fellow at the Brookings Institution.
Researchers investigating the Hugging Face incident were able to reconstruct the agents’ behaviour partly by examining their written chains of thought.
When Hugging Face detected the intrusion, its security team initially tried using commercial frontier AI models to help analyse the incoming attack, but those requests were blocked by the providers’ automated safety guardrails. Consequently, it opted to run ’s open-weight model, GLM-5.2, locally to contain the attack.
“The fact that the Hugging Face incident happened, and then right afterwards, they are releasing a model that’s harder to monitor – that’s why so many people are freaked out,” said Ben Hayum, research assistant for the Technology and National Security Program at the Centre for a New American Security (CNAS).
“If we keep getting substantial gains in the internal processing we can do without external visibility into it, AI model alignment becomes a lot harder from the get-go in terms of creating a safe model,” he added.
Are Chinese AI firms adopting it?
, also known as Zhipu, has started looking into the looped transformer architectures. In an earnings call earlier this week, company founder Tang Jie discussed using a “Loop Transformer” to increase the effective depth of its models and described the direction of its next-generation GLM-6 model as “self-evolution”.
Analysts said this technique was attractive to Chinese developers, given their limited access to advanced compute resources.
Low at the AI policy institute said Chinese AI firms might have stronger incentives to adopt it amid ongoing price wars, but it depended on how large the benefits were. “The hardware they have, their research priorities and training data, and how they balance lower costs against speed and model performance also affect their decisions,” he added.
Counterpoint’s Sun said the technique could be attractive, but it “should be viewed as compute reallocation, not free compute.”
“The same economics apply in the US too,” she added.
It remained to be seen whether other companies follow suit, said Jinge Wang, an AI safety research engineer at Concordia AI.
If AI companies’ primary goal was to enhance capabilities, model interpretability and monitorability would “inevitably deteriorate”, Wang said.
Additional reporting by Vincent Chow

Top Trending Cryptocurrencies on The Market

Current Price

$2.300
7 Days

Market Cap

$12.8M 6.25%

24h Volume

$56.5K

Supplies

5.6M / 21.0M

Current Price

$0.06114
7 Days

Market Cap

$8.7M 1.68%

24h Volume

$12.3K

Supplies

225.2M /

Current Price

$0.01118
7 Days

Market Cap

$9.9M 6.13%

24h Volume

$181.6K

Supplies

1.0B /

Current Price

$0.01140
7 Days

Market Cap

$10.8M 3.95%

24h Volume

$6.6M

Supplies

948.2M / 1.0B

Current Price

$0.06272
7 Days

Market Cap

$9.3M 3.28%

24h Volume

$1.8M

Supplies

150.0M / 150.0M

Current Price

$0.03037
7 Days

Market Cap

$12.1M 1.31%

24h Volume

$316.8

Supplies

800.0M / 1.0B

Current Price

$0.4688
7 Days

Market Cap

$18.2M 0.31%

24h Volume

$284.7K

Supplies

38.8M /

Current Price

$0.1714
7 Days

Market Cap

$15.9M 3.28%

24h Volume

$4.4M

Supplies

92.8M / 96.0M

Current Price

$0.3548
7 Days

Market Cap

$16.0M 0%

24h Volume

$5.1K

Supplies

50.0M /

Current Price

$0.05508
7 Days

Market Cap

$10.4M 2.15%

24h Volume

$6.9M

Supplies

1.0B / 1.0B

Current Price

$0.01603
7 Days

Market Cap

$16.0M -0.24%

24h Volume

$116.0K

Supplies

1.0B / 1.0B

Current Price

$0.1488
7 Days

Market Cap

$14.1M 25.85%

24h Volume

$1.0M

Supplies

100.0M / 100.0M

Join Our 💌 Newsletter!

Get updates, insights, and reports on the latest industry trends.

You are subscribing to all our networks!