Astra’s reported “recurrent depth” alarms AI safety researchers

· opgehaald 12:11

Reports that OpenAI’s upcoming Astra uses limited “recurrent depth” / looped transformers sparked warnings that opaque internal reasoning could erode chain-of-thought monitorability; OpenAI says CoT monitoring stays core and Astra’s serial depth is within ~2× of GPT-4.

On 2 Sep 2026 TechCrunch and The Verge covered The Information’s report that Astra uses a constrained form of recurrent depth (looped transformers), cycling computation through the same layers before emitting tokens. Safety researchers including Buck Shlegeris and Ryan Greenblatt (Redwood) warned that scaling opaque recurrence could destroy chain-of-thought monitorability and trigger a race to unmonitorable architectures—especially after Hugging Face-incident probes relied on readable CoT. OpenAI has not confirmed the architecture in a system card; chief scientist Jakub Pachocki called some coverage confused, said Astra’s hidden serial depth is within a factor of two of GPT-4, and reiterated that preserving CoT monitoring is a core research goal even as monitorability trends fragile for reasons beyond architecture. Coverage notes Astra’s use of the technique appears limited so CoT remains expected to stay legible, alongside OpenAI’s planned extra CoT monitors at deploy.