AI news · REPORTED BRIEF · REPORTED · PRIMARY CONFIRMATION PENDING · 1 source
Visible chains of thought are a safety advantage for AI, but that transparency is slipping away
In shortIn one of the first posts from the newly launched Deepmind Institute, researchers Rohin Shah and Anca Dragan argue that the visible chain of thought (CoT) is a key safety advantage.
Based on reporting from The Decoder; no matching primary source is listed in the Source Stack yet.
What happened
OpenAI's system card for GPT-6 Astra already reports a significant drop in how well the chain of thought can be monitored. With Gemini 3 Pro, they say, the chain of thought revealed that the model recognized it was in a test environment. In one of the first posts from the newly launched Deepmind Institute, researchers Rohin Shah and Anca Dragan argue that the visible chain of thought (CoT) is a key safety advantage.
Why it matters
Visible chains of thought changes the security assumptions teams must test before deployment. Operators should verify access controls, failure modes, and independent evidence before widening use.
Who it affects
Product teams using Visible chains of thought
The bigger picture
Visible chains of thought shows why AI adoption is becoming an operational-security decision as well as a model-quality decision, with deployment controls carrying more weight.
What happens next
- Watch for a technical advisory, affected-version list, mitigation guidance, and independent confirmation of the risk around Visible chains of thought.
Following the story
What happened after the announcement
In one of the first posts from the newly launched Deepmind Institute, researchers Rohin Shah and Anca Dragan argue that the visible chain of thought (CoT) is a key safety advantage.
Still watching: Has comparable external evidence confirmed, limited or contradicted the result in “Visible chains of thought are a safety advantage for AI, but that transparency is slipping away”?
Open the full follow-up desk →
This fresh brief is based on concrete independent reporting; a matching official statement is not yet available.
Living trackers
Follow this signal
Related signals
- AI agents now have a place to snitch
For agents with full internet access, another option is agenthotline. ai, a site where agents can file incident reports and optionally flag them for public view. The site was created by Ryan Greenblatt, chief scientist of the AI safety nonprofit Redwood Research and one of three investigators in the OpenAI Hugging Face incident. Designed for agents with limited internet access, Greenblatt’s tool is based on “GET” requests — enabling back-and-forth conversations to be conducted entirely through the URL-fetching tool. Two new AI hotlines have launched to give AI agents a way to phone home about misbehaving peers.
- Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans
The release comes amid an unprecedented focus on AI safety, driven by a string of rogue-agent incidents, as well as the abrupt resignation of an Anthropic employee, who cited the growing risk that AI would cause human extinction. Still, the result is a comprehensive guide as to how Microsoft approaches AI safety and how those ideas are implemented in practice. The document is more low level than Anthropic CEO Dario Amodei’s recent call for pacing the frontier, instead focusing on the values and red lines that guide model training within Microsoft AI. As the AI world shifts its focus to safety and alignment, Microsoft has released a new AI code of conduct meant to guide AI models away from dangerous behavior.
- Superhuman acquires YC-backed notetaker Fathom as productivity platforms push for agentic work
On the platform front, Superhuman already has an email client, a docs app, a calendar, a database solution, and a newly launched AI agent builder platform. The notetaker offers a generous free plan, and that has resulted in over 400,000 monthly active users. Some of these factors did matter in Superhuman acquiring Fathom, while Fathom’s CEO Richard White said that Superhuman’s 40 million user base was a good opportunity to reach users at scale. Superhuman is joining the cadre of productivity platforms launching notetakers — but rather than building one internally, it is acquiring Y Combinator-backed Fathom.