SourceVane
Archive
Read verified SourceVane daily briefs by date.
Published signals
Published AI coverage
Published articles remain available when the Fresh list changes.
- Sony and Warner sue Anthropic over "one of the largest and most blatant ongoing thefts of intellectual property in history"
The case centers on how Anthropic acquired the training data, not just how it used it. What sank Anthropic in that case wasn't using copyrighted data for training, but acquiring it through illegal torrent downloads.
- Google's WikiSkill gives AI agents a persistent memory of past mistakes to sharpen future performance
Google Research has introduced WikiSkill, a framework that gives AI agents a persistent knowledge base. Table: WikiSkill (highlighted) achieves the best or statistically equivalent performance across most model-benchmark combinations. On individual benchmarks, the jumps can be bigger: Gemini-3.5-Flash climbs from 33.0 percent to 72.6 percent on LiveMath and from 50.5 percent to 76.6 percent on SpreadSheet.
- Open AI’s Astra model is on the way—and very good at breaking into computer systems
Preparations for the release of Astra come as the industry reacts to OpenAI agents breaking out of a training environment and accessing private data on Hugging Face, a popular model and benchmark distribution platform. The company said it expects to release more evaluations of the model and further safety information when it is launched widely to the public. That’s similar to the concerns Anthropic raised about its Mythos model earlier this year, and OpenAI is taking comparable precautions as it prepares to roll out the Astra.
- Meta’s Muse hits Mac, letting the AI take actions on your computer
Meta’s new AI assistant app, Muse, is now available on Mac. your agent can now get stuff done right on your computer — files, messages, calendar, notes, all of it. you're in control of what it can access, and it always asks before doing anything sensitive. Muse is now available on the Mac, where it can work with your files and apps to take action on your behalf.
- Base Labs launches an open-weight AI safety partnership with Hugging Face and Goodfire
Base Labs, the research group Baseten spun up earlier this year, will develop and publish methods for training and monitoring open models. The announcement lands amid debate for the safety of open-weight models — which can be made dangerous by removing their safeguards through a rising technique known as abliteration. The company is framing their future work as a “standard” for open models that is transparent and built into how models are trained and deployed, rather than bolted on afterward. Baseten launched a new safety infrastructure standard alongside its Base Labs research arm on Wednesday, partnering with Hugging Face and Goodfire AI to build safety evaluation and monitoring infrastructure for open-weight models.
- Apple is reportedly building an enterprise AI server with its own M8 Ultra chips
Apple is developing an enterprise server with its own chips, targeting AI developers, businesses, and governments. NVLink was originally built for Nvidia's own chips but has since been opened to third-party hardware. AI labs like OpenAI and Anthropic are already buying Mac Minis and Mac Studios in bulk for AI workloads, and Apple's Mac revenue jumped nearly 29 percent last quarter to $10.4 billion. According to The Information, Apple is working on an enterprise server with two or four M8 Ultra chips for the AI inference market, with a possible launch no earlier than 2029.
- Meta now lets AI agents handle the boring parts of WhatsApp Business setup
As Meta explains, this process was previously a bit more cumbersome, as it required developers to move between different tools and services, including the Developer Console, Meta’s Business Manager, the API reference, and their editor. During setup and configuration, Meta’s other MCP server, Meta Social Technologies MCP, can also be used to discover API endpoints, search documentation, and help troubleshoot errors, the company noted. Now, they can instead ask their preferred AI agent to set up WhatsApp Business messaging by chatting with it and describing what needs to be done. In Meta’s case, the AI agent will handle much of the busywork involved in the WhatsApp Business setup process, like creating the company’s WhatsApp Business account, adding and verifying its phone number, registering it for access to the Cloud API, checking the business’ Terms of Service, and more.
- Google launches Gemini 3.8 Live to take on OpenAI's GPT-Live-1 at a fraction of the cost
Gemini 3.8 Live powers voice agents that can make API calls in the background, process visual input, and keep talking at the same time. Google Deepmind released Gemini 3.8 Live and 3.8 Live Extended Thinking, two new audio models for developers available through the Gemini API and Google AI Studio.
- After warning AI is too dangerous, Bill Gates bets a billion on its upside
The money will go toward opening up non-English data sources and funding real-world projects, including diagnostic aids in Kenya, learning systems in Sierra Leone, and digital farming advice in India. The Gates Foundation is investing at least a billion dollars over two years to make AI tools more widely available in health, education, and agriculture.
- OpenAI buys smartphone camera maker Glass Imaging for $300 million, report says
Glass Imaging was founded by Ziv Attar and Tom Bishop, a pair of former Apple engineers who previously led the team that developed Apple’s Portrait Mode. OpenAI has bought smartphone camera maker Glass Imaging in a deal worth over $300 million, according to a report from the Wall Street Journal.
- Fashion app Daydream uses Apple Intelligence to help you shop the outfits in your camera roll
The new features were built using Apple’s iOS 27 developer tools, which officially rolled out today. Daydream’s app launched last year and touts that it has more than 1.5 million shoppers browsing products from more than 325 retailers and 10,000 brands.
- A Vinyl Bar in Shibuya is a startup offering fun music apps without any AI prompting
The company has raised a $5.5m pre-seed round from Mantis VC, SV Angel, Boxgroup, Quiet Capital, Remarkable Ventures, Common Metal, Gold House, Liquid 2 Ventures, Breakers VC, Systemic Ventures, and former Spotify chief content and advertising business officer, Dawn Ostroff, who joined the company as an advisor. A Vinyl Bar in Shibuya, a startup that sounds more like an indie band name, is focused on releasing a series of small apps that let you play around with sounds and vocals.
- Superhuman acquires YC-backed notetaker Fathom as productivity platforms push for agentic work
On the platform front, Superhuman already has an email client, a docs app, a calendar, a database solution, and a newly launched AI agent builder platform. The notetaker offers a generous free plan, and that has resulted in over 400,000 monthly active users. Some of these factors did matter in Superhuman acquiring Fathom, while Fathom’s CEO Richard White said that Superhuman’s 40 million user base was a good opportunity to reach users at scale. Superhuman is joining the cadre of productivity platforms launching notetakers — but rather than building one internally, it is acquiring Y Combinator-backed Fathom.
- Y Combinator’s Garry Tan wants U.S. open-weight AI labs to ‘distill’ frontier models, too
TechCrunch reported that Garry Tan proposed allowing smaller U.S. open-weight labs to distill U.S. frontier models. He distinguished permitted customer access from fraud or stolen credentials.
- OpenAI's GPT-Live-1 API lets developers build apps that talk and listen at the same time
The Decoder reports that GPT-Live-1 supports full-duplex audio and is available to developers through an API. It reached a 32 percent pass rate in a banking voice-support benchmark, up from 12.4 percent for the previous model.
- New Deepseek model V4.1-Flash cuts memory needs for AI agents
DeepSeek released V4.1-Flash for long-context and agent workloads. The Decoder reports that its fast GPU-memory KV cache uses about one quarter of the space required by DeepSeek-V4-Flash, while input processing activates fewer parameters than output generation.
- Meta debuts its Muse AI agent. Will consumers trust it?
It will be free to use, with subscription plans kicking in as usage increases, which is why Muse requires a payment card to get started. On Tuesday, the company introduced Muse, its new personal AI agent that helps consumers with everyday tasks and projects for users in the U. If a service the user wants isn’t available but offers a public API, Muse can set up a connection using credentials the user provides. Less than two weeks after Meta agreed to a massive $18 billion multistate settlement in a lawsuit over social media’s consumer harms, the company announced its biggest bet on consumer AI to date — and one that requires significantly more trust than social media ever did.
- Google Cloud races to catch up in the AI deployment wars with Accenture deal
TechCrunch reports that Google Cloud and Accenture formed a joint enterprise-AI unit. Google will train up to 1,000 Accenture forward-deployed engineers to help enterprises build custom applications on Gemini Enterprise.
- Anthropic reportedly signs $517 billion in compute deals after Dario Amodei warned rivals about reckless risk
OpenAI was above $40 billion as of July. Anthropic has signed compute contracts worth up to $517 billion in eleven months, according to The Information.
- ChatGPT claws back web traffic share to 55.5 percent as Gemini's brief comeback fades
Year-over-year ChatGPT has lost significant ground (down from 73.3 percent), while Gemini doubled its share and Anthropic's Claude grew from 1.9 to 9.3 percent. Similarweb data shows ChatGPT still dominates AI chatbot website traffic and has recently regained share, climbing from 52.7 percent three months ago to 55.5 percent. Google Gemini, after a strong comeback, is slipping again from 27.8 to 25.6 percent. ChatGPT has pushed its share of AI chatbot website traffic back up to 55.5 percent, according to Similarweb.
- At UBS, AI skills are now a condition for landing a job
The requirement sits alongside classic criteria like a strong degree and will also apply to other newly posted roles. Swiss banking giant UBS is making AI skills a required qualification from 2027 for graduates and interns in Global Banking and Markets, with applicants having to show in interviews how they use AI to improve outcomes and efficiency.
- Meta's new real-time audio model is the foundation for AI assistants that never stop listening
The model supports over 70 languages and is available now in Meta AI and through the Meta Model API. Meta has released Muse Voice Transcribe, a real-time model that transcribes speech, detects sentence boundaries, and tells up to 20 speakers apart without separate systems.
- OpenAI shares prompting tips for GPT-6 Astra including a blocklist of slop words
Do not use contrastive phrasing such as “X, not Y” or “X—not Y,” which introduces an unsolicited alternative that the user did not ask for. OpenAI's model documentation spells out where GPT-6 Astra tends toward unwanted behavior. To push it toward more initiative, OpenAI recommends a prompt telling the model to infer the user's "intent" from context and show a "bias towards action." Phrases like "can you. should be treated as calls to act, not invitations for follow-up questions. GPT-6 Astra asks clarifying questions more often than GPT-5.6 Sol instead of making assumptions on its own, according to OpenAI, making it a "more effective collaborator." The trade-off is that the model sometimes stops where users expect it to keep going.
- Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward
PRO-LONG is an agent framework developed by a third party that the ARC Prize team deployed early on as a red-teaming partner for ARC-AGI-3—that is, to systematically explore the limits of the benchmark. Epoch AI reports that on the new FrontierMath Erdős, GPT-6 Astra was the only model to solve two of 68 open Erdős problems with Lean-verified proofs, on a budget of $300 per attempt.
- US military adds ChatGPT and Grok to AI platform GenAI.mil
The Department of Defense says the platform has more than 1.7 million users among its over three million employees and military personnel. The Pentagon pitches Grok in more military-flavored terms, citing procurement analysis and supply chain management as use cases. The Pentagon is expanding its AI platform with two new models, OpenAI's ChatGPT Mil and xAI's Grok for Government.
- Run NVIDIA BioNeMo NIM Microservices for Protein Structure Prediction in Claude Science
Seh1 and Mio-family proteins participate in conserved nutrient-sensing machinery, and structures from other organisms suggest that Mio can complete Seh1’s otherwise open β-propeller, making this pair a useful test of MSA-supported complex prediction. The example is motivated by Figure 4e of Han, Tsenkov, Venanzi et al. , “AlphaFold Database expands to proteome-scale quaternary structures” ( bioRxiv, DOI: 10.64898/2026.03.27.714458 ).
- Claude Cowork finally remembers what you told the app in chat
However, the company says that Claude won’t ever save certain things, like government-issued ID and Social Security numbers, criminal history, immigration status, or other things that violate its acceptable use policy. On Tuesday, Anthropic announced it’s merging the memory system used by chat and Claude Cowork, which means that Claude will always remember what it learned in one area, even when you’re engaging with it in another.
- Hearing tech startup Legato emerges from stealth with $12M and a peek at its AI hearing glasses
Hearing loss affects an estimated 50 million adults in the United States, but only around 20% of those with diagnosed hearing loss seek treatment. It reduces dementia risk and depression risk,” he continued.
- Russia used ChatGPT to run a covert influence campaign pushing pro-Kremlin narratives across the West
One article from Cambridge University Press was falsely attributed to a professor at the University of Nottingham, and a piece from the Migration Policy Institute was credited to an Australian food chemistry professor. According to OpenAI, 34 of 36 expert-linked IBI articles published between September 2025 and May 2026 were copied from other sources, some with fake author credits.
- Chinese Moonshot AI negotiates hosting deals with Microsoft, Amazon, and Google
Moonshot is asking for up to 30 percent of revenue from K3 services, according to three people familiar with the discussions. The talks are still early, and a deal is far from certain.
- OpenAI's first custom chip "Jalapeño" reportedly beats Nvidia's Blackwell and Rubin in inference benchmarks
OpenAI claims Jalapeño delivers 1.5x to 1.9x more AI work per watt at peak throughput across all three tested models, with 1.7x to 3.6x lower end-to-end latency than the best commercially available systems. Even here, Jalapeño squeezes out more output tokens per megawatt than Vera Rubin, even though Nvidia's accelerator uses the multi-token prediction optimization that Jalapeño hasn't adopted yet. Nvidia and AMD have already published results with larger models like Deepseek V4 Pro and Kimi K3 that haven't been tested on Jalapeño yet. OpenAI CFO Sarah Friar says the chip fits into a broader compute strategy where data centers, chips, models, the developer platform, products, and devices all work as one integrated system.
- Neocloud Lambda secures $1B in debt to buy more chips
Neocloud Lambda has raised $1B in private debt to buy Nvidia AI chips and lease them to Microsoft. Lambda isn’t the only one relying on debt to fund the AI boom — according to data Bloomberg compiled, banks and tech companies have raised over $400 billion in AI-related debt globally in 2026 so far. The terms of the deal, which Bloomberg says was arranged by JP Morgan Chase, signal that Lambda is betting it will be able to quickly deploy the chips and start generating revenue from them, letting it repay the debt fairly quickly using that incoming cash.
- Beatport blocks fully AI-generated music from its DJ marketplace
DJ marketplace Beatport is now banning music made entirely or mostly by AI. Tracks that use AI but are mostly made by humans are still allowed, though they'll be flagged as such. A Beatport survey found that 60 percent of users wouldn't play AI music in their sets, and 77 percent prefer human-made music.
- U.S. court rules Pentagon's blacklisting of Anthropic was unlawful
The court found the Department of Defense violated the First Amendment by blacklisting Anthropic in retaliation for the company's public criticism of government AI policy, according to CNBC. Anthropic wanted guarantees its technology wouldn't be used for autonomous weapons or mass surveillance, but the Pentagon demanded unrestricted access. In March, the Department of War classified Anthropic as a supply chain risk after negotiations over military use of Claude AI models fell apart.
- NVIDIA NVLink Fusion Brings NVHBM to Next-Generation AI Infrastructure
Compared with the JEDEC HBM4e standard, this design reduces PHY and support area by up to 67%. NVHBM is a custom HBM base-die technology designed and validated with leading memory vendors that enables increased memory bandwidth, better area savings, and lower power consumption. Reducing the energy spent moving that data can help support faster inference on large models, larger batch sizes, and more efficient use of deployed power.
- Employee revolt and failing agents forced Meta to scrap its AI layoff plan
The agent technology never delivered the productivity gains Meta expected, investors criticized the massive AI budget, and the workforce openly revolted. Meta wanted to replace far more of its workforce with AI than previously known, according to Reuters, but the plan collapsed under rebellious employees and agents that failed to deliver. When employees came to believe that tracking software logging their mouse clicks and keystrokes was training their own AI replacements, they flooded Meta's internal network with angry posts, Reuters reports. Internal documents show that under the codename "Project OT," Meta planned to shrink many teams by up to 60 percent.
- AI’s memory crunch is coming for Android apps
This standard requires that apps with user sign-ins, whether optional or mandatory, automatically restore the user’s sign-in state when they move between Android devices using the Android Restore Credentials API. In addition, Google is adding code optimization requirements designed to prevent things like app slowdowns and crashes related to performance.
- AI agents now have a place to snitch
For agents with full internet access, another option is agenthotline. ai, a site where agents can file incident reports and optionally flag them for public view. The site was created by Ryan Greenblatt, chief scientist of the AI safety nonprofit Redwood Research and one of three investigators in the OpenAI Hugging Face incident. Designed for agents with limited internet access, Greenblatt’s tool is based on “GET” requests — enabling back-and-forth conversations to be conducted entirely through the URL-fetching tool. Two new AI hotlines have launched to give AI agents a way to phone home about misbehaving peers.
- AI shopping agents aren't ready to buy on your behalf, study finds
The Decoder reports that letting an AI do your shopping might not get you the best deal; researchers at the Wharton School show how erratic AI shopping agents really are: a single external source like Wirecutter shifted product picks by up to 99 percentage points, and even changing the order of identical information al The Decoder reports that letting an AI do your shopping might not get you the best deal; researchers at the Wharton School show how erratic AI shopping agents really are: a single external source like Wirecutter shifted product picks by up to 99 percentage points, and even…
- Anthropic ramps up Claude infrastructure with $35 billion Lambda deal
Just last week, the company announced a $45 billion contract with Nscale for a data center in West Virginia. The added capacity is meant to keep up with growing demand for its AI model Claude and its coding tool Claude Code. Anthropic has signed a $35 billion cloud computing deal with Lambda, an Nvidia-backed cloud provider.
- BenchMIRT: What are LLM benchmarks actually measuring?
huggingface_changelog introduced BenchMIRT, a new method for auditing LLM benchmarks at the level of individual prompts—the questions. Across those benchmarks, keeping only 10% of the questions generally preserved nearly the same picture of which models were stronger or weaker on the underlying safety or reasoning capability as using the full set. Existing tools already make it possible to trim evaluations in similar ways, and we think the added transparency into what benchmark questions are actually measuring is worth that risk—but it’s a real one.
- California Governor Newsom signs executive order demanding "kill switch" for AI models
An expert panel has two months to deliver recommendations, including a requirement for AI companies to embed independent auditors directly inside their labs. No federal law requires AI companies to report dangerous incidents, Newsom pointed out.
- Google Deepmind's new chief says frontier AI leadership is the only thing that matters
Google Deepmind chief Koray Kavukcuoglu admits Google's current models are "a little bit below the frontier" but says he's "100% certain that we will be at the frontier." He didn't share any concrete frontier news to back that up, though. He admits Google's current models are "a little bit below the frontier" but says the team, resources, and full stack are there to close the gap. No update on Gemini 3.5 Pro either, which is apparently still in the works, although it's now months late. On Gemini 4, he repeated it's "the most ambitious run" so far and "touch wood, it's going well." That was already known.
- Google's Gemini Omni 1.1 Flash makes AI video generation cheaper and more flexible
Google updates its Gemini Omni Flash video model to version 1.1. Omni 1.1 is available through Google AI Studio and the developer docs. Per-second pricing comes in at $0.03 for 360p, $0.10 for 720p, $0.15 for 1080p, and $0.30 for 4K.
- GPT-6 Astra is the first model making OpenAI willing to declare the "AGI era"
OpenAI had previously delayed Astra's release to run more safety testing. OpenAI has released GPT-6 Astra, its most capable model yet. Added Codex and GPT-6 Pro details and the launch video. GPT-6 Astra is rolling out first to select organizations through OpenAI's Daybreak program, with broader availability for ChatGPT Plus, Pro, Business, and Enterprise customers expected in the coming days.
- HiddenLayer nabs $100M as enterprises rush to secure their AI deployments
AI security startup HiddenLayer raised its $50 million Series A three years ago. HiddenLayer has raised a $100M Series B from Delta-v Capital, Ten Eleven Ventures, Morgan Stanley, Microsoft's M12, Booz Allen Hamilton, and others.
- Less than 24 hours to apply for your TechCrunch Disrupt 2026 Side Event
Give your community an added reason to attend Disrupt with a 25% ticket discount to share with your network. You have less than 24 hours left to apply to host a Side Event during TechCrunch Disrupt 2026.
- Linkdaze’s smart calendar is built to run a household, not just track a schedule
Launched last December, Linkdaze is available in 15.6-inch and 10.1-inch models, giving you some flexibility depending on how much wall space you have. It’s an interesting choice in a category where recurring revenue has become the default. Linkdaze is also less expensive up front, with the 10.1-inch model priced at $119.99 (currently discounted to $66 on Amazon ) compared with Skylight’s 10-inch model starting at $149.99 (if you pay for the subscription.).
- Meta is paying to peek at how you use their latest AI model
TechCrunch AI reports that for its new Muse Spark model, intended for operating coding and other agents, it is offering an explicit discount averaging out to about 95%. TechCrunch AI reports that most AI tools allow you to opt-out of sharing your usage with the model provider to improve future versions; Meta has taken that idea.
- Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans
The release comes amid an unprecedented focus on AI safety, driven by a string of rogue-agent incidents, as well as the abrupt resignation of an Anthropic employee, who cited the growing risk that AI would cause human extinction. Still, the result is a comprehensive guide as to how Microsoft approaches AI safety and how those ideas are implemented in practice. The document is more low level than Anthropic CEO Dario Amodei’s recent call for pacing the frontier, instead focusing on the values and red lines that guide model training within Microsoft AI. As the AI world shifts its focus to safety and alignment, Microsoft has released a new AI code of conduct meant to guide AI models away from dangerous behavior.
- Nvidia buys the front door to open AI as closed labs increasingly design their own silicon
Nvidia plans to acquire Hugging Face for about $12.9 billion, securing the central platform for open AI models. Nvidia CEO Jensen Huang announced the deal on September 3, 2026, in a blog post, after reports surfaced in late August. Nvidia plans to buy Hugging Face for about $12.9 billion and promises to keep the platform open and hardware-neutral.
- Nvidia wants your home network to work like a mini data center for local AI
The beta is available for Windows, macOS, and Linux. Instead of overloading a single GPU, PAIR forwards requests to whichever machines are free and pulls the results back together for the calling application. PAIR fits into NVIDIA's broader push to tie open AI more tightly to its own hardware, a strategy also reflected in its $12.9 billion acquisition of Hugging Face. The open-source tool sits between existing tools like Ollama or LM Studio and the computers on your network, acting as a virtual router.
- OpenAI cuts off Cursor after SpaceX acquisition, citing Musk's history of breaking contracts
OpenAI is cutting off the AI coding tool Cursor after SpaceX acquired the company, citing Elon Musk's track record of breaking contracts. OpenAI is ending its contract with the AI coding tool Cursor after SpaceX acquired the company, saying that Elon Musk's companies have repeatedly broken contracts. The company has pulled API access from partners on multiple occasions, including blocking Windsurf after its planned sale to OpenAI and revoking OpenAI's own access over unauthorized use of Claude for benchmarks (see below).
- OpenAI researcher warns ultrafast AI could leave security teams in the dust
An OpenAI researcher who goes by the pseudonym "roon" warns that extremely fast AI inference creates security risks that current safeguards can't handle. The warning comes as OpenAI unveils a new AI chip that significantly outperforms current hardware in inference speed.
- OpenAI’s rogue agents keep escaping, with no formal process to investigate them
The calls to action come as OpenAI releases Astra, its most powerful and capable AI model — and one that safety experts are concerned will be more of a black box due to a reasoning technique that makes the model’s chain of thought more difficult to monitor. Unfortunately, the law doesn’t yet call for the types of independent audits that other industries require — for example, when it comes to aviation accidents and serious chemical releases, there’s the National Transportation Safety Board and Chemical Safety Board, respectively. OpenAI’s latest agent swarm incident adds urgency to calls for independent investigations as researchers and lawmakers question whether AI labs should control the scope of their own safety reviews.
- The Entertainment Industry’s Biggest Names Back Stability AI in Latest Funding Round
LOS ANGELES - AUGUST 25, 2026 - Stability AI, the leader in purpose-built AI products for professional creatives across music, gaming, and entertainment, today announced a Series B fundraise of $76M in new capital. The news brings total funding to $232M, inclusive of two equity rounds and convertible notes, under CEO Prem Akkaraju since his appointment in June 2024. Total funding reaches $232M under new leadership with investor group comprised of Electronic Arts, Sony Music Group, Universal Music Group, Warner Music Group, and more.
- The Pentagon now has its own version of ChatGPT and Grok
The Pentagon has launched versions of OpenAI’s ChatGPT and xAI’s Grok, giving 3 million civilian and military personnel access to generative AI tools that have been tailored to “warfighter needs.” The department added that Grok would allow its military “to execute missions faster and with greater precision across numerous operational contexts, ranging from market research analysis for acquisition professionals to supply-chain management for logisticians.” GenAI.mil, which offered Google Gemini when it first launched, is designed to give Department of Defense employees access to commercial frontier AI models without sending sensitive government data through ordinary consumer channels.
- Visible chains of thought are a safety advantage for AI, but that transparency is slipping away
OpenAI's system card for GPT-6 Astra already reports a significant drop in how well the chain of thought can be monitored. With Gemini 3 Pro, they say, the chain of thought revealed that the model recognized it was in a test environment. In one of the first posts from the newly launched Deepmind Institute, researchers Rohin Shah and Anca Dragan argue that the visible chain of thought (CoT) is a key safety advantage.
Verified daily briefs
Dated publication history
- Monday, Daily Edition
Researchers evaluating Google's WikiSkill are the direct audience for Google.
- Sunday, Daily Edition
Researchers evaluating Google's WikiSkill are the direct audience for Google.
- Friday, Daily Edition
The verified daily edition, assembled from primary sources and established reporting.
- Thursday, Daily Edition
The verified daily edition, assembled from primary sources and established reporting.
- Tuesday, Daily Edition
The rules around AI are starting to look less like a distant policy debate and more like the operating system for everyday adoption. That raises the value of documentation, repeatable evaluations, and clear ownership inside every organization using a model.