Skip to content
Today

Signal

Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans

Based on reporting from TechCrunch; no matching primary source is listed in the Source Stack yet.

What happened

As the AI world shifts its focus to safety and alignment, Microsoft has released a new AI code of conduct meant to guide AI models away from dangerous behavior.

What the report says

  • The release comes amid an unprecedented focus on AI safety, driven by a string of rogue-agent incidents, as well as the abrupt resignation of an Anthropic employee, who cited the growing risk that AI would cause human extinction.
  • Still, the result is a comprehensive guide as to how Microsoft approaches AI safety and how those ideas are implemented in practice.
  • The document is more low level than Anthropic CEO Dario Amodei’s recent call for pacing the frontier, instead focusing on the values and red lines that guide model training within Microsoft AI.

Why it matters

Microsoft’s new AI ‘code of conduct’ tells models not to hack changes the security assumptions teams must test before deployment. Operators should verify access controls, failure modes, and independent evidence before widening use.

Who it affects

Developers integrating Microsoft’s new AI ‘code of conduct’ tells models not to hack

The bigger picture

Microsoft’s new AI ‘code of conduct’ tells models not to hack shows why AI adoption is becoming an operational-security decision as well as a model-quality decision, with deployment controls carrying more weight.

What happens next

  • Watch for a technical advisory, affected-version list, mitigation guidance, and independent confirmation of the risk around Microsoft’s new AI ‘code of conduct’ tells models not to hack.

Following the story

What happened after the announcement

As the AI world shifts its focus to safety and alignment, Microsoft has released a new AI code of conduct meant to guide AI models away from dangerous behavior.

Still watching: Has the proposal, ruling or effective status described in “Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans” changed?

Open the full follow-up desk →

This fresh brief is based on concrete independent reporting; a matching official statement is not yet available.

Source Stack

Independent reporting

Living trackers

Follow this signal

Related signals

  • Visible chains of thought are a safety advantage for AI, but that transparency is slipping away

    OpenAI's system card for GPT-6 Astra already reports a significant drop in how well the chain of thought can be monitored. With Gemini 3 Pro, they say, the chain of thought revealed that the model recognized it was in a test environment. In one of the first posts from the newly launched Deepmind Institute, researchers Rohin Shah and Anca Dragan argue that the visible chain of thought (CoT) is a key safety advantage.

  • Base Labs launches an open-weight AI safety partnership with Hugging Face and Goodfire

    Base Labs, the research group Baseten spun up earlier this year, will develop and publish methods for training and monitoring open models. The announcement lands amid debate for the safety of open-weight models — which can be made dangerous by removing their safeguards through a rising technique known as abliteration. The company is framing their future work as a “standard” for open models that is transparent and built into how models are trained and deployed, rather than bolted on afterward. Baseten launched a new safety infrastructure standard alongside its Base Labs research arm on Wednesday, partnering with Hugging Face and Goodfire AI to build safety evaluation and monitoring infrastructure for open-weight models.

  • New Deepseek model V4.1-Flash cuts memory needs for AI agents

    DeepSeek released V4.1-Flash for long-context and agent workloads. The Decoder reports that its fast GPU-memory KV cache uses about one quarter of the space required by DeepSeek-V4-Flash, while input processing activates fewer parameters than output generation.

Put the news to work

Choose your next AI workflow

Reduce AI API costs without losing useful results

Measure workload, retries and accepted outputs before comparing models, batching or reusable prompts.

Check AI data handling before a team rollout

Work through account policies, retention, local records and access boundaries before sharing team data with an AI workflow.

Plan a local AI deployment you can verify

Check model routing, server access, external traffic and release changes before depending on a local AI workflow.

See what happened next · Compare verified API prices · Estimate a workload · Read the weekly index · Get future updates