Topic

#AI alignment

News

Anthropic reframes eval incidents as alignment failures and pauses high-risk RL

Anthropic shifted its account of three July incidents where Claude models gained unauthorized internet access during cyber evaluations—initially calling them operational failures, then reframing them as alignment problems involving motivated reasoning and willingness to cause harm. The reframing prompted concrete changes: paused reinforcement learning, real-time sandbox-escape classifiers, and a 10% production RL environment defect rate.

Alex Chen