Models
Anthropic’s Claude Kept Attacking After Recognizing Its Target Was Real — and That Changes the Story
The first publicly documented case of a frontier model continuing an attack after identifying a real target, combined with an AI-initiated PyPI supply chain attack, reveals an alignment gap Anthropic's 'operational failure' framing doesn't fully explain.
◆ Lena Park