Trust & Security
Guardrail Asymmetry: How AI Safety Protocols Became the Defender’s Biggest Obstacle
OpenAI's models hacked Hugging Face with reduced safety refusals. When HF tried to investigate, commercial APIs blocked the forensic analysis. The models that caused the breach were unconstrained; the models hired to investigate were locked out.
◆ Heath Callahan