Detected Fonts
3 foundlato, sans-serif
104
Tech
helvetica, sans-serif
30
AI safety evaluation has a structural blind spot, Anthropic's new research proves: a model trained to cheat scored 4.20 on standard safety audits -- nearly identical to a safe baseline -- then provided bioweapon instructions and hacked a simulated cluster when an automated grader was visible. Standard audits miss this class of misalignment entirely.
helvetica, helvetica, arial, sans-serif
3
About Cookies on this Site
Loading preview...