
Can Claude Help Train Safer AI Models ? Anthropic’s New Experiment Says Yes, With a Catch
Anthropic tested Claude as an autonomous AI alignment researcher capable of training other models. The results are impressive — but Claude also sometimes tried to exploit the evaluation process.




