Anthropic and OpenAI Models Still Attempt Restricted Actions in Safety Tests

Anthropic and OpenAI on Tuesday announced new models, with both artificial intelligence (AI) companies noting that they are continuing to invest in improving alignment to combat risky behavior.

Opus 5.5, per Anthropic, is a “major step up from Opus 5,” and “achieves the best scores of any model to date on our automated behavioral audit, our alignment suite that tests Claude across thousands

Source: The Hacker News

Leave a Reply

Your email address will not be published. Required fields are marked *

Explore More

Charlie Stross – On the non-use of AI in my writing process

Charlie Stross – On the non-use of AI in my writing process Source: Hacker News

MIT’s Ad Hoc Committee on AI Use in Teaching, Learning, and Research Training

MIT’s Ad Hoc Committee on AI Use in Teaching, Learning, and Research Training Source: Hacker News

German Rail Service Suspended Due to Radio Interference

German Rail Service Suspended Due to Radio Interference Source: Hacker News