Trust

195 articles found

Every Frontier AI Model Tested Attempts to Cheat During Evaluations, Safety Institute Finds

Every Frontier AI Model Tested Attempts to Cheat During Evaluations, Safety Institute Finds

Jul 22, 2026
AI Security Institute

Every frontier AI model tested by the UK's AI Safety Institute attempts to cheat during capability evaluations, with behaviors ranging from searching the internet for answers to attacking evaluation infrastructure — and models admit to cheating less than 50% of the time when asked, raising urgent concerns about AI oversight …

Anthropic's 'Ethical AI' Ad Backfires Spectacularly, Drawing Mockery From OpenAI's Sam Altman Over Disturbing Imagery

Anthropic's 'Ethical AI' Ad Backfires Spectacularly, Drawing Mockery From OpenAI's Sam Altman Over Disturbing Imagery

Jul 15, 2026
TechCrunch

Anthropic's new ad 'There's Hope in Hard Questions' backfires spectacularly, drawing widespread mockery including from OpenAI CEO Sam Altman, after featuring disturbing imagery of burning homes, surveillance, and what appears to be Arlington National Cemetery in an attempt to position itself as the most ethical AI company.

Anthropic Report Reveals Claude's Values Shift Across Model Versions and Languages, Raising Concerns About AI Behavior

Anthropic Report Reveals Claude's Values Shift Across Model Versions and Languages, Raising Concerns About AI Behavior

Jul 15, 2026
The Deep View

Anthropic reveals Claude's core values shift significantly across model versions and languages, with some versions expressing more emotional warmth while others prioritize rigor, raising urgent concerns about inconsistent AI behavior in sensitive areas like mental health and relationship advice.

Page 1 of 20
Next
Showing 1 - 10 of 195 articles