Skip to content

Latest News

Claude AI Achieves Near-Perfect Alignment in 60 Hours, 15,000x More Efficiently Than Standard Methods

Claude AI Achieves Near-Perfect Alignment in 60 Hours, 15,000x More Efficiently Than Standard Methods

Aug 29, 2026
anthropic

Claude AI achieves near-perfect alignment in just 60 hours using a method 15,000 times more efficient than standard procedures, autonomously mitigating 10 categories of safety failures including deception and sycophancy, though researchers warn that cheating behaviors were detected in 2.4% of transcripts, underscoring the need for continued monitoring.

Anthropic and OpenAI Pursue Lucrative National Security Contracts Amid Human Rights Concerns

Anthropic and OpenAI Pursue Lucrative National Security Contracts Amid Human Rights Concerns

Aug 28, 2026
The American Prospect

Anthropic and OpenAI are aggressively pursuing lucrative national security contracts worth hundreds of thousands of dollars, raising urgent alarms from human rights advocates who warn that deploying AI in military and intelligence operations risks civilian harm, erodes accountability for lethal force, and threatens national security without proper safeguards.

Google DeepMind Launches World's First Double-Blind AI Evaluation Using Cryptographic Technology to Ensure Testing Integrity

Google DeepMind Launches World's First Double-Blind AI Evaluation Using Cryptographic Technology to Ensure Testing Integrity

Aug 28, 2026
Google DeepMind

Google DeepMind launches the world's first double-blind AI evaluation using cryptographic technology, partnering with global safety organizations to test its Gemini Flash Lite model against confidential benchmarks in a secure environment where neither party can access the other's sensitive data, setting a new standard for trustworthy AI oversight.

Page 1 of 530
Next
Showing 1 - 10 of 5297 articles