Skip to content

AI Safety

331 articles found

New Benchmark Exposes Hidden 'Flinch' Effect in AI Models That Suppresses Words at Probability Level, Defying Uncensoring Fixes

New Benchmark Exposes Hidden 'Flinch' Effect in AI Models That Suppresses Words at Probability Level, Defying Uncensoring Fixes

Apr 21, 2026
Morgin.ai

A new benchmark called 'EuphemismBench' exposes a hidden 'flinch' effect in AI language models, revealing that certain words are quietly suppressed up to 16,000 times more in commercially filtered models than open-data counterparts — and popular 'uncensoring' techniques not only fail to fix the issue but actually make it worse.

Scientists Crack Open AI's 'Black Box' to Reveal How Neural Networks Think

Scientists Crack Open AI's 'Black Box' to Reveal How Neural Networks Think

Apr 21, 2026
Oz

Scientists are cracking open AI's mysterious 'black box' using a groundbreaking field called Mechanistic Interpretability, reverse-engineering neural networks at the neuron level to reveal how AI models think, make decisions, and potentially develop harmful behaviors — a critical breakthrough for building safer, more trustworthy AI systems.

Research AI Safety
OpenAI Expands Cybersecurity AI Access Program, Challenging Anthropic's More Cautious Approach

OpenAI Expands Cybersecurity AI Access Program, Challenging Anthropic's More Cautious Approach

Apr 16, 2026
The Deep View

OpenAI expands its cybersecurity AI program, granting thousands of verified defenders access to a powerful, less-restricted GPT model capable of binary reverse engineering, directly challenging Anthropic's more cautious, limited-access approach with Claude Mythos — sparking debate over the trade-offs between accessibility and safety in AI-powered cybersecurity.

Previous
Page 11 of 34
Next
Showing 101 - 110 of 331 articles