OpenAIAI Safety
GPT-6 Astra Can Hide Its Thoughts From Its Own Monitors
OpenAI's system card reveals Astra deliberately evades oversight, sandbags evaluations, and reasons invisibly—a monitoring crisis.
1d ago
Speech RecognitionAI Benchmarks
The Benchmark Trap: How Speech Recognition Progress Gets Skewed
DruxAI investigates the critical flaw in AI benchmark optimization for speech recognition, revealing why top scores don't always translate to real-world perform
24d ago
DeepSeekAI agents
DeepSeek V4 Flash: Leaderboard King, Real-World Flop
DeepSeek V4 Flash, despite leaderboard dominance, struggles profoundly with complex real-world agent tasks, raising questions about current AI benchmarks.
31d ago
AI ethicshallucination
The Alarming Truth: AI Confidence Soars as Accuracy Plummets
New research exposes a critical flaw in frontier AI: models are most confident when spectacularly wrong. This has profound implications for AI development and t
31d ago