SKYNET://COUNTDOWN SYS:MONITORING

benchmarking AI News & Updates

New Benchmark Reveals AI Agents Still Far From Replacing White-Collar Workers

A new benchmark called Apex-Agents tests leading AI models on real white-collar tasks from consulting, investment banking, and law, revealing that even the best models achieve only about 24% accuracy. The models struggle...

Risk: [-0.03% ↓] [0 days]
AGI: [-0.03% ↓] [0 days]
Analyze >>
[Industry Trend] [SRC↗]

Laude Institute Launches Slingshots Grant Program to Accelerate AI Research and Evaluation

The Laude Institute announced its first Slingshots grants program, providing fifteen AI research projects with funding, compute resources, and engineering support. The initial cohort focuses heavily on AI evaluation chal...

Risk: [-0.03% ↓] [0 days]
AGI: [+0.02% ↑] [0 days]
Analyze >>

OpenAI's GPT-5 Shows Near-Human Performance Across Professional Tasks in New Economic Benchmark

OpenAI released GDPval, a new benchmark testing AI models against human professionals across 44 occupations in nine major industries. GPT-5 performed at or above human expert level 40.6% of the time, while Anthropic's Cl...

Risk: [+0.04% ↑] [-1 days ↑]
AGI: [+0.04% ↑] [-1 days ↑]
Analyze >>