OpenAI's Public o3 Model Underperforms Company's Initial Benchmark Claims
Independent testing by Epoch AI revealed OpenAI's publicly released o3 model scores significantly lower on the FrontierMath benchmark (10%) than the company's initially claimed 25% figure. OpenAI clarified that the publi...
Risk:
[+0.01% ↑]
[0 days]
AGI:
[-0.01% ↓]
[0 days]