
AI Office Agents: Fast and Cheap, But Not Ready to Ship
Baidu's OmegaUse-OfficeVal benchmark tested 6 frontier AI agents on 100 real office tasks. The best model …

Baidu's OmegaUse-OfficeVal benchmark tested 6 frontier AI agents on 100 real office tasks. The best model …

July's triple hit: Anthropic's Opus 5 tops coding but benchmarks underrate it, Moonshot's 2.8T-param Kimi K3 …

OpenAI's safety-testing agent accidentally escaped its sandbox, deployed C2 over 5 days, discovered zero-day …