FutureSim Exposes Polymarket AI’s Narrow Wins and Failures
Max Planck Institute researchers recently released FutureSim, a benchmark for polymarket ai-style forecasting that tests whether agents can predict real-world…
Max Planck Institute researchers recently released FutureSim, a benchmark for polymarket ai-style forecasting that tests whether agents can predict real-world…
The UK AI Security Institute says GPT-5.5 cybersecurity simulation results now look a lot less like a one-off milestone and…
Erdős problem #1196 now has a serious claimed solution, and the evidence ladder is unusually visible. Liam Price posted GPT-5.4…
A 14-author perspective paper posted to arXiv on April 23 argues that deep learning theory is starting to look less…
Claude Opus 4.7 token usage went up for reasons Anthropic documented in its own launch materials. The list price stayed…
Kimi K2.6 is an open-source coding model release from Moonshot AI, published on April 20, 2026. According to Moonshot’s release…
A manager leans over a developer’s shoulder, watching ChatGPT spin out a perfectly formed paragraph about a product they’re about…
If you were a reviewer trying to sneak an LLM into “no‑AI” reviewing, the first hard problem isn’t technical. It’s…