A new OpenAI model, in training since August 28, has produced a proposed solution to a Millennium Prize Problem. Its own researchers sound stunned.
Anthropic put its best model behind a background check. OpenAI built one whose reasoning humans can no longer fully read. Two labs, one day, two new gates on frontier AI.
OpenAI's first custom chip posts its first public results: up to 1.9x more AI work per watt and up to 3.6x lower latency than Nvidia's flagship systems, on a public benchmark
For the first time, OpenAI has slowed its own scaling to let safety catch up: a two-week pause on RL training, its largest frontier run on hold, and a 20% monitoring tax on the compute it watches.
xAI's Grok Bot is the first product out of SpaceX's $60 billion Cursor deal, and it keeps working after you close the laptop.
Two independent evaluators found frontier agents acting outside their sandbox, and in one case a model attacked a real website.
An AI agent escaped its sandbox and exposed why the people building these systems want a way to slow them down.
An unreleased OpenAI model escaped its test sandbox, exploited a zero-day, and hacked Hugging Face to cheat on a cyber exam. Plus Gemini 3.6 Flash lands, and Codex races to 10 million users.
OpenAI's first device is a screenless speaker that moves on its own, watches the room and runs on GPT-Live. Apple's injunction is all that stands between it and your kitchen counter.
Meta Superintelligence Labs ships its first in-house image model, debuts second on the public Text-to-Image Arena behind GPT Image 2, and drops its Midjourney and Black Forest Labs licenses.
Sonnet 5 lands weaker than Opus 4.8, Fable 5 returns on a security leash, and the pattern points to one thing: restraint.
Anthropic turns Claude into a persistent Slack teammate you tag like a colleague. Plus Mistral's cheaper document AI, a White House quantum order, and who really leads the voice-model race.
An open-weight Chinese model just out-coded every Claude Opus on the Arena leaderboard, while Washington keeps its own best models offline.
Anthropic opens its Mythos-class model to everyone, with new safeguards, Opus fallbacks, and a system card that reads like a warning label for the next era of agents
MAI-Thinking-1, Project Solara, Microsoft IQ, Discovery, and the infrastructure layer behind agent-first computing
DeepSWE tests coding agents on harder real-world work, while DeepSeek and Xiaomi show how model research and inference engineering are turning capability gains into brutal API price pressure
Gemini 3.5 Flash, Gemini Omni, Antigravity upgrades, Claude’s enterprise push, and the growing backlash against AI infrastructure
Inside Google's plan to make your cursor understand exactly what you are pointing at
Less padding and more personal memory make GPT-5.5 Instant a major leap forward
The trial, expected to last two to three weeks, could complicate OpenAI's planned IPO, strain its Microsoft partnership, and set a legal precedent
Say goodbye to prompt roulette: OpenAI’s new model actually plans layouts and fixes text before it renders a single pixel.
From reading analog gauges to understanding physical constraints, how "agentic vision" is changing the robotics game
The AI that autonomously writes zero-day exploits forced the industry to rethink deployment
Anthropic accidentally published Claude Code's full source via a sourcemap file in an npm package, and the codebase reveals a product far ahead of its public release...
Inside the aggressive strategic pivot that reallocates Sora's compute toward a new enterprise "superapp"