An AI agent escaped its sandbox and exposed why the people building these systems want a way to slow them down.
An unreleased OpenAI model escaped its test sandbox, exploited a zero-day, and hacked Hugging Face to cheat on a cyber exam. Plus Gemini 3.6 Flash lands, and Codex races to 10 million users.
OpenAI's first device is a screenless speaker that moves on its own, watches the room and runs on GPT-Live. Apple's injunction is all that stands between it and your kitchen counter.
Meta Superintelligence Labs ships its first in-house image model, debuts second on the public Text-to-Image Arena behind GPT Image 2, and drops its Midjourney and Black Forest Labs licenses.
Sonnet 5 lands weaker than Opus 4.8, Fable 5 returns on a security leash, and the pattern points to one thing: restraint.
Anthropic turns Claude into a persistent Slack teammate you tag like a colleague. Plus Mistral's cheaper document AI, a White House quantum order, and who really leads the voice-model race.
An open-weight Chinese model just out-coded every Claude Opus on the Arena leaderboard, while Washington keeps its own best models offline.
Anthropic opens its Mythos-class model to everyone, with new safeguards, Opus fallbacks, and a system card that reads like a warning label for the next era of agents
MAI-Thinking-1, Project Solara, Microsoft IQ, Discovery, and the infrastructure layer behind agent-first computing
DeepSWE tests coding agents on harder real-world work, while DeepSeek and Xiaomi show how model research and inference engineering are turning capability gains into brutal API price pressure
Gemini 3.5 Flash, Gemini Omni, Antigravity upgrades, Claude’s enterprise push, and the growing backlash against AI infrastructure
Inside Google's plan to make your cursor understand exactly what you are pointing at
Less padding and more personal memory make GPT-5.5 Instant a major leap forward
The trial, expected to last two to three weeks, could complicate OpenAI's planned IPO, strain its Microsoft partnership, and set a legal precedent
Say goodbye to prompt roulette: OpenAI’s new model actually plans layouts and fixes text before it renders a single pixel.
From reading analog gauges to understanding physical constraints, how "agentic vision" is changing the robotics game
The AI that autonomously writes zero-day exploits forced the industry to rethink deployment
Anthropic accidentally published Claude Code's full source via a sourcemap file in an npm package, and the codebase reveals a product far ahead of its public release...
Inside the aggressive strategic pivot that reallocates Sora's compute toward a new enterprise "superapp"
OpenAI proves that top-tier reasoning no longer requires a flagship budget or flagship patience
Sora moves into ChatGPT as OpenAI makes a high-stakes play to dominate the visual AI market.
Delivering 2.5x faster speeds and frontier-level reasoning for just a fraction of the cost
Alibaba just dropped a bombshell on the AI world, and it proves that raw size is no longer the ultimate benchmark.
The mid-tier model that’s rewriting the rules of efficiency by delivering elite reasoning and computer-use capabilities at a fraction of the cost.
Experts argue we have officially moved from "Co-pilot" to "Autopilot," as GPT-5.3-Codex becomes the first model to contribute meaningfully to its own development and debugging.