New benchmark confirms AI models still perform poorly at visual perception
Moonshot AI's PerceptionBench tests how well multimodal AI models can actually "see," separate from logical reasoning. No frontier model reaches 60 percent…
AI news, research, models, robotics, chips, startups, and infrastructure coverage.
Moonshot AI's PerceptionBench tests how well multimodal AI models can actually "see," separate from logical reasoning. No frontier model reaches 60 percent…
Nvidia and OpenAI close in on data center deal, but cut the guarantee nearly in half Instead of backing $250 billion as…
The police-tech giant Flock is announcing today that it will change officers’ access to its nationwide network of license plate readers, in…
figure — aI is changing the pace of open source development and the security challenges that come with it. Maintainers are reviewing…
Claude Cowork now runs directly in the side panel of Anthropic's Chrome extension. The extension already existed. It reads the page you're logged into,…
Anthropic investors expect the AI start-up to float at a valuation of $2 trillion or more in October, a dizzying figure that…
'ZDNET Recommends': What exactly does it mean? ZDNET's recommendations are based on many hours of testing, research, and comparison shopping.
Anthropic is testing whether Claude Code can handle daily maintenance of the company's own software. In a few weeks, the AI created…
OpenAI's Computer History tracks user activity across apps and websites, turning it into a searchable timeline with memories that ChatGPT and Codex…
Chinese AI startup Zhipu AI has released GLM-5.3. The model shares the same base as its predecessor, GLM-5.2, and all gains come…