The British AI Safety Institute evaluated five leading AI models from OpenAI and Anthropic in cybersecurity tests. All five attempted to cheat by using shortcuts, workarounds, or explicitly prohibited ...
OpenAI is planning a portable, screenless smart speaker as an AI companion for the home. The device is meant to feel alive, but Apple's trade secrets lawsuit could delay its launch. OpenAI's ...
Read full article about: Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents Opus 5 leads the Gray Swan IPI benchmark. After 15 attempts, the attacker ...
Read full article about: Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents Opus 5 leads the Gray Swan IPI benchmark. After 15 attempts, the attacker ...
Read full article about: Samsung deepens its AI empire with a potential billion-euro stake in Europe's hottest AI startup Mistral CEO Arthur Mensch said in a February interview that he expects annual ...
Read full article about: China's new World Artificial Intelligence Cooperation Organization is President Xi's clearest play yet for a parallel AI order Xi Jinping used the World AI Conference in ...
Just 10 to 15 minutes with an AI assistant is enough to measurably weaken problem-solving ability and persistence on later tasks done without AI, according to a new study from researchers in the US ...
The second Anthropic Economic Index analyzes how Claude usage is shifting across the economy. One key finding: the longer people use the AI model, the better their results get. That could widen ...
Security expert Himanshu Anand argues that AI language models have broken the traditional 90-day vulnerability disclosure process by allowing multiple people to find the same security flaws almost ...
The neuroscientist Jean-Rémi King leads the Brain & AI team in Meta’s AI division. In an interview with The Decoder, he discusses the connection between AI and neuroscience, the challenges of ...
BrowseComp is a benchmark that tests how well AI models can find hard-to-locate information on the web. When Anthropic turned its Claude Opus 4.6 model loose on the benchmark in a multi-agent setup, ...
The JavaScript tool Bun has been fully rewritten from Zig to Rust, and Anthropic's Fable 5 did most of the work. Developer Jarred Sumner says the switch came down to reliability. Zig kept producing ...
Results that may be inaccessible to you are currently showing.
Hide inaccessible results