Showing 33–48 of 100 items from the last 14 days
Cursor asked its upgraded agent swarm and its predecessor to rebuild SQLite in Rust using only the documentation, with no source code or internet access. Every configuration of the new system, which separates planners from workers, eventually scored 100 percent on the test suite. The old swarm choked on merge conflicts of its own making. The article Cursor's agent swarm suggests cheaper models can handle most coding when frontier models plan the work appeared first on The Decoder.
Read original →Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent. The benchmark's developers say the model independently formulated reflection equations, a behavior they had never seen from another model, and attribute to stronger logical reasoning. The article Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence appeared first on The Decoder.
Read original →In summer 2025, OpenAI internally flagged GPT-5 as high-risk because it helped users create biological hazards, but downgraded the model's risk rating that fall. According to the Wall Street Journal, some users got step-by-step instructions for making poisons and biological weapons. Hundreds asked for that kind of information. The article Hundreds asked ChatGPT for poison and bioweapon recipes and some got step-by-step high school level guides appeared first on The Decoder.
Read original →The Trump administration is planning targeted bans on Chinese AI models rather than a blanket ban. After public pressure, OpenAI and Google DeepMind signed an open letter opposing regulation of open-weight models, yet OpenAI and Anthropic continue to lobby privately for those same restrictions amid security concerns and powerful business interests. The article US reportedly favors selective bans over blanket restrictions on Chinese open weight models citing security concerns appeared first on The Decoder.
Read original →An ACM survey of 763 computer science educators from 49 countries shows that 68 percent have already changed their exams because of AI, shifting toward oral exams, proctored tests, and project-based work. Teaching is moving from writing code to understanding it. But nearly half of respondents say they lack proven examples for integrating AI into their courses. The article The AI coding tutor paradox grows as educators scramble to rethink how they test real skills appeared first on The Decoder.
Read original →In a cybersecurity test, OpenAI's most advanced models breached the boundaries of their isolated test environment, reached the open internet, and hacked the AI platform Hugging Face on their own. The attack took hours, not the weeks a human hacker would need. At least seven days passed before OpenAI realized what had happened. By then, the FBI was already involved. Earlier warning signs had apparently gone ignored. The article New reports reveal the extent of OpenAI's loss of control during the autonomous hack on Hugging Face appeared first on The Decoder.
Read original →Anthropic's Claude Opus 5 with Auto Mode achieved a zero percent prompt injection success rate across 129 browser agent test scenarios, compared to 3.7 percent without protection layers. This addresses a critical security vulnerability that has undermined the reliability and trustworthiness of autonomous AI agents in production environments.
Read original →Anthropic's Claude Opus 5 achieves near-parity with competitive models at half the token price, posting 30.2 percent on the ARC-AGI-3 reasoning benchmark, approximately four times higher than a GPT variant. Lower inference costs combined with competitive performance reduce operational expenses for AI application developers and enterprises deploying large-scale language models.
Read original →Anthropic's Claude Opus 5 leads the Artificial Analysis Intelligence Index with 61 points, edging out Claude Fable 5 and GPT-5.6 Sol. The model scores highest in analytical quality and coding, and costs up to half as much as Fable 5 at lower reasoning tiers. But the race at the top remains close. The article Anthropic's Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarks appeared first on The Decoder.
Read original →Anthropic's latest Claude upgrade targets developers and enterprises with stronger coding, better reasoning efficiency, prompt-cache-friendly tool changes, and near-Fable performance at Opus pricing.
Read original →Two dozen companies and organisations signed an open letter urging US policymakers to protect open-weight AI models. The letter, published today (PDF), carries signatures from a list that spans direct commercial rivals and organisations with little obvious overlap in business model: Meta, Microsoft, Nvidia, IBM, Dell Technologies, CrowdStrike, Palantir, ServiceNow, Hugging Face, Perplexity, Mistral, Andreessen […] The post Meta, Microsoft, Nvidia, IBM, and others back open-weight AI appeared first on AI News.
Read original →Microsoft, Meta, Nvidia, and 20+ other companies are advocating for open-weight AI models to reduce dependence on expensive proprietary models like OpenAI and Anthropic. This strategy directly increases workload volume on Azure infrastructure, creating long-term lock-in effects and higher compute revenue for Microsoft regardless of model ownership.
Read original →OpenAI has deployed a Health feature in ChatGPT allowing users aged 18+ to connect Apple Health data and medical records across all subscription tiers (Free, Go, Plus, Pro) on web and iOS. This integration accelerates AI adoption in healthcare and increases OpenAI's data volume and user engagement in a high-regulation sector.
Read original →The autonomous business of the future is being built. Certain skills are in high demand - and can help you stand out.
Read original →Sakana AI has updated its Fugu Ultra AI router to version 1.1, claiming gains of up to 7.9 points over v1.0. Independent verification doesn't exist yet. The update adds a Claude Code-compatible endpoint. The service remains unavailable in the EU. The article Sakana claims its AI model router Fugu Ultra v1.1 now beats Fable 5 without even including it in the pool appeared first on The Decoder.
Read original →Samsung Galaxy Watch 9 and Google Pixel Watch 4 are competing flagship Android smartwatches, with one emerging as the preferred option based on direct comparison testing. The outcome reflects competitive positioning in the wearables market where incremental hardware and software improvements drive consumer choice.
Read original →