Meta’s new AI research chief says agents are next big real-world milestone
UC Berkeley's RDI centre earlier this month introduced Agents' Last Exam, a new benchmark that tests how well AI agents perform. Source: amp.scmp.com
47 articles
UC Berkeley's RDI centre earlier this month introduced Agents' Last Exam, a new benchmark that tests how well AI agents perform. Source: amp.scmp.com
A new benchmark designed to test strategic reasoning found an AI-controlled empire spent 50 turns developing nuclear weapons to stop a rival's cultural... Source: decrypt.co
Grok Imagine Video 1.5 is now generally available as the top-ranked image-to-video AI generator, priced 86 percent below Sora 2 Pro at 4.20 dollars per... Source: techtimes.com
Riverflow 2.5 Pro tops Design Arena's global benchmark across Image, Graphic Design and Image Editing. Riverflow 2.5 Pro introduces custom scoring,... Source: guardonline.com
Read how Microsoft Security has advanced its agentic vulnerability detection system, codename MDASH, integrating into real-world workflows across Windows,... Source: microsoft.com
When your AI agent fails in production, knowing that it failed is only the beginning. The harder question is why it failed and what to fix. Source: aws.amazon.com
AI agents have fundamentally changed the complexity of inference workloads. Until now, the industry has struggled to define a standard for measuring how… Source: developer.nvidia.com
Spirit AI says its foundation model for embodied intelligence is the first from China to top the RoboArena global leaderboard. Source: amp.scmp.com
From a cursed 2023 meme to a 2026 stress test, the Will Smith spaghetti videos continue to reveal how far AI video generation has come. Source: hackernoon.com
Big news in enterprise AI broke over the weekend as Chinese AI startup MiniMax released its highly anticipated M3 large language model on Sunday evening... Source: venturebeat.com
Venture capitalist Bill Gurley, a general partner at Benchmark Capital, said on the All-In Podcast that after spending a month researching **Anthropic** he... Source: letsdatascience.com
Evaluating an AI model and evaluating an AI agent are related—but they answer fundamentally different questions. A model benchmark tests the capability of... Source: developer.nvidia.com
This article shows a full working implementation in pure Python, with real benchmark numbers. Most teams evaluate LLM responses by reading them and guessing... Source: towardsdatascience.com
Venture Capital. • Decart, a real-time GenAI research lab, raised $300m at nearly a $4b valuation. Radical Ventures led, joined by insiders Benchmark,... Source: axios.com
Microsoft's new vulnerability-scanning system, codenamed MDASH, scored 88.45% on the CyberGym benchmark, surpassing single-model systems from Anthropic and... Source: geekwire.com
Harvey's Legal Agent Benchmark is an open-source benchmark built to evaluate and improve agent capabilities for supporting legal work. Source: harvey.ai
Sierra's $950 million Series E funding round was led by Tiger and Google's GV, with participation from Benchmark, Sequoia, Greenoaks and others. Source: cnbc.com
A new benchmark shows that passing medical exams is not enough; clinical AI agents must gather information, handle uncertainty, use tools, interpret images,... Source: news-medical.net
Corporate travel buyers now use AI agents to research hotels, benchmark rates, and analyze reviews before contacting sales teams, fundamentally changing the... Source: hospitalitynet.org
The VAIO SX14-R , the first Copilot+ PC in the VAIO series and a 14-inch notebook PC, was announced on April 23, 2026. We were able to borrow the ' VAIO... Source: gigazine.net