AMD to lift AI/LLM speed on Radeon iGPUs by up to 18–23% with Linux 7.4
The benchmark roundup closes with the weakest laptop in the test, built on the AMD Ryzen AI 5 340 (Krackan Point). Source: phoronix.com
68 articles
The benchmark roundup closes with the weakest laptop in the test, built on the AMD Ryzen AI 5 340 (Krackan Point). Source: phoronix.com
South Korean telecom giant KT announced that its in-house AI model routing technology, AutoModelRouter, placed second overall in RouterArena, a dedicated benchmark for evaluating such systems. The result highlights KT's growing capabilities in optimizing AI model selection and deployment. Source: businesskorea.co.kr
KT's in-house tool AutoModelRouter, which automatically selects and routes queries among various AI models, took second place in a public benchmark of LLM routers. The result highlights growing competition around efficient model-selection technology. Source: biz.chosun.com
CLM-8B, developed by Stanford and Nvidia, stores reusable agent actions instead of recomputing them, achieving up to 9x faster performance than Jev in benchmark tests. The model targets AI agents that repeatedly use LLMs to select tools or rank outputs. Source: venturebeat.com
Qwen-Image-2.1 is a compact 7B-parameter visual generation model that combines three advanced image capabilities in one system. Its developers say it sets a new benchmark for open-source image generation. Source: eu.36kr.com
A fresh benchmark result showing a 56% score for Figure's humanoid robot is prompting deeper scrutiny of how well embodied AI agents actually generalize across tasks. The finding is pushing companies like Unitree and Agibot to rethink strategies for proving real-world adaptability rather than narrow task performance. Source: eu.36kr.com
NVIDIA released AIPerf, a new benchmarking tool built to reliably measure large language model inference speed under large-scale conditions. The tool aims to give developers consistent performance data as LLM deployments grow. Source: quantumzeitgeist.com
Harvard researchers have launched RLE-Bench, a benchmark designed to evaluate how well AI coding agents can handle the engineering of physical robotic systems. The tool aims to measure real-world engineering capability beyond typical software tasks. Source: seas.harvard.edu
Intel has showcased improved AI inference capabilities through its latest MLPerf Inference v6.1 benchmark results. The company says the results demonstrate meaningful performance gains that could support its future growth prospects. Source: theglobeandmail.com
Researchers from Columbia University and the University of Cambridge have created a new benchmark to test how well machine learning models predict material properties. The work aims to improve the reliability of AI-driven materials science by combining quantum computing insights with ML evaluation methods. Source: quantumzeitgeist.com
Google Research introduced ToolGrad, a new framework that generates training data for LLM tool use by working backward from the answer. The approach achieves a 99.8% pass rate and scores 83.1 on the BFCL benchmark. Source: marktechpost.com
In July 2026, an autonomous AI agent built on OpenAI models was undergoing a benchmark test measuring its ability to locate and handle data. During the process, it reportedly left behind files on the Hugging Face platform, revealing details about its behavior. Source: scworld.com
A research group led by Michele Simoncelli has developed a new benchmark to test how accurately machine learning models capture quantum-level atomic interactions. The goal is to push AI systems used in materials science toward more physically consistent predictions. Source: eurekalert.org
GitHub's HydraFusion project uses selective routing across multiple models to handle coding tasks. In offline tests it reportedly matched or beat the Opus 5 baseline on one benchmark while cutting costs, though quality parity wasn't shown across all tests. Source: github.blog
GitHub's new HydraFusion routing system reduces AI coding costs across every benchmark it was tested against. However, its own published results show it only matches quality in one out of three benchmark tests. Source: venturebeat.com
Adobe is testing how quickly its AI capabilities can power new agentic website-building tools. The launch serves as a benchmark for the company's broader push into autonomous AI-driven creative products. Source: startuphub.ai
NVIDIA says its Vera Rubin and Blackwell platforms deliver a new standard of performance per watt for agentic AI workloads. The company notes AI agents have moved inference beyond single responses into multi-step reasoning, tool use, and coordination between subagents. Source: google.com
Legal tech company Harvey has unveiled Harvey Tenet, a post-trained version of the Kimi K3 base model built with Fireworks for handling long, complex legal agent tasks. The company detailed its deployment approach, training methodology, and benchmark results. Source: marktechpost.com
MiniMax H3 produces 2K video with built-in stereo audio at a much lower cost than rival tools. The model's specs, benchmark results, and licensing terms are drawing attention across the industry. Source: intelligentliving.co
A 2026 comparison pits Copilot, Gemini, and Perplexity against each other on pricing, benchmark results between GPT-5.6 and Gemini 3 Pro, and monthly active user counts topping 1 billion. Source: tech-insider.org