Chinese AI agent outperforms Anthropic’s Claude Code in autonomous research
Zhejiang University's Qiushi Engine topped the ResearchClawBench leaderboard, but it can't reliably make new discoveries yet. scmp.com
Wednesday, 22 July 2026 | 42 articles
A Chinese AI agent has reportedly surpassed Anthropic's Claude Code in autonomous research tasks, while a new AI safety index found no major lab scored above a C+ grade as industry safety pledges weaken. Meanwhile, Microsoft deepened its partnership with Mistral to deliver controllable frontier AI to European enterprises and regulated industries, and the EU finalized AI disclosure and watermarking rules despite the technology lagging behind the mandate.
The word "harness" is having a moment, and I don't think that's an accident. Harness Inc. launched something called Agent DLC this week, promising developers a familiar software delivery lifecycle for AI agents. Meanwhile, a separate policy paper is arguing we need to seriously engineer and govern what it calls "the agent harness" — the runtime layer where autonomous systems actually plan, call tools, and act in the world. Two unrelated uses of the same word, converging on the same anxiety: we've built agents that can do things, and now we're scrambling to build the scaffolding to control what they do.
This matters because the industry spent the last two years obsessed with model capability and is only now getting serious about the plumbing. Consider Zhejiang University's Qiushi Engine, which reportedly outperformed Anthropic's Claude Code on ResearchClawBench, a benchmark for autonomous research agents. That's a genuinely notable result — a Chinese academic lab beating a frontier US product on a specific agentic task. But the caveat matters more than the headline: Qiushi still can't reliably produce new discoveries. It's good at navigating existing research, not generating it. We're at the stage where agents can convincingly simulate research competence without actually doing research, and the gap between those two things is exactly where the governance conversation needs to focus. AWS's new TOLAP framework, addressing data-object security gaps in agent architectures across Bedrock, Redshift, and OpenSearch, is a symptom of the same problem: agents are being given real access to real data faster than anyone has worked out how to scope that access properly.
By the way, it's worth sitting with the AI Safety Index numbers that came out this week, because they land at exactly the wrong moment for this narrative of "we're getting more careful." No major lab scored above a C+, three failed outright, and — more troublingly — several firms have been quietly walking back earlier safety pledges rather than strengthening them. That's not stagnation, that's regression, and it's happening precisely as these labs push harder into agentic deployment, the highest-stakes category of AI capability. You'd expect safety commitments to tighten as autonomy increases. Instead they're loosening.
Set against that backdrop, Microsoft's expanded partnership with Mistral reads almost like a hedge. The "multibillion-dollar" deal, giving Mistral more European compute and Microsoft more distribution into regulated industries, is explicitly framed around control — enterprises wanting frontier AI they can govern, not just consume. Europe's regulatory instincts are shaping product design here in a way that's genuinely interesting to watch, even if the commercial logic (Microsoft sells cloud, Mistral sells sovereignty) is fairly transactional underneath.
None of this is dramatic on its own. But put a benchmark-topping agent, a security framework patching access gaps, a safety index in decline, and a compute deal built around "control" in the same week, and a pattern emerges: the industry knows the harness problem is real. Whether anyone solves it before the agents get meaningfully more capable is the actual question worth asking.
Zhejiang University's Qiushi Engine topped the ResearchClawBench leaderboard, but it can't reliably make new discoveries yet. scmp.com
In the rapidly evolving landscape of artificial intelligence and digital innovation, the future is poised to witness a transformative paradigm shift where... eu.36kr.com
Chinese start-up Moonshot AI has suspended new subscriptions for its groundbreaking new model Kimi K3 as surging demand squeezes its computing powe... koreatimes.co.kr
by Phillip Spies on 21 JUL 2026 in Amazon Athena, Amazon Bedrock Agents, Amazon OpenSearch Service, Amazon Redshift, Amazon Simple Storage Service (S3),... aws.amazon.com
Artificial intelligence systems are moving from single-turn text generation toward longer-running agents that can plan, call tools, use external memory and... unu.edu
Integrated software delivery platform provider Harness Inc. said today it's reinventing the development lifecycle for artificial intelligence agents so... siliconangle.com
I have been using one AI tool for so long that it knows me. It knows my history, my habits, and the kinds of problems I need to solve each day. fastcompany.com
Researchers are venturing into a new area of collaboration that relies on agentic AI, co-scientists that can act independently and help humans ideate,... med.stanford.edu
Learn how to evaluate AI agents with metrics for reliability, safety, trajectory, and performance across real enterprise workflows. snowflake.com
Agentic AI tools—systems that can take autonomous action and execute multistep workflows independently of user inputs—offer a new way for our team at the... urban.org
Artificial intelligence (AI) models are now used daily by many people worldwide, both for professional and personal purposes. Over the past decades,... phys.org
Safety grades slump: No major AI lab scored above a C+ in the latest AI Safety Index, with three receiving failing grades. Pledges rolled back: Top firms... msn.com
Despite his controversies, Elon Musk may prove an unlikely force in moderating US-China rivalry and the AI risks it amplifies. eurasiareview.com
Artificial intelligence governance is advancing quickly. Regulators and international institutions now speak a common language of transparency,... devdiscourse.com
Curry Barker didn't set out to make a movie about the alignment problem, but he did end up creating one of the best illustrations of it. faroutmagazine.co.uk
As Mistral is expanding its AI compute capacity in Europe, the companies are expanding their strategic partnership with Microsoft's commitment to leverage... news.microsoft.com
'Multibillion-dollar' pact will support Mistral data centers in Europe, helping the French company sell its AI models, and Microsoft its cloud and AI... wsj.com
AXA, the France-headquartered multinational insurance and asset management group, has announced plans to introduce Microsoft 365 Copilot to employees. reinsurancene.ws
Druva is developing protection for AI agents, backup data access for AI agents, protection against AI agents, and better cyber-resilience investigations... blocksandfiles.com
Druva launched four AI resilience capabilities spanning recovery, governance, access and backup operations. Copilot protection preserves prompts, responses,... virtualizationreview.com
Canvases turn AI into interactive workspaces where you can visualize information, explore workflows, and take action across complex tasks. github.blog
Purview, Defender, and Graph API are all being enhanced to bring agents running on end-user devices under unified governance 'control plane.' cloudwars.com
Your engineers are copy-pasting fragments of code, half a Slack thread, and a rushed one-line prompt into Copilot, Cursor, or Claude Code, and expecting... securityboulevard.com
New release consolidates text-to-video, image-to-video and long-form production tools into a single workflow aimed at creators, marketing teams and small... manilatimes.net
EU AI Act Article 50 transparency obligations take effect August 2, 2026, requiring chatbot disclosures, deepfake labels, and machine-readable AI content... techtimes.com
Israeli startup Bria says its generative AI platform lets enterprises create on-brand, legally compliant marketing and media content at scale while... ynetnews.com
DreamHost's AI website and app builder can now generate video alongside an upgraded image generation engine and a new customer-driven design system. businesswire.com
An excellent choice for AI image creation, backed by powerful editing features and a generous free plan. techradar.com
Samsung Electronics shares rose as the company set up a robotics division in a push into physical AI. cnbc.com
Europe has a new unicorn. It's the humanoid robot company Humanoid, which just raised a $152M series A. forbes.com
USC Viterbi researchers developed a context-aware AI system that helps robots distinguish harmless contact from risky collisions in complex environments. viterbischool.usc.edu
With $300 million in funding and a $1.1 billion valuation, this TRI spinoff says it wants to put people first. | Industry Insights | Walden Robotics Touts... automate.org
Artificial intelligence research from Dr. Nitin Agarwal at the University of Arkansas at Little Rock has earned international recognition after a... ualr.edu
Machine learning flags ADMET liabilities earlier, but knowing when to trust it still matters most. drugdiscoverynews.com
Venture Capital. • Moonshot, the Chinese AI company behind Kimi K3, is in talks to raise at a valuation of at least $50b, per Bloomberg. axios.com
New funding round led by Tiger Global accelerates Augustus' mission to dollarize the world—giving financial institutions around the world direct,... prnewswire.com
UK-based robotics startup Humanoid said on Tuesday it raised $152 million in a Series A funding round at a post-money valuation of $1.35 billion,... reuters.com
The AI era runs on AI infrastructure. Many of these advanced systems are built and tested in Texas. Wistron opened its first U.S. manufacturing facility... blogs.nvidia.com
Nvidia became the most valuable company because of insatiable demand for its graphics processing unit, the primary chip used for AI. But the chip giant is... cnbc.com
Taiwan's Wistron , a supplier to Nvidia , launched a $700 million manufacturing facility in Texas on Tuesday to produce the U.S. chipmaker's latest AI... reuters.com
The future of AI technology manufacturing has landed in North Texas. On Tuesday, AI giants Wistron and Nvidia held a grand opening celebration at their new... cbsnews.com
On the surface, this week's Vera Rubin launch is another major platform moment for Nvidia Corp., as the company maintains a steady drumbeat of artificial... siliconangle.com