AIskimIQ

Daily AI & tech news brief

Archive/benchmark

Tagged: Benchmark

47 articles

Importance:News

Mysterious AI Video Model 'Happy Horse 1.0' Dominates Leaderboard

A previously unknown AI video model called 'Happy Horse 1.0' has suddenly risen to the top of the Artificial Analysis Video Arena leaderboard,... Source: nationaltoday.com

Importance:News

Claude’s 20-Hour Psych Eval Fuels Model Ethics Debate

Anthropic's 20-hour AI psychiatry evaluation ignites vital Model Ethics debate, guiding cybersecurity strategy and governance decisions. Source: aicerts.ai

Importance:News

Mystery AI Video Generator Happy Horse 1.0 Reaches No. 1, Surpasses Sora, Veo

Anonymous text-to-video model leads Artificial Analysis' blind benchmark by 101 Elo points across nearly 8,000 user comparisons, surpassing all major rivals... Source: usatoday.com

Importance:News

Alibaba’s new AI video-generation model tops global ranking after debut

Alibaba Group's new artificial intelligence video-generation tool has taken the top spot in a global leaderboard that tracks AI models' abilities,... Source: msn.com

Importance:NewsLLM evaluation tools

How to choose the best LLM using R and vitals

Use the vitals package with ellmer to evaluate and compare the accuracy of LLMs, including writing evals to test local models. Large language models, LLMs. Source: infoworld.com

Importance:Launchefficient LLM model release

PrismML Releases 1-Bit Bonsai 8B Model

PrismML, a Caltech spinout, released Bonsai 8B on April 4, 2026, a 1-bit large language model that fits in 1.15 GB of memory and claims benchmark parity... Source: letsdatascience.com

Importance:Newsbenchmarking

Nebius partners with Positronic on Physical AI Leaderboard (PhAIL)

Physical AI is moving from controlled demos to real-world deployment — and that shift demands benchmarks grounded in actual operations, not lab conditions. Source: nebius.com