AIskimIQ

Daily AI & tech news brief

Archive/large language models

🧠 Large Language Models

News about foundation models, LLMs, multimodal models, benchmarks, and model releases from OpenAI, Anthropic, Google, Meta, Mistral, and others.

973 articles

Importance:Researchtraining optimization

Host Offloading Eases GPU Memory Strain in JAX-Based LLM Training

Training large language models increasingly hits GPU memory limits before compute capacity is fully used, due to the size of model weights, gradients, and optimizer states. A new approach uses host offloading in JAX to reduce these high-bandwidth memory bottlenecks. Source: developer.nvidia.com

Importance:Newsbenchmarks

Rethinking How We Measure LLM Performance

As large language models continue to spread worldwide, experts are reassessing which benchmarks actually capture meaningful performance. The piece takes stock of current LLM evaluation methods and their limitations. Source: cacm.acm.org

Importance:Researchagents

ARTEM Gives LLM Agents a Sense of Space and Time

Researchers unveiled ARTEM (Agentic Retrieval with Temporal-Episodic Memory), a hybrid agent architecture that combines LLMs with a self-organizing neural memory system. The design aims to give AI agents better spatial and temporal episodic memory for more context-aware reasoning. Source: ojs.aaai.org

Importance:Newsspace LLM deployment

NASA sends Google's Gemma language model into orbit

NASA has deployed Google's Gemma LLM in space, testing the concept amid ongoing debate over whether orbital data centers can realistically host the largest and most capable AI models. Source: aol.com

Importance:Newsrobotics AI model

TurboVLA matches 7B robot AI performance without a language model, runs at 32Hz on consumer GPU

For three years, vision-language-action (VLA) research has assumed that top-tier robot AI models need a large language model at their core. TurboVLA challenges that assumption by achieving comparable results without one. Source: techtimes.com

Importance:NewsLLM security vulnerability

A core design flaw leaves LLMs surprisingly easy to exploit

Researchers found a fundamental weakness that makes it simple to manipulate LLMs into producing harmful outputs, including instructions for sabotaging an aircraft's navigation system. Source: technologyreview.com

Importance:PolicyLLM usage policy

Debian weighs five competing proposals on AI and LLM use in the project

Debian developers last week opened discussion on a general resolution addressing how AI large language models should be used within the project, with five different proposals now on the table. Source: phoronix.com

Importance:Launchconversational AI

PolyAI unveils real-time voice model aiming for more human-like AI calls

PolyAI Ltd. announced Dialog-RSN-1, a new voice dialog AI model designed to make conversations with AI-driven phone agents sound more natural and human. Source: siliconangle.com

Importance:LaunchLLM runtime tool

JetBrains open-sources KotlinLLM, a runtime code generator

JetBrains released an experimental IntelliJ IDEA plugin that lets Kotlin/JVM projects hand off runtime logic to an LLM directly from Kotlin code. Source: infoworld.com

Importance:LaunchLLM security

Cequence AI Gateway adds LLM governance to close model access gaps

The new governance layer gives enterprises a single controlled path for every tool call, API request, agent interaction, and LLM prompt, as companies increasingly seek to manage the tools and APIs used by AI agents. Source: securityboulevard.com

Importance:Newsopen-source models

After Kimi K3's open-source release, top-tier open models look increasingly realistic

The open-sourcing of Kimi K3 suggests that AI developers' strategy of releasing highly optimized, high-performance models as open source is gaining real momentum. Source: eu.36kr.com

Importance:Researchbenchmarks

New African LLM Benchmark Stress-Tests AI Safety Across Languages

Backed by the GSMA, the African Trust & Safety LLM Challenge has released a benchmark of 4,216 verified, reproducible tests designed to probe AI safety across diverse African languages and cultural contexts. The project aims to fill gaps in safety evaluation for regions underrepresented in existing LLM benchmarks. Source: thefastmode.com

Importance:ResearchLLM applications

LLM-Guided Search Algorithm Improves Satellite Scheduling

Researchers introduced a Dual-Population Co-Evolutionary (DPEC) framework that lets an LLM-driven population of algorithms learn from expert-designed strategies to solve complex scheduling problems. The approach was applied to agile earth observation satellite scheduling, combining adaptive large neighborhood search with LLM assistance. Source: eurekalert.org

Importance:Researchedge AI

Developer Runs 28.9M-Parameter LLM on an $8 Microcontroller

GitHub user Slava S. built a 28.9-million-parameter language model capable of generating text on the ESP32-S3, a microcontroller that costs roughly $8. The project demonstrates how far small-scale LLMs can be squeezed onto cheap embedded hardware. Source: blog.adafruit.com

Importance:Launchhealthcare AI

Truveta's Language Model Extracts Cancer Staging Data from Millions of Clinical Notes

A study published in JCO Clinical Cancer Informatics shows that the Truveta Language Model can accurately pull critical cancer staging details from vast collections of unstructured clinical records. The research highlights a scalable way to turn raw medical notes into research-ready data. Source: globenewswire.com

Importance:Launchautomotive AI

OMODA Launches AI-Powered Cockpit in Southeast Asia Using ByteDance's Seed LLM

OMODA unveiled its Super AI Cockpit in Jakarta, built on ByteDance's Seed LLM, positioning it as a next-generation in-car experience aimed at younger drivers. The announcement frames the cockpit as a key part of the brand's push into AI-driven mobility in the region. Source: manilatimes.net

Importance:LaunchEU LLM and multilingual benchmark

EU Commission unveils its own LLM and benchmark for European languages

The European Commission's DG Translation has released an open-source LLM trained on multilingual data, aimed at supporting EU languages that are often underrepresented in mainstream AI models. The project also includes a benchmark to evaluate model performance across the bloc's official languages. Source: slator.com

Importance:NewsLLM training litigation and settlement

Anthropic's $1.5 billion settlement over pirated books used to train Claude

A US court has approved Anthropic's $1.5 billion settlement with a group of authors whose pirated books were used to train the Claude LLM. The case highlights growing legal risks for AI companies over how training data is sourced. Source: business-standard.com

Importance:Opinionmodel distillation and curriculum learning

Distillation explained: when training data itself becomes the teacher

As models grow capable of generating their own training curricula, data is shifting from being a static resource to acting as a medium that transmits intelligence between models. This shift reshapes how distillation techniques are understood in modern ML. Source: thesequence.substack.com

Importance:Newsencrypted AI benchmark

DESILO's THOR named reference model for encrypted AI in global FHE benchmark

South Korean company DESILO announced that its THOR system was chosen as the reference implementation for encrypted AI within a new global benchmarking suite for fully homomorphic encryption (FHE). The initiative aims to set a standard for evaluating private AI and encrypted LLM inference. Source: manilatimes.net