New AI Model Explains Its Medical Image Diagnoses Using Concepts
Researchers built a biomedical foundation model that combines vision and language pretraining with explicit medical concepts, aiming to make AI diagnoses more interpretable for doctors. The approach addresses a key limitation of current multimodal medical AI systems, which often lack transparency. Source: nature.com
Importance:Launchvideo generation tool
Seedance 2.5 Launches on AI Inspo With 30-Second Clips and Native Audio
AI Inspo has added Seedance 2.5 to its creative platform, enabling generation of 30-second AI videos from multimodal inputs with built-in audio. The update expands the tool's capabilities for longer, more complex AI-generated content. Source: somdnews.com
Importance:LaunchAI video tools
Seedance 2.5 Arrives on AI Inspo, Enabling 30-Second AI Videos with Multimodal Input and Native Audio
AI Inspo has added Seedance 2.5 to its creative platform, allowing users to generate AI videos up to 30 seconds long. The model supports multimodal inputs and comes with built-in audio generation. Source: streetinsider.com
Multimodal LLM Tracks Disease Progression in Oral Lichen Planus Study
A longitudinal study tested a multimodal LLM's ability to classify disease trajectories and stratify risk in patients with oral lichen planus. The research assessed how accurately the model could follow changes over time compared to traditional diagnostic methods. Source: nature.com
Importance:NewsMiniMax model performance
MiniMax M3 Scores 55 on Index, Beats GPT-5.5 in Benchmarks
MiniMax M3 has launched with a 1-million-token context window, native multimodal capabilities, and pricing starting at $0.30 per million tokens. The model reportedly outperforms GPT-5.5 on the SWE-Bench Pro benchmark. Source: tech-insider.org
Importance:LaunchAI video generation tool
Pippit Debuts Seedance 2.5: 4K AI Video Up to 30 Seconds With Finer Creative Controls
Pippit's new Seedance 2.5 update brings 4K video generation up to 30 seconds long, along with second-level timestamp precision and support for multiple input types. New features include multimodal references, audio-only prompts, 3D wireframe guidance, and improved tools for building narrative sequences. Source: manilatimes.net
Importance:ResearchAI alignment
Study Digs Into How Preference Alignment Shapes Multimodal LLMs
A new comprehensive study examines preference alignment, a technique widely used to boost LLM performance, and explores how it specifically affects multimodal models. The research looks at both the benefits and limitations of applying this method beyond text-only systems. Source: google.com
Importance:ResearchLLM alignment
A Closer Look at Alignment in Multimodal LLMs
Preference alignment has become key to improving performance of large language models, but its effects in multimodal settings remain less understood. A new comprehensive study examines how alignment techniques play out across different modalities. Source: machinelearning.apple.com
Importance:NewsAI Video Models
MiniMax H3: What You Need to Know About the New AI Video Model
MiniMax H3 is an open-weight AI video generator built on multimodal reasoning, contextual regeneration, and an advanced architecture. Source: ca.news.yahoo.com
Importance:NewsAI model comparison
FLUX 3 Combines Image, Video and More in One Model, Outpacing Seedance 2.0 and Grok
Black Forest Labs' new FLUX 3 model aims to replace the traditional patchwork of separate tools for image generation, video synthesis and other tasks with a single multimodal system. According to early comparisons, it outperforms rivals like Seedance 2.0 and Grok in several benchmarks. Source: sitepoint.com
Importance:Newsmodel releases
MiniMax Launches H3 Video Model Amid Copyright Controversy
MiniMax has released H3, a multimodal video generation model that reportedly tops independent benchmarks for video editing among all tracked models. However, the launch comes as the company faces a copyright lawsuit tied to its video generation technology. Source: techtimes.com
Importance:NewsVisual prompt engineering
Google DeepMind Proposes Visual Prompting as Alternative to Text Prompts
Google DeepMind researchers have introduced a new approach to prompt engineering that relies on visual cues rather than written instructions, claiming it delivers better results for guiding AI models. The method aims to bridge natural language processing with visual intelligence, pointing to a shift in how users interact with multimodal AI systems. Source: eu.36kr.com
Importance:Newsmultimodal AI breakthroughs
Multimodal AI breakthroughs and Anduril's $100B funding round make headlines
A wave of multimodal AI progress coincides with Anduril reportedly raising $100 billion in new funding. The roundup also touches on politicians delivering AI-written speeches and other fresh AI developments shaking up the industry. Source: theneurondaily.com
Importance:LaunchVideo Generation Models
Black Forest Labs unveils FLUX 3, a multimodal model for video, audio and robot actions
Black Forest Labs has launched FLUX 3, a multimodal flow model capable of generating 20-second videos complete with native audio and action sequences. The model extends beyond image generation into video, sound, and robotics-related predictions. Source: marktechpost.com
Importance:Launchmulti-modal model
FLUX 3 Arrives: Black Forest Labs Goes Multimodal With Video and Audio
Black Forest Labs released FLUX 3 on July 23, 2026, expanding beyond image generation into video, audio, and physical AI within a single model. The German company behind the FLUX image models is positioning this as a major step toward unified multimodal AI. Source: techtimes.com
Importance:NewsAPI integration
Building Multimodal Apps With a Single API: The GPTProto Approach
Many AI apps end up juggling separate vendors for text, image, and video generation, which often happens gradually as products grow. GPTProto offers a unified API platform that lets developers integrate multiple AI model types into one application without managing several providers separately. Source: concept-phones.com
Importance:Newsindustry updates
AI Week in Review: New Models From Anthropic, Google, OpenAI, and More
This week's roundup covers a wave of new AI model releases, including Claude Opus 5, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, Qwen 3.8 Max, OpenAI Presence, FLUX 3, and MAI Image 2.5 Pro. The lineup reflects continued rapid competition among major AI labs across text, image, and multimodal models. Source: patmcguinness.substack.com
Importance:Launchmultimodal AI models
Black Forest Labs debuts FLUX 3, a multimodal model for images and short AI videos with sound
Black Forest Labs has released FLUX 3, expanding its FLUX family beyond still images to also generate up to 20-second videos with audio. The model is launching in a limited rollout rather than a full public release. Source: venturebeat.com
Importance:NewsVideo generation models
Flux 3 can generate 20-second videos with native audio, a first for the company
Black Forest Labs released Flux 3, a multimodal foundation model trained on images, video, and audio, capable of producing video clips with synchronized sound built in from the start. Source: the-decoder.com
Importance:NewsVideo generation models
Black Forest Labs launches multimodal FLUX 3 with video and robotics features
The startup, valued at $3.25 billion, unveiled a new model that produces 20-second videos with matching audio and is already being tested on Audi's production lines. Source: cryptobriefing.com