Shutterstock: Is Multimodal AI a Data Licensing Problem?
AI Magazine speaks to Dan Mandell, SVP of Data Licensing & AI Services at Shutterstock, about multimodal data, creativity and Shutterstock's efforts in AI. Source: aimagazine.com
66 articles
AI Magazine speaks to Dan Mandell, SVP of Data Licensing & AI Services at Shutterstock, about multimodal data, creativity and Shutterstock's efforts in AI. Source: aimagazine.com
A new study in Nature Communications reports a multimodal large language model (MLLM) designed to assist with esophagogastroduodenoscopy (EGD)—a procedure... Source: bioengineer.org
Xiaomi on Wednesday unveiled Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive foundation model for embodied AI. The company said the. Source: technode.com
SINGAPORE, July 15, 2026 (GLOBE NEWSWIRE) -- Aurora Mobile Limited (NASDAQ: JG) (“Aurora Mobile” or the “Company”), a leading provider of custom... Source: markets.businessinsider.com
Researchers from the MIT-IBM Computing Research Lab and IBM Quantum have developed a multimodal alignment framework that maps quantum unitary operators... Source: quantumcomputingreport.com
A key feature of the Precision Nutrition and Health approach is the ability to tailor interventions to individual variability using multimodal data from... Source: nature.com
Naver turns 27 years of search into AI with tailored LLM, SLMs, multimodal The search infrastructure and know-how accumulated over the past 27 years, Source: biz.chosun.com
Take a practical look at multimodal, any-to-any systems for vision-language reasoning, speech interaction, document intelligence, real-time assistants,... Source: KDnuggets
Google's Gemini Omni Flash ranks first in Video Arena for text-to-video and image-to-video, surpassing competitors with multimodal editing capabilities. Source: cryptobriefing.com
Conducting dynamic exploration in complex and unpredictable environments, particularly in space exploration, reveals the great potential of systems based on... Source: nature.com
Alibaba's Qwen team has released Qwen3.7-Plus, a multimodal agent model that combines visual perception, GUI operation, and coding in a single agent loop. Source: the-decoder.com
The new, open NVIDIA world foundation model brings vision reasoning, multimodal generation and action prediction together to help robots,... Source: blogs.nvidia.com
Text prompts and structural scripts; Images, hand-drawn sketches, and illustrations; Existing video clips as style or structural references... Source: adgully.com
Google DeepMind unveiled Gemini Omni at Google I/O, a multimodal AI model family for video generation with implications for decentralized compute and Web3... Source: cryptobriefing.com
ByteDance released the open-source multimodal model Lance, a native unified system that handles **image and video understanding, generation, and editing**... Source: letsdatascience.com
AI leaders discussed multimodal systems, sensory computing, privacy risks, robotics, and future human-machine collaboration possibilities. Source: forbes.com
ByteDance releases Lance, a 3B native unified multimodal model for image and video understanding, generation, and editing. Source: marktechpost.com
At Google I/O 2026, the tech giant unveiled Gemini Omni, a new multimodal AI model that can create and edit videos using text, images, audio and video... Source: ndtvprofit.com
Google has unveiled Gemini Omni, a new multimodal AI model designed to generate and edit videos using combinations of text, images, audio, and video prompts... Source: indianexpress.com
Google's Gemini Omni is a new multimodal model that reasons across text, images, audio, and video to generate and edit videos through simple conversation... Source: techcrunch.com