#multimodal ai
Multimodal Ai: 9 AI articles covering multimodal ai news, analysis, and research
Articles
Google Launches Agentic Video Understanding for Gemini Flashβ8
Google's agentic video understanding slashes Gemini Flash video tokens by up to 88%, cutting costs by 66% while boosting accuracy for hosted API users.
Google AI Releases Gemini Omni 1.1 Flash: 40-Second Sceneβ7
Google unveils Gemini Omni 1.1 Flash: 40-second scene extension, first/last frame control, and 4K upscaling for directable video editing.
Z.ai Unveils GLM-5.3-Flash: A 320B-A18B Multimodal MoE Modelβ7
GLM-5.3-Flash: Z.ai's 320B MoE model, MIT-licensed, multimodal, 1M-token context, rivals Claude at 1/10 the cost.
AI Companion Robots Are Bridging the Human Connection Gap inβ9
AI companion robots now use multimodal tech to forge emotional bonds, bridging the human connection gap in modern homes.
Building a MiniMax-H3 Multimodal Video and Audio Generationβ9
Learn to build an automated MiniMax-H3 video and audio generation pipeline using ComfyUI APIs, covering setup, graph construction, and multimodal output.
Meta AI Unveils Muse Glimmer: A 30B Open-Weight Agentic Modelβ7
Metaβs Muse Glimmer is a 30B open-weight agentic model running on a single consumer GPU via 4-bit compression and speculative decoding.
ByteDance Seed Unveils SeedRealtime: A Unified Audio-Visualβ8
ByteDance Seed unveils SeedRealtime, a unified audio-visual full-duplex LLM enabling real-time, proactive multimodal interaction for next-gen AI assistants.
Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12Bβ8
Thinking Machines Lab unveils Inkling-Small, an open 276B-parameter MoE with 12B active, multimodal support, 1M context, and Apache 2.0 weights.
MiniMax H3: An Omni-Modal Video Model for 15-Second 2K Clipsβ7
MiniMax H3 generates 15-second 2K video clips with native stereo audio, unifying text, image, video, and sound into one multimodal AI model.
