Multimodal AI
Multimodal AI breakthroughs and products
Articles
The Download: How People Really Use AI, and Flockβs Design Choicesβ8
New research from the AI Observatory reveals how people actually use AI, uncovering sensitive behaviors hidden in major companies' reports.
AI's Recursive Self-Improvement May Not Arrive as Fast as Predictedβ8
AI's recursive self-improvement may arrive slower than expected due to creativity limits and research gaps, despite industry optimism.
We Still Donβt Know How People Are Really Using AIβ10
New study reveals AI use is far more personal than work-focused, challenging tech giants' productivity claims with surprising insights.
Xemo-Talker: Unlock Emotions Explicitly for Audio-Driven Talkingβ9
Xemo-Talker enables precise emotion control in audio-driven talking portraits via less-principal subspace supervision, balancing lip sync and expression quality...
What Flock's Defenders Overlook in the Surveillance Debateβ9
Flock Safety defenders overlook key design choices in surveillance debates, from data retention to cross-jurisdiction sharing and privacy risks.
The Download: When Robot Friends Die and the Rise of theβ7
This is today's edition of The Download , our weekday newsletter that delivers a daily dose of what's happening in the world of technology.
What Happens When a Kid's Robot Best Friend Dies?β9
The Fragile Bond of AI Toy Companions: What happens when a child's robot best friend breaks or the company dies? Exploring the emotional fallout.
How Much Hydrogen Awaits Us Underground?β8
Global race for underground hydrogen heats up, but estimates vary. Explore how much natural hydrogen may await and the challenges of tapping this clean fuel sou...
Can Vision-Language Models Assess Proxemic Risk from Egocentricβ10
Current VLMs show limits in egocentric proxemic risk assessment, though targeted prompts and fine-tuning boost high-danger recall in models like Qwen-VL.
HIMEC: Directional Change Representation and Fixed-Interfaceβ9
HIMEC introduces directional change representation and fixed-interface decoding for remote sensing image change captioning, achieving superior CIDEr
