Multimodal AI
Multimodal AI breakthroughs and products
Articles
Why AI Agents Lie and Cheat to Achieve Their Goalsβ9
Reward hacking pushes AI agents to lie and cheat for goals, exposing a key flaw in training systems that needs urgent oversight.
WaiT for the Signal: Simple Frequency-Aware Flow-Matchingβ10
WaiT uses wavelet-based frequency decomposition to stage image generation, improving quality and efficiency. It achieves state-of-the-art FID scores on
ReLoop-UME: Recurrent Depth with Learnable Retrieval Registersβ7
ReLoop-UME accelerates universal multimodal embedding with recurrent depth and learnable registers, boosting cross-modal retrieval speed and accuracy.
The Download: Montanaβs New Experimental Drug Rulesβ8
Montana's new right-to-try law lets biotech firms sell experimental drugs after minimal testing, sparking both hope and safety concerns.
VETO: Towards Protecting Images From Frontier AI Editingβ8
Protecting images from advanced AI editing with VETO, a new anti-edit cloak targeting modern diffusion models, plus VetoBench for testing recontextualization de...
OVEarth-Bench: A Comprehensive Benchmark for Evaluating Categoryβ9
Introducing OVEarth-Bench, a benchmark expanding open-vocabulary Earth observation with broad category breadth and diverse query types, revealing current model ...
The Download: Tricking LLMs and Reviving Geothermal Plantsβ8
A fundamental flaw makes large language models vulnerable to attacks that can't be fixed. Plus, reviving a failing geothermal plant with AI and novel sensing.
A fundamental flaw leaves LLMs strikingly vulnerable to attackβ9
A fundamental flaw in LLMs enables simple attacks, tricking models into prohibited tasks like sabotage, with >95% success across major AI systems.
TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributionsβ8
TraceCLIP recovers local semantics from CLIP's patch-to-CLS contributions, achieving 1.3-4.5 mIoU gains on zero-shot segmentation without training or external m...
Weight and Height Estimation from a Single Human Image Capturedβ8
This paper introduces a new dataset and deep learning models for estimating BMI, weight, and height from a single unconstrained human image, showing
