#clip
Clip: 5 AI articles covering clip news, analysis, and research
Articles
FailSAE: Interpretable Failure Prediction for Vision-Language⭐9
Discover how sparse autoencoders enable interpretable failure prediction in vision-language models, revealing why errors occur for risk-aware AI decisions.
Does Marginal Coverage Guarantee Class-Conditional Safety for⭐9
Marginal conformal coverage fails to ensure class-conditional safety for zero-shot VLMs under shift, with worst-class coverage near zero on ImageNet-Sketch desp...
LEGO: Leveled Language Gaussian Splatting⭐10
LEGO introduces a leveled language Gaussian splatting framework for open-vocabulary 3D scene understanding with hierarchical segmentation and spatial reasoning.
Pixel-Native RAG: A Practical Guide to Visual Document Indexing⭐10
Build a pixel-native RAG pipeline: render docs as images, tile, embed with SigLIP/CLIP, FAISS search, OCR-BM25 fusion, and Qwen3-VL grounded answers.
TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions⭐8
TraceCLIP recovers local semantics from CLIP's patch-to-CLS contributions, achieving 1.3-4.5 mIoU gains on zero-shot segmentation without training or external m...
