Articles
Detecting and Controlling Sycophancy with Cascading Linear Features⭐8
This paper introduces an iterative data pipeline that isolates cascading linear features to detect and control sycophancy in LLMs, offering more interpretable a...
CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework with Evidence-Grounded Reasoning⭐7
CaVe-VLM-CoT introduces an interpretable vision-language model framework using evidence-grounded reasoning to reduce hallucinations and improve trust in AI outp...
