#model interpretability
Model Interpretability: 2 AI articles covering model interpretability news, analysis, and research
Articles
Baseten's Base Labs Launches Open-Weight AI Safety Partnershipβ8
Baseten's Base Labs partners with Hugging Face and Goodfire to build open-weight AI safety infrastructure, tackling the rise of abliterated models and unsafe op...
A New Trick Reveals AI Modelsβ Inner Thoughtsβ8
New AI technique exposes hidden reasoning traces in models like Claude and GPT, revealing possible training links to Chinese AI systems and reshaping transparen...
