Recommended References
No video available for this topic. Explore these curated study references:
1 / 3
article
arxiv.org
Steering Language Model Refusal with Sparse Autoencoders
Explore this reference material for in-depth technical documentation and background theory on Feature Interpretability and Limitations.
Source: arxiv.orgRead Reference
Feature Interpretability and Limitations
Learning Objectives
- •Evaluating feature faithfulness
- •Addressing incomplete feature suites
Weekly Outcome
Identify and analyze safety-relevant features, and demonstrate model steering using extracted features.