EulerFold
Recommended References

No video available for this topic. Explore these curated study references:

1 / 3
Steering Language Model Refusal with Sparse Autoencoders
articlearxiv.org

Steering Language Model Refusal with Sparse Autoencoders

Explore this reference material for in-depth technical documentation and background theory on Feature Interpretability and Limitations.

Source: arxiv.orgRead Reference

Feature Interpretability and Limitations

Learning Objectives

  • Evaluating feature faithfulness
  • Addressing incomplete feature suites

Weekly Outcome

Identify and analyze safety-relevant features, and demonstrate model steering using extracted features.