Interpretable AI: Sparse Autoencoders and Feature Extraction
Learn to extract and interpret features from large language models using sparse autoencoders, focusing on scalability and safety implications.
W1
Foundations of Neural Networks and Representation Learning
Understand the linear representation hypothesis, superposition hypothesis, and dictionary learning in neural networks.
3 videos•135m
3 readings
3 topics
1 homework
References
Week 1: Foundations of Neural Networks and Representation Learning
Week 2: Sparse Autoencoders and Feature Extraction
Week 3: Scaling Sparse Autoencoders to Large Models
Week 4: Safety-Relevant Feature Analysis and Model Steering
Rate this course
Community Insights
0Join the discussion
Sign in to share your thoughts and technical insights.
Loading insights...


