Loading...

Interpretable AI: Sparse Autoencoders and Feature Extraction

Learn to extract and interpret features from large language models using sparse autoencoders, focusing on scalability and safety implications.

Created Bysankalpsankalp
4 weeks
0 Learners
Aug 21
to start learning
Curriculum

Share:

W1

Foundations of Neural Networks and Representation Learning

Understand the linear representation hypothesis, superposition hypothesis, and dictionary learning in neural networks.

3 videos135m
3 readings
3 topics
1 homework
Learn

Topics

1.1
Linear Representation Hypothesis
Feature learning & the linear representation hypothesis for steering & monitoring LLMs
29 minutes
1.2
Superposition Hypothesis
Superposition in LLM Feature Representations | Boluwatife Ben-Adeola | Conf42 LLMs 2024
47 minutes
1.3
Dictionary Learning
Collaborative Dictionary Learning from Big, Distributed Data
59 minutes
W2

Sparse Autoencoders and Feature Extraction

Implement and train sparse autoencoders to extract interpretable features from model activations.

3 videos108m
3 readings
3 topics
1 homework
Learn
W3

Scaling Sparse Autoencoders to Large Models

Apply scaling laws to train sparse autoencoders on large language models and analyze feature properties.

3 videos67m
3 readings
3 topics
1 homework
Learn
W4

Safety-Relevant Feature Analysis and Model Steering

Identify and analyze safety-relevant features, and demonstrate model steering using extracted features.

3 videos94m
3 readings
3 topics
1 homework
Learn
Rate this course
0.0
0 reviews

Community Insights

0

Join the discussion

Sign in to share your thoughts and technical insights.

Loading insights...