Conceptual
Login

About Look inside a language model and judge whether it is safe: interpretability, alignment, and AI control

Copied

Learn how researchers look inside a large language model to find the features behind its answers, and how they test whether it is faking alignment. You will be able to explain sparse autoencoders, circuit tracing, AI control and jailbreak classifiers, and weigh claims about model welfare.

What you will learn

No introduction video available

About Zoe Graystone

Z

Guide profile coming soon.