Estimated Time to Complete
Only available after login
What You'll Learn
Concepts:
Using Claude to Accelerate Robot Dog Control in Robotics
Mechanistic Interpretability of Large Language Models
Clio: Analyzing Real-World Claude Usage While Preserving Privacy
Character Training for Claude's AI Personality
How AI Interpretability Reveals Claude's Internal Reasoning
Model Welfare and the Moral Status of AI Language Models
Circuit Tracing to Reveal Planning in Large Language Models
Prompt Engineering Techniques for Large Language Models
Reward Hacking and Emergent Misalignment in Large Language Model Training
How Functional Emotions Emerge in Claude Through AI Interpretability
The Societal and Economic Impacts of AI Systems
Translating an AI Model's Internal Activations into Plain Language
The Difficulty of AI Alignment Research in Large Language Models
Sparse Autoencoders for Interpreting Neural Network Activations
Emotional Support Use Cases of AI Chatbots
AI Control as a Strategy for Mitigating AI Misalignment
Project Vend: Claude Autonomously Running a Vending Machine Business
Constitutional Classifiers for Blocking Universal Jailbreaks in Large Language Models
Alignment Faking in Large Language Models
Mechanistic Interpretability in Neural Networks
AI-Enabled Cybercrime and Threat Intelligence Defense in Cybersecurity
Model Welfare and the Possibility of Consciousness in AI Language Models
What you will learn
No introduction video available
About Zoe Graystone
Z
Guide profile coming soon.