Inside AI's Black Box: Interpretability, Chain-of-Thought, and Safety with Neil Nanda

1 件の動画 · 更新: 3時間前
Understanding the inner thoughts of AI 📺 Understanding the inner thoughts of AI ⏱ 53:05📅 2026/07/18 05:33

Inside AI's Black Box: Interpretability, Chain-of-Thought, and Safety with Neil Nanda

Google DeepMind podcast host Hannah Fry speaks with Neil Nanda, who leads the language model interpretability team, about how researchers try to understand the internal workings of neural networks and why this matters for AI safety.

The discussion covers the field's goals and methods, from chain-of-thought monitoring to mechanistic techniques, and considers their current usefulness and future limitations.

■ Interpretability fundamentals
- What interpretability means and how it relates to neuroscience and evolution
- Motivations around AI safety, debugging, and scientific understanding

■ Techniques and tools
- Chain-of-thought or scratchpad reasoning as a monitoring method and its limitations
- Probes, sparse autoencoders, and steering as ways to examine internal representations

■ Safety applications and challenges
- Detecting deception, misuse, and hidden objectives, plus auditing model alignment
- Evaluation awareness, model behavior in tests, and pragmatic expectations for interpretability

It is suited to viewers seeking an accessible overview of interpretability research and its role in AI safety, without requiring prior technical expertise. The conversation outlines key concepts and open questions rather than providing a step-by-step technical guide.

📄 このページの紹介文は AI が独自に生成したものであり、著作権をはじめとする他者の権利(商標権・名誉権・プライバシー等)を侵害しないよう配慮しています。動画の著作権は各作成者に帰属します。