All news and announcements
| Jul 2026 | Invited speaker at EIML@ICML 2026 (2nd Workshop on Epistemic Intelligence in Machine Learning). Talk: “Understanding model behaviour in the age of AGI.” |
|---|---|
| Jul 2026 | Teaching at two summer schools: invited lecture at the ML Summer School on Reliability and Safety, Kraków (Jul 1–4); and at the Oxford Machine Learning Summer School (OxML 2026), Track 02: Representation Learning & Generative AI (Jul 15–18). |
| Jun 2026 | Several upcoming invited talks: debating AI governance at the Oxford Union (Connected Life Summit, Jun 25); invited talk at ETH Zürich on automated interpretability (Jun 25, remote); speaking at BLISS (Berlin Learning and Intelligent Systems Society) Speaker Series, Berlin (Jun 30); and a talk at the Foresight Secure & Sovereign AI Workshop, Berlin (Jul 18–19). |
| May 2026 | Paper accepted at ACL 2026: Make Mechanistic Interpretability Auditable. |
| May 2026 | 7 papers accepted at ICML 2026, including 2 Spotlights: There Are Futures That Benchmark-Driven AI Cannot See and Don’t Just “Fix it in Post”: A Science of AI Must Study Learning Dynamics. |
| Apr 2026 | New paper: in Science Robotics: Beyond Alignment: Why Robotic Foundation Models Need Context-Aware Safety. |
| Apr 2026 | Gave an invited talk at the Barcelona Supercomputing Centre (Severo Ochoa Research Seminar) and an upcoming seminar at the University of Cambridge (Language Technology Lab) in May. |
| Apr 2026 | Featured in The Independent and the Irish Independent on AI safety and existential risk. |
| Feb 2026 | 2 papers accepted at ICLR 2026. |
| Oct 2025 | Gave a talk at Sogang University |
| Oct 2025 | Gave a talk at the Seoul AI Safety & Security Forum |
| Sep 2025 | 3 papers accepted at NeurIPS 2025. |
| Aug 2025 | 4 papers accepted at EMNLP 2025. |
| Jul 2025 | Speaking at the Human-aligned AI Summer School in Prague, July 22-25, 2025. |
| Jul 2025 | Panelist at the Actionable Interpretability workshop at ICML 2025, Vancouver. |
| Jul 2025 | Speaking at Post-AGI Civilizational Equilibria in Vancouver, co-located with ICML 2025. |
| Mar 2025 | Invited talk on open problems in machine unlearning for AI safety at Ploutos. |
| Mar 2025 | Invited to lecture on AI Safety and Alignment at Oxford Machine Learning Summer School in August. |
| Feb 2025 | I am serving as an Area Chair for ACL 2025. |
| Jan 2025 | Our paper Towards interpreting visual information processing in vision-language models accepted at ICLR 2025. |
| Nov 2024 | Excited to speak about unlearning and Safety at the How To Evaluate AI Privacy Tutorial at NeurIPS 2024. |
| Sep 2024 | Our paper Interpreting LFPs in Large Language Models has been accepted at NeurIPS 2024. |
| Sep 2024 | Two papers accepted at EMNLP 2024. |
| Aug 2024 | Presenting our work on LLMs Relearn Removed Concepts at ACL 2024. ! |
| Aug 2024 | Co-organised the first workshop on Mechanistic Interpretability at ICML 2024. |
| Aug 2024 | Presented 2 main confrence and 1 workshop paper at ICML 2024. |
| Jun 2024 | New paper: Sycophancy to Subterfuge: Investigating Reward Tampering in Language Models |
| May 2024 | Our paper on how LLMs relearn removed concepts has been accepted at ACL 2024. |
| May 2024 | I gave a talk about Machine Unlearning at Foresights AGI workshop |
| May 2024 | Two papers accepted at ICML 2024 and excited to be co-organising the first MI workshop in Vienna : ICML 2024 |
| Apr 2024 | Serving as a Programme Committee at ECAI 2024 : 27th European Conference on Artificial Intelligence |
| Apr 2024 | Invited speaker at AI Safety Colloquium at KIAST, South Korea : Introduction to AI safety: Can we remove undesired behaviour from AI. |
| Mar 2024 | Gave a keynote at EACL 2024 Personalization Workshop in Malta : PERSONALIZE @ EACL 2024 |
| Jan 2024 | Our paper has been acceoted at ICLR 2024. See you In Vienna : Understanding Addition In Transformers |
| Jan 2024 | New paper: : Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training |
| Jan 2024 | New paper: : Can language models relearn removed concepts. |
| Dec 2023 | Presented Measuring Value Alignment at NeurIPS 2023, New Orleans : Presentation |