All news and announcements

Jul 2026 Invited speaker at EIML@ICML 2026 (2nd Workshop on Epistemic Intelligence in Machine Learning). Talk: “Understanding model behaviour in the age of AGI.”
Jul 2026 Teaching at two summer schools: invited lecture at the ML Summer School on Reliability and Safety, Kraków (Jul 1–4); and at the Oxford Machine Learning Summer School (OxML 2026), Track 02: Representation Learning & Generative AI (Jul 15–18).
Jun 2026 Several upcoming invited talks: debating AI governance at the Oxford Union (Connected Life Summit, Jun 25); invited talk at ETH Zürich on automated interpretability (Jun 25, remote); speaking at BLISS (Berlin Learning and Intelligent Systems Society) Speaker Series, Berlin (Jun 30); and a talk at the Foresight Secure & Sovereign AI Workshop, Berlin (Jul 18–19).
May 2026 Paper accepted at ACL 2026: Make Mechanistic Interpretability Auditable.
May 2026 7 papers accepted at ICML 2026, including 2 Spotlights: There Are Futures That Benchmark-Driven AI Cannot See and Don’t Just “Fix it in Post”: A Science of AI Must Study Learning Dynamics.
Apr 2026 New paper: in Science Robotics: Beyond Alignment: Why Robotic Foundation Models Need Context-Aware Safety.
Apr 2026 Gave an invited talk at the Barcelona Supercomputing Centre (Severo Ochoa Research Seminar) and an upcoming seminar at the University of Cambridge (Language Technology Lab) in May.
Apr 2026 Featured in The Independent and the Irish Independent on AI safety and existential risk.
Feb 2026 2 papers accepted at ICLR 2026.
Oct 2025 Gave a talk at Sogang University
Oct 2025 Gave a talk at the Seoul AI Safety & Security Forum
Sep 2025 3 papers accepted at NeurIPS 2025.
Aug 2025 4 papers accepted at EMNLP 2025.
Jul 2025 Speaking at the Human-aligned AI Summer School in Prague, July 22-25, 2025.
Jul 2025 Panelist at the Actionable Interpretability workshop at ICML 2025, Vancouver.
Jul 2025 Speaking at Post-AGI Civilizational Equilibria in Vancouver, co-located with ICML 2025.
Mar 2025 Invited talk on open problems in machine unlearning for AI safety at Ploutos.
Mar 2025 Invited to lecture on AI Safety and Alignment at Oxford Machine Learning Summer School in August.
Feb 2025 I am serving as an Area Chair for ACL 2025.
Jan 2025 Our paper Towards interpreting visual information processing in vision-language models accepted at ICLR 2025.
Nov 2024 Excited to speak about unlearning and Safety at the How To Evaluate AI Privacy Tutorial at NeurIPS 2024.
Sep 2024 Our paper Interpreting LFPs in Large Language Models has been accepted at NeurIPS 2024.
Sep 2024 Two papers accepted at EMNLP 2024.
Aug 2024 Presenting our work on LLMs Relearn Removed Concepts at ACL 2024. !
Aug 2024 Co-organised the first workshop on Mechanistic Interpretability at ICML 2024.
Aug 2024 Presented 2 main confrence and 1 workshop paper at ICML 2024.
Jun 2024 New paper: Sycophancy to Subterfuge: Investigating Reward Tampering in Language Models
May 2024 Our paper on how LLMs relearn removed concepts has been accepted at ACL 2024.
May 2024 I gave a talk about Machine Unlearning at Foresights AGI workshop
May 2024 Two papers accepted at ICML 2024 and excited to be co-organising the first MI workshop in Vienna : ICML 2024
Apr 2024 Serving as a Programme Committee at ECAI 2024 : 27th European Conference on Artificial Intelligence
Apr 2024 Invited speaker at AI Safety Colloquium at KIAST, South Korea : Introduction to AI safety: Can we remove undesired behaviour from AI.
Mar 2024 Gave a keynote at EACL 2024 Personalization Workshop in Malta : PERSONALIZE @ EACL 2024
Jan 2024 Our paper has been acceoted at ICLR 2024. See you In Vienna : Understanding Addition In Transformers
Jan 2024 New paper: : Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Jan 2024 New paper: : Can language models relearn removed concepts.
Dec 2023 Presented Measuring Value Alignment at NeurIPS 2023, New Orleans : Presentation