My research focuses on understanding, evaluating, and controlling advanced AI systems. I am particularly interested in failures that ordinary evaluations can miss: models that retain behaviours we thought we had removed, appear safe during testing but behave differently in practice, or give confident answers that are not supported by what they actually processed.

I also work on connecting technical AI safety research to governance. Interpretability and evaluations are most useful when they produce evidence that people can act on: evidence that allows auditors to test safety claims, regulators to set meaningful standards, and institutions to determine responsibility when systems fail. More broadly, I study how increasingly capable AI systems could weaken human agency, and what technical and institutional safeguards could help keep people in control.

Some of the questions I am interested in include:

  • When a model behaves unexpectedly, how can we identify the internal mechanisms responsible?
  • How can we detect deception, reward manipulation, or other dangerous behaviours that ordinary evaluations may miss?
  • How can we tell whether a safety intervention changed a model, rather than merely making an unwanted behaviour harder to observe?
  • What evidence would allow independent auditors and regulators to verify claims about an AI system’s safety?
  • How can increasingly capable AI systems be developed and governed without gradually shifting important decisions away from people?

My work has been published at NeurIPS, ICML, ICLR, ACL, EMNLP, and FAccT. For more detail, see my Research Agenda and recent papers.