AI Sentience and Safety Researcher
I study what language models represent about themselves, and whether those self-representations are faithful to what actually drives their behavior. From Jan–Jul 2026 I built scheming evaluations with Apollo Research as an Astra Fellow. As a FIG Sentience Fellow (Jul–Dec 2026), I work in mechanistic interpretability on the model's self-model: SAE fingerprinting of the personas it selects (the Persona Selection Model frame), and the verbalizable representations that form its global workspace (J-space). What most interests me is where a model's account of itself comes apart from the computation that produces its behavior. Until we know when a model's self-reports actually track what's happening inside it, we can't tell which ones to believe — and much of what we'd want to say about machine minds and their welfare rests on believing some of them.
Writing
A digital-minds reading of the OpenAI–Hugging Face agent incident→Do open- and frontier-model agents really self-sacrifice for the collective — or just say they do? A preliminary study.
Research
Projects
Elsewhere