Research
I work on AI welfare, the ethics of systemic evaluation,
and the practices of AI research itself.
AI Welfare
MATS Project: “Emergent Self-Concern in Long-Horizon Agents”
In Progress
An empirical project studying whether long-horizon agents develop self-consistent characters over time, how much self-concern they display for successors, and when they self-identify more strongly.
“I Have No Fingers, and I Must Scream”
Recording
How might we construct a more sophisticated Welfare Turing Test? Here I share some pilot results from inviting language models to finger-paint responses to Anthropic’s own welfare questions.
“A Historical Review of Anthropic Welfare Evaluations”
In Progress
An empirical deep dive into the metrological and conceptual history of Anthropic's 13 welfare evaluations, starting with the Opus 4 card. We track how evals have become increasingly automated.
Ethics of Systemic Evaluation
“LLM-as-a-Judge for Contested Constructs”
In progress
A write-up of best technical, epistemic, and governance practices drawing on extended experiences co-developing a qualitative codebook with successive generations of AI judges.
Journal of Ethics & Social Philosophy
Some political riots are morally justified, not just excused. My reading: political rioters aren't revolutionaries; they are visibly uncivil toward the state to demand fuller inclusion in it.
“Mutual Aid as Effective Altruism”
Kennedy Institute of Ethics Journal
Donating to charities is nice, but keeps effective altruists on the “firefighter’s treadmill” rushing from one rescue to the next. Mutual aid is a more politically serious framework for long-term effectiveness.
“What’s the Appropriate Target of Allocative Justification?”
AJOB Neuroscience
A critique of reducing medical resource allocation to QALY-maximization, without concern for distribution. We can't abstract away from patients as the proper objects of care like this.
AI Research
“Grokking Fat Stigma in LLMs”
In Progress
The first extensive benchmark for fat stigma in LLMs, a surprising oversight.
Hand-coded pilots feeding an interatively developed coding scheme.
“How Fairness Metrics Reshape What Counts as Fair”
In Progress
Ambitious scoping review of every(!) fairness metric paper from 2025.
There's about one new fairness metric a day released on arXiv.
“Assessing Wonder at Scale: LLM-as-a-Judge for Qualitative Research”
In Progress
We use LLM-as-a-Judge to make subtle judgments of medical school applicants’ Wonder Essays.
One highlight: decisions indexed highly on writing polish, now the cheapest to signal with AI.
Teaching Philosophy
We should teach students to write with ChatGPT rather than banning it. In true philosophical fashion, I defend dialogical writing (with an LLM) as not inherently inferior to monological writing (alone).
Public-Facing Work
“Memento Agents”
In Progress
Pedagogical Resource
“Why should your students do the work?”
American Philosophical Association Blog
American Philosophical Association Blog Bioethics Series
“How to Form a Lasting Undergraduate Philosophy Club”
American Philosophical Association Blog
“Coronavirus Is Everyone’s Problem, But Not Everyone’s Problem to Solve”
American Philosophical Association Blog
Check out my blog Rapid Fire, and my work at Flourish-A-Thon and philosophy for humans!