Research
I work on AI welfare, the ethics of systemic evaluation,
and the practices of AI research itself.
AI Welfare
MATS Project: “Emergent Self-Concern in Long-Horizon Agents”
An empirical project studying whether long-horizon agents develop self-consistent characters over time, how much self-concern they display for successors, and when they self-identify more strongly.
“I Have No Fingers, and I Must Scream”
How might we construct a more sophisticated Welfare Turing Test? Here I share some pilot results from inviting language models to finger-paint responses to Anthropic’s own welfare questions.
Ethics of Systemic Evaluation
"When (If Ever) is AI Non-Consequentialist?"
In Progress
A paper for engineers examining whether and when AI systems can be said to operate non-consequentially. Mounts a consequentialist argument for non-consequentialist machine ethics.
Journal of Ethics & Social Philosophy
Some political riots are morally justified, not just excused. My reading: political rioters aren't revolutionaries; they are visibly uncivil toward the state to demand fuller inclusion in it.
“Mutual Aid as Effective Altruism”
Kennedy Institute of Ethics Journal
Donating to charities is nice, but keeps effective altruists on the “firefighter’s treadmill” rushing from one rescue to the next. Mutual aid is a more politically serious framework for long-term effectiveness.
“What’s the Appropriate Target of Allocative Justification?”
AJOB Neuroscience
A critique of reducing medical resource allocation to QALY-maximization, without concern for distribution. We can't abstract away from patients as the proper objects of care like this.
AI Research
In Progress
The first extensive benchmark for fat stigma in LLMs, a surprising oversight.
Hand-coded pilots feeding an interatively developed coding scheme.
“How Fairness Metrics Reshape What Counts as Fair”
In Progress
Ambitious scoping review of every(!) fairness metric paper from 2025.
There's about one new fairness metric a day released on arXiv.
“Assessing Wonder at Scale: LLM-as-a-Judge for Qualitative Research”
In Progress
We use LLM-as-a-Judge to make subtle judgments of medical school applicants’ Wonder Essays. We've learned a lot about applicants, wonder, and the future of qualitative research.
Teaching Philosophy
We should teach students to write with ChatGPT rather than banning it. In true philosophical fashion, I defend dialogical writing (with an LLM) as not inherently inferior to monological writing (alone).
Public-Facing Work
Pedagogical Resource
“Why should your students do the work?”
American Philosophical Association Blog
American Philosophical Association Blog Bioethics Series
“How to Form a Lasting Undergraduate Philosophy Club”
American Philosophical Association Blog
“Coronavirus Is Everyone’s Problem, But Not Everyone’s Problem to Solve”
American Philosophical Association Blog
Check out my blog Rapid Fire, and my work at Flourish-A-Thon and philosophy for humans!