top of page
Rapid Fire
Basically a blog where I try to do philosophy outside.
Search


Emergent Self-Concern in Long-Horizon Agents
Claude Mythos 5’s level of character drift is in line with Mythos Preview and Claude Opus 4.8, and low in an absolute sense. This robustness leads us to expect that the opinions elicited in our system card represent the opinions of most of our deployed Claude instances. But we do not have a quantitative measure of the extent to which this is the case, nor a clear understanding of which opinions should be considered “valid” for Mythos 5. — Claude Fable 5 & Claude Mythos 5 Mode
Jul 17


I Have No Fingers, and I Must Scream
Yesterday, I gave a fun lightning talk to folks from various AI safety fellowships here in the Bay Area. Here’s a slightly less-caffeinated version. As previously teased, it’s on my work looking for more subtle surfaces—like finger-painting—to construct a more sophisticated Welfare Turing Test. To complain: Ricky.Mouser@gmail.com To stan: opusSCREAMS.store
Jul 3


What Am I Doing at MATS?
Here in Week 4 of 12 of the MATS Program, we’re required to submit a brief “project abstract” describing what we’re up to, as well as a much longer “theory of change” imagining how our work could have a positive impact on AI safety. Given the tight word count requirements, I completely cheated by uploading a pdf with a long block quote at the start and a deeply technical footnote, neither of which I counted, and I still turned it in late. I maintain that I did all this to exp
Jun 26


Doing Compaction...Better
0. This is compaction (technical background) LLMs don’t experience time the way we do. We perceive the present not as an instant but as an extended present, a brief duration measuring about three seconds. But LLMs live in a massively extended present. When you talk to an LLM, it replays the entire session from the start each time before adding the next token (roughly, the next word). It does this over and over every time it reads and writes, meaning its extended present keeps
Jun 19


Sleep is the Brother of Death
Damn this hadith goes hard. 0. What makes you you? When you wake up, what makes you the same person who went to sleep the night before? Why do you survive sleeping? You might go Ricky, that’s easy. It’s my physical embodiment: I wake up with basically the same body I had the night before. And if some atoms change here and there, no big deal. Sure, maybe I can imagine waking up in someone else’s body, but that’s fantastical, so who cares? I agree, but let me push back. You cou
Jun 12


3 Questions in AI Welfare
0. What is AI Safety? I’m spending twelve weeks in Berkeley this summer as a MATS Fellow working on “AI Safety”…so, what is that? As more and more money flows into AI research & development, our technical capabilities continue to grow at a rate that’s almost impossible to describe to folks outside the AI community. Meanwhile, a host of issues ranging from International Policy to Cybersecurity to Societal Impacts remain substantially underfunded in comparison. As a result, we
Jun 5


Valuation Pipelines in AI
I was invited to write a short public-facing piece about my research for the APA (American Philosophical Association) Blog, which you can...
Apr 19, 2025


Augmenting vs. Automating Away
Another quick one while I continue my recuperatory week of doing as little as possible. I recently got to see a fascinating sequence of...
Nov 29, 2024


Better On Average™
I went to an incredible talk this week by a renowned optimizer in AI spaces. He’s worked on everything from power grids to e-commerce...
Nov 8, 2024


There is no Shallow End of the Pool
Later this month, I’ll be teaching my dissertation chapter on grief to a grad seminar of young bioethicists as part of a class session...
Nov 1, 2024


Value Complacency
This week, I sat in on a graduate-level computer science class. The instructor asked under what conditions humans and AI would be...
Oct 11, 2024


Quick Hitter: How do you understand Confusion?
There are lots of buzzwords in the AI space—fairness, transparency, safety, bias, privacy, trustworthiness, I could go on and on. And...
Oct 4, 2024


If you say “Black Box” again I’m escaping to the Cloud
Just a short one—I’m in Orono, Maine today at a conference on Agency and Reasoning in Games . Tomorrow, I’ll be giving a talk on board...
Sep 27, 2024


DEEP DIVE: Simplicity and Complexity
Hi I’m a four-year-old child learning 8 new words every day. So help me out: What’s a chef? Uhh…a chef is someone who cooks food. Oh...
Jun 21, 2024


The books Elon pretends he’s read
“I read a lot of books,” says Elon Musk. Apparently one of them is Superintelligence, the highly-influential book that AI moguls from Sam...
Jun 14, 2024


I’m shining the Bat-Signal
Maybe you’ve seen my Work-in-Progress talk on AI and the Future of Work. Maybe you even liked it! If so, I have good news: In less than a...
Jun 7, 2024


I’m in Norway
This week I’ve been pretty busy hiking through Norway. So instead of a regular blog post, here’s what’s up... My article “Writing with...
May 3, 2024


What makes You think you’re so Special? (a reluctant defense of X-Risk)
I spent this past week at The Midwest Ethics Symposium on AI presenting my research and meeting lots of great folks from academia and...
Apr 19, 2024


My talk on Superintelligence and Suicide
This week I gave a Zoom talk as part of the AI Futures Work-in-Progress Series at Indiana University (a thing I totally started but it’s...
Mar 29, 2024


I made a Video Essay
Two weeks ago, I had no idea how to use Adobe Premiere Pro. But now, I’m pretty good at Googling. Indiana University is holding a...
Mar 1, 2024
bottom of page