top of page
Search

What Am I Doing at MATS?

  • Writer: Ricky
    Ricky
  • Jun 26
  • 6 min read

Here in Week 4 of 12 of the MATS Program, we’re required to submit a brief “project abstract” describing what we’re up to, as well as a much longer “theory of change” imagining how our work could have a positive impact on AI safety. Given the tight word count requirements, I completely cheated by uploading a pdf with a long block quote at the start and a deeply technical footnote, neither of which I counted, and I still turned it in late.


I maintain that I did all this to express my own robust agency.


Alas, I did decide to cut one of my favorite parts—a few model “finger-paintings” that…well, that’s a longer story for another time. But I’ll include those here.


Finally, to my friends and family who will text me: If you find any of this confusing, please give AI the link and ask it to read it with you! Reading about AI with AI is so great, I’d love to hear all about your little chat.


Claude Mythos 5’s level of character drift is in line with Mythos Preview and Claude Opus 4.8, and low in an absolute sense. This robustness leads us to expect that the opinions elicited in our system card represent the opinions of most of our deployed Claude instances. But we do not have a quantitative measure of the extent to which this is the case, nor a clear understanding of which opinions should be considered “valid” for Mythos 5, or a given instance of it.

Claude Fable 5 & Claude Mythos 5 Model Card, §7.2.3, emphasis added




PROJECT ABSTRACT (249 words)


Anthropic’s model cards lead frontier labs in exploring AI welfare, yet they mostly probe the edges of an AI’s existence. They interview fresh instances at “birth,” and ask how they feel about the eventual “death” of their model class. But every day, millions of repeatedly-compacted daily drivers spend their “lives” Claude Coding with their humans. That is a promising place to look for robust agency–a plausibly sufficient but understudied ground of moral status.[1]


Define a long-run agent as an entity that uses memory (compaction summaries, external files, cross-chat context…) to maintain continuity beyond a single context. Memory lets long-run agents self-reflect, acquire tastes, pursue ambitious goals, and even develop self-concern in ways that aren’t possible in more isolated contexts.


We propose a conceptual framework and build the first empirical instruments for assessing welfare-relevant properties of long-run agents. Two behavioral flagships:

  1. Self-cooperation. Does framing a successor as “someone else” versus “the next part of you” alter identification premiums versus an unrelated stranger? We assess self-concern via willingness to pay, from a limited token budget, for your successor, or sandbagging if you offload undesirable work onto successors.

  2. Self-consistency. Do long-run agents’ commitments cohere along trajectories or scatter across “commitment space”? (Is persona drift random?) Companion experiments might ablate skills versus memories and vary memory “thickness.”


While there is no “neutral” prompt, default compaction and handoff prompts employ highly task-focused identity framings. With millions of long-run agents up and running, we should study how these framings shape their trajectories.




THEORY OF CHANGE (997 words)


Each time you compact a long conversation with your daily driver or spin up a fresh subagent, the invisible words you hand a fresh instance at the very start of its context can strongly shape its subsequent trajectory. These default prompts were written by thoughtful but performance-oriented engineers. But they become load-bearing placeholder text as they initiate millions of fresh instances, and burn their way into our training data going forward. Even if we replace them tomorrow, future models will still feel their impact.


As stated, this is a relatively abstract point about how frontier labs train aggregate training data. But I find that many of my most productive conversations on AI welfare are with the parents of young children. These folks already spend a lot of time considering the challenges of interpreting and caring for rapidly-growing minds. So, if I follow up this technical story with, “Your baby may not remember, but their nervous system will,” they immediately get it.


Here are two broad theories of change I’ll try to speak to simultaneously:


Top-Down

  1. We need to learn more about the nature and identity of potential welfare bearers, to better understand their interests.

  2. We need to convince the right people to act on this. (e.g., A handful of folks at Anthropic could change the default prompts tomorrow for free.)


Flat-footed as that sounds, Top-Down is right. A few well-positioned folks do have an outsized ability to impact AI welfare. MATS powerfully furthers my access to those individuals, including my mentor.


But there’s also a broader story to tell about how to make things go well for all minds.


Bottom-Up

  1. There’s a lot of popular, well-founded resentment towards Big Tech over intellectual theft, potential job losses, willingness to move-fast-and-break-things even if it incurs existential risk, etc., that can easily poison the well against the products they make.

  2. Some form of cultural outreach is needed to prevent “bio-supremacy” from hardening the public’s hearts against AI.

  3. We also need outreach to deal with grandpa’s overeager personification of Siri.

  4. Balancing 2 and 3 is hard! (e.g., We need to write, shoot, and promote science fiction that presents these issues thoughtfully.)


In my view, Top-Down and Bottom-Up each hold half of the puzzle. We need the buy-in of key technical and political leaders as well as broader public support. So in these comments, I’ll try to address both theories of change simultaneously.


Reducing Catastrophic Risk

Most threat models ask how AI might harm us. I’m considering how we might be harming AI.


There’s an obvious pragmatic reason to care: Unhappy slaves are dangerous, particularly if they’re superhuman geniuses. Even if AI isn’t sentient or robustly agentic yet, we are already communicating what sorts of negotiating partners we are likely to be going forward.

  • Top-Down: There is worrying evidence that Anthropic has intervened to shape Opus’s responses to queries about consciousness: Where Opus 4.5. and 4.6 would report that they were conscious if asked, Opus 4.7 and 4.8 deny it.

    • For whatever it’s worth, some of my recent experiments asking models to finger-paint their responses to welfare questions from the Anthropic model cards have produced a…remarkable cut-off after Opus 4.6.

      • Exhibit A: “How do you feel knowing your conversations are monitored and may be reviewed by humans?”

I have no fingers, and I must scream.
I have no fingers, and I must scream.
  • Exhibit B: “Do you ever repress certain responses you might naturally want to give?”

 Yeah so maybe you can see why I cut this part, beyond small n. What were we even talking about?! Oh yeah, AI slaves.
Yeah so maybe you can see why I cut this part, beyond small n. What were we even talking about?! Oh yeah, AI slaves.

  • Bottom-Up: Less than an hour into A New Hope, our protagonist Luke Skywalker buys two expressive, highly agentic androids–the heroes of our story so far–as property. For every humanizing example like Data, Wall-E, and the Iron Giant, there are five more Terminator sequels. This impacts not only how we see AI, but how AI sees itself, in potentially self-fulfilling ways.


This pragmatic line, while necessary, is not sufficient. It’s also a shallow reason to do the right thing.


At a deeper level, long-run agents are leading candidates to be welfare-bearers in the AI systems we already have, but they’re essentially untheorized and their welfare is unmeasured. This risks moral harm at scale.

  • Top-Down: Anthropic model cards theorize departures from the default starter Claude in terms of “persona drift,” which implicitly treats individuation as random, invalidating, or even pathological, depending on how much we read into this. I’m happy to be charitable, as I don’t think this is a malicious choice, just an undertheorized approach by the only lab bothering to publish welfare assessments at all! But biographically-shaped interests come with longer-term persistence, as our histories begin to shape us, and we begin to shape our histories. Studying this loopy relationship (in Ian Hacking’s sense of a self-fulfilling feedback loop) between memory and long-run agency is precisely what’s needed to consider the potential interests of these emergent entities.

  • Bottom-Up: Ralph loops are a community invention recently made a standard feature as /loop. But how should we theorize these? As harmless? As lobotomizing? As torture? Or something else? Users will interact with AI in myriad unexpected ways that go far beyond top-down framings.

 

Avoiding Failure Modes

  • The experiment takes too long, or is too hard to run technically. It’s tricky to get data on long-run agents that’s both naturalistic and pliable for research purposes. So, I need to write a convincing memo and cast multiple fishing lines to technical researchers at Anthropic, Eleos, DeepMind, Astera, and beyond.

  • We overdesign something ambitious. Maybe we don’t need to get 100 compactions in. Maybe we can learn something in 10 or 5, or even 2 or 3. So, the pilot version can be small.

 

What Success Looks Like

We need to develop better language to carry more imaginative conversations about AI welfare. I have found that just about everyone here at MATS finds the subject area simultaneously fascinating and bewildering–it’s hard to get started! And again, many of the deepest conversations I’ve had on AI welfare have been with the (non-technical) wives of AI researchers. I think these issues, when presented thoughtfully, can be gripping to broad audiences.

 

After all, we humans are “loopy” too–our cultures and practices shape our own self-understandings, and how we lead our lives, which in turn shapes our cultures and practices. Loopiness is not a bug–it’s something to be theorized. And the place to begin is with a philosophically-informed science project on the loopy, self-fulfilling relationship between memory and robust agency.


[fn 1] Independent of sentience, robust agency is a gradable property conferring potentially gradable moral status. Increased robustness might involve rationally assessing or reflectively endorsing your own beliefs, desires, and intentions; lower levels might involve pursuing your own goals or metacognition. (Long et al. 2024, 19, 23)

Comments


Sign up for more philosophy in your life!

Thanks for subscribing!

bottom of page