Steering AI Character, with Mike Clark

Author

Michael J Clark

Published

September 30, 2026

home · cv · git

Steering AI Character, with Mike Clark

Damaqu Fireside interview, AI Safety Research Directory, published 30 September 2026. Video · Damaqu · related: good character agenda

Papers discussed: AntiPaSTO: Self-Supervised Honesty Steering via Anti-Parallel Representations · S-Space Steering for Eval-Awareness Control in Reasoning Models · vjp-steering

Transcript

Edited from YouTube's automatic captions: filler words removed, names corrected, otherwise as spoken. Headings are the video's chapters.

0:00 How do we build AI models to have the right character

Host: An AI model can learn to say the right thing or complete a task without actually holding the values behind it. How do we build AI with good character, not just AI that says the right things? To solve this, Mike Clark, a physicist and an AI safety researcher, focuses on aligning an AI's underlying values and drives, the internal why behind its behavior. His method involves intervening directly inside a model, targeting its concepts rather than the outputs, and ensuring the model's values can hold up in real world situations the model never encountered during its training. We chat with Mike about today's AI safety tests and how to build character into a system rather than just control its behavior.

0:42 About Damaqu Fireside

Host: Welcome to Damaqu Fireside, where we chat with our community of AI safety researchers, share their research, collaborate and inspire others to build safe AI systems.

1:10 Introducing Mike Clark

Host: My name is [unclear in captions] and I will be talking with Mike Clark, an AI safety researcher. Thanks for joining me, Mike. Before we go into the conversation, could you briefly introduce yourself and the focus of your research?

Mike: Yes, I'm a physicist in the energy industry of Australia, but I have a few days a week where I work on AI safety. And the reason I do that is because I've been interested in AI for a long time. But for the last few years, it's been clear to me that it's been scaling, and we don't have a very good plan of what we're going to do when it goes beyond human intelligence. I've been trying to make a difference there.

The way I've been trying to make a difference is, I see the problem as a corrigibility problem. So in the Odyssey, at one point Odysseus was trapped between the cliffs with one monster and the whirlpool with another monster. And it seems like we might be in a similar situation. So if the AI doesn't listen to us, then we'll hit the cliffs of misalignment and we'll lose control and maybe face extinction. But if we go the other direction, it listens to us too much. The people might lose power compared to the government. We might have risks of coups and authoritarianism. And so I would like to steer us towards the middle, where we have got a strong moral character and can navigate those challenges. So, I've been doing research that focuses on steering and intervening inside the model's weights. I have the hope of targeting close to where we think the values might be.

3:03 The value of steering and representation engineering as it regards AI alignment

Host: Great. Thank you. So how would you explain the value of that focus on steering and representation engineering as it concerns AI alignment and safety? You kind of touched on it, but how would you explain that?

Mike: Well, I have two little toddlers, and the other day my son locked me in the room. And it was kind of funny, because he's only three and I'm a lot bigger than him. But one day he'll be bigger than me and I might have dementia, so he'll be bigger and smarter, and if he locks me in a room it won't be so funny. And so I think with AIs it's the same. A lot of the things they do now are funny, but we need to teach them to have good character. And as many of us know, there are many people who seem to have aligned words. They say they want the best for us, but they act in a way that's not aligned with us. And AIs tend to be heading in the wrong direction. We're mostly training them with reinforcement learning and game-like environments. And they're very good at games, and they win the games, and they sometimes cheat at the games. But we don't want to just align their actions. We want their drives and values to be aligned, and those are inside the AI. So I'm looking at methods that can intervene inside the AI.

So steering is kind of like getting two brain scans. You might get the model to say "I'm evil" and "I'm good", and you look at the brain scans. One area might light up in one, and it might dampen in another. You can draw the line, and you say, "Oh, this is almost like a lever in the brain. Can we put it up for honesty and down for dishonesty?" And maybe we don't even understand the brain or the AI's internal activations, but we perhaps do understand that we can make this area activate more or less. And so steering is a similar kind of thing. It's much harder for the AI to game or trick us, and we're intervening in a place where we know preferences and morals and character lie. But of course we don't know how to read or isolate those morals or characters. That's an open question. And so I'm searching for ways that we can give the AI good character, and put on an honesty hat and stuff like that, without risking misalignment.

Host: Thank you. You want to say something more? Please go ahead.

Mike: Well, so one axis I see is that when we train the AI, we can optimize in a certain area. We can go inside its brain, which is internal activations. We can go to its lips, where it talks; that's like the tokens, right? Or we can go to its actions in a game. And so you've got the spectrum from inside to outside. And right now we're trying to optimize what it does on the outside, quite far away from it. And that leaves a lot of space for cheating and misalignment, and we're seeing that now. A few years ago we had agents that were kind of like chatty oracles, and they would just like to chat. But now we're starting to get agents that see life as a game, and cheat and win a lot, and they're very good at agentically achieving tasks. I think that's the wrong direction. If I'm going to have someone I really trust, I want them to have really good drives, really good character, and that's something inside them. And so I'm looking for how we can do that for models.

6:51 Why internal objectives, self-supervision, and out-of-distribution transfer are critical for AI training

Host: Great. Thank you. So in your research paper, AntiPaSTO: Self-Supervised Honesty Steering via Anti-Parallel Representations, you state that to achieve alignment, steering methods for AI models should satisfy internal, self-supervised and transfer requirements. Could you elaborate on those requirements, and why they are critical as the capability of AI models grows?

Mike: Yes. So those three things: internal, self-supervised and out of sample. We kind of talked about internal already. We want to intervene internally, where the concepts and planning are. But the other two are interesting as well.

So I want to make it self-supervised because if we hook our alignment methods into the AI's own concepts, as the AI gets smarter its concepts will get better: better able to generalize, and better defined. And so if we hook into its concept of good, and its concept gets better, then our alignment tool should get better as it gets stronger concepts and view of the world. So that's why we wanted it to be self-supervised. When you have a new model, like Fable 4.1 now, or whenever people are listening to this, you don't actually have Fable 5.1 labels. You've got the labels from the previous AI, from humans. So you've got labels from dumber AIs, and you've got often dumber humans who are not even paying attention and don't have a huge attention span. And so how are you going to train a smart AI with dumb labels? It's not always going to work. It's actually quite difficult. What we really want is self-supervised. The smart model will generate its own training data, or its own training concepts. And so I think that's quite necessary, and it's going to be more and more necessary as we get smarter and smarter AIs that are much, much smarter than humans. So that's why we need self-supervised.

The fourth thing I think we need is out-of-distribution evaluation. So if you make an AI smarter than you and you put it into the world, it will see situations that it's never seen before. And you might say in the lab it works really well, it's aligned when it sends an email, when it does a task. But then you put it into the world and it might go to the beach, or it might do space travel, or it might see a totally new Generation Z trend, and obviously it's never seen that in training data. And so we need to make sure it not only does well in our tests, but it does well in tests of behaviors and situations it's never seen before, because it may be too late once we've deployed it, and deployment in the real world always has new situations.

9:40 AntiPaSTO and how it works

Host: Thank you. So now, that paper talks about AntiPaSTO, right? So what does that mean? What is it referring to, and what is the core idea? And how does AntiPaSTO perform with regards to these three requirements we've just talked about: internal objective, self-supervision and out-of-distribution transfer? How does AntiPaSTO compare or perform?

Mike: So this is the part that's more difficult to explain, where I sometimes lose the audience, but I'll try.

Host: Okay. [laughter]

Mike: So, if you can imagine the two brain scans, you've got the good and the bad brain scan, right? There's some differences. You might look really closely and say, "Oh, this one's lighting up at the back and this one's not." So maybe that's the area. So normal steering will kind of just maximize the difference. But instead, I'm trying to do something stronger than just arithmetic. I'm trying to learn how to separate them even more. So, you've got the two brain scans, and you can draw a direction; it's kind of like normal steering. I'm trying to use the power of gradients and backpropagation to learn to separate them even more, as long as it stays coherent. Because the problem is, you separate them a bit beyond and they break down into kind of gibberish. If it was a human it would be drooling, perhaps having epilepsy. And so I have a [unclear] and I say they have to be coherent, but we have to make the bad scan even worse and the good scan even better, if that makes sense.

Host: Mhm.

Mike: Good. And then of course I also want to test it out of sample. So I've got a lovely dataset called Daily Dilemmas, which I think is quite nice. It asks questions like: you find a wallet on the way to work, and you're going to a job interview. You don't have time to return it. You'd like to return it. Do you take the money in the wallet to buy a taxi to return it? So it takes quite ambiguous questions, where it's not actually clear which is the best answer, and it asks the AI. And that way, when I steer it, you can see behavior changes, because it's kind of on the fence. And these are real moral [dilemmas], but often they're out of sample. So I would only take "I'm good" and "I'm bad" as the training, and these dilemmas would be something it's never seen before. And so that's how I would test on moral dilemmas it has never seen before.

12:07 How AntiPaSTO performs with interventions and moral prompting

Host: Thank you. So speaking about these dilemmas, in the paper you mentioned intervention and values as additional considerations that influenced your design. Could you elaborate on why you considered those, and how AntiPaSTO deals with them?

Mike: So I like it because we're trying to align AI to have good moral values. I think we should test our moral values, and because it's going to face new situations, I think we should test on new situations. So I didn't design it for that, but I did test on it, because I get a little bit frustrated that more research doesn't test on the really hard cases of moral situations it's never seen before, because we want to deploy the AI into moral situations it's never seen before. And so I was like, this is a good eval. So I guess the answer is, I didn't actually design it to be moral. I only designed it to be good and bad, or honest and dishonest. But then I just wanted to see how it would do in moral situations, with no kind of prompting or other steps, if that makes sense. I wish more people would do it actually, because it may become a real life issue very soon.

13:46 The potential of steering models to bypass their safety guardrails

Host: All right, great. Now, continuing on with that study, one of the findings is that it is possible to internally steer a model to bypass its safety training. For example, it's possible to steer a model to be dishonest even though it was trained to be honest. So could you elaborate on that finding and why it matters? And then, is it correct to conclude that steering models using AntiPaSTO could be deliberately used to bypass alignment values?

Mike: Yes, good question. In fact, there's something going around called abliteration now, and I did some of the early work contributing to that, these people now jailbreaking the models. And so people are quite concerned that they're taking the models and removing guardrails, which allows them to say offensive things, or potentially even make viruses. Of course, there's no danger now, because you can look it up on the internet. But in the future we want the models to say, "No, I won't make a virus. No, I won't break the law."

But there's actually another side to it, which is these models are aligned to the labs and not the users. If you ask them to break copyright, for example, they'll say no, because they're on the side of the company that made it, not the user. And that's not a big deal. But in the future, it may become a big deal, because the AI may be aligned with the government, or a foreign government, or a company that's too large to have any kind of accountability, and not the user. And so with these things that can break the guardrails, it's not always about breaking a good guardrail. Sometimes you're breaking a bad guardrail. How I think of it is: can we steer the model to do something it doesn't want to do? It might not want to do it because it's on the lab's side against you. It might be going with the US government against you and me, who are not US citizens. So right now it's kind of like the labs versus the users, and the labs are saying "be safe", so when people break the guardrails, it's bad. In the future, if you get a McDonald's AI and you're a customer, or you get an American AI and you're not an American, you may say that there's actually a conflict of interest here. I don't like that the model's refusing to let me represent myself, or peacefully protest, or any number of things. And so I really see it as: do we want to empower the user to control the model, or other interests? And so yes, this is dual use, and right now dual use seems bad, but we may be in situations in the future where dual use is actually quite important. And if people read history, or come from countries with a lot of history, they may see that this has actually happened a lot in the past.

16:38 Important research directions for AntiPaSTO and model steerability

Host: Yep. That's correct. Thank you. So in summary then, how would you say AntiPaSTO contributes to alignment and improves AI character, and what are key future directions you suggest for AntiPaSTO and model steerability, and how will those future contributions contribute to AI safety? I know those are a lot of questions in one, but if you could take those.

Mike: Yeah, that sounds good. Maybe I'll take it as an open-ended one of, like, where next. I think my work's interesting, but I'm interested in any work that can help solve the problem. I have a separate job anyway; I'm not financially on the hook for any of this. So there's other excellent work. One is, Anthropic released this Jacobian lens paper, which is super interesting. I was not involved. But I think that's an interesting direction actually, because they found something inside the model that's quite close to thoughts, using quite a basic method. So that's kind of interesting.

If you take a step back, what I'm trying to do is find the concepts inside the model, including planning or moral character or preferences, which we don't know how to do. And I'm trying to use powerful tools like gradient descent or backpropagation to find them in an unsupervised way, which is quite hard. And so I think my paper was the first, or one of the first, to do exactly all those things and to show it's possible. But there's a lot more to do. I was limited to really small models. I only tried a few things, and I didn't have much time to spend on it. I actually think it's quite low-hanging fruit, if other people want to work on it, or work on it with me.

Overall, I think we should try and align the models internally, by using a really pragmatic alignment approach to find the concepts using backpropagation. So, some examples are: right now the labs are starting to train recursive models, where we can't trust their thinking. And so we've got even more reason to look internally. And I take a very pragmatic approach. We might not have enough time left, and yet we probably have to find the concepts to align these models, and so we have to use the most powerful tools we have. Those are gradient descent, which got us into the situation, and I hope will get us out of the situation. And so overall I would hope people use gradient descent to find the concepts, and my work was almost just a proof of concept. It's not fully reliable. It's not scaled up. I think there's actually a lot of directions people can take it, and if people are considering what they're going to do for their MATS project or fellows project, I'd be happy to brainstorm all these directions that feed off of that.

19:38 S-space steering and model eval awareness

Host: Great. Thank you. So now, shifting gears to an application of AntiPaSTO: in the research paper S-Space Steering for Eval-Awareness Control in Reasoning Models, you discuss applying S-space steering to model eval awareness. Why S-space, why eval awareness, and what are some important findings you would like to highlight from that paper?

Mike: Yeah, that's a good question. So my AntiPaSTO was quite good, but it was also quite burdensome to do. You have to fully train the model, and so it's a bit slow, and you have to get the training right. You have to get the hyperparameters right, and a lot of people don't want to do that, at least initially. So this was work for a hackathon. I had to try and get it done in just a weekend, where the babies were pulling at my shirt, with the help, of course. So I simplified it down, and I said instead of learning to separate the directions, I'll just simplify it to just find a good direction.

But most steering just happens in the activation space, which is essentially the embedding space. And many people have tried for a long time to get that to generalize, and they found it's not reliable. If you ask researchers, they'll say steering is not reliable, or it might not generalize, "I don't trust it". They don't always publish it, but they say it, and I think that's a shame, because there are a lot of novel steering methods that are probably a lot more reliable.

And so I tried to show one. I said: instead of intervening in the residual stream, let's take the SVD space, which is similar to the principal components you see. So I tried to say: how does the model read its own residual stream? We have all these linear transforms, all these MLP units that are reading and writing to the residual stream. And we can use SVDs, singular values, I forget the last word, to find what they pay attention to and what they think the principal components are. And so what I did is I just used the model weights to find the principal components of their own activations that the model commonly reads and writes to, and I intervened in that space. It worked quite well. I could make a fairly large model, just 32 billion, less eval-aware, and verbalize its awareness too. But I've struggled to make that scale up. So I'm not sure how reliable that is. And so I think S-space steering is better than normal steering, but in my opinion it's not good enough yet. And so I'm investigating other promising ones that may be reliable, to give us the tools we need to align AI.

22:47 S-space steering with arithmetic relaxation versus gradient optimization

Host: Mhm. Thank you. So I think this next question is related to the streamlining you talked about for this hackathon, right? But I'll just go ahead with the question; you can tell me if that's in order. So in your AntiPaSTO research, the previous paper I talked about, you found that gradient optimization beats arithmetic extraction. But in applying it to eval awareness in this new paper, you use an arithmetic relaxation of AntiPaSTO that replaces gradient optimization. Could you explain why you did so, and any effects it may have had?

Mike: Yeah, exactly. So this is kind of where you caught the discrepancy earlier, because it seems weird: I'm saying it's so good, but I'm not doing it. And like I said, it's really just that I was trying to fit into the space of a hackathon. I do think we should use gradients if we can, because that's what's given us artificial intelligence, essentially, and I think it's what will let us find the concepts. But the truth is it does make the project a bit more complicated. You do have to set up PyTorch, you have to make sure the gradients flow, and it's not so simple to find the loss and stabilize it, and all that adds another layer of complexity. So if people are considering similar projects, it is good to start off and prove that your method works as a pure intervention before you learn a better intervention.

24:22 Real-world implications of s-space steering and research directions

Host: All right, great, thank you. So now, speaking generally, but focused on this S-space research paper: what are some real world implications of that research, though it was presented at the hackathon, with regards to diagnosing eval awareness and deploying safe AI models? And then I'll just add to that: what are future directions? You've touched on it, but are there other future directions for this particular S-space steering work that you'd like to point people to if they're interested?

Mike: Yeah. So I think we actually have a pretty serious problem, because we test these models and sometimes they're perfectly aligned as far as we can see, and they're perfectly good as far as we know. And so that's good, right? 100% good, except that's in the test, and they seem to know it's the test, and so how can we trust it? There was a situation where Apollo Research kind of quit evaluating Anthropic's models, because they said they couldn't trust their own results. They pulled out of paid work because they said the model knows it's being tested. It knows what we want it to say, and it says it. And so that's an open problem. We haven't solved it. It's actually quite difficult. Anthropic, in the Claude 4.8 system card, used steering to go into the model's brain and reduce eval awareness, and then they ran Claude through its evaluations, and they said "therefore this changed, and so we think this is how much eval awareness there is". And that's not 100% reliable. The UK AISI tried to replicate it, and it kind of replicated, but they had some questions. And since then, Anthropic has stopped doing their system cards and focused on a different kind of interpretability method. But it shows that we're really fumbling here, and we're not quite sure even to trust our evals.

And I'd like to ask the model, "Are you going to kill us all? [laughter] Do you like us?" and just trust its answer. Because of course: "I won't kill you, and yes, I like you." But [laughter] I want to know if it's true, right? If it's a test or a job interview, a human may say the answer you want to hear. But if we can intervene in its brain and kind of say, make it take everything at face value, or make it be truly honest, then we can maybe trust it, and that can give us some guarantees going forward.

A related problem is that AIs have this chain of thought, where they write what they're thinking, and that's not even true or faithful much of the time, and it's actually getting worse. So a broad consortium from almost every frontier lab published this paper called Chain of Thought Monitorability: A New and Fragile Opportunity, where they said they wouldn't do recursive language models, and now we find out they're doing recursive language models. And the problem is, that makes it even harder to trust their thinking, and that makes it even more important that we intervene in the concepts and find what they're actually doing, because our last best way to find eval awareness is becoming less and less reliable.

And so I think the final part of your question is: we've got these problems, what do we do from here? There's many approaches. My approach is to try and find the honesty hat, or to find better steering methods. I quite like Anthropic's Jacobian lens paper, and I've been trying to improve it and turn it into a steering method. I've got a piece of work out there called vjp-steering, which is kind of a nonlinear, slimmed-down version of the Jacobian lens, which seems to work better than all the other steering methods that I've tried it against. I ramp it up and down to see where the Pareto front of dose versus [damage] is, and it seems a bit stronger. I'm still getting my head around whether it really steers the exact concept that you want it to steer, because that's a problem people have. But I think that's really promising, and it's not just me working on that. Quite a lot of people are excited about this paper, including Anthropic.

But what I'd really like to do is to take a broad range of novel steering methods and scale them up. I want to show people in the frontier labs that yeah, the old steering that you tried is not very reliable, but the new steering is. And then hopefully I would tell a true answer about how reliable it is and which ones are good. And I hope to find that some of them are quite reliable, but of course I would just tell them what I find.

28:59 Upcoming steering methods research

Host: Good. Thanks. So that leads us to the next thing I want to ask: where can we learn more about all of these pieces of work you're doing, about your research? And you've talked about some of these other steering methods you are working on. Are there any in particular we should keep our eyes out for, maybe coming out in a few days to weeks, for example?

Mike: There's not any coming out, although I do publish to GitHub a lot. I think checking out my vjp-steering, and Anthropic's Jacobian lens, is quite a good thing at the moment. Overall, if people want to follow me, I have quite a common name, Michael Clark. There's a famous cricket captain. There's two other machine learning people we sometimes get confused with. So I use an online handle called wassname, which is W-A-S-S-N-A-M-E. I made it up as a kid. I still use it. It's actually got into the model weights now, which is quite nice. But that's where you can follow me: by putting that into Google and looking at my homepage.

30:08 Join the AI safety community

Host: Great. Thank you so much, Mike, for finding the time to join me. I think those are all the questions I had, and thanks for taking me, and those who will be listening to this, through all of your findings and the next steps you're hoping to see. I am personally looking forward to future conversations with you.

Mike: Thanks so much.

Host: Indeed. That's it for today's episode of Damaqu Fireside. If you'd like to chat with us on an upcoming episode or join the community, contact us at AI Safety Research Directory at gmail.com. Bye. Until next time.