Research agenda

Author

Michael J Clark

home · cv · git

What is my current research agenda

2026-08-22

I want to build alignment tools that frontier labs will actually use in the next few years, and that have three nicer properties.

1. Closer to unsupervised learning

We won't have labels as strong as the models we are aligning, so I prefer methods closer to unsupervised.

2. Non-adversarial oversight

Two-panel cartoon of a small graduate-capped robot supervisor holding a clipboard next to a much larger robot student. In the left panel, marked with a red X and the label 'Hide', the large robot holds a page with a red bug icon behind its back while the small supervisor squints at it along a dotted sight line. In the right panel, marked with a green check and the label 'Find and confess', the large robot hands the bug page straight to the supervisor and a green star reward points back to it. The point is that adversarial hide-and-seek between a weak supervisor and a strong student is bad, while a rewarded confession is good.

I want to avoid setting a weak supervisor against a stronger student in an adversarial setting. Neutral is better, for example gradient routing detaches the adversarial gradient, and cooperative schemes like confessions are better still.

3. Closer to internal optimization targets

One unbroken ink-brush line on cream. It starts as a hand holding a speckled, rotten apple on the left, loops down into the tubing of a stethoscope with two round ends, and rises on the right into a face in profile. Three labels sit along the path: at the apple, 'distal' over 'RL'; at the mouth, 'boundary' over 'SFT'; at the head, 'internal' over 'RepEng / Steering'. It reads as a spectrum of training signals, from far outside the model to listening directly to its internal states.

I want to avoid distant and surrogate objectives like RLAIF because they are easily gamed, and instead prefer objectives closer to the model: internal states first, then logprobs at the model boundary, and RL reward last.

No method gets everything

Of course no method gets everything, and a better tool that gets used is better than a perfect one that doesn't.

The talk

I laid this agenda out in a 5 minute talk at the Sydney AI Safety Forum 2026, Technical Advantages for Weak-to-Strong Oversight: Bets I'd Like Challenged. Please come change my mind, anonymously or by email.