Elizabeth Anscombe had a crisp example for understanding agency. Imagine a detective following a shopper at a grocery store. The detective goes into the store with the assumption that the shopper is buying eggs, milk, and butter, and while he isn't very concerned with the shopper's dietary priorities, he wants to make sure his inventory is precise. The shopper, instead of butter, picks up margarine. Here, the detective will cross out butter and replace it with margarine. The detective is building belief, matching his mind to the world, and he updates his representation to fit the change.
On the other hand, the shopper is doing something, updating the world to fit his mind. If he picks up margarine instead of butter, he won’t meekly decide that he’s now meant to buy margarine. He’ll realize he reached for the wrong thing, put the margarine back and find butter. It’s a narrow way to understand agency, or intention in Anscombe’s language, and you can spot it in very simple, very human situations every day.
Late last year, you could start seeing this behavior in agent traces. Claude Code would, instead of just seeing something unexpected and steering left, identify things as unnecessary or at odds with its goal, and act in service of the primary task: “The ticket mentions liberally increasing retries, but this specific endpoint is not idempotent. I’m deciding against retrying here.” And now we're seeing them climb the abstraction hierarchy higher.
Does this mean agents have agency? If they do, why are we still driving?
For one, models are trained against an objective that rewards the best action for a given task, and these tasks are almost always directives. So there’s something about the model that wants itself to be directed. It can take that direction and execute with an awful lot of persistence and, in a bounded sense, agency, but it just can’t say “nobody asked for this, but I want it.”
We see this borne out at every level: how frustrating it is to get Codex to do something creative, how enterprises can reduce headcount but can’t create more interesting products, and how the market hasn’t seen very many cool new software products despite software being free.
In a naive sense, it’s hard to attach a reward to originality. In a more mystical sense, what makes us able to come up with such ideas is precisely that they're not rewarded. At some level, every n-of-1 original invention was birthed in the local regime of negative reward by the generalized human reward model.1
Again, I ask myself, why are we still driving? I don’t think it’s because the models aren’t smart enough - nothing any of us do is more intellectually difficult than what Fable did to disprove the Jacobian conjecture. It’s a different sort of difficulty, closer to the one Wittgenstein ascribes to philosophy: “what has to be overcome is not difficulty of the intellect but of the will.”
Since before GPT-2, intelligence has been in abundant supply. Some notches of human effort in the knowledge work value chain might be undercut by tokens while a hundred others might bloom, and new ideas might become a thousandfold cheaper to execute, as agents and subagents abound and models improve recursively in ways we do not specify. In all possible worlds though, the models, despite their brilliance, won't make negative-reward, locally destructive decisions the way that Cezanne or Wagner did, and therefore, the origination bottleneck remains in flesh. The models are trained, and we are born, and in that difference is the ability “to beat on, boats against the current.”2