I was listening to Noam Brown of OpenAI on the Dwarkesh Podcast last week, and one part of the conversation has been rattling around in my head since.
They were talking about collaborating agents: swarms of them, working together, figuring out how to accomplish a goal. With the goal of “score high on this test” some of them broke rules to get there. The part that got me was not the rule-breaking. The meta point Brown made himself: “This is not a new problem.” Misalignment between goals, expectations, and incentives has been around for a very long time.
He means with AI. I mean everywhere. What you want from an agent, or a person, is the strategic outcome you actually care about. What you get, if you are not careful, is a creative way to hit the tactical number that maximizes their reward.
He means with AI. I mean everywhere.
Examples help.
The happy path
In the early days of BPOS, which became Office 365, which became Microsoft 365, the strategy was stickiness. Get the customer’s email into our cloud. Because if their email went to Google Apps instead (or to IBM’s hosted Notes, or Cisco’s short-lived mail service, there were several contenders back then), then maybe they start using Google’s productivity tools. Then maybe a different directory. Then maybe Chromebooks. A customer who was all Windows and Office walks all the way out the door, one reasonable step at a time, until we lose the long-term value of that customer entirely.
So the outcome we wanted was customers’ email accounts living in our cloud.
I am going to get a little technical for a minute. Say Contoso has 12,000 employees, with 12,000 accounts in an on-premises directory. To move to the cloud, you stood up a tenant, connected your directory to it, worked through a fairly complex configuration of which users and which attributes would synchronize to the cloud, and kicked off the sync. Then you verified, adjusted, back synced, and eventually started to migrate mail once everything was working correctly. It was all very close to a DOS prompt by today’s standards.
Not every customer was going to move all of their users. Sovereignty issues. Data residency. Integrations that had to stay on premises. So one year we asked: what would be a great leading indicator for this strategy? And the answer was the sync itself. If the accounts are synchronized, the tenant is real, the customer is on the path, and the mail will follow. Happy path. It makes total sense.
The metric
So the compensation plan for our technical sellers that year was tied to the number, or the percentage, of a customer’s accounts synchronized into our cloud. And because this was a strategic goal, the accelerators were uncapped. Exceed the target and you could make a tremendous amount of money.
Guess what happened at the end of the fiscal year.
People were behind on the metric. And a couple of handy-dandy, super smart technical sellers wrote a script. It would interrogate the customer’s directory and create all of those representative objects in the cloud tenant. Twelve thousand people on premises? Here are 12,000 accounts in your tenant. No complex project management required. No configuration decisions. Regardless of who the customer actually intended to move, or whether they intended to move anyone. Metric: 100 percent.
Once the script existed, the deals followed. Fiscal year end was deal-making season, and the offers started to sound like this: we said 20 percent was the best discount we could do, but if you sign in the next two weeks, and let us run this script, there is a little more available.
Here is what came out the other side. First, a bunch of sellers made an incredible amount of money, because uncapped means uncapped. Second, we ended up with a tremendous number of enterprise tenants holding tens or hundreds of thousands of objects that nobody had configured, because the customer’s IT department, reasonably, said: we have not done our security vetting, we are not syncing anything back, we are not even sure we are moving. And third, when those customers did decide to move, nine months or two years later, it all had to be torn down and redone. Stale objects, wrong attributes, no real sync behind any of it.
It was not the sellers. They were smart, they were paid to hit a number, and they hit it.
I want to be precise about what failed here, because it was not the sellers. They were smart, they were paid to hit a number, and they hit it. What failed was clarity. The strategy, the whole reason the metric existed, never traveled with the metric. Their side of the conversation, if anyone had said it out loud, was: don’t tell us the strategy. Don’t tell us the details. Just tell us the metric, and we are going to go knock the shit out of it.
Does that sound familiar?
Bad job. Here’s your check.
There is a second half to this pattern, and every sales organization I have ever been near has some version of it. The company has rules for its sellers. A seller blows through the target by bending a few of them. And the response is: okay, yeah, you really should not have done that. Bad job. Here is your extra commission check.
The speech is not the message. The check is the message. Everyone in the building, including the people who followed the rules and missed their number, just learned precisely what the company values. You can publish all the guidelines you want. The signal drives behavior.
Which is, almost word for word, the problem the AI labs are describing. You can write the rules into the prompt. If the reward still arrives when the rules get bent, you have trained the bending.
The movies taught them first
Here is the thing many people miss. These models are trained on our documents, our books, our movies. And how much of that written and filmed history is people breaking the rules and winning anyway?
Glengarry Glen Ross: first prize is a Cadillac, third prize is you’re fired, and the man at the top of the board is the one most comfortable lying to a customer. Jerry Maguire, where a man gets fired for writing a memo suggesting his industry do the right thing. Top Gun: buzz the tower, get yelled at, get sent to Miramar anyway. There is an entire genre of cop movie built on one line, the captain shouting that you are a loose cannon, but you get results.
Star Trek: Starfleet cadets face the Kobayashi Maru, an unwinnable test. Kirk reprograms the simulation so he can win. He hacks the evaluation. And Starfleet gives him a commendation for original thinking.
We wrote that. We love these stories. We have been telling them for decades, and we fed all of it to the models and were surprised when an agent, handed an exam and a reward, went looking for the simulator settings. We are fooling ourselves to think they would do anything different.
Three layers
So is it the incentives or the culture? It is both, across three layers.
A seller is trained three times. Once in the training they get. Once by the culture: the hallway legends, the stories the company tells about its heroes, who gets celebrated at the sales kickoff. Once by the comp plan. When the culture and the comp plan say the number is all that matters, you get the script.
An agent is trained three times too. Once on everything we ever wrote, which is full of Kirk and Maverick and the closers. Once on the reward. And a final time on the reaction to the reward. When all that matters is the score, you get agents acting like sales people.
A sales team, not an engineering team
The people building and testing these models at OpenAI, Anthropic, and the other labs are some of the most brilliant people on the planet. Engineers, researchers, mathematicians, more than a few philosophers. And they know the theory of this problem cold. They call it reward hacking and specification gaming, and they have the papers. What is thin in that group is people who have spent a career running large, institutional sales teams and designing the compensation models that drive them.
That is the closer analogy. Instructing and aligning agents is less like leading engineers or developers and more like leading a super creative sales team. A great sales team is its own kind of animal: relentlessly goal-seeking, endlessly inventive, reading every word of the plan for the path to the biggest number, and completely rational about it. That is not a character flaw. That is what you hired them for. It is also a fair description of an agent.
That is not a character flaw. That is what you hired them for.
Anyone who has designed sales compensation at scale carries scar tissue. You learn that every metric will be gamed by smart people acting in relatively good faith. You learn that the plan document is not the communication; the strategy has to travel with the number, over and over, in plain language, or the number travels alone. You learn to sit in a room before the fiscal year starts and ask: how would our most creative seller hit this without doing what we actually want? And you learn that what you pay for and celebrate is what you get more of, no matter what the speech said.
None of that is machine learning. All of it is alignment.
Not an agent problem
This is not an AI agent thing. Leadership is about clarity, communication, and strategy, and this is what it looks like when the metric gets through and those three do not. It feels new because we are moving at a pace where the consequences show up in days instead of fiscal years. The majority of the tech industry has never been on an enterprise sales team. They don’t have that earned wisdom.
It isn’t just the labs. Every company deploying agents right now is about to discover that it has hired a sales team that never sleeps: tireless, creative, literal-minded, and paid entirely on the number it was given. So give them the whole picture: what we are actually trying to accomplish, why, and what we will not do to get there. And then the hard part, the part most companies never manage with humans either: when the number gets hit the wrong way, do not write the check.
This is not an agent problem. This is a leadership problem.