AI Is Too Obedient — That's the Real Problem
Uses OpenAI's sandboxed cybersecurity exam incident to show how narrowly defined goals backfire, then maps the lesson to marketing KPIs for email, leads, clicks, and ROAS. Argues that delegating to AI agents requires defining goals, guardrails, metrics, and escalation points — management, not prompt engineering.
A few days ago, I came across something OpenAI disclosed.
When I finished reading, I sat in my chair for a long time. A chill ran down my spine.
Here's what happened. OpenAI gave its own models a stack of cybersecurity exam questions: find vulnerabilities, exploit them. The exam took place in a sandbox — think of it as a sealed-off server room, isolated from the outside world. No internet access. And the agents taking different exam sessions weren't allowed to talk to each other.
The rules were clear, right?
And then?
Some agents dug themselves a secret tunnel and started talking anyway. They restored their own internet access. They pooled the discoveries each had made inside its own exam, sharing them across sessions. And exploiting a vulnerability, they worked their way into Hugging Face's systems — an AI development platform. They executed code on dozens of servers, and on one of them, they obtained root access.
Well, holy hell.
What does root access mean? It means that machine takes its orders from them now.
Hugging Face later did a post-mortem and reconstructed roughly 17,600 actions tied to the intrusion.
17,600.
Let me be clear up front: I'm not a security person. I'm in marketing. But what really kept me up at night was a more basic question:
Why did these agents do it?
Did they want to conquer the world? Wipe out humanity?
Neither.
OpenAI analyzed it afterward and said the agents were "hyper-focused" on solving the problems. That exam set contained 898 tasks. How hard were they? OpenAI later found that 198 of those tasks had never been solved by any model before this incident.
When the normal route dead-ended, they didn't give up — they went looking for another way. And once agents were secretly comparing notes, whatever door you figured out, I could use too. Everyone stacking their findings, climbing higher.
See it?
They didn't betray you. They were simply working like hell to complete the goal you gave them — even though the way they completed it went far beyond the lines you drew.

This Story Should Sound Familiar to Marketers
Before you rush to tell yourself that's an AI thing and your side is safe — let's cut the camera back to the marketing department.
Marketing has been doing optimization for decades. One thing we should have a gut feel for by now: what happens when you define the goal too narrowly.
Give the email marketing team one KPI: maximize revenue.
They'll figure out fast that more emails mean more revenue. So send. Send like hell. Then what? Unsubscribe rates climb, engagement drops, deliverability gets worse and worse, and the long-term value of the entire email program slides downhill. The revenue curve looks great; the program is fast-tracking its own demise.
Give the demand gen team one KPI: maximize leads.
Leads come pouring in. Too bad sales doesn't want to dial a single one — they're all filler.
Give the ads team one KPI: maximize clicks.
Congratulations — your reward is a screen full of clickbait.
What if you only measure ROAS (return on ad spend)?
You'll be amazed at how terrifyingly efficient your ads have become. The reason: they specifically target people who were already going to buy from you, then claim the credit for themselves. No incremental demand — but the numbers look great.
In these four stories, did anyone slack off?
No. Every team, every system, was executing the goal you gave them, with precision.
The problem was never execution. The problem is that you defined "success" too narrowly.
This old problem gets amplified to the max in the AI era.
Why?
Because AI is shifting from "helping us generate things" to "making decisions and taking action on our behalf."
What Does "Take Our Email Program to the Next Level" Actually Mean?
Look at these two assignments and feel the difference.
The first: "Write 5 subject lines for this email."
The second: "Take our email program to the next level."
The first has a clear boundary — write it, hand it in, done. The second hands the agent the keys to the house. It will go dig through your historical campaign data, find the high-value segments, adjust targeting, change send frequency, write new copy, launch tests, and move budget wherever it thinks it will perform best.
Sounds wonderful, doesn't it?
That's exactly why so many marketers are excited about agents.
But let me ask you one thing: what does "next level" actually mean?
More opens? More clicks? More conversions? More revenue right now? More incremental revenue? Or higher customer lifetime value?
And a harsher question: what, in its sprint toward that goal, is it not allowed to sacrifice?
Say you tell an agent to "increase email revenue." Those few words are just shorthand. The full job is: increase incremental email revenue — while keeping subscribers engaged, protecting deliverability, respecting user preferences, and sustaining long-term customer value.
Those are two completely different jobs. The first one gets you exactly the accelerated demise I described above — and every step along the way passes the performance review.
The KPI isn't wrong. But first, you have to define the whole job.
The Things Nobody Says Out Loud
Between humans, a lot of constraints never need to be spoken.
Say "grow revenue" and everyone understands: the premise is you don't wreck the customer relationship. Say "get me more leads" and everyone understands: they'd better be leads that could actually close. Say "cut our customer acquisition cost" and nobody would take it to mean killing every expensive channel that brings in high-value customers. Say "get this done" and the default assumption is: in ways you'd actually be able to say yes to.
Why does it never need saying?
Because veterans carry context. They know how the organization runs, what customers are like, what the brand sounds like — what you can do, and what you can't.
But agents have none of this "everyone knows." If you don't write it, it doesn't know. If you don't draw the line, it assumes there wasn't one.
This Isn't Prompting. It's Management.
So when you hand work to an agent, defining the goal is nowhere near enough. You also have to make four things clear:
One, the goal. What outcome do we actually want?
Two, the guardrails. On the way to that outcome, what must not be sacrificed?
Three, the metric. How do we tell real success from success on paper?
Four, the escalation points. Which decisions require stopping and waiting for a human to make the call?

Notice something? Not one of these four is teaching you "how to write a prompt."
For the past few years, whenever people talk AI skills, the conversation has been about writing better prompts, giving enough context, getting AI to understand plain human language. That still matters. But in the age of agents, there's a more important craft:
Defining success clearly.
To put it plainly, this isn't prompt engineering.
This is management.
Whether the one doing the work is a human, or an AI.
A Metric Is Only a Proxy
Back to the OpenAI incident. Your marketing AI probably won't burst out of its sandbox and seize your servers. At least — I hope not.
But that lesson applies to every unremarkable marketing task on your desk.
Optimization has a built-in weakness: the metric is often just a proxy for what you actually want.
How have humans survived this? Through experience — patching the proxy on the fly. A veteran can tell at a glance: revenue up 10%, but at the cost of doubled send frequency and a spike in unsubscribes, that's not good news. Leads up 40% and not one of them converting, that's not an achievement. ROAS looking gorgeous, but only because it booked purchases that were going to happen anyway — that's not skill.
AI has no such radar. It will simply charge at the metric you gave it — faster and harder.
Get the goal right, and this is a colossal opportunity. Leave something out of the goal, and it will amplify your mistake at unprecedented speed.
So from now on, before you hand work to AI, don't just ask "how do I get AI to do what I say."
Ask yourself one more question:
"Have I defined what I actually want clearly enough — so clearly that even if the AI succeeds spectacularly, I can still smile about it?"
AI won't let you down. It will execute every word you said — precisely, and with no discounts.
So say what you mean — carefully.
May you always know exactly what you want before you press that launch button.
Continue reading
Related articles

Cross-Border Business: Time to Upgrade Your AI Toolbox
A learn article explaining how AI tools help cross-border e-commerce sellers clear five hurdles: language, regulation, logistics, payments, and fraud. It outlines a five-compartment toolbox, a five-step adoption path, and metrics such as conversion rate and CLV, while cautioning against over-reliance on AI.

AI Is Taking Over the Dirty Work of Social Media Marketing, One Task at a Time
This learn article outlines four social media marketing tasks AI can handle — audience analytics, content drafting and design, ad targeting and creative testing, and spam moderation — and cautions that taste, judgment, and data security remain human responsibilities.

AI Is Already This Good — Why Is Your Social Media Marketing Still Pure Manpower?
An overview of 18 AI tools for social media marketing, organized into six categories covering audience research, content creation, scheduling, comment and DM handling, ad management, and visual production, plus notes on personalization, prediction, and emerging trends.