Skip to content
    Back to articles

    How to Calibrate AI Controls to Risk

    September 19, 2026 7 min read
    Share
    Overhead desk showing light and heavy review trays, illustrating how to calibrate ai controls to risk

    Somewhere in your organization right now, a marketer is waiting forty minutes for sign-off on an AI-drafted internal email that nobody outside the team will ever see. Down the hall, a first draft of external creative went out under someone's name with no second look at all. Same tools, same leadership approval, wildly different exposure. Nobody designed it this way on purpose.

    This piece is about why that mismatch happens and what fixes it: controls sized to what a mistake would cost.

    How much human review does a given AI task need

    A task needs enough human review to match what happens if it is wrong, no more and no less. Low-stakes, reversible, low-visibility work needs a light touch. Anything that is hard to unwind, seen by a client or regulator, or attached to a real decision needs a human who can be named as accountable for the outcome.

    Most teams do not ask this question at all. They ask a different one: was this made with AI? That question feels like governance but it is not, because it sorts work by how it was produced instead of by what is at stake. An AI-drafted internal Slack message and an AI-drafted client-facing proposal get treated the same way if the rule is about the tool rather than the task.

    The fix starts with naming, for each recurring category of AI-assisted work, what a bad outcome looks like and who would have to clean it up. A rough tagline gets rewritten. A misquoted statistic in a client deck gets forwarded, screenshotted, and remembered for years. Those are not the same risk, and they should never carry the same checkpoint.

    Reversibility is the variable that matters

    Can this be fixed after the fact, or is it out the door the moment it is sent? A social post can be deleted. A number cited in a board memo cannot be un-seen. Sort your AI-touched work by that question before you sort it by anything else.

    Why our AI governance feels too heavy for some tasks and too loose for others

    Governance feels mismatched because it was built once, in a hurry, for the loudest fear in the room rather than for the actual range of tasks now touching AI. One blanket policy gets stretched over work with very different consequences, so it is simultaneously too slow for the safe stuff and too thin for the risky stuff.

    This usually traces back to how the policy got written. Leadership approved the tools, someone in legal or compliance drafted a single set of guardrails under time pressure, and that document then had to cover everything from internal brainstorming to client-facing deliverables. A single lane trying to carry every kind of traffic will always feel wrong to somebody.

    The practical answer is to build swim lanes instead of one rule. Low-risk, high-reversibility work gets a fast lane with light or no review. Anything touching a client, a regulator, or a public audience gets a slower lane with a named reviewer and a real checkpoint. Teams stop feeling micromanaged on the easy stuff and stop improvising on the hard stuff, because the lane tells them which mode they are in before they start.

    Guardrails need clarity more than they need volume

    People do not need every edge case pre-solved. They need to know where they can experiment freely, where caution is required, and who owns the judgment call when something falls in between. That clarity is worth more than a longer policy document.

    How do we decide which AI outputs need approval before they go out

    An output needs approval before release if getting it wrong would be expensive, hard to reverse, or visible to someone outside the team who did the work. Everything else can move on trust, spot checks, and after-the-fact review rather than a pre-publication gate.

    A useful way to sort this is to ask three questions about any given piece of AI-assisted work: who sees it, what does it cost to be wrong, and can it be corrected quickly if it is. Client-facing copy, anything with numbers attached, and anything tied to a legal or financial commitment usually fails all three and needs a human check before it ships. Internal drafts, early-stage ideation, and low-stakes variations usually pass all three and do not.

    Research on AI performance across different task types makes a related point worth building into this decision: AI does not fail evenly across tasks, it fails unevenly, doing well in some areas and badly in adjacent ones that look similar on the surface. That argues for review gates placed at the boundary of what a given AI use case is good at, not at the boundary of what feels comfortable to a nervous manager.

    Why is adoption stalling even though leadership approved the AI tools

    Adoption stalls when the tools were approved but the decision rights around them were never redesigned, so people are left guessing what they are allowed to do without asking permission each time. Approval at the top does not automatically produce clarity in the middle, and clarity is what drives use.

    This is the messy middle showing up in a governance costume. Leadership announced the tools, ran a training session, and moved on, assuming adoption would follow. Meanwhile the person doing the work has no idea whether a given task needs sign-off, whose name is on the outcome if it goes wrong, or whether asking that question will make them look like they do not trust the new system. So they either freeze, over-check everything to be safe, or do their own thing and hope nobody notices.

    That last pattern, shadow AI use that never surfaces in an official adoption number, is not evidence of bad intent. It is a signal that the sanctioned path felt slower or riskier than going around it. Treat it as data about where your controls are miscalibrated rather than as a compliance problem to stamp out. If you want a clearer picture of where your own organization sits on this, the AI Profit Readiness Assessment is built to surface exactly this kind of mismatch between what was approved and what people feel permitted to do.

    What does appropriate AI oversight mean for marketing work

    Appropriate oversight for marketing work means the review a piece gets is proportional to its reach and its reversibility, not to whether AI touched it. A social caption and a paid campaign built on the same AI-assisted process deserve very different levels of scrutiny before launch.

    Marketing carries a particular version of this problem because so much of the output is externally visible and hard to fully unwind once it is live. A tone mismatch in a client-facing deck is not a small thing; it can erode trust that took years to build, which is exactly the kind of risk that experienced creatives tend to spot faster than anyone else in the room. Their hesitation is frequently ground truth about where quality is at risk, not resistance to the technology itself.

    The practical move is to separate production risk from distribution risk. Drafting, ideating, and internal iteration can run fast and loose because nothing has shipped yet. The moment something is scheduled to go external, whether that is an ad, a client-facing document, or a public statement, it should pass through a named reviewer who owns the call. For teams trying to redesign this workflow properly rather than patch it, the AI Profit Sprint works through exactly this kind of swim-lane design with the people who will be living inside it.

    Watch for the overconfident adopter

    The person moving fastest with AI is not automatically the one taking the least risk. Overconfident use of a tool in an area it is not strong in often needs more review, not less, even when that person is your most enthusiastic user.

    If your governance still treats every AI-touched deliverable the same way, that is worth putting in front of someone before the next client-facing mistake makes the decision for you. A short conversation is often enough to see where the mismatch sits: book time with Average Robot and bring the specific workflow that has been bothering you.

    Take it with you

    Download this as a PDF

    A clean, branded version to read offline or share with your team.

    Frequently Asked Questions

    Usage describes how much AI touches a workflow. Risk describes what happens if that output is wrong, who sees it, and how hard it is to reverse. Two tasks can have identical AI usage and completely different risk, which is why controls tied to usage alone tend to misfire.

    Weekly newsletter

    With People

    Every Tuesday: the people side of making AI pay.

    • One story from the week's news about AI at work, read through our book, The Elephant in the Algorithm.
    • A few links worth your time, and most weeks a podcast conversation.
    Privacy Policy
    Share