
That gap — between "the steps ran" and "the intended outcome was achieved" — is the subject of a recent explainer on executable operational specifications, which argues that automation needs a companion layer: a machine-readable description of intent that can be checked against evidence, independent of whatever tool actually performed the work.
Execution Is Not the Same Question as Intent
Think about what a typical deployment pipeline actually proves. It builds, tests, packages, and ships an application, then reports success. What it has verified is that a sequence of configured steps completed without error. What it usually has not verified is whether the resulting system matches what anyone actually wanted: the correct version running, enough healthy replicas, acceptable error rates, a rollout confined to the right region.
Those two things — execution success and operational conformance — sound similar but are logically distinct. A script can be flawless and the outcome can still be wrong, because the script only knows how to act, not what "correct" means in this specific case.
Part of the reason this gap persists is that operational knowledge in most organizations is scattered rather than centralized. The number of replicas that should exist might live in a Kubernetes manifest. The acceptable error rate might live in a dashboard threshold. The rollback trigger might live in a senior engineer’s memory or an old incident ticket. Each artifact captures a fragment of intent, but no single artifact says, in one place, "this is what we’re trying to achieve, these are the limits, and this is how we’ll know if we succeeded." When that unifying statement is missing, the tool that executes the work becomes the de facto definition of correctness — and it becomes very hard to ask whether the tool behaved well, because there’s nothing independent to check it against.
Making Intent Something a Machine Can Check
An executable operational specification tries to close that gap by writing the objective itself in a form that can be evaluated later against observed evidence: what should happen, what constraints must hold, what evidence proves it, and how conformance is judged. It doesn’t replace the deployment script or the infrastructure tool — it sits next to them as an independent reference point. This is conceptually close to ideas already familiar in software engineering: an interface separates a caller from an implementation, SQL separates a query from how the database physically stores data, and "desired state" systems in infrastructure automation continuously reconcile what’s running against what’s declared. An operational specification applies that same separation to the broader question of whether an operation — not just a configuration file — achieved its purpose.
The practical value shows up the moment something goes only partly wrong. Instead of a flat "deployment failed," a system checked against a specification can report that execution completed, but the operation still failed because, say, available replicas dropped to two when three were required. That is a far more useful failure message than a boolean.
It’s worth being precise about what this buys you and what it doesn’t. A dashboard telling you the error rate is 1.4% is an observation — a fact about the present. It says nothing about whether 1.4% is fine or a violation, because that judgment depends on an expectation that lives outside the metric. A specification supplies that expectation. Put differently: observability tells you what a system is doing; a specification tells you what it should be doing. You need both, and neither substitutes for the other.
| Layer | What it answers | What evidence it uses | What it cannot do alone |
|---|---|---|---|
| Execution | Did the steps run to completion? | Pipeline logs, exit codes, task status | Cannot say whether the outcome was correct |
| Observability | What is the system doing right now? | Metrics, traces, dashboards, alerts | Cannot say whether that state is acceptable |
| Specification | What should the system be doing, and did it? | Declared constraints evaluated against evidence | Cannot generate evidence or run the operation itself |
From Intent to Verification: The Missing Chain
Laid out as a sequence, the relationship between the people setting goals, the specification, the executor, and the resulting judgment looks like a short chain rather than a single leap from "click deploy" to "success":
flowchart TD A[Human / organizational intent] --> B[Operational specification] B --> C[Executor: pipeline, script, or agent] C --> D[Observed evidence] D --> E[Conformance evaluation]
Every link in that chain matters. Skip the specification, and evidence has nothing to be measured against. Skip the evidence, and the specification is just an aspiration. The chain is also where executor independence becomes valuable: because the specification names the outcome rather than the mechanism, the same objective — deploy this version, keep replicas above three, keep latency under a threshold — can be handed to a Kubernetes controller, a managed cloud platform, or, increasingly, an AI agent, without rewriting what "success" means each time.
Why This Matters More Once Agents Choose Their Own Steps
That last executor is the one that changes the stakes. A traditional automation script follows a fixed, predictable path, so verifying it is largely a matter of checking whether it followed the script. An AI agent given an operational goal — restore checkout availability, resolve an incident, provision an environment — may choose a different sequence of actions each time: inspect logs, restart a service, adjust a feature flag, reroute traffic, in whatever order it judges effective. You can no longer verify the agent by comparing its actions to one canonical script, because there isn’t one. What you can still verify is whether the outcome satisfies the stated constraints, regardless of the path taken to get there.
This is precisely the tension showing up in AI safety research on agent evaluation, though in a different setting. One research framework for testing frontier AI models, described as AutoControl Arena, builds evaluation environments on what its authors call "logic-narrative decoupling": deterministic parts of the environment — file systems, databases, permission checks — are grounded in executable code, while open-ended, unpredictable parts, like how a simulated colleague responds, are left to a language model to generate. The stated reason is that letting a model freely generate the state of the world invites "logic hallucination" and untrustworthy tests, whereas anchoring the parts that must stay consistent in real code lets researchers evaluate behavior against a state they can actually trust. The paper is about evaluating AI risk in test environments, not about running production operations — but the underlying instinct is the same one animating operational specifications: separate the part that must remain a fixed, checkable fact from the part that is allowed to vary, and evaluate against the fixed part.
Industry conversations around software delivery describe a related shift, often summarized as moving attention from implementation to intent — from "how did the code do it" toward "what was the code supposed to achieve." It’s a framing rather than a settled methodology, but it echoes the same underlying instinct: the thing worth preserving across tool changes is the goal, not the mechanism.
What a Specification Cannot Fix
None of this converts operations into something that runs itself safely. A specification is only as good as the judgment behind it — a badly chosen constraint, a metric nobody validated, or an incomplete evidence requirement will produce false confidence just as easily as a missing spec produces none at all. It doesn’t resolve organizational disagreement about what the right outcome even is, and it doesn’t replace security review, ownership, or human sign-off. Real incidents involving autonomous coding agents illustrate this starkly: postmortem work in this area increasingly asks not just what broke, but whether an agent operated under declared constraints, whether a human approval gate actually existed and was enforced, and whether the deployment record can even prove who — or what — made the decision. A stated constraint that is never enforced at the moment of action is not meaningfully different from no constraint at all; the specification has to be checked, not merely written down.
The Real Shift
The value of an executable operational specification isn’t that it makes automation infallible — it’s that it gives failure a vocabulary. Instead of "deployment failed," teams can say which constraint failed, against what evidence, compared to what was intended. As execution keeps getting faster and more autonomous, that vocabulary matters more, not less. The right question to ask of any automated system, human-written script or AI agent alike, is no longer just "did it run?" It’s whether anyone can show, with evidence, that it ended up where it was supposed to be — and whether a human is still positioned to judge if that was the right place to end up.


