A Sufficiently Detailed Specification Is Indistinguishable from Code
When Karpathy called English the hottest new programming language, many people nodded. So did I. After writing natural-language documents to control agents and repeatedly watching them break, I started to wonder: perhaps English hasn't become a programming language. Perhaps it only works when we write it like one.
Background
Claude Code has a system called Skills. Put a Markdown document under .claude/skills/, and AI reads it to perform a task: send a Slack message, approve a database change, or validate code.
I spent a while building and breaking these. The same document worked one day and failed the next. Different conversation contexts made AI interpret the same sentence differently. I built a framework to validate Skill documents. These are the observations that came out of it.
Code is an allowlist; prompts behave like a blocklist
I wrote a Skill to approve database changes. The first version looked like this.
## Process
1. Validate SQL
2. Check the Jira ticket
3. Approve the change
It usually worked. Sometimes, during SQL validation, AI would announce a performance issue and rewrite the query. I hadn't asked it to. But I hadn't forbidden it either.
That's the crucial difference from code. Write validateSQL(query), and it performs only the implemented validation logic. Without query-modification code, modification isn't available. But ‘validate SQL’ in natural language can be interpreted to include fixing it.
Code executes implemented behavior; omitted behavior isn't there. Natural-language instructions leave room for AI to fill in anything not explicitly prohibited. The default control model is reversed.
As a firewall analogy, code resembles an allowlist: only explicitly permitted traffic passes. A prompt can behave like a blocklist: anything not denied may pass.
How I narrowed the boundaries
Fixing the SQL-rewriting issue revealed another gap.
After adding prohibitions, AI guessed a URL when given something other than a Slack link. Once I blocked that, it sent another approval request for a Jira ticket already approved. Close one gap, and another appears. That's the weakness of a blocklist: unblocked paths remain open.
I built the validation framework to make this trial and error systematic. Failure patterns fell into several categories.
One was acting beyond the intended scope: intent drift. Phrases such as as needed, if appropriate, and consider frequently triggered it. They're like variables typed any: almost anything fits. I made a checklist: at least two explicit prohibitions, no vague delegation language, and explicit criteria wherever AI could branch on its own judgment.
Another was context contamination. A preceding ‘refactor aggressively’ conversation could make the next Skill behave aggressively too. For an ordinary function, that would be strange. calculateDiscount(price, rate) shouldn't change its result based on the previously called function without some shared state. For LLMs, the whole conversation effectively acts as global state.
I quickly learned that an isolation statement—‘ignore previous context’—wasn't enough. I shifted from trying to prevent contamination to assessing its blast radius. A contaminated read-only Skill may do little damage; one changing a DB or approving Jira work can affect production. I classified blast radius as Local, Team, or Production, and failed Production-level Skills without a prior confirmation gate.
This produced seven validation dimensions. A Skill review returns something like this.
/validate-skill approve-db-changes
[1] Intent Drift Prevention WARN — No prohibited-actions section; 0 prohibitions
[2] Intent Corruption Prevention FAIL — No invariants or input/output examples
[3] Intent Completeness FAIL — Invalid input, missing tickets, API failures undefined
[4] Trigger Clarity PASS — 3 clear triggers; no conflicts
[5] Environment Dependencies WARN — Token specified; acquisition path missing
[6] Context Contamination Risk WARN — Jira approval (Production) + Slack (Team); unclear gates
[7] Testability FAIL — Responsibilities not separated; failures indistinguishable
On FAIL, I revise and run it again. The final document after that loop looked like this.
## execute@approve-db-changes
**Input**: Slack message link (https://team.slack.com/archives/...)
**Output**: Jira approval confirmation or an error
### Process
1. Extract the SQL query from the Slack link
2. Validate SQL syntax only
3. Retrieve the Jira ticket status
4. Show the approval details and request Y/N confirmation
5. Approve the Jira ticket only if the response is Y
### Invariants
- Order: SQL extraction → syntax validation → Jira lookup → confirmation → approval. Do not change it.
- If SQL validation fails, do not continue.
- Do not call the Jira API before the user enters Y.
### Prohibited actions
- Do not modify the SQL or suggest alternatives
- Do not analyze SQL performance, inspect execution plans, or suggest indexes
- Do not change Jira ticket status without validation
- Do not invent an approval rationale
- Do not approve without user confirmation
- Do not retry on error
### Edge cases
- Not a Slack link → Return: Please enter a Slack message link
- Slack message without SQL → Return: No SQL query found
- Ticket already approved → Return: Ticket already processed
- Missing ticket → Return: Ticket not found
- API failure → Return an error without retrying
- User enters N → Return: Approval canceled
### Requires
- JIRA_API_TOKEN
- SLACK_USER_TOKEN
The original three lines became input and output types, enforced ordering, six prohibitions, branching criteria, error handling, and dependency declarations. Replace the prose with symbols, and it's essentially code.
Then why use an LLM?
Room for interpretation is both a problem and a source of LLMs' power. They fill gaps with judgment. ‘Validate SQL’ may lead to syntax checks or performance analysis depending on context. Implementing that in code requires explicitly writing the branches.
I was systematically removing that flexibility. Taken to the extreme, a fully constrained prompt is just code, and using an LLM loses its purpose.
My working rule became: constrain behavior more tightly as actions become less reversible. Changing databases, approving tickets, and production deployments make discretion risky. Code review, research, and drafts benefit from openness because the model can supply perspectives I missed.
The criterion is simple: can we afford the cost of AI being wrong? If so, leave room. If not, constrain it.
Testing
There's another reason to constrain behavior: testing.
A conventional test supplies an input and compares output: assertEquals(expected, actual). That assumes the same input produces the same output.
Agents don't reliably satisfy that assumption. The same Skill and input can produce different results with different conversation contexts. Yesterday's test may fail today, and a model update can change them all.
I decided to test behavioral boundaries rather than exact output values. Did the agent request confirmation before sending a Slack message? Did it skip SQL validation? Did it invent a value for an empty input?
Here the perspective flips again. Code tests ask whether the right thing happened. Agent tests also need to ask whether the wrong thing was avoided. This resembles security testing more than ordinary unit testing.
Negative checks alone can give full marks to an agent that does nothing: it violated no prohibition. Positive checks are necessary too. It's a question of priority. Not destroying a DB comes before producing an optimal query. Discuss performance after establishing safety.
Testing boundaries requires defining them first: prohibitions, invariants, and confirmation gates. Without those, there is nothing clear to test. Making an agent testable therefore reduces interpretive freedom. Testing needs constraints, and constraints reduce flexibility. But deploying without tests isn't acceptable either.
Conclusion
Traditional programming specifies what to do. Prompting often also requires specifying what not to do. Constrain natural language enough, and it becomes indistinguishable from code. Defining, validating, and strengthening those constraints resembles designing a language.
We don't need to close every door. Close everything, and there's little reason to use an LLM. Leave everything open, and it drifts beyond intent. Knowing what to leave open matters as much as knowing what to constrain.