← all posts

Refactoring an LLM Pipeline: Finding Where Tokens Leak

I have a pipeline that reads English technical articles and turns them into Korean introduction cards. It checks feeds from 178 vendors, selects useful material, writes cards, and serves them automatically. The output looks like this.

Illustration: Refactoring an LLM Pipeline: Finding Where Tokens Leak

Writing and assembly needed no human review. But I still had to inspect every selection and card. Is that automation, or has the work merely moved one step sideways?

I took the pipeline apart, starting with a simple question: why does one card cost so many tokens? There were 929 sources awaiting triage. Writing all of them would take over 40 million tokens. Even a small per-card saving would matter.

I went looking for savings and found two different problems: one in four generated cards was silently disappearing from the usable output, and I'd been optimizing only one of three cost dimensions. Neither was what I set out to find. Neither would have surfaced without measurement.

Let's start with the results.

Measured results: tokens per article 72,136 to 45,788 (-37%); tool calls 13 to 5; duration 282s to 121s; corruption 25% to 0; triage agreement 80% to 93%. Refactoring also revealed that 25% of written cards were silently disappearing.

A note on how I worked: this refactoring was done with a Claude Code session. The session performed most execution and measurement; I chose directions and questioned results. ‘I’ here records that joint work. Both the session's mistakes and mine appear below. I haven't edited either out. All reported figures are measured.


Overall structure

Four-stage pipeline: ingestion, triage, writing, assembly; output files represent state, with humans at review points

Four stages, only two using LLMs.

Layer Actor Work
[1] Ingestion Scripts, no LLM Check feeds, download and extract sources
[2] Triage Haiku, batched Apply gates and decide bring/skip
[3] Writing Sonnet, one agent per article Read the source and write a card
[4] Assembly Scripts, no LLM Upsert the DB and export serving files

Core design: file existence is state

There is no separate progress record. Each article has a directory; the files in it determine its state.

content/<id>/
  meta.json → original.md → triage.json → card.md → inclusion in cards.json = served
 (ingestion)   (source)      (triage)     (writing)          (serving)

Without a progress ledger, recorded progress can't diverge from files. After interruption, processing resumes where the next file is missing.

That principle broke twice during this work. More on that below.

Why this structure: paths I didn't take

Four layers and two LLM stages weren't predetermined. They came from painful experience.

One agent can fetch feeds, select articles, and write cards in a demo. But if it dies at item 40 in a batch of 60, progress living only in its context goes with it. No disk checkpoint means no resume. Separate stages contain failure to one article and let us restart at that boundary.

Ingestion and assembly don't require judgment. Extracting feed URLs and inserting cards into a database shouldn't need a model. Adding one makes the work expensive and nondeterministic, undermining reproduction and debugging. I reserved models for selection and writing and made the rest scripted.

I also avoided a workflow framework. State here is simply which files exist; many frameworks add a checkpoint database or execution graph. That creates two sources of truth and reconciliation code when they diverge. I removed state instead of adding it. Inspection is ls; resumption is ‘process the missing next file.’ That simplicity mattered more than framework convenience.

Layers also have different resource profiles: network I/O for ingestion, cheap batched inference for triage, isolated expensive inference for writing, and deterministic assembly. Combining them can bind cheap work to expensive resources.


[1] Ingestion: the source is the foundation

Scripts inspect 178 vendors' feeds, discover URLs, and extract source articles into Markdown. No LLM is involved.

One rule: no source, no card. Writers cannot fetch the web again, because changing source material would make repeated runs incomparable.

This stage determines downstream quality. Extracted site menus at the start of articles caused triage to reject good material. That's the next section.

Another issue: 22 of 33 served cards have no preserved source. They predate source storage, so rewriting and fact comparison aren't possible. I exposed that with a ‘missing source’ badge in the admin interface.


[2] Triage: selection is the product

A low acceptance rate isn't inherently bad. Not covering everything defines the product. An article must pass five gates to receive bring.

Gate
G1 Form: an article grounded in facts or concepts?
G2 Usefulness: substance usable in practice?
G3 Discovery: something readers may not have known?
G4 Graspability: digestible in a short read?
G5 Duplication: overlap with existing cards?

I still had to judge whether those decisions were correct. I reviewed 40 articles to investigate.

Eight rounds, 40 articles

Rather than build first and measure later, I ran a feedback loop immediately.

  • Each round used one article from each of five different vendors to reduce source bias.
  • The session judged, I reviewed, disagreements produced feedback, and the next round used five unseen articles.

Revised criteria were applied to new articles, separating feedback material from subsequent evaluation.

Eight rounds, 40 articles: agreement increased from 80% to 93% after switching from 60-line excerpts to full-text reading; most errors involved reading scope

Classifying five disagreements

Among the first 25 articles, five decisions differed from mine. Three weren't rubric problems.

Article Decision Reality
Brendan Gregg — No More Blue Fridays Extraction failed Intact body, 7,206 characters
Slack Engineering — Agent context No substance Intact body, 21,758 characters
Google SRE Book — Eliminating Toil Extraction failed Intact body, 13,726 characters

All three were exactly the kind of sources we wanted, rejected for the same reason. The judge was instructed to read only the first 60 lines, which these sites filled with menus or chapter contents. It saw navigation and declared extraction failure. It followed the instruction correctly; the instruction was the trap.

Treating all three as rubric feedback would have added rules to the wrong gates while leaving the real cause untouched.

First fix: the extractor. It failed

I expanded content-container candidates, removed elements whose class matched sidebar|toc|related, and removed link-dense <ul> elements.

Unit tests passed. Menus disappeared while prose lists remained.

On real sites, improvement was only 1.3%.

Site Menu lines among the first 60
Brendan Gregg 29 → 30
Slack 24 → 24
Google SRE Book 26 → 26

The actual HTML didn't match what I'd imagined. Those sites didn't implement their menus as <ul class="sidebar">. I reverted the change.

Something more embarrassing emerged: our claim that the body began at line 159 was based on the > source: URL we prepend. The detector counted lines over 80 characters as prose. We'd measured the problem with a broken ruler, then built a solution on top of it.

Remove the 60-line limit

If extraction wasn't the answer, the reading limit was next. original.md contained the full article, but the judge read only 60 lines to save costs. I'd never actually measured those costs.

Here was the estimate:

Reading scope Judgment input, 60-article batch
First 60 lines 44K tokens
Full text 314K tokens, 7.1×

Sevenfold sounds large, but compare it with a wrong decision. A wrong bring wastes about 45K writing tokens; a wrong skip loses a valuable source. Avoiding six wrong decisions in 60 articles would recoup the added input. Our observed disagreement rate was 20%, or twelve per 60.

I set a 400-line cap per article rather than an unlimited read—an outlier was 340 KB—and reduced batches from 60 to 30.

Rejudgment: recovering the three articles

Article 60-line decision After broader reading
Brendan Gregg SKIP: extraction failed BRING: eBPF verification mechanisms
Slack Engineering SKIP: no substance BRING: multi-agent context
Google SRE Book SKIP: extraction failed SKIP: insufficient discovery under G3

The third matters especially. It was still skipped, but because its content was familiar, not unreadable. Judgment finally rested on the article itself.

Agreement also improved.

Condition Sample Agreement
60-line limit, R1–5 25 80%
Broader reading, R6–8 15 93%

Actual cost grew much less than estimated: about 40–46K to 54–58K per round, roughly 25%. The judge read only what it needed within the 400-line cap. The sevenfold estimate was wrong too.

Three things the loop accomplished

We hadn't changed a single rubric line, yet the loop had already paid off three ways.

First, it found a silent failure. The judge returned ‘extraction failed’ normally, with no error or warning. Without human inspection, those high-quality sources would have kept disappearing.

Three of 40 experimental articles, 7.5%, suffered this issue. Extrapolated to 929 queued articles, roughly 70 might be rejected for the same reason. They weren't random sources, either: long navigation disproportionately affected older, well-organized technical sites.

Second, it prevented a bad rubric revision. I drafted feedback to reject an AWS Builders' Library caching article as generic advice without numbers. The session challenged me:

‘Is that criterion right? Principles and advice should pass when they fit our purpose.’

It was right. The rubric already accepted explanations of classic concepts and why they work. My new criterion was wrong. Adding ‘skip without numbers’ could have filtered out classic system-design articles that the product explicitly includes.

Feedback-loop risk isn't just overfitting. The feedback itself can be wrong, and this time the human supplied the mistake.

Third, the overfitting guard worked. The revision tool refuses fewer than three feedback items. At two, it wouldn't run. Manufacturing a third item would have bypassed our own safeguard.

That guard came from an earlier incident: writing policy v0.1 overfit one example, making every card follow ‘odd pattern → paradox → solution.’ One or two examples can distort a whole policy.

Then we revised—and measured the effect

At three feedback items, the gate opened. The revision tool added one line to G1.

+ skip: reactions to others' articles or announcements—extended quotation or
+ commentary on news/policy. Substance remains in the original; a card becomes a summary of a summary.

Two examples, charity and cloudflare, supported the change; an event recap was deferred for insufficient evidence.

The evaluation was:

  • Rejudge the two supporting examples: does the new rule catch them?
  • Judge three new articles: does it wrongly reject unrelated material?
Article Before revision After
charity: commentary on Mat Duggan BRING SKIP
cloudflare: commentary on an executive order BRING SKIP
Three new articles — BRING 2 / SKIP 1, as expected

The decisions explicitly cited the new rule.

‘A reaction quoting Mat Duggan extensively. The substance is in the original source; the card would summarize a summary.’

‘Policy commentary interpreting a White House executive order through Cloudflare's position. The substance is the policy itself.’

Language from my comment returned as a judgment rationale. The loop had completed a cycle.

There is a limitation: the two changed decisions were the examples used to revise the rule, effectively evaluating on training material. A genuinely new reaction article would be stronger evidence; none appeared among these three new articles.

The defensible claim is limited: the rule behaved as intended on those examples, with no regression in this small check. How well it absorbs editorial judgment needs more rounds.

The conclusion of this experiment

The loop's main success wasn't rewriting a rubric. It was discovering that the rubric wasn't what needed fixing.

Adding rules for every disagreement would have been natural. It would have added three clauses, left the reading limit untouched, and kept agreement around 80%. Classifying causes instead led to a one-line change and 93% agreement in the later sample.

The loop's value was telling us where to change, not merely changing criteria.


[3] Writing: craft is the moat

Each article gets one subagent to read its source and write a card. That isolates failure: one failed article doesn't take others down, and only that article needs rerunning.

This stage consumed the most tokens, so it got the most attention.

A cheaper model still reads the same input

Sonnet handled writing. I first considered switching to cheaper Haiku, but measured the workload before doing so.

Average source: 13,331 bytes (median 10,440; p90 25,422)
   ↓
Output body: 2,000 characters

This was read-heavy: input was five to ten times output. A cheaper model still reads the same 13 KB. The token price falls, not the token count.

I ran Haiku several times to test it. The observations shaped everything that followed.

Constraint Result
Use a polite conversational Korean register Failed: only 1 of 3 cards
Bullet format Succeeded: 100%
Do not calculate new numbers Succeeded
No length cap Overshot: 2,000 characters

The pattern: Haiku followed prohibitions and structural slots better than sustained writing-style instructions.

A writing register must be reapplied sentence by sentence; maintaining it over hundreds of sentences was difficult. A structural constraint, once set, tended to persist.

I split the task into three calls: extract → lay out → restyle. One job per call stabilized structure in 3/3 runs and reached 100% style conversion.

But the writing wasn't good.

Here are translated versions of sentences from the two models on the same source.

Sonnet: They don't trust the cache and ask the origin every time, while making actual data transfers almost unnecessary.

Haiku: A 304 Not Modified response means the artifact hasn't changed, so it is served directly from the proxy cache.

Same fact, but the first explains why the design is clever; the second lists behavior. Haiku also left ‘artifact’ untranslated twenty times, while Sonnet explained it once as a build output.

I kept Sonnet. Writing quality was the product's moat; sacrificing it wasn't optimization. I changed direction and searched for waste instead.


Change the measurement: file size wasn't the answer

I had been counting file size: reducing a policy from 11.7 KB to 5.5 KB, for example. Total tokens barely moved.

I inspected actual usage in the agent's jsonl execution logs. Each request records:

  • cache_creation_input_tokens: newly charged input.
  • cache_read_input_tokens: cached input, over ten times cheaper.
  • output_tokens: output.

A trap: content blocks sharing the same message.id repeat usage records. Without deduplication, I overcounted fivefold. That fooled me once.

Then I ran a decisive experiment: launch an agent that does nothing.

Prompt: ‘Reply only ok. Do not use tools.’
Result: 25,559 fresh input tokens

Receiving two characters, ok, cost 25,559 input tokens. The system prompt, tool schemas, and project configuration arrived as soon as the agent launched.

Of 48,932 input tokens per writing task, 25,559 (52%) are fixed overhead

Fixed overhead accounted for 52% of writing cost. The six KB I'd carefully removed from the policy were only a small part of that total.

First lesson: inspect billed usage, not file size, and measure fixed overhead with an empty task. If overhead is half the cost, shaving small documents won't move the total much.


The real lever: exact prefix reuse

I couldn't eliminate the 25.6K launch overhead, but I could serve it from cache.

Two measurements:

Prompt Fresh input Cache reads
Differs by even one byte 27,223 15,783
Exactly identical 0 43,006

The cached prefix needed an exact match up to the breakpoint. My prompt put the assigned article path at the very beginning.

<dir> = content/cloudflare--hyper-bug

That line differed for every article, invalidating reuse afterward. I paid again for a 43,006-token prefix per article.

A one-byte prompt difference invalidates the entire 43,006-token prefix before the cache breakpoint

The solution was structural: move the variable out of the prompt and into a file.

content/_state/write-queue/
  01.todo   → content/cloudflare--hyper-bug
  02.todo   → content/uber--modernizing-artifact-storage

A writer claims a .todo file by moving it to .taken with mv, then processes the path inside. The pipeline uses the rename as its claim operation. Every writer receives identical instructions.

At normalized pricing, the equivalent cost fell from 72,136 to 45,788 tokens, a 37% reduction.

One trap remained: starting all workers together makes them cold before the first fills the cache. I ran one first to warm it, then launched the others in parallel.

Another finding: changing parameters such as effort invalidated reuse, dropping cache_read from 15,783 to zero. Don't mistake that one-time transition cost for steady-state option cost in an A/B test. I made that mistake again later.


Tool round trips: thirteen calls became five

Usage inspection revealed thirteen tool calls during rewriting.

Write card.json
  → error: File has not been read yet
Read card.json      ← forced to read the previous version
Edit card.json

The harness prevents writing an existing file before reading it. Rewriting became Write failure → Read → Edit. Edit carries both old and new text, effectively paying for the body twice.

The fix was one step: remove the previous output before rewriting.

An unexpected finding followed: the forced old-version read contaminated the new card.

Version Title, translated
Original They send a request every time, but don't download
First rewrite Moving from legacy storage to managed SaaS plus a validation proxy reduced egress 99%
Second rewrite A validation proxy reduced SaaS artifact-store egress by over 99%

The second rewrite copied the first instead of resembling the original. ‘Write afresh’ didn't prevent anchoring when the old result entered context. The Uber card shown above is this article; the surviving title was ultimately the original one.

Deletion addressed quality as well as tokens.

I found another redundant sequence: despite an explicit field contract, the agent inspected other cards via find → Read → cat to check formatting. One instruction against using other cards as examples removed three calls.

Thirteen calls became five—four parallel reads and one write, the theoretical minimum here—and time fell from 282 to 121 seconds.


Then I found one in four outputs was unusable

While inspecting outputs from option experiments, I noticed something wrong.

json.decoder.JSONDecodeError: Invalid control character at line 5 column 675

Of eight outputs inspected, two—25%—were broken. It happened with both Sonnet and Haiku, independent of the option being tested.

The body mixes multiline Markdown, Korean text, and quotation marks. Embedding it in a JSON string requires escaping everything, and the model didn't do that consistently.

What happened afterward was worse.

def read_json(path):
    try:
        return json.load(...)
    except (FileNotFoundError, json.JSONDecodeError, ...):
        return None      # the silent failure is here

The pipeline used file existence as state: a card file meant written, no file meant waiting. But an existing unreadable file was treated like a missing file, pushing it back to waiting.

A 60K-token job could exist on disk yet vanish from the admin interface. Nobody would know.

Swallowed JSON parse failures reset the derived state, silently hiding completed writing

Why change the format instead of adding validation?

I could parse-check after writing and retry failures, or choose a format that didn't require escaping the body.

Validation reduces failures rather than removing the failure mode, and retries cost more. Haiku's self-validation loop reached sixteen tool calls.

I changed the format. Put the body where body text belongs, and it no longer needs JSON-string escaping.

---
slug: cloudflare-hyper-flush-race
title: The logs said everything was fine; only the kernel knew
tags: concurrency, networking, observability
---

## 200 OK, but only half the photo arrived

Quotation marks and line breaks in the body need no JSON escaping.

That removes this particular escaping failure structurally. I also exposed parse exceptions as flags in the admin interface instead of swallowing them.

The format change saved essentially no tokens: 66.1K versus 65.5K for the same article. It was a reliability change, not a cost optimization.


This time the session was wrong—and I almost believed it

The first effort: low test showed a 40% increase. The session rejected it, explaining that weaker reasoning produced trial-and-error loops. It sounded plausible, and I almost accepted it.

But lower reasoning effort consuming more tokens seemed odd. I asked whether the prompt needed adapting. Reexamining the experiment revealed two errors.

We hadn't changed the prompt. The lower-effort mode still had a self-gating instruction that encouraged validation loops rather than explicitly avoiding them.

We also counted cache invalidation as an option failure. Although we knew parameter changes invalidated the prefix, we charged that one-time transition against normal operations.

Removing those effects produced a 10% reduction instead.

But the rerun uncovered the 25% output corruption. The supposedly wasteful self-debugging loop had actually been catching that problem.

I still rejected the option. The 10% difference came from one fewer request; raw tokens were effectively the same. It was the right conclusion reached for the wrong reason, and the rerun found something much more important.

When results conflict with expectations, inspect the experiment first. The model supplied a plausible explanation, and the person nearly stopped questioning it.


Writing-stage results

Before and after: tokens fell 37%, from 72,136 to 45,788; tool calls 13 to 5; corruption 25% to 0

Quality held on the compared article: all nine key facts covered, with zero numbers absent from the source.

I also found the floor. The remaining roughly 45K breaks down like this:

Component Tokens or equivalent cost Can it shrink?
Prefix 43K, equivalent to 4.3K through caching Already about 90% discounted
Fresh input 18K: source reading and returned output Less source reading harms coverage
Output 10K: reasoning and card Lower settings harm quality

The last two trade directly against quality. That was the floor for per-item optimization.


[4] Assembly: no LLM

Scripts collect cards, upsert the DB, and export serving files. Neither a person nor an LLM is involved.

I fixed one defect: a rewritten card with a new slug left its old version served too. The DB used slugs as primary keys without linking cards back to sources. An internal tracking field now identifies and removes the source's old row.


The human's role: absorb review into criteria

Cost isn't one-dimensional

Near the end, I added a requirement: comments on selection decisions should accumulate into criteria revisions. It sounds like a separate feature, but it addresses the same cost problem.

Total cost = cost per item × number of items × rework factor

I'd focused only on cost per item. Reducing 45K to 28K helps once, but doesn't change which work happens. It's a constant-factor improvement.

Total cost equals unit cost times volume times rework rate; unit-cost improvements are fixed factors, better decisions compound

I measured the other two dimensions.

294 decisions: bring 78 / skip 216
  Acceptance rate: 26.5%
  Bring precision: 71.8% — accepted cards not rejected afterward

Of 78 accepted articles, 22—28%—were rejected or deferred after writing. We'd already spent roughly 45K each, about one million tokens on rework.

That's costlier than a triage rejection. A mistaken skip wastes judgment effort; a mistaken bring also wastes the full writing cost.

Improved judgment compounds. Fixing one type of mistaken acceptance prevents future work of that type. Per-item savings happen once; better selection accumulates value.


Keep humans in the loop without calling them every time

I built a feedback loop. Treating it only as a token-saving mechanism would miss the design. Its purpose is to accumulate human judgment; savings are a result.

Human comments accumulate in files; a manually triggered skill revises criteria in a human-on-the-loop workflow

Several details mattered.

Comments sit beside decisions in triage, writing, and serving lists. They accumulate in per-article files, not a new log table that would duplicate state.

The screen determines the feedback stage automatically: triage from selection, write from serving. Asking users to select it creates avoidable mistakes.

One article can receive both kinds. Correct selection but poor writing is common; one comment type per article couldn't express that.

A human explicitly triggers revision. Comments don't automatically change criteria, and fewer than three comments are rejected to reduce overfitting. That guard came from the real v0.1 incident where one example made all writing converge on the same structure.

I separated revision tools for selection gates and writing craft—style, structure, and voice. Combining them bloats the prompt and mixes different jobs, contradicting the measured benefit of one job per call.


Is a revision an improvement? Measure first

How would we distinguish improvement from damage?

If revised criteria reduce precision, we should roll back. Without measurement, we only have shifting preferences.

Instrumentation needs to precede revision. Add it later, and there's no baseline.

The earlier work taught exactly that: file sizes misled me; usage records revealed the real cost. To avoid repeating it, I needed metrics first.

I chose three metrics and recorded a baseline.

Measured baseline: triage pass rate 26.5%, bring precision 71.8%, writing acceptance 42.3%

Bring precision is central. A low acceptance rate isn't automatically bad; filtering is the job. Low bring precision means paying to write and then discarding the result—the most expensive failure.

Snapshots go into an append-only file. Overwriting a baseline defeats measurement. Revision tools automatically capture before-and-after snapshots, so records remain even if a person forgets.


Conclusion

On measurement

  1. Inspect billed usage, not file sizes. Fresh input and cache reads have very different prices.
  2. Measure fixed cost with an empty task before optimizing smaller components.
  3. Question experimental design when results contradict expectations. Models can generate plausible explanations too.
  4. Record a baseline before changing the system.

On LLM pipelines

  1. Exact prefix reuse matters more than shaving a small file. Move variables out of shared prompts.
  2. Measure waste before downgrading the model. Tool round trips alone fell from thirteen to five.
  3. Avoid asking LLMs to escape long bodies into JSON strings; two of eight outputs broke here.
  4. Don't swallow parse failures. except: return None can silently erase usable work.

On cost structure

  1. Per-item optimization is a constant factor; better judgment prevents repeated unnecessary work.
  2. Writing and then rejecting is expensive. A false acceptance pays the entire downstream cost.

I started with token optimization. The two findings that mattered most weren't optimizations: one in four outputs had been unusable without me knowing, and I'd been working along only one cost axis.

Neither would have surfaced without measurement. Neither was what I'd set out to find.