← all posts

Choosing Transaction Boundaries: Narrowing Them with the Outbox Pattern

0. Test environment and metrics

  • Load generator: k6 (ramping-vus).
  • Database: MySQL, InnoDB.
  • ORM: Spring Data JPA.
  • Connection pool: HikariCP, maximumPoolSize=10.
  • Metrics:
    • API latency: p50, p95, p99, max.
    • HikariCP active and pending counts.

1. The problem

The reward-earning API, POST /rewards/earn, records a completed order in our database, writes an audit log, and synchronizes with a partner API.

Raising the load to 50 VUs caused response times to spike.

Average latency was roughly 6.2 seconds, with 3,510 requests processed.

Soon after load began, active connections reached the maximum of ten. Pending requests then accumulated, peaking at 35.

That suggested the application was holding connections too long, not simply that the DB was slow.


2. Why I originally designed it this way

Looking back, putting external HTTP calls inside a transaction seems questionable. At the time, the reasoning felt sensible.

  • If rewards are earned, the audit log and partner synchronization should succeed too.
  • If something fails midway, rolling everything back seems safer.

It looked like one business workflow, so I wrapped it in one transaction.

But a transaction doesn't make every step of a logical workflow atomic. It also determines how long connections and locks remain held.


3. Transactions lived too long

Two common causes of pool exhaustion are:

  1. Slow queries holding connections for a long time.
  2. Fast queries inside broad transactions, delaying commit and connection return.

The second looked overwhelmingly likely here. Pending requests accumulated quickly, and requests appeared to stall in one particular part of the flow.


4. External I/O inside the transaction

This was the original code.

@Transactional
fun earnReward(request: EarnRewardRequest): EarnRewardResponse {

    // 1) Acquire an exclusive lock
    val summary = userRewardSummaryRepository
        .findByUserIdAndYearMonthForUpdate(userId, yearMonth)

    // 2) Update internal state
    summary.totalAmount += request.amount

    // 3) Write the ledger
    val ledger = rewardLedgerRepository.save(
        RewardLedger(...)
    )

    // 4) Write the audit log
    auditLogRepository.save(
        AuditLog(...)
    )

    // 5) Synchronize with partner over HTTP
    partnerApiClient.syncRewardPoints(
        userId = request.userId,
        amount = request.amount,
        ledgerId = ledger.id
    )
}

The partner call was synchronous.

@Component
class PartnerApiClient(
    private val webClient: WebClient
) {
    fun syncRewardPoints(userId: Long, amount: Long, ledgerId: Long) {
        webClient.post()
            .uri("/api/v1/points/sync")
            .bodyValue(
                SyncRequest(
                    userId = userId,
                    amount = amount,
                    idempotencyKey = ledgerId.toString()
                )
            )
            .retrieve()
            .bodyToMono(Void::class.java)
            .block() // Wait for the response
    }
}

The key is .block().

  • The thread waits for the network response.
  • The transaction stays open.
  • The connection and exclusive lock remain held.

External network latency becomes connection-holding time and lock-holding time.


5. Estimating the upper bound

With a pool of ten, if an external call adds an average of 100 ms to each transaction:

Throughput upper bound ≈ poolSize / holdingTime
                = 10 / 0.1
                = 100 tx/s

It's a simple estimate with important implications.

  • Only ten requests can hold connections; the rest wait.
  • The pending queue grows and p95/p99 climb.

6. Lock waits tie up more connections

SELECT ... FOR UPDATE makes it worse.

  • Request A holds the exclusive lock while waiting for the partner.
  • Requests B through N wait for the same row's lock.
  • Those requests wait inside transactions, holding their own connections.

When one userId becomes a hot key:

  • A lock queue becomes a connection-holding queue.
  • The entire pool is quickly consumed.
  • Requests for unrelated users are affected too.

Connection scarcity was a visible consequence.

Broad transactions, lock contention, and external I/O combined to exhaust the pool.


7. The wrong atomicity boundary

Ask which operations actually need to commit together:

Operation Atomic for internal consistency? Benefit of inclusion
summary UPDATE + ledger INSERT ✅ High
audit INSERT ❌ Low
Partner API call ❌—not atomic with our DB None

Putting partner synchronization inside a database transaction doesn't make it atomic with that database.

  • External call succeeds, internal DB rolls back: partner credits rewards, internal state doesn't.
  • External call fails, internal DB commits: internal rewards exist, partner state doesn't reflect them.

Consistency with an external system requires idempotency, retries, and asynchronous processing—not just a local transaction.


8. Why Outbox instead of @Async?

There are several ways to move the call outside the transaction. The simplest is @Async.

Option A: @Async

Advantages:

  • Easy implementation and immediate latency improvement.

Disadvantages:

  • Work can be lost when the process restarts.
  • Retries, status tracking, and operational recovery need their own implementation.

Option B: Transactional Outbox

Advantages:

  • Database changes and the record of work to perform commit together.
  • Explicit PENDING, DONE, and DEAD_LETTER states support operations.
  • Pending events survive restarts and can be retried.

Disadvantages:

  • Polling adds delay rather than immediate partner updates.
  • The outbox needs cleanup, indexes, and monitoring.
  • At-least-once delivery requires idempotency.

Losing work was hard to accept, and traceability mattered more. I chose Option B.


9. Narrower transactions with an Outbox

The main transaction handles internal consistency only.

@Transactional  // Shorter connection occupancy
fun earnReward(request: EarnRewardRequest): EarnRewardResponse {

    val summary = userRewardSummaryRepository
        .findByUserIdAndYearMonthForUpdate(userId, yearMonth)

    summary.totalAmount += request.amount

    val ledger = rewardLedgerRepository.save(...)

    // Record deferred work in the same transaction
    outboxEventRepository.save(
        OutboxEvent(
            type = "REWARD_EARNED",
            payload = toJson(ledger),
            status = "PENDING",
            retryCount = 0
        )
    )
}

A separate scheduler processes the work.

@Scheduled(fixedDelay = 500)
fun processOutbox() {
    outboxEventRepository.findPending(MAX_RETRY).forEach { event ->
        try {
            val payload = parse(event.payload)

            partnerApiClient.syncRewardPoints(
                userId = payload.userId,
                amount = payload.amount,
                ledgerId = payload.ledgerId
            )

            auditLogRepository.save(...)
            event.status = "DONE"
        } catch (ex: Exception) {
            event.retryCount++
            if (event.retryCount >= MAX_RETRY) event.status = "DEAD_LETTER"
        }
        outboxEventRepository.save(event)
    }
}

10. Trade-offs accepted

10.1. Define delay instead of promising immediacy

With a 500 ms polling interval, partner synchronization is no longer immediate.

  • Average polling wait: about 250 ms.
  • Up to about 500 ms of polling wait, plus processing and queueing time.

I chose a bounded polling interval over immediate execution. That is a requirements decision; processing and queueing still add time.

10.2. Exactly-once doesn't appear automatically

Outbox delivery is generally at least once. Duplicate delivery is possible, so the partner API needs an idempotency key.

  • Use ledgerId as the idempotency key.
  • Repeated requests with the same key must not credit rewards twice.

10.3. The Outbox is another operational component

Introducing it creates these responsibilities:

  • Monitor and alert on accumulating PENDING events.
  • Clean up DONE rows according to retention policy.
  • Tune polling indexes and batch sizes.
  • Handle DEAD_LETTER events through investigation and manual reprocessing.

11. Results

Metric Before After Change
avg 624.1ms 15.9ms -97%
p50 477.7ms 4.7ms -99%
p95 1,483.2ms 58.7ms -96%
p99 1,746.3ms 239.2ms -86%
max 2,734.7ms 1,309.0ms -52%
Requests processed 3,510 21,725 +6.2×

Average response time fell sharply, and requests processed in the test increased roughly sixfold.

  • Active connections reached their maximum later and stayed there for less time.
  • Pending counts fell by more than tenfold.

This didn't require query tuning or a broad architectural rewrite.

The improvement came from separating network waiting from transaction resources: connections and locks.


12. Key takeaways

  • A transaction boundary is a resource-holding interval, not just a code block. Business consistency determines what belongs together.
  • External calls inside transactions turn network latency into DB resource occupancy.
  • An Outbox isn't magic. It's a basic mechanism for durable asynchronous work.
  • You pay for it with delayed consistency, operational work, and idempotency design.