Skip to main content

Command Palette

Search for a command to run...

Your Worker Read the Past. Then It Saved It.

Replication lag doesn't just confuse users. It lets background jobs write stale data that never heals.

Updated
•7 min read•View as Markdown
Your Worker Read the Past. Then It Saved It.
J
I'm a software engineer who spends most days building systems that solve real problems. When I'm not shipping code, I'm either untangling a tricky problem or writing about what I learned doing it. Currently exploring AI on the side.

A customer changes her email address. The settings page says "Saved," and a refresh shows the new address. Two weeks later, receipts still go to the old one.

The database is correct. The email service isn't. The sync job that connects them ran a second after she clicked Save and logged success. It just read her record from a moment before the change, and saved that moment where nothing would revisit it.

That's the replication-lag bug nobody writes about: not the stale page a user sees, but the stale read your own code writes down.

Replication lag in one paragraph

A read replica is a copy of your primary database that replays every change the primary makes. Writes go to the primary; most reads go to the replica. The replica is always slightly behind: milliseconds usually, seconds under heavy write load, sometimes longer. That gap is replication lag.

The famous version heals. The quiet version doesn't.

The well-known case: a user saves a profile, the next page reads the replica, and shows the old value. They refresh, the replica has caught up, and it's fixed. Nothing wrong was stored.

Now replace the user with a background job:

  1. The app commits the new email to the primary.

  2. It queues sync user 42.

  3. An idle worker picks it up within milliseconds.

  4. It loads user 42 from the replica, which hasn't replayed the change.

  5. It sends the old email to the email service, logs success, and completes.

Moments later the replica catches up. The database is right everywhere. The email service stays wrong, because the only event that could have fixed it was already processed. No exception, no retry, no alert.

A stale read shown to a person is a display problem. A stale read handed to your code is a data problem.

A public GitLab issue from May 2026 documents this exact shape in their search indexing: a reindex triggered by an update read the record from a lagging replica, wrote the pre-change version into the search index, and cleared the pending entry with no retry. Search showed the old value for close to two hours, until the item was edited again.

Why after_commit doesn't protect you

The louder cousin of this bug: you dispatch a job inside a transaction, the worker runs before the commit, and you get "record not found." The standard fix is dispatching after commit (after_commit in Rails, afterCommit in Laravel, and the pattern Sidekiq's troubleshooting guide recommends).

That closes the race between the commit and the message. It doesn't close the race between the commit and the replica. Once you add read replicas, the same race moves one hop down and goes silent.

Framework stickiness doesn't help either. Laravel's sticky option, for example, sends reads to the primary after a write only within the same request cycle. A queued job runs later in another process and doesn't inherit it.

Why it only happens in production

  • Dev and tests: one database, no replica, no race.

  • Staging: often no replica, or an idle one with near-zero lag.

  • Production: idle workers grab jobs instantly and lag grows with write volume, so the bug peaks during imports, spikes and migrations, exactly when odd data gets ignored.

It's intermittent. Most jobs arrive after the replica catches up. The occasional early one writes the past.

The mental model: every message races the data it announces

A job that says "go look at user 42" travels through your queue. The change it refers to travels through replication. Nothing guarantees which arrives first.

Any message that says "go look" instead of "here it is" is betting the data got there first.

The same race shows up with webhooks (a partner calls your API right after your "updated" event, and the API reads a replica) and with cache refills after invalidation (the refill reads a lagging replica and caches the old value with a fresh expiry).

Four fixes, simplest to most thorough

1. Send the facts, not just the ID

Put the new values in the message, plus a version number so workers can discard messages older than what they've already applied. No read, no stale read.

2. Carry the commit position

In Postgres, record the primary's WAL position when queuing, and check the replica has replayed that far before reading:

-- on the primary, when queuing the job
SELECT pg_current_wal_lsn();            -- e.g. 0/3A2F1C8

-- on the replica, before the worker reads
SELECT pg_last_wal_replay_lsn() >= '0/3A2F1C8'::pg_lsn;
-- false: the replica hasn't seen the commit yet. Wait, or read the primary.

GitLab's Sidekiq setup works this way: jobs are enqueued with the database position, and each worker declares its data-consistency level. The preferred level guarantees the replica has caught up to the enqueue point, or falls back to the primary.

3. If a read decides a write, read the primary

Syncing, indexing, sending, totalling: read from the primary. Replicas are for pages a human can refresh. In Laravel that's one call:

$user = User::onWriteConnection()->find($this->userId);

4. Don't complete the job until you've seen the version

public function handle()
{
    $user = User::find($this->userId);      // may hit a replica

    if ($user->version < $this->version) {
        $this->release(2);                  // replica is behind: retry shortly
        return;
    }

    SearchIndex::put($user);
}

Use an integer version or a high-precision timestamp. A timestamp rounded to the second breaks when two updates land in the same second.

Key Takeaways

  • Replication lag is harmless when a person reads the stale value and refreshes. It's harmful when a job reads it and writes it somewhere else.

  • after_commit fixes the race to the commit, not the race to the replica.

  • The bug hides in dev (no replica) and peaks under production write load.

  • Any "go look" message races the data it points to.

  • Fixes: send the data, carry the commit position, read the primary for read-then-write jobs, or verify the version before completing.

FAQ

What is replication lag in a database?

Replication lag is the delay between a change committing on the primary database and that change being applied on a read replica. During that window, reads from the replica return older data.

Why does my background job read stale data after I save a record?

Because the job runs quickly after the commit, but reads from a read replica that hasn't applied the change yet. The job is correct about when to run; it's reading from a copy that's behind.

Does dispatching jobs after commit prevent stale reads?

No. Dispatching after commit ensures the data exists on the primary. If the worker reads from a replica, it can still see the pre-change row until replication catches up.

What is read-your-writes consistency?

It's a guarantee that after you write something, your own subsequent reads will see that write. Replica setups break it by default unless reads are pinned to the primary or checked against a commit position.

How do I make a worker read from the primary in Laravel?

Use Model::onWriteConnection() for the query, for example User::onWriteConnection()->find($id). Laravel's sticky option only applies within the same request cycle, so it doesn't cover queued jobs.

How can I check whether a Postgres replica has caught up?

Record pg_current_wal_lsn() on the primary after the write, then compare it with pg_last_wal_replay_lsn() on the replica. If the replica's position is lower, it hasn't applied your change yet.

The bottom line

Replication lag is the price of replicas, and usually a cheap one. It becomes expensive the moment your code treats a replica's answer as the truth and writes it somewhere that never refreshes.

Replicas are for reading. Truth is for writing. Don't let a replica decide what gets written.

Adam Jaber is a software engineer who writes Simply Explained: complex topics, made simple, no jargon, no hype.

The Systems Behind the Software

Part 1 of 9

Plain-English deep dives into the backend and system-design ideas that quietly run everything — caching, idempotency, databases, distributed systems, and the subtle bugs the obvious approach never prevents. No jargon, no hype. Each piece starts with a real production problem, builds the intuition, then shows how to actually get it right.

Up next

There's No Safe Order for Two Writes

Why saving data and then calling an email, payment, or storage API breaks in production, and the one-write fix that works even in a monolith.

More from this blog

S

Simply Explained

65 posts

Complex topics in AI and software, made simple — no jargon, no hype.

I'm Adam Jaber, a software engineer writing the clear explanations I wish existed. Powerful technology becomes useful the moment you actually understand it.

Three tracks: AI for Humans (how AI works and affects your life), Building with AI (prompting, agents, RAG, for developers), and The Systems Behind the Software (backend design, explained through everyday bugs).

New here? Start with the pinned "Start Here" guide.