Skip to main content

Command Palette

Search for a command to run...

The Sum Is the Bug: Why Accurate Software Estimates Make Late Projects

Every task passes its estimate check. The project fails anyway. Here's the reproduction, the root cause, and the fix.

Updated
•9 min read•View as Markdown
The Sum Is the Bug: Why Accurate Software Estimates Make Late Projects
J
I'm a software engineer who spends most days building systems that solve real problems. When I'm not shipping code, I'm either untangling a tricky problem or writing about what I learned doing it. Currently exploring AI on the side.

Here's a bug report every engineer has filed in their head at least once:

Expected: project done in 20 days (10 tasks × 2 days). Actual: 26 days. Notes: 8 of 10 tasks landed within hours of their estimate. No one slacked. No estimate was padded or lowballed.

Every unit passes. The integration fails. That's not a people problem. It's a math problem, and like most math problems in software, you can reproduce it in a few lines of code.

Key Takeaways

  • Your gut estimates the typical case (the median) of a task, and it's usually good at it.

  • Task durations are right-skewed: a task can finish a little early but can run very late. So the average cost of a task is higher than its typical cost.

  • Adding typical-case estimates undershoots the total. In a simple model where every estimate is exactly right, a 20-day plan holds only about 1 time in 4.

  • The overrun concentrates in the few tasks you understand least. In the model, 2 unfamiliar tasks out of 10 carry roughly three-quarters of the expected delay.

  • Padding every task fails because early finishes rarely get passed on. Pool one visible buffer instead.

  • The fix: run a Tail Audit, give two numbers per task, and touch the riskiest tasks first.

Reproduce it

Ten tasks. Each is estimated at exactly its true median of 2 days, so every estimate is, by construction, correct. Eight tasks are familiar work with a narrow spread. Two are unfamiliar (a new library, an external API) with a wide spread.

import numpy as np

rng = np.random.default_rng(42)
RUNS, GUESS = 100_000, 2.0           # every task estimated at its true median
spreads = [0.4] * 8 + [1.2] * 2      # 8 familiar tasks, 2 unfamiliar

tasks = np.array([GUESS * np.exp(rng.normal(0, s, RUNS)) for s in spreads])
total = tasks.sum(axis=0)
plan = GUESS * len(spreads)          # 20 days

print("on time:", (total <= plan).mean())
print("median:", np.median(total), "mean:", total.mean(),
      "85th pct:", np.percentile(total, 85))

Output, rounded:

on time:  ~0.26     (about 1 project in 4)
median:   ~23 days
mean:     ~26 days
85th pct: ~32 days

A plan built from ten correct estimates ships on time about a quarter of the time. Its typical finish is about 23 days, its average about 26, and an 85%-confident date would be around 32.

Model caveat: this is a toy, not field data. The log-normal shape and the spread values (0.4 and 1.2) are assumptions; swap in your own. The exact numbers will move. The direction holds whenever tasks can overrun by much more than they can underrun.

Root cause: task durations are lopsided

A 2-day task has a floor. On a perfect day you finish in one. You can't finish in minus three.

It has no ceiling. The library can't do the one thing you need. The migration locks a table that's busier than anyone thought. A test fails in CI one run in five. The vendor API documents behavior it doesn't have. Any of these turns 2 days into 8.

When outcomes are lopsided, the median (what your gut estimates) and the mean (what the plan actually pays on average) come apart. Engineer Erik Bernhardsson made this argument in 2019 with a statistical model of estimate-versus-actual data: people seem to estimate the median well, and the mean is what hurts.

Plans add estimates. But medians don't add up to the median of the total, and they fall even further below the mean of the total. Lucky tasks can't pay back unlucky ones, because the luck is capped on one side and unbounded on the other.

Where the overrun actually comes from

Break the expected overrun down per task:

Task type Count Expected overrun each Share of total overrun
Familiar 8 ~0.2 days ~24%
Unfamiliar 2 ~2.1 days ~76%

Two tasks out of ten decide the date. That reframes the whole skill. Sharper estimates on familiar work barely move the total. The schedule is set by the tasks you understand least, which are the ones people are most tempted to hand-wave.

This is also why "just break it into smaller tasks" only half works. Decomposition shrinks the spread of known work. Splitting an unfamiliar task doesn't remove the unknown; it hides it inside one of the pieces, usually the one called "integrate with X."

Why per-task padding is a bad patch

Padding every estimate looks like a fix. Eliyahu Goldratt's 1997 business novel Critical Chain explains why it isn't:

  • Parkinson's law: work expands to fill the time available.

  • Student syndrome: with slack available, people start later, so the margin is spent before trouble arrives.

  • Early finishes don't get passed on: finish a 3-day task in 2 and the spare day turns into polish or other work. It never reaches the next task.

Model that last effect directly: treat any task that finishes early as finishing exactly on estimate. Now the project only hits its date if every task lands at or under estimate. With 10 independent tasks at even odds each, that's 0.5¹⁰, or 1 in 1,024. The simulated average finish moves from about 26 days to nearly 29.

Goldratt's alternative still holds: strip hidden padding from individual tasks and pool it into one visible project buffer. Shared slack only gets spent where trouble actually happens, and you can watch it drain.

The fix: find the tail, then plan around it

The Tail Audit (3 questions per task)

  1. Has someone on the team done this exact thing, in this codebase, recently? Not "something like it."

  2. Can it finish without waiting on anyone or anything you don't control? Another team, a vendor API, an approval, a maintenance window.

  3. Will you recognize "done" when you see it? Vague acceptance criteria let tasks grow indefinitely.

Any "no" marks a tail task. Two or more, and the estimate is a guess wearing a number. Say so.

Then the five moves

  1. Run the Tail Audit before you commit to anything.

  2. Give two numbers: "likely 3 days, pretty sure within 6." The gap is the signal. Narrow means known work; wide means risk.

  3. Touch the tail first. Order work by uncertainty, not ease. A throwaway call against the real API or a migration dry run on a copy of production data on day one turns a late surprise into an early decision.

  4. Pool the buffer, out loud. One visible buffer, sized by how many tail tasks you found. Track buffer used against work done.

  5. Keep a ratio log: task, estimate, actual, kind of work. You're not trying to guess closer; you're learning how wide your tail is per kind of task.

Why it matters past the sprint

People don't remember that eight of your ten estimates were close. They remember the slipped date, and whether they heard about it early. Reporting the tail ("done by the 14th if the payments integration behaves; I'll know by Wednesday") turns a broken promise into an updated forecast.

The engineers trusted with deadlines are rarely the best guessers. They name the tail early and go poke at it first, which is exactly the kind of judgment the job is shifting toward. One catch: when early de-risking works, it can look like you worried about nothing, because prevention erases its own evidence. Writing the two numbers down before you start is the receipt.

FAQ

Why are software estimates always wrong?

Individual estimates often aren't wrong. Engineers tend to estimate the typical case well. Projects run late because task durations are right-skewed (they can run far longer than they can run short), so adding typical-case estimates produces a total below what the project will usually take.

What is the difference between median and mean in software estimation?

The median is the typical duration: half the time you finish faster, half slower. The mean is the average including rare disasters. For skewed task durations, the mean is higher than the median, and project plans that add medians undershoot the average total.

Does breaking work into smaller tasks improve estimates?

Partly. Smaller pieces of well-understood work have smaller spreads, so decomposition helps there. But splitting an unfamiliar task doesn't remove its unknowns; they usually end up concentrated in one of the pieces, such as the integration step.

Should I add padding to every task estimate?

Usually not. Per-task padding tends to be consumed by Parkinson's law and student syndrome, and early finishes rarely get passed on. A single visible project buffer, as proposed in Goldratt's Critical Chain method, protects the date more reliably.

How do I identify which tasks will blow up my estimate?

Ask three questions per task: has the team done this exact thing in this codebase recently, can it finish without waiting on anything outside your control, and is "done" clearly defined. Any "no" marks a high-risk tail task to schedule and investigate first.

How should I communicate an uncertain estimate to my manager?

Give two numbers (a likely date and a pretty-sure date), name the specific risk that separates them, and attach a checkpoint when you'll know more. That gives planners a forecast they can act on instead of a single number that silently hides the risk.

The bottom line

Late projects are rarely made of bad estimates. They're made of good typical-case estimates, added together, with the uncertainty stripped out at every step. Padding hides it. Smaller tasks move it around. Finding the two or three tasks that decide the date, and touching them first, is what changes it.

The date is decided by the tasks you understand least. Find them first.

Adam Jaber is a software engineer who writes Simply Explained: complex topics, made simple. No jargon, no hype.

More from this blog

S

Simply Explained

70 posts

Complex topics in AI and software, made simple — no jargon, no hype.

I'm Adam Jaber, a software engineer writing the clear explanations I wish existed. Powerful technology becomes useful the moment you actually understand it.

Three tracks: AI for Humans (how AI works and affects your life), Building with AI (prompting, agents, RAG, for developers), and The Systems Behind the Software (backend design, explained through everyday bugs).

New here? Start with the pinned "Start Here" guide.