Proving re-implemented business logic with a parallel run

How to be sure a rule you rebuilt from a specification behaves like the one it replaces, when the specification is incomplete and the original is unreadable. Where to run the comparison, what to compare, how to read the differences, and what to do when the old system turns out to be wrong.

14 min read · Updated

01The problem tests cannot solve

Sooner or later a migration hands you a rule you cannot read. It lives in COBOL on a mainframe, or in a stored procedure nobody has opened since the person who wrote it left, or in a spreadsheet that finance treats as scripture. You are given a specification and asked to build against it.

The specification will be wrong. Not maliciously and often not obviously: it describes what the rule was intended to do, and what runs in production is what it actually does, after twenty years of amendments, exceptions and fixes that were never written down.

This is why a test suite is not enough. Tests written from the specification encode the specification. If the specification and the system disagree, your tests will agree with the specification and pass, and you will discover the difference when a customer's invoice is wrong. The gap you need to find is precisely the gap your tests cannot see, because both are derived from the same incomplete document.

The only source of truth about what a system does is the system. Not its documentation, not the people who maintain it, and not the specification written to replace it.

02The idea, in one line

Run both implementations against the same real inputs, compare the outputs, and treat every difference as a question rather than a failure.

That framing matters. A difference is not automatically a bug in the new code. It might be, but it might equally be a defect in the old system that has been quietly running for years, or an ambiguity the specification never resolved, or an artefact of the comparison itself. Sorting differences into those categories is most of the work, and it is why the process needs a person rather than a build pipeline.

The pattern is sometimes called shadowing or dark launching. GitHub's Scientist library popularised it for refactoring, and the ideas port cleanly to any language: run both, return the old result, record the comparison.

03Where to run it

Three options, in increasing order of confidence and cost.

Offline replay
Capture a window of production inputs, replay both implementations over them in a batch, and diff. Cheapest and safest, and where you should start. Its weakness is that captured inputs are historical: they cannot contain a case that has not happened yet.
Shadow in production
Both implementations run on live traffic. The old result is returned to the user; the new result is computed and discarded, with the comparison recorded. Highest confidence, because the inputs are real and current. Costs compute, and the shadow must be genuinely incapable of affecting the real path.
Shadow on a sample
Shadow, but on a percentage of traffic. The usual compromise when the calculation is expensive. Sample by input characteristics rather than at random, or you will get a great deal of evidence about the common path and none about the rare one.

Start with replay to clear the obvious differences cheaply, then shadow to find what replay could not. Going straight to shadow means triaging hundreds of known differences under production conditions.

04The rule that must not be broken

The shadow implementation must not be able to change anything. Not the database, not a queue, not a file, not an email. Every write path in the new code has to be inert while it is shadowing.

This sounds obvious and it is the most common way the technique causes an outage. The new implementation is a faithful re-implementation, which means it faithfully re-implements the part that posts a ledger entry, and now every transaction is posted twice.

  • Run the shadow against a read-only connection, so the database refuses a write rather than relying on the code not to attempt one
  • Stub or disable outbound side effects: mail, queues, webhooks, file writes
  • Wrap the shadow so an exception inside it can never propagate into the real request
  • Bound its runtime, so a slow shadow degrades the comparison rather than the user's response time

Enforce this at the boundary, not by review. A read-only connection cannot be forgotten in a hurry; a code comment saying do not write here can.

05What to compare, and how

Outputs are rarely a single number, and naive equality produces a difference set full of noise that hides the real findings. Normalise before comparing, and be explicit about what normalisation you are applying, because each one is a small decision about what you are prepared not to notice.

  • Money is the usual source of false differences: scale, rounding mode and the sign of zero all differ between platforms
  • Decimal and floating point will disagree eventually. If the legacy used fixed-point arithmetic, match it rather than converting
  • Collection ordering is meaningful surprisingly often. Sort before comparing, but check first that the order was not itself part of the answer
  • A tolerance band is sometimes justified and always dangerous. A one-cent tolerance hides a one-cent bug that appears on every line of a million-line file
Normalising before comparison
compare(expected, actual):
    # Decisions, each of which hides something. Make them deliberately.
    normalise money      -> fixed scale, half-even rounding, no -0
    normalise dates      -> single timezone, drop sub-second precision
    normalise collections-> sort by a stable key before comparing
    normalise absent     -> decide once whether null and empty are equal
    ignore               -> generated ids, timestamps, run metadata

    if normalised(expected) == normalised(actual): return SAME
    return difference(expected, actual, inputs)

06Reading the difference set

Every difference falls into one of five categories. Naming them is what turns an intimidating list into a work queue.

New implementation is wrong
The common case early on, and the easy one. Fix the code and re-run.
Comparison is wrong
Non-determinism, ordering, rounding or precision that normalisation should have handled. Fix the harness, not the code, and be suspicious: a comparison bug you fix by loosening the comparison is how real differences get hidden.
Specification is ambiguous
Both results are defensible readings of the document. This is not an engineering decision. Take it to whoever owns the rule and record the answer, because it will be asked again.
Legacy is wrong
The old system has been producing the wrong answer, possibly for years. See the next section. This is the finding that changes the shape of the project.
Legacy is deliberately odd
The old behaviour looks wrong and is correct, because of a regulation, a contract, or a decision made in 2009 that still binds. Preserve it, and write down why, or someone will remove it as a bug next year.

Track the category on every difference. The mix tells you where you are: mostly category one means you are early, mostly categories three and four means you are nearly finished and the remaining work is not yours.

07When the old system turns out to be wrong

This will happen, and it is the moment the project stops being technical.

Somewhere in a rule that has run for fifteen years there is a case handled incorrectly. Nobody noticed because the case is rare, or because the error is small, or because the downstream systems have been quietly compensating for it. Your new implementation, built from the specification, produces the correct answer, and it therefore disagrees.

You cannot simply ship the correct answer. Historical data was produced by the wrong one. Reports reconcile against it. Somebody may have been invoiced on it. Correcting the rule and correcting the history are two different projects, and only one of them is yours.

  • Take it to the business immediately, with the input, both outputs and an estimate of how often the case occurs
  • Let them decide whether the new implementation matches the old behaviour or corrects it. Both are legitimate; only they can choose
  • If they choose to correct it, that is a scoped piece of work with its own risk, not a line in your migration
  • If they choose to preserve the defect, implement it deliberately, comment it with the decision and the date, and add a test that fails if someone fixes it by accident

Framing matters here. You have not found a bug in their system; you have found a question only they can answer. Presented as the former, the conversation goes badly and slowly.

08Knowing when you have run enough

Volume is not coverage. Ten million invoices from a quiet month can exercise fewer branches than two hundred chosen carefully, and the cases that break a re-implementation are almost never the common ones.

Track coverage by input characteristic rather than by row count, and go looking for the shapes the ordinary traffic does not contain.

  • Period boundaries: month end, quarter end, year end, and the day either side of each
  • Leap years, and February in general
  • Negative amounts, zero amounts, refunds, cancellations and reversals
  • The oldest records in the system, which often predate a schema change or a rule change
  • Records touched by a migration or a manual correction, which may not satisfy the invariants the code assumes
  • The largest and smallest values present, where precision and overflow behaviour diverge
  • Anything the legacy code has an explicit branch for. Read it for the branches even if you cannot read it for the logic

That last point is the most useful. You may not be able to follow the legacy implementation's arithmetic, but you can usually see where it makes a decision, and each of those is a case your input set must contain.

09Exit criteria

Stop when the difference set is empty, or when every difference remaining in it is a category three, four or five that somebody outside the engineering team has signed off in writing.

Not before, and equally not indefinitely. A parallel run that has been green for a full business cycle is telling you something; one that has been running for six months because nobody wants to make the decision is telling you something else.

  • The run has covered at least one complete business cycle, which for anything financial means a month end
  • The difference set is empty, or fully triaged and signed off
  • The rare cases from the previous section have appeared in the input set, rather than merely being absent
  • The comparison harness itself has been reviewed, because a comparison that never fails may be broken rather than reassuring

10What actually goes wrong

The shadow wrote something
Discussed above, and worth repeating because it is the one that causes an incident rather than a delay.
Non-determinism swamped the signal
Timestamps, generated identifiers and unordered collections produce thousands of differences and the real ones are lost in them. Normalise aggressively before the first real run, or nobody will read the output twice.
Nobody owned the triage
The harness runs, the differences accumulate, and no one is responsible for categorising them. A parallel run with an unread difference set is theatre. Name the owner before turning it on.
The comparison was loosened to make it green
Under deadline, a tolerance widens and the run goes green. This is the failure mode that produces confidence without correctness, and it is unrecoverable in the sense that you will not know it happened.
It never ran on real traffic
Replay against a captured window is a good start and it cannot contain what has not happened. If the calculation matters, get it in front of live inputs before you cut over.

11What it costs, and why it is worth it

A parallel run adds a harness, a comparison, somewhere to store differences, and a person to read them. On a migration of any size that is days of work, not weeks, and most of it is reusable across every rule you have to move.

What it buys is the ability to say, with evidence, that the new implementation agrees with the old one across real inputs including the awkward ones. That sentence is the difference between a cutover the business approves and a cutover the business tolerates, and on a rule that touches money it is the only honest basis for going live.

Also in guides

This is the method we use, published in full.

If you would rather not run it yourself, the two-week assessment produces the sequence for your specific system, and the plan is yours whoever executes it.