Building a Rules Engine You Can Actually Maintain

The worst rules engine I ever inherited had 1,847 rules in a single database table, evaluated top to bottom, and the order they ran in was determined by the...

Originally published onanselmfowel.com

The worst rules engine I ever inherited had 1,847 rules in a single database table, evaluated top to bottom, and the order they ran in was determined by the auto-increment primary key. Someone had discovered years earlier that if you needed a rule to win a conflict, you deleted it and re-inserted it so it got a fresh, higher ID. That was the documented procedure. It was written down in a Confluence page titled "How to make a rule stick."

I have spent a good chunk of my career either building rules engines or apologizing for them. Fraud scoring, credit decisioning, fee calculation, KYC routing, dunning logic. They all start the same way, as a reasonable idea, and a surprising number of them end the same way too, as a haunted forest nobody wants to walk into after dark. So this is what I have learned about keeping one alive past its second birthday.

Why you reach for one in the first place

The honest reason teams build a rules engine is that the business wants to change behavior without waiting for a deploy. A risk analyst notices that a particular merchant category is spiking with chargebacks on Friday nights, and they want to tighten the limit today, not next sprint. That is a legitimate need. Compliance requirements shift, promotional logic changes weekly, and pretending all of this belongs in a quarterly release train is how you end up with shadow spreadsheets driving production decisions.

But be clear-eyed about the trade you are making. You are pulling logic out of your codebase, where it has tests, version control, and a review process, and putting it somewhere softer. Every rules engine is a small bet that the flexibility is worth the loss of those guardrails. My rule of thumb is simple: if the logic changes less than once a quarter, it does not belong in a rules engine. It belongs in code, with a pull request and a reviewer who has to think about it.

The two jobs a rules engine actually has

People talk about rules engines as if the hard part is evaluating rules. It is not. Evaluating a boolean expression against a fact object is a solved problem and always has been. The hard parts are the two jobs nobody puts in the design doc: making a decision explainable after the fact, and letting a human change a rule safely.

In regulated fintech, the first one is not optional. When a customer disputes a declined application, or a regulator asks why a transaction was flagged, "the engine said so" is not an answer that keeps you out of trouble. You need to reconstruct, months later, exactly which rules fired, what data they saw, and what version of the ruleset was live at 14:32 on a Tuesday in March. If your engine cannot do that, it is not a rules engine. It is a random number generator with good PR.

Write down every decision, not every rule

The single highest-leverage thing I have ever added to a rules engine is a decision log. Not a log of rule changes, though you need that too. A log of decisions: one row per evaluation, capturing the input facts, the ruleset version, the rules that matched, and the final outcome. Immutable, append-only, retained for as long as your regulator says plus a comfortable margin.

This sounds obvious and almost nobody does it properly on the first try. They log the outcome but not the inputs, so you can see that an application was declined but not what the engine believed about the applicant. Or they log the inputs by reference to mutable records, so when you go back to investigate, the customer's address has since changed and the reconstruction is a lie. Snapshot the facts as they were at evaluation time.

public sealed record DecisionRecord(
    Guid DecisionId,
    string RulesetVersion,
    DateTimeOffset EvaluatedAt,
    string FactsSnapshot,        // serialized copy of inputs, not a foreign key
    IReadOnlyList<string> MatchedRuleIds,
    Outcome Outcome,
    string Explanation);         // human-readable, generated at decision time

// The explanation is built while the context is hot.
// Do not try to reconstruct "why" six months later from raw facts.
var explanation = matched.Count == 0
    ? "No rules matched; default outcome applied."
    : string.Join("; ", matched.Select(r => $"{r.Id}: {r.Because}"));

Generating the human-readable explanation at decision time, while all the context is in memory, is the trick most teams miss. Six months later, reconstructing "why" from raw facts is archaeology. Do it while the pot is still warm and you will thank yourself during your first audit.

Decide how conflicts resolve before you have any

Two rules will eventually disagree. One says approve, one says decline. The question of who wins is the single most important design decision in the whole system, and it is the one most teams answer by accident, the way that haunted table answered it with primary keys.

Pick an explicit strategy and make it boring. The options are not exotic:

  • First match wins with an explicit, human-visible ordering. Simple, predictable, and my default for decisioning flows.
  • Most specific match wins, where specificity is a number you can actually compute and show, not a vibe.
  • Salience or priority, an integer on each rule. Honest, but priority inflation is real; give yourself gaps like BASIC line numbers so you can insert between them.
  • Accumulate then decide, where rules contribute to a score and a final threshold makes the call. This is what most fraud scoring actually wants.

What matters is not which one you pick. What matters is that a person looking at two rules can tell you, without running the engine, which one will win and why. If they cannot, you have built a system that surprises its own operators, and surprise is the enemy of a maintainable rules engine.

The DSL temptation

At some point somebody will suggest building a domain-specific language so business users can write their own rules. I have done this. I have watched other people do this. It is one of those ideas that is correct in a slide deck and expensive in reality.

Every DSL grows until it is a bad, undocumented programming language, and then you are maintaining a compiler you never meant to write.

The version that actually works is narrow on purpose. Give users a constrained set of conditions over a well-defined fact schema, expose it through a UI with dropdowns and typed fields rather than a free-text box, and validate hard at save time. The moment someone asks for loops, variables, or the ability to call out to another rule, push back and ask what real problem they are solving. Usually it is one specific case that deserves code, not a general escape hatch that turns your engine into a Turing tarpit maintained by people who do not know they are programming.

Test the ruleset like it is code, because it is

Here is a hill I will die on: a ruleset is production logic and it deserves the same testing discipline as the application around it. The fact that a risk analyst can edit it in a web form does not make it configuration. It makes it code with a friendlier editor.

The practice that saved us the most grief was a golden-file regression suite. We kept a corpus of a few thousand real, anonymized past decisions, and any proposed rule change was replayed against all of them before it could go live. The diff was the review artifact. "This change flips 340 previously-approved applications to declined" is a sentence that stops a bad edit cold, and it is far more useful than any amount of staring at the rule text. When we introduced that gate, our rate of "emergency rule rollback" dropped from roughly one a fortnight to maybe one a quarter.

Rulesets deploy; they do not just save

A rule change is a production change. Treat the act of publishing a ruleset with the same seriousness you treat a deploy, because that is exactly what it is. That means versioning, staged rollout, and a rollback that takes seconds and not a frantic re-insert.

Version the whole ruleset as an immutable unit, not each rule independently. When you promote version 47 to production, the decision log records "47" and you can always ask what version 47 contained. Never mutate a published version in place; the day you let someone hotfix a live ruleset without minting a new version is the day your decision log starts lying, and a decision log that lies is worse than none at all because you will trust it. Keep the last known-good version one click away and make sure the on-call engineer has actually practiced the rollback.

Watch what the engine does, not just whether it runs

Standard uptime monitoring will tell you the service is up. It will not tell you that a well-meaning analyst just published a rule that quietly declines every applicant from a particular postcode. For that you need to watch outcomes, not health checks.

The metrics that earn their keep are outcome distributions over time and per-rule fire rates. If your approval rate drops eight points in an hour, alarms should go off regardless of whether every rule technically evaluated without error. Track how often each rule actually matches, too. In that inherited system with 1,847 rules, we instrumented fire rates and found that 61 percent of the rules had not matched a single transaction in over a year. Dead rules are not harmless. They are landmines that make the whole thing harder to reason about, and they are cover for the handful of rules that are quietly doing something wrong.

Anselm Fowel, CTO and fintech architect
Anselm Fowel — CTO & fintech architect

Conclusion

If I could tattoo one idea onto every team that builds one of these, it would be this: a rules engine does not let you escape engineering discipline, it only moves where you have to spend it. You hand the trigger to someone who is not an engineer, so the scaffolding around them has to be stronger, not weaker. People get that backwards constantly. They see a web form and assume it buys them less rigor, when the honest version demands more.

So here is the one thing I would actually do differently if I were starting the next one tomorrow: build the audit trail before you write the second rule. Not the tenth, not the hundredth, the second. The decision log, the immutable versioning, the replay harness, all of it feels absurd when you have a table with four rows in it. It stops feeling absurd the first time a regulator emails you a transaction ID from nine months ago and expects a straight answer. You will never regret being able to explain what the engine did. You will regret every single day you cannot.

Chat with us