Amazon runs a thing called the Weekly Business Review, a standing meeting where a fixed deck of operational metrics gets walked, line by line, every week. I first met it secondhand, from an ex-Amazon PM who joined a payments team I ran, and my honest first reaction was that it sounded like a bureaucratic tax. Sixty slides of graphs. Who has time.
Six months later I had built one for my own engineering org and I would not give it up. Not because it made us look good in front of the board, but because it was the first time the whole team was arguing about the same numbers on the same cadence, instead of each of us carrying a private, slightly-wrong mental model of how the systems were actually behaving.
What a WBR actually is, and is not
A Weekly Business Review is a recurring meeting built around a fixed set of metrics that you look at every single week, in the same order, with the same definitions. The deck barely changes. That sameness is the whole point. When the shape of the report is constant, your eye is drawn to the one line that moved, and the conversation goes straight to the anomaly instead of relitigating what to measure.
It is not a status meeting. It is not people reading out what they did last week. It is not a demo, a sprint review, or a place to plan the next quarter. Those are all fine meetings; they are just different meetings. The WBR looks at outputs and outcomes of the running business, not at individual effort. If someone shows up wanting to tell you how hard they worked, the metric either moved or it didn't, and the graph already told you.
The distinction matters because engineering leaders love to smuggle a roadmap discussion into any meeting with a projector. Resist it. The moment your WBR becomes a planning session, it stops being a mirror and becomes a sales pitch, and you lose the one forum where the numbers get to speak without spin.
Why engineering specifically needs one
Sales has always had a weekly number. Finance closes the month. Ops lives on dashboards. Engineering, meanwhile, tends to run on vibes and Jira burndown charts that nobody trusts. We are strangely allergic to looking at our own operational data on a schedule, even though we generate more telemetry than any other function in the company.
The teams I have seen struggle most were not the ones with bad engineers. They were the ones where nobody could tell you, without opening three tools and guessing, what last week's p99 latency was on the checkout path, or how many deploys got rolled back, or whether the on-call queue was getting worse. The data existed. It was just never assembled into one place at one time so a human could reason about the trend.
A dashboard nobody is required to read is a graveyard for good instrumentation. The WBR is the standing appointment that forces the reading.
The metrics I actually track
Every org's deck is different, but the spine is remarkably stable. I want a mix of reliability, delivery, and cost, because those three pull against each other and the tension is where the interesting conversations live. Here is roughly what sits in mine for a payments platform:
- Availability and error rate per critical path, with the payment authorization flow always first because that is the one that costs us money and reputation by the minute.
- Latency at p50, p95, and p99 for those same paths. p50 is for morale; p99 is where the truth lives.
- Deployment frequency and change-failure rate, the two DORA metrics I trust most, tracked as a four-week rolling average so a single bad week does not send everyone into a panic.
- On-call load: number of pages, how many were actionable, and how many fired between midnight and 6am. That last one is a leading indicator of attrition and I take it personally.
- Open incident count by severity, and mean time to recovery, which I care about far more than mean time between failures.
- Cloud spend against forecast, broken out by the two or three services that always misbehave.
Notice what is missing. No story points. No velocity. No lines of code. Those measure activity, and activity is not the business. The WBR is about whether the machine is running well and getting better, not whether the people are busy. They are always busy.
Show the trend, not the snapshot
The single most valuable thing the Amazon format taught me was to plot the trailing six weeks and the same week last year, side by side, on every graph. A single number is nearly useless. 99.9% availability sounds fine until you notice it was 99.99% the four weeks before. That missing nine is roughly 40 extra minutes of downtime a month, and something quietly broke to spend it.
Year-over-year is the one that catches slow rot. Latency creeps up two percent a month and no weekly snapshot will ever alarm you, but the line against last July makes it undeniable. I once had a service whose cold-start time had roughly doubled over eleven months, and not one person had noticed because every week looked like the week before. The YoY panel in the first WBR we ran surfaced it in about four seconds of silence.
The mechanics that make it survive
The format is boring on purpose and the discipline is everything. We meet Monday at 9am for exactly one hour. The deck is auto-generated by a job that runs Sunday night, so nobody spends Friday afternoon massaging numbers, and nobody can quietly reframe a bad week. The metric owner narrates their section in two minutes, then we go to questions.
The rule I enforce hardest: no surprises in the room from the person presenting, but plenty of surprises welcome from the data. If a metric moved, the owner should already know why or already be finding out. What we do not do is solve the problem live. When something is clearly broken we name it, assign one person, and give it a follow-up slot. Trying to debug in a room of twelve people is how a one-hour meeting becomes a two-hour one, and the two-hour version gets cancelled within a month.
I also killed the attendance list twice before it stuck. The first version had thirty people and became theater. Now it is the engineering leads, the on-call lead for the week, and whichever product counterpart owns the paths we are reviewing. Small enough that silence is uncomfortable and everyone has to actually look.
What it does to the culture
The interesting effects were the ones I did not plan for. Once people knew the change-failure rate would be on a screen every Monday in front of their peers, the quality of pre-deploy checks improved without me sending a single memo. Not because anyone was shamed, but because the number became real and shared. Peer visibility does quiet work that top-down mandates never manage.
It also gave junior engineers a map of the business they otherwise would not get. A second-year engineer sitting in that room learns which paths actually matter, what good looks like, and how the parts connect. I have watched people go from writing code against a ticket to reasoning about the platform, and a lot of that happened in the WBR by osmosis. That is worth the hour on its own.
The ways it goes wrong
I have broken this meeting in most of the available ways, so let me save you some time. The first failure is metric sprawl. You add a graph every time something goes wrong until the deck is ninety pages and nobody can hold it in their head. Prune ruthlessly. If a metric has not prompted an action or a question in three months, cut it.
The second is gaming. The moment a metric becomes a target that affects how people are perceived, they optimize the metric instead of the outcome. Deploy frequency is a classic trap; reward it naively and you get a flood of trivial one-line deploys that inflate the number and prove nothing. Watch pairs of metrics that constrain each other, and treat any single line moving in isolation with suspicion.
The third, and the quietest killer, is letting it drift into a blame session. The first time someone gets grilled for a red number, everyone else learns to hide theirs. I would rather have an honest ugly graph than a pretty dishonest one, and I say that out loud, often, because people need to hear it more than once to believe it.
How to start without a big program
Do not build the ninety-slide version first. That is the mistake that guarantees the meeting dies before it earns trust. Pick five metrics you already have data for, put them on five graphs with a six-week trailing line, and meet for thirty minutes next Monday. That is the entire minimum viable version, and it is enough to start the flywheel.
The deck will grow itself. Something will break, you will wish you had been watching a number, and you will add it. Something will turn out to be noise and you will drop it. After about eight weeks you will have a deck that fits your org rather than one copied from a blog post, mine included. The cadence is the asset; the specific graphs are just this quarter's version of it.

Conclusion
The thing I underestimated most was how much a WBR changes what leadership even means day to day. When the operational truth is on a screen every Monday, my job stops being the person who collects status and becomes the person who asks better questions about a shared picture. That is a much better job. Build the boring meeting, protect it from becoming interesting in the wrong ways, and let the graphs argue on your behalf.