2 min lesson
Anchor it with real numbers
Rebuild the parts of "Anchor it with real numbers", then say why each one matters.
Step 1 of 3
Anchor it with real numbersQuantify outcomes
- Latency
- p50 and p99 before vs after, with the budget you were holding to.
- Cost
- Monthly spend or unit cost moved and the lever - Spot, right-sizing, killed waste.
- Availability
- Nines before vs after or failover time cut from minutes to seconds.
- Scale held
- Peak RPS, DAU, region or cluster count the design carried.
- Your decision
- The one architectural call that was yours and the alternative you rejected.
If a row is blank, that's a gap to fill before the loop, not a number to invent.
Then have a peer attack it. Every major choice gets a “why not X?” and you answer without flinching.
Learn more
Advanced table
Strong answers cite the constraint and the rejected option
- They ask
- Why EKS not self-managed k8s?
- Weak answer
- “It's the standard / it's what we used.”
- Strong answer
- Control-plane ops cost vs team size, IRSA for IAM and the upgrade burden we didn't want to own.
- They ask
- Why active-passive not active-active?
- Weak answer
- “Active-active is too complex.”
- Strong answer
- The cost of double capacity vs our actual failover SLO; minutes of RTO were acceptable for that tier.
- They ask
- Why Terraform not Pulumi?
- Weak answer
- “More people know Terraform.”
- Strong answer
- State and module maturity, drift detection in our pipeline and the existing module library we'd reuse.
| They ask | Weak answer | Strong answer |
|---|---|---|
| Why EKS not self-managed k8s? | “It's the standard / it's what we used.” | Control-plane ops cost vs team size, IRSA for IAM and the upgrade burden we didn't want to own. |
| Why active-passive not active-active? | “Active-active is too complex.” | The cost of double capacity vs our actual failover SLO; minutes of RTO were acceptable for that tier. |
| Why Terraform not Pulumi? | “More people know Terraform.” | State and module maturity, drift detection in our pipeline and the existing module library we'd reuse. |
Strong answers cite the constraint and the rejected option; weak answers cite the convention.
Practice going three “why” levels deep on your single biggest decision. Why service mesh? mTLS and traffic splitting. Why Istio over Linkerd? The sidecar overhead vs feature need. Why eat the sidecar cost at all instead of sidecarless? The third level is where memorized answers run out and real ownership shows. Rehearse until you don't hand-wave at level three.
“I moved us from active-passive to active-active across two regions because our failover RTO of four minutes was burning the SLO during AZ events. It roughly doubled steady-state cost, which I justified against the revenue at risk per minute of downtime. If I redid it, I'd have shipped global load balancing first and split traffic gradually instead of cutting over.”