Dogfood: making our systems more resilient
Move infra before the public site depends on a single box. Dual masters, proxies, and expendable nodes beat heroics.
Dogfooding: run the architecture you recommend. Early infra is cheap and coupled; dependencies accumulate; the public site is often the last thing migrated because everything else hangs off it.
What broke
A shared web host going unresponsive was predictable and mostly outside our control — but coupling made the move hard. After cutting over, traffic metrics rose sharply: downtime and glitches had been a silent tax.
Target shape
- Dual MySQL masters across separate facilities (MMM-style tooling was a poor fit for that topology — design the failover yourself)
- Proxy and light frontends (HAProxy, lighttpd era) in front of expendable app nodes
- Mix of rented VMs and owned machines; mail and edge services split early
Failures should not become major incidents. Resilience is incremental: split services, rehearse failover, measure before and after. Keeping mishaps secret helps neither your team nor anyone learning from public postmortems.