Dogfood: making our systems more resilient

Move infra before the public site depends on a single box. Dual masters, proxies, and expendable nodes beat heroics.

Dogfooding: run the architecture you recommend. Early infra is cheap and coupled; dependencies accumulate; the public site is often the last thing migrated because everything else hangs off it.

What broke

A shared web host going unresponsive was predictable and mostly outside our control — but coupling made the move hard. After cutting over, traffic metrics rose sharply: downtime and glitches had been a silent tax.

Target shape

  • Dual MySQL masters across separate facilities (MMM-style tooling was a poor fit for that topology — design the failover yourself)
  • Proxy and light frontends (HAProxy, lighttpd era) in front of expendable app nodes
  • Mix of rented VMs and owned machines; mail and edge services split early

Failures should not become major incidents. Resilience is incremental: split services, rehearse failover, measure before and after. Keeping mishaps secret helps neither your team nor anyone learning from public postmortems.