Zero-trust on a slide is easy. Zero-trust on 12,000 production workloads — without an outage window — is engineering.
The Three Phases
1. Identity Before Encryption
Before turning on mTLS, every workload must have a verifiable identity. We use SPIFFE/SPIRE to issue short-lived SVIDs tied to Kubernetes service accounts and node attestation.
2. Permissive, Then Strict
Roll out mTLS in PERMISSIVE mode first. Watch the dashboards. Find the legacy clients still speaking plaintext. Then — and only then — flip to STRICT.
3. Policy-as-Code From Day Zero
Every authorization rule lives in Git as an OPA or Kyverno policy. Reviewed, versioned, rolled back like any other code.
Operational Realities
- Sidecar overhead: budget 50–80 m of CPU per pod
- Certificate rotation: every 60 minutes, automated
- Egress: explicit allowlists or you've replaced one perimeter with another
Result
99.99% of internal traffic is mTLS-encrypted with workload identity. Lateral movement, even after a container compromise, is a non-event.