Site Reliability Engineer · Senior / Staff

Systems I run grow 1,000x without a rebuild or a hero.

Scale stays survivable when it stays boring. I build capacity ahead of growth, observability everyone reads, and runbooks anyone on the rotation can follow.

Reliability with receipts

Scale, visibility, and what happens when the pager goes off.

Petabytes
at Protectwise, acquired by Verizon

Thousands of Cassandra nodes, petabytes in S3, and a $10M+/yr AWS environment run entirely as code, holding steady under a product that shipped dozens of times a day and could not miss data.

1,000x
growth on one platform at Cloaked

The EKS platform I built carried the product from hundreds of users to hundreds of thousands without a rebuild. I was the one-person ops team behind it, handling capacity and cost under SOC 2, ISO 27001, and ISO 27701.

PCI
incident response the auditors signed off on

At Cardfree I ran a large-scale PCI-compliant C# installation in AWS and replaced ad-hoc Slack firefighting with a structured incident-response program. It satisfied the auditors without burying the engineers in paperwork.

My phone
is the pager for my own house

My homelab serves production traffic from a k3s cluster. Certs renew themselves, backups run on schedule, every change ships from git, and Prometheus pages me only when something needs a human.

For the keyword scanners: Prometheus, Grafana, Loki, PagerDuty, CloudWatch, Cassandra, Kafka, S3, EKS, k3s, Flux CD, Terraform, Chef, Packer. Every one of them attached to a system above or on the resume.

The day-2 receipts

Quarterly to on demand

Took CyberGRX from quarterly deploy events to shipping whenever anyone wanted, with a bad deploy undone in under a minute, by writing a Go operator that ran blue/green rollouts as a custom resource and kept the old version warm behind the new one.

On call without the author

Made the rotation runnable by whoever joined it last, so pages close without a call to me, by putting Prometheus, Grafana, and Loki on everything I run and linking every alert to the graph that fired it and the runbook that closes it.

Zero hand-built hosts

Took hand-built servers out of Cardfree’s PCI environment, every Windows and Linux host behind the C# and Ruby services coming from one Packer build, by building a hybrid AMI platform where a drifted host got replaced from the image rather than repaired.

“Jake and I ran the AWS behind a petabyte-scale security platform. Thousands of Cassandra nodes, Kafka pipelines, a $10M-a-year bill, all of it managed with Chef and deliberately boring tech. I’d sign up to do it with him again tomorrow.”

Alex Artigues, DevOps Engineer, Protectwise

Sound like the role you’re hiring for?

Email me and let’s find out. My resume digs into the details if you need more convincing.

Or browse my projects & ventures. Screening with an AI assistant? Point it at ai.jakegaylor.com.