Engineering
Keeping critical systems fast and dependable
I've spent 17 years in infrastructure and site reliability engineering. Today I'm a principal site reliability engineer, leading platform and infrastructure work for machine learning systems that make decisions inside payment networks, where a few milliseconds matter.
Career
From radio masts to real time machine learning
-
ν
Wireless networks
I started out building infrastructure for 3G and 4G wireless networks, including a data centre built from scratch.
-
Σ
Systems and DevOps
Systems administration and DevOps for large enterprises, automating how software is built, deployed and run.
-
μ
Machine learning platforms in payments
About 8 years running real time machine learning inference at financial scale. I now lead platform and infrastructure engineering for a machine learning product in the critical path of payment authorisation.
What I focus on
Reliability as an engineering discipline
Availability
Service level objectives, availability modelling and designing systems that degrade gracefully instead of failing all at once.
Latency
Keeping machine learning inference inside tight time budgets, measured by tail percentiles, not averages.
Observability
Metrics, traces and logs with OpenTelemetry, eBPF and Grafana, so problems show up before customers notice them.
Compliance grade infrastructure
Platforms built to meet the security and audit standards of the payments industry.
Where it meets research
The same skills, pointed at markets
My research runs on infrastructure I built and operate myself, from streaming ingestion and time series storage to monitoring down to disk write latency. The habits of production engineering, such as validating data, failing safely and measuring everything, turn out to matter just as much in research. See Research and Projects.