Probal Bose

Engineering

Keeping critical systems fast and dependable

I've spent 17 years in infrastructure and site reliability engineering. Today I'm a principal site reliability engineer, leading platform and infrastructure work for machine learning systems that make decisions inside payment networks, where a few milliseconds matter.

Career

From radio masts to real time machine learning

  1. ν

    Wireless networks

    I started out building infrastructure for 3G and 4G wireless networks, including a data centre built from scratch.

  2. Σ

    Systems and DevOps

    Systems administration and DevOps for large enterprises, automating how software is built, deployed and run.

  3. μ

    Machine learning platforms in payments

    About 8 years running real time machine learning inference at financial scale. I now lead platform and infrastructure engineering for a machine learning product in the critical path of payment authorisation.

What I focus on

Reliability as an engineering discipline

α

Availability

Service level objectives, availability modelling and designing systems that degrade gracefully instead of failing all at once.

τ

Latency

Keeping machine learning inference inside tight time budgets, measured by tail percentiles, not averages.

θ

Observability

Metrics, traces and logs with OpenTelemetry, eBPF and Grafana, so problems show up before customers notice them.

κ

Compliance grade infrastructure

Platforms built to meet the security and audit standards of the payments industry.

Where it meets research

The same skills, pointed at markets

My research runs on infrastructure I built and operate myself, from streaming ingestion and time series storage to monitoring down to disk write latency. The habits of production engineering, such as validating data, failing safely and measuring everything, turn out to matter just as much in research. See Research and Projects.