SRE & Platform Engineering Leader with 13+ years of experience. I lead teams that keep mission-critical infrastructure running for 7M+ users, and I write code in Rust, Go, and Python to automate the toil away.
I didn't begin my career writing software. I began by solving customer problems — moved into technical support, then infrastructure operations, incident management, and platform engineering.
That journey gave me something many engineers never experience: a deep understanding of how systems actually fail in production.
Every system I build is designed from that perspective — to reduce operational complexity, automate decision-making, and make distributed systems easier to understand and operate at scale.
Production-grade tools and platforms solving real SRE challenges — from distributed log aggregation to governed AIOps incident triage.
A lightweight, fault-tolerant, production-grade distributed log aggregation and real-time streaming system. Built to understand the mechanics behind ELK/Kafka at a code level.
Cloud-native, Rust-powered distributed workflow orchestrator with HA scheduling, DAG execution, transactional outbox pattern, and NATS JetStream. A modern alternative to legacy schedulers like AutoSys.
AI Infrastructure Intelligence Platform to observe, optimize, and automate AI workloads across GPUs, Kubernetes, and distributed training pipelines.
Client-facing POC combining AI-assisted P1/P2 analysis with L2/L3 approval workflows, SOP-driven runbooks, dry-run/execution separation, and full audit traceability for safe AIOps adoption.
AI-powered incident investigation tool that correlates distributed logs, metrics, and traces to automate Root Cause Analysis and accelerate MTTR during major incidents.
Production-quality RAG application using 100% open-source and locally-hosted AI models. Enables secure document Q&A with grounded citations — zero data leaves your infrastructure.
Generates realistic synthetic incident alert datasets to accelerate ML anomaly detection model training and eliminate manual data-gathering toil for data science teams.
Machine Learning POC for automated runbook execution with Human-in-the-Loop confidence thresholds. Enables safe, governed auto-remediation of known anomalies.
Custom dashboard monitoring Blue Prism production workflows, providing real-time visibility into digital worker health to prevent business SLA breaches.
13+ years of progressive responsibility — from solving customer problems face-to-face to leading reliability operations for millions of users.
Leading a 10-member SRE team supporting 7M+ end users across hybrid-cloud. Defined SLIs/SLOs, drove MTTR reduction by 20%, and pioneered governed AIOps adoption for mission-critical payment gateway APIs.
Enterprise IT operations and multi-cloud infrastructure monitoring. Developed Python automation for provisioning, log analysis, and incident response workflows.
Advanced enterprise hardware and software support. Automated diagnostic workflows and earned Customer Champion recognition from Director-level leadership.
L1/L2 IT support operations covering incident triage, escalation, and resolution. Progressively took on automation initiatives and earned the ENU Excellence Award for a Windows 10 migration project.
First point of contact for customer escalations. Built the foundation for customer-centric thinking and operational discipline that defines my SRE philosophy today.
I'm interested in conversations about SRE leadership, distributed systems, AIOps, platform engineering, and ambitious infrastructure challenges.