Available for SRE Leadership Roles

I build resilient systems that scale beyond human operators.

SRE & Platform Engineering Leader with 13+ years of experience. I lead teams that keep mission-critical infrastructure running for 7M+ users, and I write code in Rust, Go, and Python to automate the toil away.

7M+
End Users Served
6,000+
Virtual Machines
95%
SLA Achievement
10
Engineers Led

I didn't begin my career writing software. I began by solving customer problems — moved into technical support, then infrastructure operations, incident management, and platform engineering.

That journey gave me something many engineers never experience: a deep understanding of how systems actually fail in production.

Every system I build is designed from that perspective — to reduce operational complexity, automate decision-making, and make distributed systems easier to understand and operate at scale.

Systems I've Built

Production-grade tools and platforms solving real SRE challenges — from distributed log aggregation to governed AIOps incident triage.

Distributed Systems

Barnacles

A lightweight, fault-tolerant, production-grade distributed log aggregation and real-time streaming system. Built to understand the mechanics behind ELK/Kafka at a code level.

Go Kafka Distributed Systems
Platform Engineering

FlowForge

Cloud-native, Rust-powered distributed workflow orchestrator with HA scheduling, DAG execution, transactional outbox pattern, and NATS JetStream. A modern alternative to legacy schedulers like AutoSys.

Rust NATS DAG
AI Infrastructure

Obsygnal-Pulse

AI Infrastructure Intelligence Platform to observe, optimize, and automate AI workloads across GPUs, Kubernetes, and distributed training pipelines.

Go Kubernetes GPU Observability
Governed AIOps

HITL Incident Triage

Client-facing POC combining AI-assisted P1/P2 analysis with L2/L3 approval workflows, SOP-driven runbooks, dry-run/execution separation, and full audit traceability for safe AIOps adoption.

Python LLMs HITL
Observability

AI SRE Copilot

AI-powered incident investigation tool that correlates distributed logs, metrics, and traces to automate Root Cause Analysis and accelerate MTTR during major incidents.

Python Streamlit RAG
Secure AI

OpenNoteBook

Production-quality RAG application using 100% open-source and locally-hosted AI models. Enables secure document Q&A with grounded citations — zero data leaves your infrastructure.

Python Local LLMs RAG
ML Infrastructure

Synthetic Data Generator

Generates realistic synthetic incident alert datasets to accelerate ML anomaly detection model training and eliminate manual data-gathering toil for data science teams.

Python Streamlit ML
Self-Healing

ML Self-Healing (HITL)

Machine Learning POC for automated runbook execution with Human-in-the-Loop confidence thresholds. Enables safe, governed auto-remediation of known anomalies.

Python Automation SRE
Business Reliability

RPA Observability Dashboard

Custom dashboard monitoring Blue Prism production workflows, providing real-time visibility into digital worker health to prevent business SLA breaches.

Python Streamlit RPA

From Customer Support to SRE Leadership

13+ years of progressive responsibility — from solving customer problems face-to-face to leading reliability operations for millions of users.

Apr 2023 – Present

Technical Lead — Resolve Tech Solutions

Leading a 10-member SRE team supporting 7M+ end users across hybrid-cloud. Defined SLIs/SLOs, drove MTTR reduction by 20%, and pioneered governed AIOps adoption for mission-critical payment gateway APIs.

Aug 2020 – Mar 2023

Senior Engineer — Wisemen Consulting

Enterprise IT operations and multi-cloud infrastructure monitoring. Developed Python automation for provisioning, log analysis, and incident response workflows.

Aug 2019 – Aug 2020

Technical Support — Dell Technologies

Advanced enterprise hardware and software support. Automated diagnostic workflows and earned Customer Champion recognition from Director-level leadership.

Mar 2013 – Aug 2019

Service Desk Analyst — Wipro Technologies

L1/L2 IT support operations covering incident triage, escalation, and resolution. Progressively took on automation initiatives and earned the ENU Excellence Award for a Windows 10 migration project.

Jun 2012 – Feb 2013

Front Desk Executive — Smart Tech Solutions

First point of contact for customer escalations. Built the foundation for customer-centric thinking and operational discipline that defines my SRE philosophy today.

Skills & Technologies

SRE & Reliability

SLI/SLO/SLA Error Budgets Toil Reduction Blameless Postmortems Runbook Automation Chaos Engineering Disaster Recovery

Cloud & Infrastructure

AWS GCP Azure Kubernetes Terraform Docker

Observability

Dynatrace Prometheus Grafana OpenTelemetry CloudWatch Azure Monitor

Programming

Go Rust Python Bash SQL

Distributed Systems

Apache Kafka NATS JetStream gRPC PostgreSQL TimescaleDB

AI & AIOps

LLMs RAG Vector Search Anomaly Detection Self-Healing

Certifications

☁️
Google Cloud — Professional ML Engineer
📊
Dynatrace Verified — Essentials
Confluent — Apache Kafka Fundamentals
🔧
Platform Engineering — Observability
🛡️
Coursera — Google SRE Culture
Coursera — GKE Foundations
☁️
Coursera — GCP Core Infrastructure
🚀
Simplilearn — DevOps Postgraduate

Let's Build Something Resilient

I'm interested in conversations about SRE leadership, distributed systems, AIOps, platform engineering, and ambitious infrastructure challenges.