Production Cluster Migration
Migrated production workloads with zero downtime, including sandbox (sbx) migration and infra-switch AP usage with controlled traffic movement.
I design, build, and operate scalable, resilient infrastructure across AWS and GCP. From production migrations to incident response, I turn complex systems into reliable, observable, and cost-efficient platforms.
GCPScalable. Global. Open.
AWSSecure. Resilient. Innovate.
AWS
Google Cloud
Kubernetes
Terraform
Docker
Linux
Prometheus
Grafana
Kafka
PostgreSQL
MySQL
Redis
ClickHouse
OthersI like systems that fail quietly and recover loudly — clear alerts, fast rollbacks, and dashboards that tell the truth. My work sits at the intersection of infrastructure, automation, and incident response: turning fragile, manual processes into resilient, observable platforms that let teams ship with confidence. The best infrastructure is invisible — it just works, and when it doesn't, it tells you exactly why.
3+ years operating and scaling infrastructure across fintech and enterprise environments.
"Work 10x and Grow 10x."
The principle that's guided every role — take ownership beyond the job description, and the growth follows.
B.E. Computer Science & Engineering
P.A. College of Engineering and Technology · 2018 – 2022
The disciplines I focus on when designing and operating production infrastructure.
Provisioning, IAM, networking, and cost optimization across multi-cloud environments.
Designing and operating production-grade Kubernetes clusters at scale.
Building metrics, logging, and tracing pipelines that surface signal, not noise.
Running and tuning databases and caches under real production load.
Automating build, test, and deploy pipelines for fast, safe releases.
Incident response, on-call rotations, postmortems, and SLO-driven engineering.
A few highlights from my SRE journey — solving real problems at scale.
Migrated production workloads with zero downtime, including sandbox (sbx) migration and infra-switch AP usage with controlled traffic movement.
Designed and operated VictoriaMetrics, vmagent and Grafana for high-cardinality metrics across multiple clusters, enabling better visibility and faster incident resolution.
Handled migrations from RDS to Aurora, created read replicas, optimized Aurora I/O, and managed Redis, MySQL, Scylla and ClickHouse at scale.
A generic, illustrative reference architecture — not a real production system — showing how I'd typically structure a resilient, observable service.
Real experiences, deep dives, and practical takeaways from operating systems at scale.
How we identified and resolved a sudden latency increase in our EKS cluster.
Right-sizing, automated scaling, and intelligent shutdowns.
Building actionable dashboards and meaningful alerts.