Short bio
15+ years of experience designing and operating large-scale distributed systems and latency-sensitive production platforms. Expertise in multi-region Kubernetes infrastructure, hybrid cloud environments, and production reliability engineering, with a strong focus on scalability, fault tolerance, observability, and operational automation. Deep experience building and operating hybrid bare-metal and cloud platforms (AWS, GCP), including Kubernetes infrastructure, traffic routing, CI/CD systems, and production observability stacks supporting globally distributed workloads. Strong background in incident response, Linux systems, networking, distributed systems troubleshooting, and infrastructure automation using Terraform and Python. Focused on end-to-end ownership of production infrastructure platforms while partnering closely with engineering teams to improve reliability, security, deployment consistency, and operational efficiency.
Tech stack
Work experience
- Senior DevOps / Infrastructure Engineer (SRE)2023 β 20263 y.
Designed and operated a hybrid cloud and bare-metal Kubernetes platform supporting latency-sensitive multiplayer services (~20K concurrent users) across North America, Europe, and Asia.
Responsibilities:- Built cross-environment autoscaling systems using Agones, dynamically scaling workloads between on-premises infrastructure and cloud based on demand and performance.
- Automated provisioning of 100+ bare-metal servers using Terraform, PXE bootstrapping, and KVM virtualization.
- Operated and modified Akamai CDN properties and Edge DNS configurations, including production updates and TLS certificate renewals supporting global traffic.
- Developed Python-based automation extending cluster autoscaling workflows to dynamically provision Tencent Cloud Elastic IPs (EIPs) from dedicated bandwidth pools during scaling events.
- Designed and operated self-hosted GitHub Actions runners, enabling environment-specific dependency management and improving CI/CD reliability and reproducibility across deployment pipelines.
- Led Elasticsearch cluster upgrades and migrations including automated backup and recovery strategies.
- Engineered PostgreSQL high-availability architecture with replication, connection pooling (PgBouncer), load balancing, and disaster recovery automation.
- Led incident response and debugging across Kubernetes, networking, and distributed systems in production.
- Operated production F5 BIG-IP infrastructure including SSL/TLS certificate management, iRules-based traffic control, firmware upgrades, and zero-downtime application patching.
- Designed and maintained observability platforms (Prometheus, Grafana, Loki, ELK, Splunk), including metrics, alerting, and production logging at scale.
- Senior DevOps / Site Reliability Engineer2020 β 20222 y.
Operated infrastructure supporting large-scale real-time advertising delivery systems with strict latency and availability requirements.
Responsibilities:- Operated and scaled distributed microservices platform running on a company EKS (~20 services, hundreds of pods across multiple clusters).
- Implemented Akamai DataStream ingestion pipeline into cloud storage, enabling large-scale collection and processing of edge traffic logs for analytics and observability.
- Operated and supported Kafka-backed telemetry and distributed messaging systems in production environments, including topic configuration management and operational troubleshooting.
- Led infrastructure initiatives including IAM migration, automated key rotation, Lacework infra, and Redshift private endpoints.
- Conducted infrastructure POC migrating persistent storage from EBS to EFS across workloads spanning 900+ pods to evaluate performance for latency-sensitive services.
- Contributed to observability platform migration from SignalFX to Prometheus and Grafana.
- Participated in follow-the-sun on-call rotations, responding to incidents in latency-sensitive distributed systems.
- Mentored DevOps engineers across North America, Europe, and Asia.
- DevOps Engineer2018 β 20202 y.
Acted as DevOps coach promoting infrastructure-as-code and CI/CD practices across engineering teams.
Responsibilities:- Designed CI/CD pipelines for ~20 microservices across four deployment environments.
- Led Kubernetes workload migration from Azure AKS to AWS EKS.
- Managed infrastructure using Terraform, Helm, Kustomize, and Ansible while defining reliability and security standards.
- Unix / Systems Administrator2011 β 20187 y.
Managed Linux infrastructure across fleets of 1000+ servers spanning multiple data centers.
Responsibilities:- Automated provisioning, monitoring, and operational reliability processes supporting large production environments.

