Automate the toil
Replace manual runbooks with GitOps, pipelines, and IaC that scale with the team.
I own the full delivery lifecycle for 52+ services across 5 Kubernetes clusters — CI/CD, GitOps, observability, security, and cost, all built to hold up at 3 a.m.
Two years making 3 a.m. pages rare: GitOps that self-heals, deploys nobody has to babysit.
I build and run cloud-native delivery platforms on AWS and Kubernetes — owning Dockerization, Helm, CI/CD, GitOps, secrets, observability, security & compliance, cost, and incident response for 52+ production services across Python, Java, and Node.js, with a backend foundation in Java and Spring Boot.
The throughline: automate the toil, secure by default, and own it end to end in production.
Replace manual runbooks with GitOps, pipelines, and IaC that scale with the team.
Shift-left scanning, Vault-managed secrets, and compliance baked into the delivery path.
From root-causing OOM crashes to leading incident response and cutting MTTR.
Mar 2025 → present
Solely own the full delivery lifecycle for 52+ production services across 5 Kubernetes clusters (Python, Java, Node.js) — Dockerization, Helm, CI/CD, and secrets. Drove org-wide GitOps with ArgoCD, migrated secrets to a HashiCorp Vault cluster, architected observability (Prometheus, Grafana, Loki, OpenTelemetry, Coralogix), hardened security & compliance (Trivy, Veracode, AWS WAF, SOC 2), cut ~$4.8K/mo in AWS spend, and led incident response to keep MTTR low.
Aug 2024 → Feb 2025 · 7m
Optimized API performance by profiling and eliminating bottlenecks with YourKit, cut data storage footprint 13% by removing duplicate Postgres audit entries, and drove unit-test coverage past 93% across a multi-module Spring Boot application.
Production ownership, squashed into numbers.
Built over two years solving delivery, observability, and security challenges across 52+ services and 5 Kubernetes clusters.
1stack:2 cloud: [aws]3 containers: [docker, kubernetes, eks, helm]4 delivery: [argocd, jenkins, github-actions]5 iac: [terraform, ansible]6 observability: [prometheus, grafana, loki, alertmanager, opentelemetry, coralogix]7 data: [kafka, msk, postgres, redis]8 security: [vault, trivy, veracode, aws-waf, soc2]9 scripting: [python, bash]1011mode: automate_the_repetitive # gitops • secure by default
ArgoCD, Jenkins, GitHub Actions, Terraform, Ansible
AWS, Kubernetes, Amazon EKS, Docker, Helm
Prometheus, Grafana, Loki, OpenTelemetry, Vault, Trivy, AWS WAF
Solely own the full delivery lifecycle for 52+ production services across 5 Kubernetes clusters (Python, Java, Node.js) — Dockerization, Helm charts, manifests, CI/CD wiring, secrets, and observability, end to end.
Drove org-wide GitOps by migrating every service to ArgoCD — automated sync, self-heal, drift detection, and one-click rollbacks with full deployment audit history.
Migrated secrets from Ansible playbooks to a HashiCorp Vault cluster using AppRole auth in a hub-and-spoke model — centralizing storage, access control, and rotation across every environment.
Architected end-to-end monitoring with the kube-prometheus-stack (Prometheus, Grafana, Loki, Alertmanager) and OpenTelemetry tracing; unified logs, metrics & traces in Coralogix, whose CSPM module surfaced 60+ security findings.
Engineered CI/CD for Python, Java & Node.js running unit, integration, and Playwright tests at PR time; parallelized Jenkins stages to accelerate builds 27%, and killed Docker Hub 429s with an ECR pull-through cache.
Embedded Trivy image scanning and Veracode SAST/DAST across every pipeline (shift-left DevSecOps), remediating 25+ CVEs to harden supply-chain security.
Solely drove infrastructure SOC 2 Type 1 & 2 — remediating weak TLS/CBC ciphers flagged by testssl.sh — and deployed AWS WAF at the edge, blocking 44K+ malicious requests/month.
Cut ~$4,800/month in AWS spend — Athena partitioning (99% cut), single-AZ non-prod RDS/EKS, and automated off-hours shutdowns — with org-wide tagging and billing alerts.
Migrated all 52+ services from AMD64 to ARM64 (AWS Graviton) with multi-platform Docker builds, and architected a fully isolated AWS environment (dedicated VPC, EKS, subnets, IAM, RBAC) for a client's network-segregation requirements.
Killed recurring production OOM crashes by provisioning PVCs to capture JVM heap dumps on crash, and led incident response — storage shortages, bastion crashes, zero-downtime node-group migrations, subnet/IP exhaustion — cutting MTTR.