nitesh@prod :~$ deploy --env production --portfolio
› initializing build environment…
0%
DevOps Engineer  · ap-south-1 / Bengaluru

NITESH KUMAR

I own the full delivery lifecycle for 52+ services across 5 Kubernetes clusters — CI/CD, GitOps, observability, security, and cost, all built to hold up at 3 a.m.

status open to opportunities
region ap-south-1 · Bengaluru
since 2024 · 2y+ uptime
local 00:00:00
scroll
DevOps Engineering AWS Kubernetes Amazon EKS Docker Helm ArgoCD GitOps Terraform Ansible Infrastructure as Code Jenkins GitHub Actions CI/CD Pipelines Prometheus Grafana Loki Alertmanager OpenTelemetry Coralogix Observability HashiCorp Vault Trivy Veracode AWS WAF DevSecOps SOC 2 Apache Kafka Amazon MSK AWS EC2 AWS RDS AWS ECR AWS Athena VPC IAM Cost Optimization Incident Response Python Bash
01 — whoami

Automation isn't a task.
It's the job.

Two years making 3 a.m. pages rare: GitOps that self-heals, deploys nobody has to babysit.

I build and run cloud-native delivery platforms on AWS and Kubernetes — owning Dockerization, Helm, CI/CD, GitOps, secrets, observability, security & compliance, cost, and incident response for 52+ production services across Python, Java, and Node.js, with a backend foundation in Java and Spring Boot.

The throughline: automate the toil, secure by default, and own it end to end in production.

01

Automate the toil

Replace manual runbooks with GitOps, pipelines, and IaC that scale with the team.

02

Secure by default

Shift-left scanning, Vault-managed secrets, and compliance baked into the delivery path.

03

Own it in prod

From root-causing OOM crashes to leading incident response and cutting MTTR.

02 — git log --career

Deployment history.

release/2025.03 — HEAD

Lumber

Mar 2025 → present

DevOps Engineer now

Solely own the full delivery lifecycle for 52+ production services across 5 Kubernetes clusters (Python, Java, Node.js) — Dockerization, Helm, CI/CD, and secrets. Drove org-wide GitOps with ArgoCD, migrated secrets to a HashiCorp Vault cluster, architected observability (Prometheus, Grafana, Loki, OpenTelemetry, Coralogix), hardened security & compliance (Trivy, Veracode, AWS WAF, SOC 2), cut ~$4.8K/mo in AWS spend, and led incident response to keep MTTR low.

Kubernetes ArgoCD AWS Terraform Vault Prometheus OpenTelemetry Jenkins Trivy AWS WAF
release/2024.08

Lumber

Aug 2024 → Feb 2025 · 7m

Software Development Engineer — Intern

Optimized API performance by profiling and eliminating bottlenecks with YourKit, cut data storage footprint 13% by removing duplicate Postgres audit entries, and drove unit-test coverage past 93% across a multi-module Spring Boot application.

Java Spring Boot PostgreSQL JMeter YourKit
02a — git diff --stat

Two years, diffed.

Production ownership, squashed into numbers.

52+ Production services owned
27% Faster CI builds
$4.8K Monthly AWS spend cut
44K+ Malicious requests blocked / mo
03 — cat stack.yaml

The toolchain, declared.

production toolkit

Built over two years solving delivery, observability, and security challenges across 52+ services and 5 Kubernetes clusters.

infra/stack.yaml · tracked · drift checked hover a card ↓
1stack:2  cloud:          [aws]3  containers:     [docker, kubernetes, eks, helm]4  delivery:       [argocd, jenkins, github-actions]5  iac:            [terraform, ansible]6  observability: [prometheus, grafana, loki, alertmanager, opentelemetry, coralogix]7  data:           [kafka, msk, postgres, redis]8  security:       [vault, trivy, veracode, aws-waf, soc2]9  scripting:      [python, bash]1011mode: automate_the_repetitive  # gitops • secure by default
ship CI/CD & GitOps

ArgoCD, Jenkins, GitHub Actions, Terraform, Ansible

run Cloud & Containers

AWS, Kubernetes, Amazon EKS, Docker, Helm

guard Observability & Security

Prometheus, Grafana, Loki, OpenTelemetry, Vault, Trivy, AWS WAF

04 — ls -la credentials/

Proof, not promises.

certifications
CNCF Certified Kubernetes Administrator (CKA) The Linux Foundation
impact across the stack

Solely own the full delivery lifecycle for 52+ production services across 5 Kubernetes clusters (Python, Java, Node.js) — Dockerization, Helm charts, manifests, CI/CD wiring, secrets, and observability, end to end.

Platform Ownership 52+ services · 5 clusters · 3 stacks

Drove org-wide GitOps by migrating every service to ArgoCD — automated sync, self-heal, drift detection, and one-click rollbacks with full deployment audit history.

GitOps at Scale ArgoCD · self-heal · rollbacks

Migrated secrets from Ansible playbooks to a HashiCorp Vault cluster using AppRole auth in a hub-and-spoke model — centralizing storage, access control, and rotation across every environment.

Secrets Management HashiCorp Vault · AppRole

Architected end-to-end monitoring with the kube-prometheus-stack (Prometheus, Grafana, Loki, Alertmanager) and OpenTelemetry tracing; unified logs, metrics & traces in Coralogix, whose CSPM module surfaced 60+ security findings.

Observability & Tracing Prometheus · OTel · Coralogix

Engineered CI/CD for Python, Java & Node.js running unit, integration, and Playwright tests at PR time; parallelized Jenkins stages to accelerate builds 27%, and killed Docker Hub 429s with an ECR pull-through cache.

CI/CD Engineering Jenkins · 27% faster · Playwright

Embedded Trivy image scanning and Veracode SAST/DAST across every pipeline (shift-left DevSecOps), remediating 25+ CVEs to harden supply-chain security.

DevSecOps Trivy · Veracode · 25+ CVEs

Solely drove infrastructure SOC 2 Type 1 & 2 — remediating weak TLS/CBC ciphers flagged by testssl.sh — and deployed AWS WAF at the edge, blocking 44K+ malicious requests/month.

Compliance & Edge Security SOC 2 · AWS WAF · TLS hardening

Cut ~$4,800/month in AWS spend — Athena partitioning (99% cut), single-AZ non-prod RDS/EKS, and automated off-hours shutdowns — with org-wide tagging and billing alerts.

Cost Optimization ~$4.8K/mo · Athena · tagging

Migrated all 52+ services from AMD64 to ARM64 (AWS Graviton) with multi-platform Docker builds, and architected a fully isolated AWS environment (dedicated VPC, EKS, subnets, IAM, RBAC) for a client's network-segregation requirements.

Migrations & Isolation Graviton · dedicated VPC/EKS

Killed recurring production OOM crashes by provisioning PVCs to capture JVM heap dumps on crash, and led incident response — storage shortages, bastion crashes, zero-downtime node-group migrations, subnet/IP exhaustion — cutting MTTR.

Reliability & Incident Response OOM fixes · zero-downtime · MTTR
education
Bachelor of Technology
Computer Science & Engineering · CGPA 8.1
HKBK College of Engineering
2020 → 2024
05 — ping nitesh

Got something
worth automating?

One conversation about your pipelines, platforms, or on-call pain, and how to make deploys boring again.

base: Bengaluru, Karnataka, IN 12.9716° N · 77.5946° E ↗ tz: IST · UTC+05:30