Cloud Operations Engineer
Building reliable systems
where cloud meets scale.
I design, automate, troubleshoot, and operate production infrastructure across AWS, Azure, and GCP, with a focus on Kubernetes, observability, distributed systems, and practical AI for operations.
PDF · updated September 2026
Selected work
Production problems.
Measurable outcomes.
Case studies from reliability engineering, platform modernization, and distributed systems, plus the side projects I build when the on-call phone is quiet.
AI for operations · SRE
AI incident-response assistant
An agentic workflow that isolates an alert's timeframe, enriches the Jira ticket with the matching logs, traces, and metrics, then runs a RAG pipeline over internal runbooks to hand on-call evidence-backed remediation steps instead of a blank dashboard.
Observability · SRE
Centralized observability platform
Built the team's first centralized open-source observability platform across three monitoring clusters, collecting six telemetry types from spoke environments and introducing SLI/SLO-driven reliability management.
Kubernetes · Databases · FinOps
MongoDB modernization
Owned the Kubernetes side of a migration from standalone database hosts and Enterprise MongoDB to a simplified Kubernetes-managed Percona Server for MongoDB architecture.
AWS · Event-driven systems
AWS Migration Hub workflows
Built backend services and event-driven workflows for customer migration tracking, including GDPR-compliant account deletion and dead-letter queue handling.
Cassandra · Performance
Distributed database troubleshooting
Investigated compaction behavior, JVM heap dumps, and replication topology for enterprise Cassandra deployments, then guided architecture changes to remove performance bottlenecks.
Side projects
Things I build to learn something specific.
Computer Vision · On-device AI
In Progress
Cricket DRS
An on-device video review app for run-outs and stumpings. The pipeline combines rolling video capture, crease detection, a surface-adaptive classical vision path, and a YOLOv11-nano model prepared for Core ML inference.
Built a training dataset of ~2,000 images from five public sources and redesigned the detector after real-ground testing showed that one global extraction threshold would not generalize.
Experience
From production incidents
to platform architecture.
My work sits at the intersection of infrastructure, software engineering, and reliability, from customer-facing escalations to multi-cloud platform operations.
Cloud Operations Engineer
Cumulocity GmbH · IoT Platform
Operate 54 multi-cloud Kubernetes clusters across 22 enterprise customers and 10 public platform environments. Built an agentic, RAG-based incident-response assistant and SLO-driven observability, run FluxCD GitOps delivery, and modernized the database tier.
System Development Engineer
Amazon Web Services
Built backend services and event-driven workflows for AWS Migration Hub using serverless and containerized AWS services, distributed tracing, and progressive delivery.
Technical Support Engineer · Enterprise Escalations
DataStax
Owned high-severity Cassandra and DataStax Enterprise escalations across Kubernetes and standalone deployments, from diagnosis through remediation and postmortems.
DevOps Engineer
InterGlobe Aviation Ltd.
Automated hybrid-cloud provisioning and application delivery, built CI/CD pipelines, and created monitoring and operations automation with Python and Bash.
Engineering toolbox
Tools are useful.
Judgment matters more.
I use the stack that fits the failure mode, scale, and operational constraints, rather than treating technology names as a sticker collection.
Cloud
AWS · Azure · GCP
EKS, EC2, RDS, S3, DynamoDB, Lambda, SQS/SNS, Step Functions, CloudWatch, X-Ray; AKS, GKE, Cloud Run, Pub/Sub, Vertex AI, networking, IAM, KMSKubernetes, IaC & Delivery
Kubernetes · Terraform · Helm
Docker, Kustomize, FluxCD, GitHub Actions, Jenkins, SOPS, CloudFormation, Ansible, Atlantis, progressive deliveryAI & Platform Automation
LangGraph · LangChain · FastAPI
RAG pipelines, Chroma DB, vector databases, multi-agent workflows, Jira APIObservability
Grafana ecosystem · Prometheus
Alloy, Mimir, Loki, Tempo, Pyroscope, Alertmanager, SLOs, incident responseSoftware & Data
Python · Bash · Java
MongoDB, Cassandra, DynamoDB, Kafka, Pulsar, REST APIs, event-driven systemsCustomer & Reliability
Architecture · Escalations · RCA
Technical demos, proofs of concept, cost reviews, postmortems, on-call operationsCertifications
Validated depth across cloud, Kubernetes, and ML.
Active
Currently validPreviously held
Lapsed 2024 · badges still verifiable
About
I like systems that explain themselves.
I’m Ravi, a cloud operations engineer with a software-engineering background. I enjoy work where reliability is observable, automation removes toil, and architecture decisions can be connected to concrete operational outcomes.
My recent focus is multi-cloud Kubernetes, GitOps delivery, SLO-driven observability, database modernization, and distributed-system troubleshooting, plus the AI incident-response assistant in my work above.
Beyond engineering
Curiosity leaks out of the terminal.
Cricket gave me a computer-vision project. Kayaking and skydiving mostly give me better stories.


