Résumé  ·  September 2026

Rashi Kaushik

DevOps & Site Reliability Engineer

DevOps and Site Reliability engineer with 4 yrs building and running production systems on Kubernetes across AWS, Azure and OCI. Infrastructure as code, CI/CD, observability and incident response — with a track record of automating 65% of alert processing, cutting alert noise 40%, reducing mean time to restore 25%, and halving recurring incidents.

Experience

Jul 2022 — present
4 yrs

DevOps / Site Reliability Engineer

Accenture  ·  Gurugram, India
  • Automated infrastructure provisioning, configuration management and operational workflows with Terraform, Ansible, Python, Bash and PowerShell — cutting manual effort and making deployments repeatable.
  • Built and optimised CI/CD pipelines covering build, test, deployment and release across development, staging and production environments.
  • Designed, deployed and supported highly available production infrastructure across AWS, Azure, OCI, Linux, Docker and Kubernetes.
  • Ran containerised workloads on Kubernetes and Docker — deployment troubleshooting, capacity planning, performance tuning and day-to-day production support.
  • Built and maintained observability platforms on Prometheus, Grafana, ELK, Zabbix and CloudWatch, including centralised logging and dashboards.
  • Developed automation for monitoring, alert management, incident response, health checks and operational remediation.
  • Led production incident management, root cause analysis, problem management and service restoration, reducing mean time to restore.
  • Defined and reported against SLIs, SLOs and availability targets, with alerting and log analytics tuned to them.
  • Held 24×7 on-call for critical production systems — triage, escalation management and post-incident review.
  • Worked across AWS EC2, EKS, Lambda, S3, IAM, VPC, SNS and CloudWatch to deploy, secure and monitor enterprise workloads.

Projects

2024
~3 months

Cloud-Native Observability Platform

Prometheus · Grafana · ELK · Kubernetes · AWS · Terraform · Python

Built a cloud-native observability platform covering containerised applications and infrastructure across AWS and Kubernetes. Prometheus, Grafana and ELK provide centralised metrics, logs, dashboards and alerting, with Terraform and Python automation making the platform and its operational workflows repeatable.

2024
~2 months

AI-Assisted Alert Correlation Engine

Python · GenAI · Prometheus · ELK · REST APIs

Developed an alert correlation engine in Python that aggregates, deduplicates, enriches and prioritises alerts from ELK and Prometheus. A generative model writes incident context and summaries on top, so on-call opens a grouped and ranked view instead of a raw notification feed — accelerating triage and cutting alert fatigue.

Skills

Cloud & containers

  • AWS
  • Azure
  • OCI
  • Kubernetes
  • Docker
  • Linux
  • EC2
  • EKS
  • Lambda
  • S3
  • IAM
  • VPC
  • SNS

Infrastructure as code

  • Terraform
  • Ansible
  • Puppet
  • Azure Functions

Observability

  • Prometheus
  • Grafana
  • ELK Stack
  • Datadog
  • Zabbix
  • CloudWatch

CI/CD

  • GitHub Actions
  • GitLab CI
  • Git

Languages

  • Python
  • Bash
  • PowerShell
  • REST APIs

Reliability practice

  • Incident management
  • Root cause analysis
  • SLI / SLO
  • On-call
  • Capacity planning
  • ServiceNow

Education

2022

B.Tech, Computer Science & Engineering

Maharshi Dayanand University, Rohtak

Grade 8.0 / 10