Résumé · September 2026
Rashi Kaushik
DevOps & Site Reliability Engineer
DevOps and Site Reliability engineer with 4 yrs building and running production systems on Kubernetes across AWS, Azure and OCI. Infrastructure as code, CI/CD, observability and incident response — with a track record of automating 65% of alert processing, cutting alert noise 40%, reducing mean time to restore 25%, and halving recurring incidents.
Experience
4 yrs
DevOps / Site Reliability Engineer
Accenture · Gurugram, India- Automated infrastructure provisioning, configuration management and operational workflows with Terraform, Ansible, Python, Bash and PowerShell — cutting manual effort and making deployments repeatable.
- Built and optimised CI/CD pipelines covering build, test, deployment and release across development, staging and production environments.
- Designed, deployed and supported highly available production infrastructure across AWS, Azure, OCI, Linux, Docker and Kubernetes.
- Ran containerised workloads on Kubernetes and Docker — deployment troubleshooting, capacity planning, performance tuning and day-to-day production support.
- Built and maintained observability platforms on Prometheus, Grafana, ELK, Zabbix and CloudWatch, including centralised logging and dashboards.
- Developed automation for monitoring, alert management, incident response, health checks and operational remediation.
- Led production incident management, root cause analysis, problem management and service restoration, reducing mean time to restore.
- Defined and reported against SLIs, SLOs and availability targets, with alerting and log analytics tuned to them.
- Held 24×7 on-call for critical production systems — triage, escalation management and post-incident review.
- Worked across AWS EC2, EKS, Lambda, S3, IAM, VPC, SNS and CloudWatch to deploy, secure and monitor enterprise workloads.
Projects
~3 months
Cloud-Native Observability Platform
Prometheus · Grafana · ELK · Kubernetes · AWS · Terraform · PythonBuilt a cloud-native observability platform covering containerised applications and infrastructure across AWS and Kubernetes. Prometheus, Grafana and ELK provide centralised metrics, logs, dashboards and alerting, with Terraform and Python automation making the platform and its operational workflows repeatable.
~2 months
AI-Assisted Alert Correlation Engine
Python · GenAI · Prometheus · ELK · REST APIsDeveloped an alert correlation engine in Python that aggregates, deduplicates, enriches and prioritises alerts from ELK and Prometheus. A generative model writes incident context and summaries on top, so on-call opens a grouped and ranked view instead of a raw notification feed — accelerating triage and cutting alert fatigue.
Skills
Cloud & containers
- AWS
- Azure
- OCI
- Kubernetes
- Docker
- Linux
- EC2
- EKS
- Lambda
- S3
- IAM
- VPC
- SNS
Infrastructure as code
- Terraform
- Ansible
- Puppet
- Azure Functions
Observability
- Prometheus
- Grafana
- ELK Stack
- Datadog
- Zabbix
- CloudWatch
CI/CD
- GitHub Actions
- GitLab CI
- Git
Languages
- Python
- Bash
- PowerShell
- REST APIs
Reliability practice
- Incident management
- Root cause analysis
- SLI / SLO
- On-call
- Capacity planning
- ServiceNow
Education
B.Tech, Computer Science & Engineering
Maharshi Dayanand University, RohtakGrade 8.0 / 10