What I look after
04 domainsPlatform & IaC
Environments that rebuild from code rather than from memory. Terraform, Ansible and Kubernetes across AWS, Azure and OCI, wired into CI/CD.
Terraform · Ansible · Kubernetes · AWS · Azure · OCI
Automation & AI
Python and generative models turning an alert flood into a ranked, summarised, already-triaged incident.
65% of triage automated Python · GenAI · REST APIs · AWS Lambda
Observability
Metrics, logs and dashboards that make production legible. Prometheus, Grafana and ELK across Kubernetes and multiple clouds.
−40% alert noise Prometheus · Grafana · ELK · Zabbix · Datadog
Incident Response
Four years of 24×7 on-call. Triage, escalation, root cause analysis, and the postmortems that stop a problem coming back.
−25% mean time to restore SLI / SLO · RCA · ServiceNow · On-call
Toolkit
36 tools & practicesCloud & containers
- AWS
- Azure
- OCI
- Kubernetes
- Docker
- Linux
- EC2
- EKS
- Lambda
- S3
- IAM
- VPC
- SNS
Infrastructure as code
- Terraform
- Ansible
- Puppet
- Azure Functions
Observability
- Prometheus
- Grafana
- ELK Stack
- Datadog
- Zabbix
- CloudWatch
CI/CD
- GitHub Actions
- GitLab CI
- Git
Languages
- Python
- Bash
- PowerShell
- REST APIs
Reliability practice
- Incident management
- Root cause analysis
- SLI / SLO
- On-call
- Capacity planning
- ServiceNow
Let's talk.
Open to DevOps, Platform and Site Reliability roles. Email is the quickest way to reach me.