Curriculum Vitae
Senior Site Reliability Engineer and Multi-Cloud DevOps Consultant with 10+ years of experience designing, automating, and managing cloud infrastructure across AWS, Azure, and Google Cloud environments.
Strong expertise in Kubernetes (EKS, AKS, GKE), Terraform, GitLab CI/CD, Azure DevOps, Jenkins, GitHub Actions, Argo CD, Docker, Linux Administration, Shell Scripting, and Python Automation.
Experienced in building highly available and scalable cloud platforms, implementing Infrastructure as Code, automating deployment pipelines, managing production workloads, incident management, monitoring, and reliability engineering.
Hands-on experience supporting mission-critical applications, maintaining SLA/SLO objectives, performing root cause analysis, cloud cost optimization, and implementing cloud security best practices.
TECHNICAL SKILLS
Cloud Platforms
AWS, Microsoft Azure, Google Cloud Platform
Containers & Orchestration
Docker, Kubernetes, EKS, AKS, GKE, Helm, Argo CD
CI/CD Tools
GitLab CI/CD, Azure DevOps, Jenkins, GitHub Actions
Infrastructure as Code
Terraform, Ansible
Monitoring & Observability
Prometheus, Grafana, ELK Stack, Datadog, New Relic, Grafana Loki
Programming & Automation
Python, Shell Scripting
Operating Systems
Linux (RHEL, Ubuntu), Windows
Databases
MySQL, MS SQL Server, MongoDB
Senior Site Reliability Engineer / Multi-Cloud DevOps Consultant
Jan 2019 – Present
Project4: Multi-Cloud Platform Engineering & Site Reliability Engineering
Duration: Jul 2025 – Present
Environment
AWS, Azure, Google Cloud, Kubernetes, Terraform, GitLab CI/CD, GitHub Actions, Argo CD, Python, Prometheus, Grafana
Responsibilities
- Designed and managed cloud infrastructure across AWS, Azure, and Google Cloud environments.
- Built reusable Terraform modules and Infrastructure as Code standards for enterprise cloud platforms.
- Implemented CI/CD pipelines using GitLab CI/CD, GitHub Actions, and Argo CD.
- Managed Kubernetes platforms including EKS, AKS, and GKE clusters.
- Developed Python automation scripts for cloud operations and infrastructure management.
- Implemented monitoring and observability solutions using Prometheus, Grafana, Loki, and Datadog.
- Performed production support, incident management, root cause analysis, and reliability improvements.
- Collaborated with development teams to implement secure and scalable deployment solutions.
- Worked on cloud governance, security, performance optimization, and cost optimization initiatives.
- Provided technical consulting and architecture guidance across multiple customer environments.
Project3: Azure Cloud Operations & Site Reliability Engineering
Duration: Jan 2022 – Jun 2025
Environment
Azure, AKS, Azure DevOps, Terraform, Docker, Kubernetes, Shell Scripting, Prometheus, Grafana
Responsibilities
- Managed production-grade Azure Kubernetes Service (AKS) clusters.
- Designed and maintained CI/CD pipelines using Azure DevOps.
- Implemented Infrastructure as Code using Terraform.
- Managed Azure services including Virtual Machines, Virtual Networks, Network Security Groups, Azure Container Registry, App Services, Front Door, Key Vault, Azure Monitor, Storage Accounts, and Application Gateway.
- Implemented high availability, disaster recovery, and scalability solutions.
- Performed Kubernetes upgrades, patching, cluster maintenance, and security hardening.
- Automated operational activities using Shell scripting.
- Monitored applications and infrastructure using Prometheus and Grafana.
- Participated in incident response, RCA preparation, and SLA/SLO management activities.
- Supported production environments and platform reliability initiatives.
Project2: Google Cloud Platform DevOps Automation
Duration: Jan 2021 – Dec 2021
Environment
Google Cloud Platform, GKE, Cloud Build, Artifact Registry, Terraform, Docker, Jenkins
Responsibilities
- Managed Google Kubernetes Engine (GKE) clusters and containerized applications.
- Implemented CI/CD pipelines using Jenkins and Google Cloud Build.
- Automated infrastructure provisioning using Terraform.
- Managed Google Cloud services including Compute Engine, Cloud Storage, Cloud SQL, IAM, BigQuery, Cloud Run, and Artifact Registry.
- Supported application deployment, monitoring, and operational activities.
- Worked with development teams to improve deployment automation and release management.
- Implemented cloud-native deployment practices and Kubernetes-based application hosting.
Project1: AWS Infrastructure Automation & DevOps
Duration: Jan 2019 – Dec 2020
Environment
AWS, Jenkins, Docker, Terraform, Linux, Ansible, Shell Scripting
Responsibilities
- Managed AWS cloud infrastructure and DevOps operations.
- Built CI/CD pipelines using Jenkins for automated application deployments.
- Implemented Infrastructure as Code using Terraform.
- Managed AWS services including EC2, S3, IAM, VPC, Route53, Application Load Balancers, Auto Scaling Groups, RDS, ECR, ECS, EKS, CloudWatch, Lambda, SNS, SQS, and Systems Manager.
- Automated server administration and operational tasks using Shell scripting and Ansible.
- Managed Linux servers, application deployments, and troubleshooting activities.
- Configured monitoring, logging, backup, and disaster recovery solutions.
- Supported production systems and performed incident resolution activities.