**Please strictly adhere to the following resume naming convention:ALL CAPS, NO SPACES BETWEEN UNDERSCORESPTN_US_GBAMSREQID_CandidateBeelineIDExample: PTN_US_9999999_SKIPJOHNSON0413: - MSP Owner: Michelle LeeLocation: Dallas, TX / HybridDuration: 6 monthsGBaMS ReqID: 10933590 We are looking for a highly experienced Senior Site Reliability Engineer to design, operate, automate, and continuously improve highly reliable, scalable, secure, and observable cloud GCP. The ideal candidate will bring deep hands-on experience in SRE practices, multi-cloud infrastructure, Kubernetes platforms, CI/CD engineering, infrastructure automation, observability, disaster recovery, incident response, and reliability governance. This role is suited for an engineer with strong experience in SLI/SLO management, error budget governance, cloud migration, platform modernization, production support, and automation-led toil reduction. The candidate will be responsible for improving platform availability, deployment reliability, incident response maturity, disaster recovery readiness, observability coverage, and operational efficiency across enterprise cloud environments. The role requires close collaboration with development, security, infrastructure, platform, and business teams to ensure business-critical applications meet defined reliability, performance, compliance, and scalability objectives.Key Responsibilities• Site Reliability Engineering and Reliability Governance (SLIs, SLOs, error budgets, reliability reviews, and blameless postmortem practices, MTTD, MTTR, recurring incidents)• Multi-Cloud Infrastructure Engineering• Kubernetes, Containers, and Platform Engineering (Kubernetes, EKS, AKS, GKE, ECS, and Docker-based platforms)• Infrastructure as Code and Automation (Terraform)• CI/CD and Release Reliability (GitHub Actions, blue-green, canary, rolling deployments, automated rollback, deployment validation, and automated testing)• Observability, Monitoring, and Logging (Prometheus, Grafana)• Disaster Recovery, High Availability, and Resilience• Security, Compliance, and Cloud Governance• Linux Systems Administration and Production SupportRequired Qualifications• 10+ years of experience in Site Reliability Engineering, DevOps, cloud infrastructure, Linux administration, production operations, or platform engineering.• Strong hands-on experience designing, operating, and supporting cloud infrastructure across AWS, Azure, and GCP.• Deep experience with Kubernetes platforms such as EKS, AKS, GKE, and containerization using Docker.• Strong Infrastructure as Code experience using Terraform, AWS CloudFormation, ARM Templates, and Bicep.• Experience building and maintaining CI/CD pipelines using Jenkins, Azure DevOps, GitLab CI, GitHub Actions, and AWS CodePipeline.• Strong observability experience with Prometheus, Grafana, ELK Stack, OpenSearch, CloudWatch, Azure Monitor, Log Analytics, Application Insights, and GCP Cloud Monitoring.• Experience with disaster recovery, high availability, backup automation, multi-region failover, and recovery validation.• Hands-on scripting and automation experience using Python, Bash, PowerShell, and Ansible.• Strong understanding of cloud networking, including VPC, VNet, subnets, NAT gateways, load balancers, VPN, Direct Connect, ExpressRoute, PrivateLink, Private Endpoints, Route 53, and Transit Gateway.• Strong Linux systems administration experience across enterprise production environments.Preferred Qualifications• Experience implementing GitOps using ArgoCD, Helm, and Kustomize.• Experience with service mesh and API traffic management using Istio Service Mesh and API Gateway.• Experience supporting regulated enterprise environments in healthcare, financial services, banking, insurance, or similarly controlled domains. The resume includes experience across financial, healthcare, insurance, retail, and banking clients.• Experience with cloud cost optimization using AWS Cost Explorer, Azure Advisor, Azure Cost Management, Budgets, Trusted Advisor, and Auto Scaling.• Experience with security and governance tooling such as AWS IAM, AWS Config, GuardDuty, CloudTrail, Azure Policy, Defender for Cloud, Azure Security Center, Kubernetes RBAC, Secrets Manager, and Azure Key Vault.• Experience authoring runbooks, DR playbooks, operational procedures, architecture diagrams, and incident response documentation.Certifications Preferred• Google Cloud Certified Cloud EngineerSkills: Digital : Python ~ Digital : Kubernetes ~ MySQLExperience Required: 8-10