Site Reliability Engineer - Terraform, Ansible, Python, AWS, GCP, GitOps, Infrastructure a
Site Reliability Engineer - Terraform, Ansible, Python, AWS, GCP, GitOps, Infrastructure as Code (IaC), Grafana, Datadog, Dynatrace, Site Reliability Engineering (SRE). Site Reliability Engineer (Cloud Platform) About the Role Join our Platform Engineering team to improve the reliability, automation, and performance of cloud-based services. You''ll apply Site Reliability Engineering (SRE) practices to enhance observability, automate operations, and support highly available production platforms. Key Responsibilities Drive SRE best practices across cloud infrastructure and platform operations. Improve monitoring, alerting, SLAs/SLOs/SLIs, and platform observability. Automate operational tasks using Infrastructure as Code and GitOps. Develop and maintain automation using Terraform, Ansible, and Scripting. Support production systems, participate in an on-call rota, and lead incident resolution and continuous improvement. About You Experience supporting cloud infrastructure in a production environment. Knowledge of SRE principles, incident management, and root cause analysis. Strong Scripting skills (Python, Ansible, or PowerShell). Experience with AWS, GCP, or similar cloud platforms, Infrastructure as Code, and GitOps. Familiarity with observability tools such as Grafana, Datadog, or Dynatrace. Strong problem-solving, communication, and stakeholder management skills. London/Hybrid/Permanent By applying to this job you are sending us your CV, which may ..... full job details .....
Other jobs of interest...
Perform a fresh search...
-
Create your ideal job search criteria by
completing our quick and simple form and
receive daily job alerts tailored to you!