Cloud Operations - Service Reliability Engineer
What you will do The Service Reliability Engineer is accountable for improving the reliability, observability and operational resilience of cloud-hosted services. The role focuses on monitoring, early issue identification, cloud engineering and automation, using tools and practices such as Bicep, Azure DevOps, GitHub and Ansible to support consistent, repeatable and well-governed service operation.The role involves:Monitoring, observability and alerting across cloud infrastructure, platform services and supported application environments;Cloud engineering and automation, including Infrastructure as Code, deployment pipelines, configuration management and standards-led delivery; andPromoting service resiliency through proactive issue identification, operational insight, automation and continuous improvement.Improve the reliability, resilience and performance of cloud-hosted services through monitoring, observability and automation.Design and enhance monitoring, alerting and operational dashboards to provide real-time insight into service health and performance.Support and develop Azure cloud platforms, working across infrastructure, networking, identity and platform services.Implement Infrastructure as Code and automation solutions using tools such as Bicep, Azure DevOps, GitHub and Ansible.Investigate incidents, identify root causes and drive continuous service improvements through automation and operational excellence.Collaborate with engineering, infrastructure and support ..... full job details .....
Perform a fresh search...
-
Create your ideal job search criteria by
completing our quick and simple form and
receive daily job alerts tailored to you!