Senior Site Reliability Engineer
Overview The Senior Site Reliability Engineer is responsible for the reliability, availability, scalability, and operational excellence of our SaaS platform. This role combines software engineering, cloud infrastructure, automation, and operational leadership to build resilient systems that enable rapid product delivery. ResponsibilitiesOwn the reliability and performance of production services.Define and operate SLIs, SLOs, and error budgets with engineering teams.Lead incident response and drive blameless postmortems and continuous improvement.Automate operational processes and reduce manual toil through engineering.Build and operate cloud-native platforms using Azure, Kubernetes, and Infrastructure as Code.Develop observability through effective monitoring, alerting, and telemetry.Mentor engineers and promote reliability best practices across the organisation.ExperienceProven experience as a Senior or experienced Site Reliability Engineer with a software engineering background.Ability to diagnose and make safe changes to PHP and Java or .NET applications.Experience operating large-scale production SaaS systems.Strong knowledge of SRE principles, incident management, observability, and operational excellence.Hands-on experience with Azure, Kubernetes, Infrastructure as Code, and monitoring platforms such as Prometheus, Grafana, or Datadog.Experience influencing engineering teams and driving reliability improvements through collaboration and technical leadership.Key ..... full job details .....
Other jobs of interest...
Perform a fresh search...
-
Create your ideal job search criteria by
completing our quick and simple form and
receive daily job alerts tailored to you!