Senior Specialist Engineer (Specialist Site Reliability Engineer SRE)
This role involves applying engineering principles to manage and improve cloud infrastructure and operational systems within the UK Health Security Agency. The postholder works closely with software engineering, DevOps, and infrastructure teams to automate workflows, ensure system scalability, and respond to production incidents. Day-to-day responsibilities include monitoring system performance, reducing manual toil through Infrastructure as Code, and helping to establish and track Service Level Objectives. The work supports critical health security operations by ensuring that technical services remain stable, reliable, and performant.
What they're looking for
- Experience as a Site Reliability Engineer, DevOps Engineer, or similar role
- Coding skills in Python, PowerShell, or Bash
- Understanding of Linux/Unix, Windows, networking, and distributed systems
- Experience with observability and monitoring tools like Prometheus or Datadog
- Understanding of infrastructure automation tools such as Terraform or Ansible
Nationality requirements
This job is broadly open to UK nationals, Republic of Ireland nationals, Commonwealth citizens with work rights, EU/EEA/Swiss nationals with EUSS status, and Turkish nationals under the Ankara Agreement.
Selection process
Interview dates are to be confirmed.
Apply on Civil Service Jobs ↗Originally listed on Civil Service Jobs.