Site Reliability Engineering

Site Reliability Engineer (On-site)

Darktrace

Greater Cambridge AreaFull-Time · On-siteCompetitive Salary

We’re looking for a Site Reliability Engineer (SRE) to bring deep expertise in a key reliability domain and help shape the future of our platform reliability strategy.

About Darktrace

Darktrace is a global leader in AI cybersecurity, providing the essential cybersecurity platform to secure organizations today and for an ever-changing future. Darktrace AI learns from each business's unique data in real time, detecting threats and intervening against attacks with precision and speed. We are a diverse and inclusive team of over 2,400 employees, each playing a crucial role in protecting nearly 10,000 organizations and communities worldwide from known, unknown, and novel cyber-threats.

About The Role

You’ll act as the go-to authority in your area of specialism, working across teams to embed best practices, solve complex reliability challenges, and improve system resilience at scale.

Responsibilities

  • Act as the subject matter expert in your chosen reliability domain
  • Define and implement standards, frameworks, and best practices across SRE, Platform Engineering, and DevSecOps
  • Stay current with industry trends and bring innovative ideas into the organisation
  • Design and implement solutions to complex, cross-cutting reliability challenges
  • Build tooling, automation, and frameworks to improve system resilience and scalability
  • Lead deep-dive investigations into systemic issues and drive long-term fixes
  • Partner with Platform Engineering to ensure your domain is embedded within the internal developer platform
  • Collaborate with DevSecOps to integrate security, compliance, and resilience practices
  • Contribute to cross-team initiatives that improve reliability across the stack
  • Play a key role in incident response, particularly within your specialism
  • Contribute to on-call rotations and continuous improvement of operational processes
  • Develop runbooks, documentation, and training materials to support teams

Requirements

  • Proven experience in Site Reliability Engineering, DevOps, or infrastructure engineering
  • Deep expertise in at least one of the following areas:
  • Observability & monitoring (metrics, logging, distributed tracing)
  • Performance engineering & capacity planning
  • Data infrastructure reliability (databases, streaming, pipelines)
  • Security-focused SRE (hardening, compliance automation, secrets management)
  • Network reliability & traffic management
  • Strong programming skills (e.g. Go, Python, or similar)
  • Experience with cloud platforms (AWS, GCP, Azure) and Kubernetes
  • Strong communication skills, with the ability to explain complex technical concepts clearly
  • Self-driven with the ability to identify and prioritise high-impact work independently

Desirables

  • Experience building internal developer platforms or tooling
  • Contributions to open-source, technical blogs, or public speaking
  • Experience working in regulated environments
  • Familiarity with SLO frameworks and error budget management
  • Relevant certifications in your specialist domain

Benefits

  • 23 days’ holiday + all public holidays, rising to 25 days after 2 years of service.
  • Additional day off for your birthday.
  • Private medical insurance which covers you, your cohabiting partner and children.
  • Life insurance of 4 times your base salary.
  • Salary sacrifice pension scheme.
  • Enhanced family leave.
  • Confidential Employee Assistance Program.
  • Cycle to work scheme.

More Roles

Other roles you might like

Cookie Notice

Cookies and similar technologies

We use cookies and similar technologies to keep this site working and to ensure the best possible experience for all our users. Kindly confirm your consent to our use of them.

Learn More