Site Reliability Engineering
Site Reliability Engineer(Senior)
SecondWind by Joveo
We are hiring a Site Reliability Engineer to keep production reliable, observable, and fast.
About SecondWind by Joveo
SecondWind is a global community to discuss the work, learn what is changing, and discover where your talent-acquisition experience can go next.
About The Role
The role combines software engineering with operations automating away toil and treating reliability as a first-class engineering problem.
Responsibilities
- Design and operate systems for high availability and performance.
- Build and maintain observability tooling (logging, metrics, tracing).
- Define and track SLOs, SLIs, and error budgets.
- Lead incident response and post-mortem reviews.
- Automate operational toil through tooling and platform improvements.
- Partner with application teams on production readiness.
Requirements
- 4+ years in SRE, DevOps, or infrastructure engineering.
- Strong scripting and software engineering skills (Python, Go, or similar).
- Deep experience with cloud platforms (AWS, GCP, Azure).
- Hands-on with Kubernetes, Terraform, and observability platforms.
- Experience leading incident response in production environments.
- Strong understanding of distributed systems.
Benefits
- Fully remote, flexible work hours.
- Performance-based bonus structure.
- Annual learning & development stipend.
- Health and wellness benefits (varies by location).
- Opportunity to work on high-scale, real-world impact projects
More Roles
Other roles you might like
Site Reliability Engineering
Site Reliability Engineer (On-site)
Darktrace
We’re looking for a Site Reliability Engineer (SRE) to bring deep expertise in a key reliability domain and help shape the future of our platform reliability strategy.
Platform Engineering
Automation Platform Engineer [AQ-20432]
Aquent
We are seeking a talented individual to step into a pivotal role where your expertise will directly shape the future of automation and digital enablement. You will be instrumental in ensuring our platforms are not just functional, but secure, stable, and highly performant, empowering delivery teams to innovate and scale.