Platform Engineering
Observability Platform Engineer
Era4
As an Observability Platform Engineer, you will design, implement, and operate Era4’s enterprise observability platform across its AI infrastructure environment. You will build and maintain highly available telemetry, monitoring, alerting, and dashboarding capabilities using the Grafana ecosystem, Kubernetes, OpenTelemetry, and Infrastructure-as-Code.
About Era4
Era4 delivers resilient, lower carbon, high-performance AI infrastructure. We are enabling the UK’s AI transformation by converting our brownfield sites into data centres for the new industrial era, creating ultra-fast AI capacity while supporting regeneration and bringing new investment into the communities around us. Era4 is the new name of Carbon3.ai since 4 March 2026.
About The Role
- You will be working across compute, GPU, storage, networking, and application platforms.
- You will establish scalable observability standards, automate platform operations, support incident detection and root-cause analysis, and ensure the platform is secure, resilient, and maintainable as Era4 scales.
Responsibilities
- Design and implement highly available observability services across multiple co-location and production sites.
- Define platform standards for telemetry collection, labelling, metadata enrichment, retention policies, and data governance.
- Implement multi-tenant observability controls and tenant isolation strategies
- Configure telemetry ingestion pipelines for metrics, logs, and future distributed tracing workloads.
- Configure and maintain object-storage-backed telemetry platforms for long-term retention and scalability.
- Deploy and manage Grafana Alloy collectors across Kubernetes clusters, Linux hosts, network infrastructure, storage platforms, and hardware management systems.
- Integrate telemetry from Kubernetes, GPU infrastructure, HPE hardware, storage platforms, network devices, and cloud-native services.
- Develop and maintain observability integrations using OpenTelemetry standards and protocols.
- Establish onboarding processes for new platforms, applications, and infrastructure services.
- Collaborate with application teams to define observability requirements and future tracing adoption strategies.
- Design and implement alerting frameworks using recording rules, AlertManager, and operational best practices.
- Develop operational dashboards and service health views for infrastructure, platform, and application services.
- Support integration of observability events with ITSM and incident-management platforms.
- Define SLIs, SLOs, alert thresholds, and operational KPIs.
- Continuously improve platform observability, incident detection, and root-cause analysis capabilities.
- Implement Infrastructure-as-Code and GitOps practices for observability platform deployment and configuration management.
- Develop automation for dashboard provisioning, alert deployment, tenant onboarding, and telemetry configuration.
- Design and validate disaster recovery, resilience, and failover capabilities across observability services.
- Contribute to platform security, compliance, and operational governance initiatives.
- Work with operational teams to ensure observability services remain reliable, scalable, and maintainable.
Requirements
- Significant experience implementing and operating enterprise observability or monitoring platforms.
- In depth understanding of metrics, logs, traces, OpenTelemetry, and modern observability principles.
- Grafana ecosystem technologies - Grafana, Prometheus, Grafana Mimir, Grafana Loki, Grafana Tempo, Grafana Alloy.
- Designing Kubernetes-native solutions and operating distributed platforms at scale.
- Linux systems administration and cloud-native infrastructure.
- Infrastructure-as-Code and GitOps approaches (preferably including Ansible).
- Automation and operational tooling using Python and/or Go.
- Technical architecture, operational documentation, and deployment designs.
- Object storage technologies and distributed data platforms.
More Roles
Other roles you might like
Platform Engineering
Platform Support Engineer (Eng I)
Kraken
Platform Support Engineering is a new team within Kraken’s Platform department. It will be the primary contact point for most day-to-day engineering interactions between the company’s 2,000 staff and the 120 engineers in the Platform department. The team will handle a high volume of tickets each month, supporting the business across a myriad of clients in the EU, North America, and Australia-Pacific regions on a 24/5 basis.
Platform Engineering
Platform Engineer (Hybrid)
9fin
We are looking for a highly-motivated and driven platform engineer, with experience in enabling a DevOps culture within tech teams. In this role, your primary responsibility will be to provide a world-class developer experience to other engineers.