We are seeking a Platform Engineer to automate software delivery and production operations across data pipelines, servers, and application services. The role spans infrastructure provisioning, change management, deployment through development, staging, and production, observability, and reliability engineering.
You will build the systems and automation that help engineering teams deliver changes safely and on-call responders detect, diagnose, and recover from failures. This role does not participate in the on-call rota.
Key Responsibilities
•
Design and manage cloud infrastructure using infrastructure as code, including Terraform.
•
Build and maintain CI/CD pipelines that test and promote changes through development, staging, and production, with appropriate approval controls, canary deployments where appropriate, telemetry-based observation periods, and rollback capabilities.
•
Implement developer platform capabilities, templates, and self-service infrastructure.
Build actionable monitoring and alerting, including alert routing and escalation, test that alerts reach the right responders, and measure incident response times using alerting and incident-management data.
•
Implement OpenTelemetry instrumentation, traces, logs, and metrics to measure service reliability, service response times, and performance against agreed service-level objectives (SLOs) and service-level agreements (SLAs).
•
Develop and test automation for provisioning and updating environments, virtual machines, and application services.
•
Build and test diagnostic tools, runbooks, and recovery automation with on-call responders so they can restore service efficiently.
•
Improve platform security, reliability, scalability, and operational efficiency.
•
Work with application and data engineers to apply deployment and reliability practices across services and data pipelines.
•
Use AI-assisted development tools such as Codex to build automation, and review and test the resulting code.
Required Skills & Experience
•
Strong practical experience with infrastructure as code, including Terraform.
•
Experience building, testing, and operating CI/CD pipelines for controlled promotion across environments, deployment monitoring, and rollback.
•
Experience operating cloud-native environments.
•
Practical experience applying Site Reliability Engineering (SRE) principles to production systems, including reliability measurement, actionable alerting, and recovery automation.
•
Experience with observability platforms, monitoring, and alerting frameworks.
•
Strong infrastructure automation and scripting skills, including automated testing of infrastructure and deployment changes.
•
Understanding of platform security, governance, and operational controls.
•
Experience diagnosing production failures and building tools that help engineering teams and incident responders resolve them.
•
Ability to understand the purpose and wider context of a requirement, explain assumptions and trade-offs, and judge when to proceed, clarify, or challenge the proposed approach.
Preferred Qualifications
•
Experience with OpenTelemetry.
•
Experience supporting Databricks and AI/ML platforms.
•
Familiarity with Kubernetes, container platforms, and cloud automation services.
•
Experience implementing canary deployments and evaluating telemetry before wider rollout.
•
Knowledge of developer platform engineering practices.
•
Experience using AI coding agents in software development.