Location: Birmingham
Rate: Up to £575/day – Inside IR35
Contract opportunity – Initially 6 months
Hybrid working: 2 days on-site per week
We’re recruiting a Lead Site Reliability & Observability Engineer to lead the implementation of Datadog across Azure and Cloudflare.
You’ll build synthetic monitoring and automated production validation for critical APIs, integrations and customer journeys, identifying issues within minutes of a software release.
The role
Own and evolve the Datadog observability platform.
Design synthetic monitoring for critical API and browser workflows.
Integrate monitoring, testing and release validation into Azure DevOps and GitHub pipelines.
Develop monitoring-as-code and testing-as-code using Terraform.
Create actionable dashboards, SLOs, SLIs, alerts and anomaly detection.
Drive reliability improvements, performance investigations and root-cause analysis.
What we’re looking for
Strong hands-on Datadog experience across Synthetic Monitoring, APM, RUM, Log Management, SLOs and alerting.
Deep Azure experience and experience integrating Cloudflare services.
Experience operating large-scale production environments.
Strong CI/CD experience with Azure DevOps and GitHub.
API, integration and browser-based testing expertise.
Terraform experience and a strong understanding of distributed systems and microservices.
A background in SRE or Platform Engineering leadership.
Datadog or Azure certifications and Cloudflare administration experience would be advantageous.
If you are interested in this role and would like to hear more, please apply for the opportunity with an updated CV and contact information
Salary description
£500.00 - £575.00 per day
