Location: Hybrid at Newcastle (3 days/week)OverviewWe are seeking an experienced L3 Live Support Engineer to provide advanced technical support and operational stability across critical production systems. The successful candidate will act as the highest level of technical escalation, owning and driving the investigation, diagnosis, and resolution of complex incidents while ensuring minimal business impact. This role requires strong analytical skills, excellent stakeholder management, and the ability to collaborate across multiple technology and business teams.
Key ResponsibilitiesIncident Management & Resolution
Own and manage the investigation and resolution of live production incidents and service requests.
Act as the primary technical escalation point for complex issues impacting business-critical services.
Execute technical recovery actions during incidents to restore service as quickly as possible.
Participate in Major Incident and High-Priority Incident response activities.Technical Troubleshooting
Diagnose and resolve application, infrastructure, integration, database, and data-related issues across supported products and platforms.
Analyse logs, system metrics, monitoring tools, and application behaviour to identify root causes and corrective actions.
Maintain a deep understanding of product architecture, business processes, and operational dependencies.Problem Management & Continuous Improvement
Conduct root cause analysis (RCA) and contribute to Problem Management processes.
Participate in post-incident reviews and recommend long-term preventative solutions.
Drive continuous service improvement through automation, process optimisation, tooling enhancements, and reduction of recurring incidents.
Monitor service health and proactively identify risks, trends, and emerging issues.
Develop and maintain operational documentation, runbooks, knowledge articles, and support procedures.
Ensure monitoring, alerting, and support processes remain effective and fit for purpose.
Work closely with Delivery, Development, DevOps, Infrastructure, Enterprise Service Management, Operational Teams, and third-party suppliers to resolve incidents effectively.
Communicate clearly with technical and non-technical stakeholders during incident resolution.
Support service reviews and contribute to operational governance activities.On-Call Support
Participate in High-Priority Incident response and on-call/out-of-hours support rotas where required.Technical Skills
Strong experience supporting enterprise applications in a production environment.
Proven expertise in incident management, problem management, and root cause analysis.
Experience with application support, infrastructure troubleshooting, and system integrations.
Knowledge of SQL and database troubleshooting.
Experience with monitoring and observability tools such as Splunk, Dynatrace, AppDynamics, Grafana, Azure Monitor, or equivalent.
Understanding of cloud platforms (Azure, AWS, or GCP).
Familiarity with APIs, middleware, and integration technologies.
Knowledge of ITIL principles and service management best practices.
Excellent problem-solving and analytical abilities.
Strong communication and stakeholder management skills.
Ability to perform under pressure during critical incidents.
Strong collaboration and teamwork mindset.
Ability to prioritise workload effectively in a fast-paced environment.
ITIL Foundation Certification.
Experience working within Agile and DevOps environments.
Exposure to automation and scripting technologies (PowerShell, Python, Bash, etc.).
Experience supporting cloud-native applications and microservices architectures
Salary description
£37000.00 - £55000.00 per year
