Site Reliability Engineering Specialist
bt
Job Description
Provide end-to-end SRE ownership for the Global Fabric service, ensuring platform reliability, performance, resilience, and operational excellence.
• Own incidents across the customer journey, from detection through to resolution, working with the appropriate ASGs and support teams.
• Act as the primary operational escalation point for CF-related incidents participating in an on-call rota.
• Escalate complex technical issues appropriately and coordinate resolution activities
• Manage incidents through ServiceNow and track defects and improvements through Jira.
• Perform root cause analysis and implement preventative actions to prevent recurrence along with providing documentation and knowledge transfer sessions.
• Design, implement, operate, and continuously improve observability and monitoring solutions using Dynatrace.
• Drive automation initiatives using Ansible, scripting, CI/CD pipelines, and GitOps practices to reduce operational toil and improve service reliability.
• Define, measure, and report service health metrics, SLIs, SLOs, error budgets, and operational performance dashboards.
• Fulfil operational service requests in line with agreed SLAs.
• Support onboarding of new customers and operational readiness activities.
• Analyse platform, network, and service trends to identify reliability and optimisation opportunities.