Site Reliability Engineering Specialist
bt
Job Description
- Executes the implementation of new software development life cycle automation tools, frameworks, and code pipelines (continuous integration/continuous delivery pipelines whilst executing best practices with a focus on the re-use of application code, demonstrates consistent software delivery practices and produces continuous integration/continuous delivery platform solutions using Amazon Web Services cloud, infrastructure as code (IaC), GitOps, and container technologies
- Coordinates a diverse team and creates the initial test schedule to deliver all aspects of testing to time, budget and quality targets, ensuring producing outlines of solutions and defining depth of testing required
- Executes the implementation of automation technologies to ensure repeatability, eliminating toil, reducing mean time to detection and resolution and repair services
- Proactively identifies and manages risk through regular assessment and diligent execution of controls and mitigations, proactively raising any concerns
- Leads scale testing to measure, tune and optimise system performance
- Executes metric/monitoring analysis that creates stability, security, and performance improvements
- Designs, analyses, develops and troubleshoots highly distributed large-scale production systems spanning on-prem and cloud-based hosting
- Executes approaches that scale systems sustainably through mechanisms like automation and evolves systems by pushing for changes that improve reliability and velocity
- Writes and delivers infrastructure as code software to improve the availability, scalability, latency, and efficiency of services
- Implements robust monitoring and alerting systems and performs root cause analysis and post-mortems with an eye towards future prevention
- Inspects queue and support processing to ensure early warning of support issues
- Executes retrospective and preventive actions after each high severity production incident
- Analyses complex systems from a reliability and resilience perspective and identifies sources of instability in distributed systems
- Champions, continuously develops and shares with team knowledge on emerging trends and changes in site reliability engineering best practices and industry standards
- Mentors other site reliability engineers, helping to improve the team’s abilities by acting as a technical resource
- Uses the network of site reliability engineers, removing BTs organisational boundaries to deliver improvements that are in synergy with initiatives being driven by other SREs.
Skill Required for the Job
- A degree in IT, Maths or Science
- A deep understanding of full stack monitoring solutions such as Dynatrace to ensure current end to end performance and trends of owned CDO Applications
- Strong proficiency in one or more programming languages (e.g. Java, Python).
- Experience with cloud platforms (AWS, Azure, or GCP).
- Solid understanding of software architecture, design patterns, and microservices.
- Familiarity with CI/CD tools and DevOps practices.
- High levels of quality presentation and reporting capabilities to collate output from Managed Service Partners.
- Resilience to ensure support teams are engaged 24x7x365 to support priority incident resolution.
- Ability to adapt to latest industry trends
- CI/CD/CT Pipeline management
- Micro-Service functionality
- Business Process Improvement
- Growth mindset