Site Reliability Engineer – Windows
ubs
Job Description
Design, implement, and maintain highly available and fault-tolerant systems in a financial environment.
· Define and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs) to ensure system reliability and customer satisfaction.
· Passionately identify, measure, and reduce TOIL, with a proactive approach to eliminating repetitive manual tasks through automation.
· Lead incident response, post-mortems, and root cause analysis for production issues.
· Collaborate with development teams to embed reliability into the software development lifecycle.
· Integrate with observability platforms (e.g., Prometheus, Grafana, ELK, Datadog) to ensure end-to-end visibility of systems and services.
What We’re Looking For
We’re seeking a seasoned Windows SRE with a strong foundation in reliability operations and a passion for building robust, scalable systems. The ideal candidate will bring a mix of technical depth, operational maturity, and a proactive mindset.
✅ Essential Experience & Skills
· Proven expertise in Site Reliability Engineering, with a background in software engineering, infrastructure, or operations.
· Hands-on experience with cloud platforms (e.g. Azure Devops), operating systems (e.g. Windows 2019+), and networking fundamentals.
· Solid understanding of networking and storage technologies (e.g. NFS, SAN, NAS).
· Strong working knowledge of authentication and naming services (e.g. DNS, LDAP, Kerberos, Centrify).
· Proficiency in scripting and automation (e.g., Python, PowerShell ).
· Azure ADO pipeline and Azure control plane.
· Practical experience with infrastructure as code tools (e.g., Terraform, GitLab).
· Demonstrated ability to define and manage SLIs, SLOs, SLAs, and to systematically reduce TOIL.
· Ability to integrate with observability platforms to ensure system visibility.
· Microsoft Failover clustering and DFS architecture .
· A metrics- and automation-driven mindset, with a strong focus on measurable reliability.
· Calm under pressure, especially during incidents and outages, with a structured approach to incident response and post-mortems.
· Strong collaboration and communication skills, with the ability to work across engineering and business teams.
· A proactive, ownership-driven attitude, always seeking opportunities to improve systems and processes.
✨ Desirable Additions
· Experience with chaos engineering, resilience testing, or disaster recovery planning.
· Familiarity with financial transaction systems, real-time data pipelines, or core banking platforms.
· An understanding of CI/CD pipelines, containerization (AKS), and orchestration (Kubernetes)
· You’re curious to explore how AI can improve how we build, deliver, and optimize workflows. You do this with sound judgment – validating outputs and aligning with policies, risk standards, and ethical use.