Site Reliability Engineer – Windows

ubs

pune NM Years Exp Posted 34d ago

Job Description

Design, implement, and maintain highly available and fault-tolerant systems in a financial environment.

· Define and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs) to ensure system reliability and customer satisfaction.

· Passionately identify, measure, and reduce TOIL, with a proactive approach to eliminating repetitive manual tasks through automation.

· Lead incident response, post-mortems, and root cause analysis for production issues.

· Collaborate with development teams to embed reliability into the software development lifecycle.

· Integrate with observability platforms (e.g., Prometheus, Grafana, ELK, Datadog) to ensure end-to-end visibility of systems and services.


What We’re Looking For


We’re seeking a seasoned Windows SRE with a strong foundation in reliability operations and a passion for building robust, scalable systems. The ideal candidate will bring a mix of technical depth, operational maturity, and a proactive mindset.


✅ Essential Experience & Skills

· Proven expertise in Site Reliability Engineering, with a background in software engineering, infrastructure, or operations.

· Hands-on experience with cloud platforms (e.g. Azure Devops), operating systems (e.g. Windows 2019+), and networking fundamentals.

· Solid understanding of networking and storage technologies (e.g. NFS, SAN, NAS).

· Strong working knowledge of authentication and naming services (e.g. DNS, LDAP, Kerberos, Centrify).

· Proficiency in scripting and automation (e.g., Python, PowerShell ).

· Azure ADO pipeline and Azure control plane.

· Practical experience with infrastructure as code tools (e.g., Terraform, GitLab).

· Demonstrated ability to define and manage SLIs, SLOs, SLAs, and to systematically reduce TOIL.

· Ability to integrate with observability platforms to ensure system visibility.

· Microsoft Failover clustering and DFS architecture .

· A metrics- and automation-driven mindset, with a strong focus on measurable reliability.

· Calm under pressure, especially during incidents and outages, with a structured approach to incident response and post-mortems.

· Strong collaboration and communication skills, with the ability to work across engineering and business teams.

· A proactive, ownership-driven attitude, always seeking opportunities to improve systems and processes.

✨ Desirable Additions

· Experience with chaos engineering, resilience testing, or disaster recovery planning.

· Familiarity with financial transaction systems, real-time data pipelines, or core banking platforms.

· An understanding of CI/CD pipelines, containerization (AKS), and orchestration (Kubernetes)

· You’re curious to explore how AI can improve how we build, deliver, and optimize workflows. You do this with sound judgment – validating outputs and aligning with policies, risk standards, and ethical use.

Similar Openings for You