Lead Site Reliability Engineer/ Expert
sita
Job Description
-
Analyze production incidents using logs, metrics, and traces to identify impacted application code and execution paths.
-
Diagnose system issues and determine whether the root cause is related to application code, configuration, Kubernetes, or infrastructure.
-
Troubleshoot Kubernetes workloads, including runtime behavior, networking, health probes, and failure scenarios.
-
Serve as the technical escalation point during critical incidents, providing clear and timely guidance.
-
Lead root cause analysis (RCA) and drive permanent corrective actions to improve reliability.
-
Enhance observability and alerting to improve issue detection and resolution.
-
Partner with Development and Platform teams to resolve systemic issues and strengthen service reliability.
-
Automate repetitive operational tasks and promote engineering best practices.
-
Assess the impact of deployments on production environments through CI/CD pipeline expertise.
-
Monitor system performance and reliability, identifying opportunities to improve resilience and reduce outages.
-