Lead Site Reliability Engineer/ Expert

sita

Bengaluru 8 Years Exp Posted 12d ago

Job Description

  • Analyze production incidents using logs, metrics, and traces to identify impacted application code and execution paths.

  • Diagnose system issues and determine whether the root cause is related to application code, configuration, Kubernetes, or infrastructure.

  • Troubleshoot Kubernetes workloads, including runtime behavior, networking, health probes, and failure scenarios.

  • Serve as the technical escalation point during critical incidents, providing clear and timely guidance.

  • Lead root cause analysis (RCA) and drive permanent corrective actions to improve reliability.

  • Enhance observability and alerting to improve issue detection and resolution.

  • Partner with Development and Platform teams to resolve systemic issues and strengthen service reliability.

  • Automate repetitive operational tasks and promote engineering best practices.

  • Assess the impact of deployments on production environments through CI/CD pipeline expertise.

    • Monitor system performance and reliability, identifying opportunities to improve resilience and reduce outages.

Similar Openings for You