Assoc, Production Eng, WRB Tech

standardchartered

Chennai, Tamil Nadu, India 10 Years Exp Posted 4d ago

Job Description

•    Ensure maximum service quality and production stability through rapid and effective response to technical incidents, while driving continual service improvement through trend analysis, problem management, and proactive identification of improvement opportunities.
•    Manage technical recovery and service restoration for High Severity Incidents (CERT/MIM), service outages, and medium/high severity incidents, providing end-to-end support and implementing timely resolutions within agreed SLAs.
•    Lead and coordinate incident management activities, including stakeholder communications, escalation management, and technical bridge facilitation during critical incidents.
•    Perform root cause analysis (RCA) for High Severity and recurring incidents, ensuring corrective and preventive actions are identified, tracked, and implemented to closure.
•    Own the operational stability, availability, and performance of production systems, directing second- and third-level support teams for problem diagnosis and resolution in accordance with agreed SLAs and OLAs.
•    Manage production changes, releases, deployments, and rollouts with zero or minimal impact to business services. Ensure comprehensive implementation, validation, rollback, contingency, and communication plans are in place for all production activities.
•    Review and assess the impact of dependent changes across applications, infrastructure, databases, middleware, cloud platforms, and networks to minimize production risk.
•    Drive proactive monitoring and event management by identifying opportunities for automation, alert optimization, early issue detection, and operational efficiency improvements.

 

•    Support capacity management, resiliency testing, disaster recovery (DR), and business continuity planning (BCP) activities to ensure operational readiness.
•    Create, maintain, and continuously improve Production Engineering documentation, operational procedures, runbooks, recovery guides, knowledge articles, and contingency plans.
•    Ensure adherence to operational governance, security, compliance, audit, and risk management requirements across supported services.
•    Provide inputs to the PE Manager for operational dashboards and service reviews, including incident trends, problem trends, availability metrics, service improvement plans (SIPs), RCA action tracking, and platform health indicators.
•    Collaborate with application development, infrastructure, security, architecture, and business teams to improve reliability, reduce technical debt, and enhance service resilience.
•    Participate in and support cross-training, knowledge transfer, mentoring, and capability-building activities within the Production Engineering organization.
•    Identify and drive opportunities for automation, shift-left initiatives, operational simplification, and reduction of manual effort to improve service reliability and support efficiency.
•    Act as a technical lead during major incidents, complex problem investigations, production go-lives, and critical business events, providing technical guidance and decision-making support to internal and external stakeholders.

Similar Openings for You