Lead Site Reliability Engineer
fisglobal
Job Description
- Defines meaningful SLIs, SLOs for owned services; tracks error budgets and burn rates.
- Advanced scripting, CI/CD, container orchestration (Docker/Kubernetes)
- Forecasts capacity for owned services; plans for growth
- Pattern Analysis and Correlation skillset
What you bring:
Must Have
- Strong Knowledge in Performance Monitoring Tools like Dynatrace, Splunk and ability to create Dashboards, Views and Alerts
- Knowledge in OS, Network, Middleware, Database, SSL, Load Balancer
- Must have knowledge in scripting language (Unix Shell, Windows Scripting, Python, Java, .NET…)
- Hands on with at least one of (Unix/Windows Scripting, Python, Java, C++, C#)
- Ability to Automate repetitive tasks - (Scripting, RPA, Power Automate, UiPath etc.)
- Ability to implement AI models in Problem Solving and reduce MTTR
- Experience with Virtualization / Containerization in a production environment (e.g. Kubernetes, etc.);
- Engineering Mindset with End to End view and provide solutions
- Passion for problem solving with strong analytical capabilities