Senior Site Reliability Engineer
cisco
Job Description
Your Impact
- Own and drive AWS cost optimization and efficiency initiatives across ThousandEyes infrastructure.
- Analyze cloud usage across compute, storage, databases, observability, networking, and data platform workloads to identify savings opportunities.
- Partner with engineering teams to improve application and infrastructure performance while reducing cloud wastage.
- Drive FinOps practices including cost visibility, tagging hygiene, showback/chargeback, budgeting, forecasting, anomaly detection, and cost allocation.
- Improve infrastructure efficiency through better resource utilization, autoscaling, right-sizing, and capacity planning.
- Build automation, dashboards, reports, and guardrails to improve cost governance and operational visibility.
- Support ThousandEyes cost reviews, OKR tracking, leadership updates, and stakeholder communications.
- Collaborate with product, finance, engineering, and platform teams to align cost optimization with business priorities.
- Identify and reduce underutilized infrastructure, idle resources, over-provisioned workloads, and inefficient service usage.
- Improve observability, alerting, and monitoring for cost, performance, and platform health indicators.
- Influence engineering teams to adopt cost-aware architecture, reliable design patterns, and performance-efficient implementation practices.
- Provide senior-level technical leadership, mentor engineers, and independently drive cross-functional initiatives to closure.
- Balance reliability, scalability, performance, and cost efficiency in all platform decisions.
Minimum Qualifications
- Bachelor’s degree or higher in Engineering, Computer Science, or equivalent practical experience.
- 8–12 years of relevant experience in Site Reliability Engineering, Cloud Infrastructure, Platform Engineering, DevOps, Production Engineering, or Performance Engineering.
- Strong experience with AWS services such as EC2, S3, RDS, EMR, Lambda, CloudWatch, OpenSearch, ElastiCache, IAM, VPC, and related cloud-native services.
- Practical experience in cloud cost optimization, FinOps, AWS billing analysis, cost allocation, tagging, budget tracking, forecasting, and cost governance.
- Experience with performance analysis, capacity planning, infrastructure optimization, and reliability improvements for large-scale cloud platforms.
- Experience with infrastructure-as-code and automation tools such as Terraform, CloudFormation, Puppet, Ansible, or similar.
- Strong scripting or programming skills in Python, Go, Shell, or similar languages for automation, reporting, and operational tooling.
- Experience with observability and monitoring platforms such as ThousandEyes, CloudWatch, Prometheus, Grafana, Splunk, OpenSearch, Datadog, or similar.
- Strong Linux systems knowledge and understanding of distributed systems.
- Experience in incident management, production support, reliability engineering, and operational excellence.
- Ability to analyze large-scale infrastructure, performance, and cost data and convert findings into actionable engineering recommendations.
- Strong communication skills with the ability to present technical, performance, and cost insights to engineering teams, finance, leadership, and multi-functional stakeholders.
- Experience working in Agile/Scrum environments and managing priorities across multiple teams.