Site Reliability / DevOps Engineer

siemens-energy

Remote NM Years Exp Posted 14d ago

Job Description

  • Support daily operations and incident handling in close collaboration with the operations team
  • Improve troubleshooting workflows and reduce time to detect and resolve issues
  • Manage deployments and patching across cloud environments
  • Enhance monitoring, alerting, and early warning systems using OpenSearch and OpenTelemetry to improve SLOs
  • Create deployment and incident reports
  • Drive improvements in platform stability, release efficiency, and operational maturity
  • Support and optimize Kubernetes-based platforms and containerized workloads
  • Collaborate with DevOps engineers to implement process improvements and platform enhancements

What You Bring

  • Experience in Site Reliability Engineering / DevOps / Cloud Operations
  • Strong knowledge of: Kubernetes and container platforms, Virtualization technologies, CI/CD pipelines (GitLab), Terraform (Infrastructure as Code), Apache Kafka
  • Experience with monitoring and observability (OpenSearch, OpenTelemetry)
  • Solid understanding of incident management and cloud operations best practices
  • Experience in large-scale or customer-facing cloud environments
  • Strong analytical and problem-solving skills
  • Ability to work in cross-functional teams

 

Similar Openings for You