Software Engineer SRE

netapp

Bengaluru, India 5 Years Exp Posted 27d ago

Job Description

  • 5+ years in SRE, DevOps, Platform Engineering, or Infrastructure Engineering roles
  • Extensive experience with Linux (RHEL/CentOS), including shells, filesystems, kernel tuning, networking, and performance optimization
  • Deep expertise with Kubernetes at scale, including cluster administration, troubleshooting, networking, storage, RBAC, and lifecycle management (on-premises and Rancher Kubernetes)
  • Hands-on experience operating GPU workloads on Kubernetes, including NVIDIA GPU Operator, device plugins, scheduling, and resource management
  • Strong experience managing Confluent Kafka in production, including operations, monitoring, performance tuning, and disaster recovery
  • Experience operating Apache Spark and/or Dremio, including cluster management, job scheduling, scaling, and performance optimization
  • Proficiency in Infrastructure as Code using Terraform, Helm, and GitOps workflows with ArgoCD/FluxCD
  • Proficiency in scripting and automation using Shell, Ansible, and Python, with a strong automation-first mindset
  • Experience with scheduling and orchestration tools such as cron jobs and Apache Airflow
  • Deep familiarity with monitoring and observability tools, including Dynatrace, Grafana, and Prometheus
  • Solid understanding of SQL and NoSQL databases, including operations, backup, and monitoring
  • Experience designing and maintaining CI/CD pipelines and release processes
  • Expertise in AWS cloud platforms and hybrid-cloud integration
  • Strong systems thinking, with an understanding of how infrastructure design choices impact failure modes, scalability, and recovery
  • Strong incident management skills and post-mortem facilitation experience
  • Excellent written communication skills for design documents, runbooks, post-mortems, and operational documentation

Nice to Have

  • Knowledge of Generative AI tools and frameworks, including the application of AI-based predictive analytics and automation in infrastructure operations
  • Familiarity with ML platforms such as Kubeflow, MLflow, and Ray, as well as AI/ML training infrastructure
  • Experience with Kafka Streams, ksqlDB, or Apache Flink

Education

  • 5-8 years of relevant experience.
    • Bachelor of Science Degree in Computer Science, Electrical Engineering, or a related field; a Master’s Degree is preferred. 

Similar Openings for You