AI Platform Engineer

ebayinc

Bangalore 5 Years Exp Posted 1d ago

Job Description

  • Design and build Kubernetes Custom Resource Definitions (CRDs) and operators for ML workloads, GPU node pools, and RayService CRDs.
  • Architect and operate Kubernetes networking layers including Gateway API, Service Load Balancers, and API gateways for hybrid cloud connectivity.
  • Design, deploy, and operate multi-NIC Kubernetes clusters with RDMA-enabled networking.
  • Implement and optimize RDMA networking using GPUDirect RDMA, RoCE, and InfiniBand for distributed GPU workloads.
  • Build and operate a custom AI-aware GPU scheduler with topology-aware placement, preemption, and GPU defragmentation.
  • Design and manage GPU pool management systems across on-premise and cloud-burst GPU environments.
  • Deploy and operate KubeRay infrastructure for distributed Ray clusters supporting training and inference workloads.
  • Implement cloud bursting and spot or preemptible GPU scheduling to improve utilization.
  • Automate GPU asset provisioning, node configuration, and cluster lifecycle management using Infrastructure-as-Code and GitOps.
  • Implement networking policies, multi-tenant isolation, RBAC, and security controls across the Kubernetes environment.
  • Build observability for GPU utilization, NCCL communication, scheduler decisions, and network throughput.
  • Collaborate with ML Platform, AI Research, and Networking teams to optimize infrastructure for training and online inference.

     

What We’re Looking For

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
  • 5+ years of experience building distributed systems or infrastructure platforms with deep Kubernetes expertise.
  • Strong programming skills in Go and/or Python.
  • Familiarity with Kubernetes controller development frameworks such as Kubebuilder, Operator SDK, or controller-runtime.
  • Deep understanding of Kubernetes internals including the API server, etcd, scheduler, controller manager, and kubelet.
  • Hands-on experience designing Kubernetes networking including Gateway API, CNI plugins, service load balancing, and hybrid cloud architectures.
  • Experience designing and operating multi-NIC Kubernetes clusters using NVIDIA Network Operator, SR-IOV device plugins, or equivalent tooling.
  • Strong understanding of RDMA networking protocols and NCCL configuration for distributed GPU workloads.
  • Experience with Kubernetes GPU scheduler frameworks and GPU pool management, including MIG partitioning and preemption policies.
  • Hands-on experience deploying and operating KubeRay for distributed Ray workloads.
  • Experience with GPU asset lifecycle management, bare-metal provisioning automation, and GitOps-based CD tooling such as Argo CD.
  • Familiarity with GPU technologies including NVIDIA CUDA, NVLink, NVSwitch, device plugins, and the GPU Operator ecosystem.
  • Experience with observability tooling such as Prometheus, Grafana, and OpenTelemetry.
    • Strong debugging and performance optimization skills across GPU driver stacks, RDMA networking, and distributed Kubernetes infrastructure.

Similar Openings for You