AI Platform Engineer
ebayinc
Job Description
- Design and build Kubernetes Custom Resource Definitions (CRDs) and operators for ML workloads, GPU node pools, and RayService CRDs.
- Architect and operate Kubernetes networking layers including Gateway API, Service Load Balancers, and API gateways for hybrid cloud connectivity.
- Design, deploy, and operate multi-NIC Kubernetes clusters with RDMA-enabled networking.
- Implement and optimize RDMA networking using GPUDirect RDMA, RoCE, and InfiniBand for distributed GPU workloads.
- Build and operate a custom AI-aware GPU scheduler with topology-aware placement, preemption, and GPU defragmentation.
- Design and manage GPU pool management systems across on-premise and cloud-burst GPU environments.
- Deploy and operate KubeRay infrastructure for distributed Ray clusters supporting training and inference workloads.
- Implement cloud bursting and spot or preemptible GPU scheduling to improve utilization.
- Automate GPU asset provisioning, node configuration, and cluster lifecycle management using Infrastructure-as-Code and GitOps.
- Implement networking policies, multi-tenant isolation, RBAC, and security controls across the Kubernetes environment.
- Build observability for GPU utilization, NCCL communication, scheduler decisions, and network throughput.
- Collaborate with ML Platform, AI Research, and Networking teams to optimize infrastructure for training and online inference.
What We’re Looking For
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
- 5+ years of experience building distributed systems or infrastructure platforms with deep Kubernetes expertise.
- Strong programming skills in Go and/or Python.
- Familiarity with Kubernetes controller development frameworks such as Kubebuilder, Operator SDK, or controller-runtime.
- Deep understanding of Kubernetes internals including the API server, etcd, scheduler, controller manager, and kubelet.
- Hands-on experience designing Kubernetes networking including Gateway API, CNI plugins, service load balancing, and hybrid cloud architectures.
- Experience designing and operating multi-NIC Kubernetes clusters using NVIDIA Network Operator, SR-IOV device plugins, or equivalent tooling.
- Strong understanding of RDMA networking protocols and NCCL configuration for distributed GPU workloads.
- Experience with Kubernetes GPU scheduler frameworks and GPU pool management, including MIG partitioning and preemption policies.
- Hands-on experience deploying and operating KubeRay for distributed Ray workloads.
- Experience with GPU asset lifecycle management, bare-metal provisioning automation, and GitOps-based CD tooling such as Argo CD.
- Familiarity with GPU technologies including NVIDIA CUDA, NVLink, NVSwitch, device plugins, and the GPU Operator ecosystem.
- Experience with observability tooling such as Prometheus, Grafana, and OpenTelemetry.
- Strong debugging and performance optimization skills across GPU driver stacks, RDMA networking, and distributed Kubernetes infrastructure.