AI Platform
ebayinc
Job Description
-
Own production inference: Take models from handoff to production-grade serving, including release engineering, capacity planning, cost optimization, and incident response.
-
Tune inference performance: reduce end-to-end latency and increase throughput across real production traffic patterns.
-
Optimize runtimes and servers: Scale inference across heterogeneous GPU fleets; optimize stacks such as vLLM, Triton, and related components (e.g., schedulers, KV cache, batching, memory).
-
Benchmark and measure: Build benchmarking suites, metrics, and tooling to quantify latency, throughput, GPU utilization, memory, and cost.
-
Reliability and observability: Improve monitoring, tracing, and alerting; participate in incident response and postmortems to harden systems.
-
Apply and ship new optimizations: Evaluate research and implement pragmatic inference optimizations (e.g., quantization, paging, kernel/runtimes improvements).
-
Partner with cross-functional teams: Work with data science and product teams to translate business requirements into performance and availability SLOs.
-