AI Platform

ebayinc

Bengaluru 5 Years Exp Posted 3d ago

Job Description

  • Own production inference: Take models from handoff to production-grade serving, including release engineering, capacity planning, cost optimization, and incident response.

  • Tune inference performance: reduce end-to-end latency and increase throughput across real production traffic patterns.

  • Optimize runtimes and servers: Scale inference across heterogeneous GPU fleets; optimize stacks such as vLLM, Triton, and related components (e.g., schedulers, KV cache, batching, memory).

  • Benchmark and measure: Build benchmarking suites, metrics, and tooling to quantify latency, throughput, GPU utilization, memory, and cost.

  • Reliability and observability: Improve monitoring, tracing, and alerting; participate in incident response and postmortems to harden systems.

  • Apply and ship new optimizations: Evaluate research and implement pragmatic inference optimizations (e.g., quantization, paging, kernel/runtimes improvements).

    • Partner with cross-functional teams: Work with data science and product teams to translate business requirements into performance and availability SLOs.

Similar Openings for You