Lead AI/ML Software Use Cases Validation Engineer
amd
Job Description
- Lead validation and quality ownership of AI/ML compute stacks on Ubuntu and Yocto
- Define validation strategy, test architecture, and coverage across functional, performance, stress, regression, and scalability testing
- Own and drive the defect lifecycle, including triage, root cause analysis, and closure
- Validate end‑to‑end AI pipelines, including:
- Model training, conversion, and optimization (e.g., PyTorch → ONNX)
- Kernel execution, memory transfers, and inference accuracy
- Define and execute AI benchmarking and profiling strategies for training and inference workloads
- Analyze compute, memory, and latency bottlenecks and drive system‑level and model‑level optimizations
- Validate AI frameworks and runtimes (PyTorch, TensorFlow, ONNX Runtime)
- Execute and optimize workloads on ROCm/HIP, CUDA, OpenCL, and heterogeneous accelerators
- Design and own Python‑based automation frameworks for validation, benchmarking, and reporting
- Drive improvements in validation scalability, efficiency, and performance coverage
- Collaborate closely with compiler, runtime, driver, and hardware teams
- Provide technical leadership and mentorship to senior and junior engineers
- Communicate validation status, performance metrics, and quality risks to stakeholders
Required Skills & Qualifications
Technical
- 8–12 years of experience in AI/ML software validation or performance engineering
- Strong expertise in Python scripting and test automation
- Strong ML fundamentals including deep learning and LLMs
- Hands‑on experience with ROCm validation, performance profiling, and optimization
- Experience with HIP, CUDA, OpenCL, and TensorFlow/PyTorch integrations
- Proven experience validating end‑to‑end AI pipelines
- Strong Linux expertise (Ubuntu, Yocto)