GPU AI Performance Engineer
qualcomm
Job Description
-
Analyze and evaluate GPU architecture/microarchitecture for performance optimizations of AI workloads
-
Work across the AI stack, from model graphs and inference runtimes down to GPU kernels and compiler IR
-
Collaborate with hardware, software, and ML teams to identify performance bottlenecks and propose architectural or algorithmic optimizations
-
Analyze AI workload characteristics and correlate their behavior across different GPU generations
-
Contribute to architectural trade-off studies and influence GPU roadmap decisions with data-driven insights
Preferred Skills:
-
Strong understanding of CPU/GPU architecture
-
Experience in Python, C++, and ML frameworks (e.g., TensorFlow, PyTorch)
-
Skills: C/C++ Programming Language, Scripting (Python/Perl), Assembly, Verilog/SystemVerilog
-
Familiarity with AI inference runtimes and deployment stacks for model compilation, optimization, and execution on GPU/accelerator platforms
-
Good understanding of common neural network layers and operations, including what they do and how they affect model behavior and performance
Nice to have:
-
Experience with LLM inference engines such as llama.cpp
-
Exposure to Triton, TTIR, TTGIR, or GPU compiler pipelines
-
Knowledge of quantization formats and tradeoffs
-
Understanding of MoE architectures
-
Experience with GPU driver and compiler development
-
Experience with OpenCL or Cuda development
Minimum Qualifications:
• Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 6+ years of Systems Engineering or related work experience.
OR
Master's degree in Engineering, Information Systems, Computer Science, or related field and 5+ years of Systems Engineering or related work experience.
OR
PhD in Engineering, Information Systems, Computer Science, or related field and 4+ years of Systems Engineering or related work experience.