Compiler Expert Triton Performance Optimization Expert

HUAWEI WARSAW 2026-09-22

Preferred Experience:



  1. Strong expertise or hands-on experience in at least one of the following: CUDA, Triton, PTX, or LLVM IR.

  2. Experience in compiler optimizations, backend code generation, or register allocation.

  3. Extensive experience in developing and optimizing high-performance GPU or NPU kernels.

  4. Familiarity with CUDA or similar parallel programming environments.

  5. Experience in open-source development and collaborative engineering practices.

  6. Strong system design and software engineering skills, with the ability to balance performance, maintainability, and generality in complex systems.

  7. Strong interpersonal skills with the ability to lead and motivate teams effectively.modeling.

  8. Excellent system design and code engineering skills, with the ability to balance performance, maintainability, and generality in complex systems.


Education:



  • Master’s or Ph.D. degree in Computer Architecture, Compiler Design, High Performance Computing, or a related field.

Full time office work in Wola, Warsaw



  • Private healthcare package. We offer premium private healthcare package for our employees.

  • Sport Cards. Our employees can choose from many options within sport subscriptions and sport associations.

  • Benefit Platform. You can choose your benefits on our Benefit Platform e.g.: cinema/theater tickets and discounts, shopping cards and many more.

  • Special discounts for employees. We cooperate with various local companies to offer unique promotions only for our Employees.

  • Office massages. It is a 15-minute chair-based.

,[Lead the development of Triton compiler and kernel performance optimization frameworks, driving high-performance implementations of deep learning operators on NPUs. , Design and implement highly optimized Triton kernels for core operators such as Attention, MatMul, LayerNorm, Convolution, and Softmax. , Improve and extend the Triton compilation pipeline, including optimization passes and code generation quality. , Optimize memory access patterns and parallel scheduling strategies, with deep understanding of performance factors such as memory footprint and instruction scheduling. , Collaborate closely with NPU architecture and performance teams to co-design performance-critical features. , Track and evaluate state-of-the-art AI compiler technologies (e.g., MLIR, Apache TVM, Hidet, CUTLASS), and introduce innovative designs to improve overall system performance. , Build and lead a Triton compiler optimization team; communicate results with cross-functional teams and leadership; contribute to open-source ecosystems such as Triton and LLVM. ] Requirements: Degree, CUDA, GPU, Triton, LLVM IR Additionally: Private healthcare, Sport subscription, Benefit Platform, Benefit Platform..