GPU Software Development Engineer at Advanced Micro Devices Inc

Oregon, Oregon, USA -

Full Time

Start Date

Immediate

Expiry Date

24 Aug, 25

Salary

0.0

Posted On

25 May, 25

Experience

0 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

Skills

Python, Parallel Computing, Machine Learning, Cuda

Industry

Computer Software/Engineering

Description

REQUIRED EXPERIENCE:

2+ years of experience in GPU kernel development for machine learning (ROCm or CUDA).
Proficiency in C/C++ and Python, with experience in performance-critical programming.
Strong understanding of ML frameworks (PyTorch, TensorFlow) and GPU-accelerated libraries.
Basic knowledge of modern AI technologies (LLMs, transformers, inference optimization).
Familiarity with parallel computing, memory optimization, and hardware architectures.
Problem-solving skills and ability to work in a fast-paced environment.

PREFERRED EXPERIENCE:

Direct experience with AMD ROCm development (HIP, MIOpen, Composable Kernel).
Knowledge of LLM-specific optimizations (e.g., FlashAttention, PagedAttention in vLLM).
Experience with distributed training/inference or model compression techniques.
Contributions to open-source ML projects or GPU compute libraries.

Responsibilities

WHAT YOU DO AT AMD CHANGES EVERYTHING

We care deeply about transforming lives with AMD technology to enrich our industry, our communities, and the world. Our mission is to build great products that accelerate next-generation computing experiences – the building blocks for the data center, artificial intelligence, PCs, gaming and embedded. Underpinning our mission is the AMD culture. We push the limits of innovation to solve the world’s most important challenges. We strive for execution excellence while being direct, humble, collaborative, and inclusive of diverse perspectives.
AMD together we advance_
Responsibilities:

THE ROLE:

We are seeking a talented Machine Learning Kernel Developer to design, develop, and optimize low-level machine learning kernels for AMD GPUs using the ROCm software stack. In this role, you will work on high-impact projects to accelerate AI frameworks and libraries, with a focus on emerging technologies like Large Language Models (LLMs) and other generative AI workloads.

KEY RESPONSIBILITIES:

Design and implement highly optimized ML kernels (e.g., matrix operations, attention mechanisms) for AMD GPUs using ROCm.
Profile, debug, and tune kernel performance to maximize hardware utilization for AI workloads.
Collaborate with ML researchers and framework developers to integrate kernels into AI frameworks (e.g., PyTorch, TensorFlow) and inference engines (e.g., vLLM).
Contribute to the ROCm software stack by identifying and resolving bottlenecks in libraries like MIOpen, HIP, or Composable Kernel.
Stay updated on the latest AI/ML trends (LLMs, quantization, distributed inference) and apply them to kernel development.
Document and communicate technical designs, benchmarks, and best practices.
Troubleshoot and resolve issues related to GPU compatibility, performance, and scalability.