Staff AMD GPU Performance Engineer, Inference Acceleration

Screened
Penarth, Wales
Posted 1 week ago
Apply Now

About the role

Inferact is seeking an AMD GPU performance engineer to advance vLLM as a premier inference engine on AMD accelerators. You will build and optimize AMD GPU backends, kernels, and benchmarking infrastructure using ROCm, HIP, Triton, CK, and AITER.

You will work at the boundary of inference systems, kernels, compilers, and hardware, improving attention, GEMM, sampling, KV cache, and other communication-heavy paths to deliver fast, scalable inference and maintainable backend.

#J-18808-Ljbffr

About this listing

Screened by Joboru

This role passed our automated spam and quality filters and was active in our feed when last checked. Joboru is an aggregator — here is how we screen listings. If anything looks off, tell us.