TPU Performance Engineer - vLLM Inference Backends

Screened
Penarth, Wales
Posted 1 week ago
Apply Now

About the role

Inferact Singapore PTE. LTD. is seeking a TPU performance engineer to make vLLM a first-class inference engine on Google TPUs.

You will build and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling so vLLM can deliver frontier inference performance on TPU hardware. You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving production-relevant model serving on TPU

#J-18808-Ljbffr

About this listing

Screened by Joboru

This role passed our automated spam and quality filters and was active in our feed when last checked. Joboru is an aggregator — here is how we screen listings. If anything looks off, tell us.