Senior ML Systems Engineer - GPU & HPC Optimizations

Screened
Boston
Posted 1 week ago
Apply Now

About the role

Proxima is seeking a Principal ML Performance Engineer to optimize training and inference for state-of-the-art models. You will profile PyTorch, write custom kernels in CUDA/Triton, and leverage compilers like torch.compile, TensorRT, and XLA to maximize throughput on large GPU clusters.

You\'ll scale distributed training across 32–64 nodes on GCP, manage memory scaling for large complexes, and build benchmarks and profiling tools for the research team.

#J-18808-Ljbffr

About this listing

Screened by Joboru

This role passed our automated spam and quality filters and was active in our feed when last checked. Joboru is an aggregator — here is how we screen listings. If anything looks off, tell us.