About the role
Proxima is seeking a Principal ML Performance Engineer to optimize training and inference for state-of-the-art models. You will profile PyTorch, write custom kernels in CUDA/Triton, and leverage compilers like torch.compile, TensorRT, and XLA to maximize throughput on large GPU clusters.
You\'ll scale distributed training across 32–64 nodes on GCP, manage memory scaling for large complexes, and build benchmarks and profiling tools for the research team.
#J-18808-LjbffrAbout this listing
This role passed our automated spam and quality filters and was active in our feed when last checked. Joboru is an aggregator — here is how we screen listings. If anything looks off, tell us.
Similar jobs you may like
Plant Equipment Trainer
1 day agoTalent Finder
Software Engineering Manager
1 day agoHalian Technology Limited
OpenShift Engineer
1 day agoTeksystems
Senior Systems Engineer
1 day agoEclectic Recruitment Ltd
Senior Safety Engineer
1 day agoMeridian Business Support
SRE Engineer
1 day agoTeksystems
CMM Programmer / Inspector
1 day agoProdrive
Activities Coordinator
1 day agoCare UK
CMM Programmer (Composites)
1 day agoThe Collective Network