Senior / Staff ML Training Optimization Engineer

Waabi · Remote US & Canada

ExclusiveRemoteFull-timeSeniorPublished May 8, 2026

Apply directly on Waabi’s careers site — no account needed.

About the role

Waabi, founded by AI visionary Raquel Urtasun, is the leader in Physical AI. With a world-class team, we're unlocking the next era of autonomous transportation with technology that's powering commercial autonomous trucks and robotaxis. Waabi is backed by and partners with world leaders in AI, automotive, logistics, and deep tech.

With offices in Toronto, San Francisco, Dallas, and Pittsburgh, Waabi is growing quickly and looking for diverse, innovative and collaborative candidates who want to impact the world in a positive way. To learn more visit: www.waabi.ai

You will...
- Build standardized distributed training frameworks for research and production, drive our training towards new levels of stability and efficiency.
- Comprehensively profile model runtime and memory to pinpoint performance bottlenecks.
- Identify and evaluate emerging technologies that can be adopted into Waabi’s training and inference frameworks. Examples include designing new CUDA kernels, quantization-aware training and inference, and compilation/deployment techniques.
- Work with researchers and ML engineers on best-practices for optimal resource usage.
- Create and improve tooling and dashboards to ensure broad adoption of your work.
Qualifications:
- MS/PhD or Bachelors degree with a minimum of 4 years of industry experience in Computer Science, Robotics and/or similar technical field(s) of study.
- Solid coding proficiency in a variety of coding languages including Python, C++ or Rust.
- Experience in deep learning frameworks such as PyTorch or Jax.
- Skilled in profiling CPU and GPU code using tools such as PyTorch Profiler and NVIDIA Nsight.
- Open-minded and collaborative team player with willingness to help others.
- Passionate about self-driving technologies, solving hard problems, and creating innovative solutions.
Bonus/nice to have:
- Experience in identifying when custom CUDA kernels are needed, and implementing them.
- Experience in Bazel in a monorepo environment, and integrating third party packages into dev environments.
- Experience with Kubernetes-based training platforms.

Skills

  • Python
  • C++
  • Rust
  • PyTorch
  • Kubernetes
  • Docker

Never be applicant #200 again

Every job here is indexed straight from company career pages — often hours after it opens, before it reaches the big boards. Create a free account and get your best matches in a twice-daily digest.

  • Your best matches, twice a day
  • No duplicates, no ghost jobs, no recruiter spam
  • Every job free to browse — pay only when you apply
Get my matched jobs

Free account — no card required

93 123 live jobs · 17 643 companies tracked · 82 added today

Similar jobs