Member of Technical Staff - Infrastructure

Gimlet · San Francisco, CA

ExclusiveRemoteFull-time$150,000 – $350,000Published Jun 12, 2026

Apply directly on Gimlet’s careers site — no account needed.

About the role

About Us

Gimlet is building the first multi-silicon neocloud designed for fast, efficient inference.

As AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough. The challenge is making increasingly diverse compute work together.

Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency. Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.

We work with foundation labs, hyperscalers, and AI-native companies to power production workloads at massive scale and help define the infrastructure layer for the future of AI.

About this Role

We are looking for an Infrastructure Platform Engineer to design, build, and operate the cluster infrastructure behind Gimlet's heterogeneous AI cloud.

In this role, you will build the platform that brings new hardware online, provisions clusters, manages capacity, and keeps production inference systems running reliably at scale. You'll work across bare metal, Linux, Kubernetes and cluster schedulers, high-speed networking, observability, and automation to ensure AI workloads can execute efficiently in production.

Unlike traditional cloud platforms built around a single hardware ecosystem, Gimlet's infrastructure spans multiple accelerator vendors and architectures. You'll build the operational systems that abstract this complexity, allowing new silicon to become production-ready quickly while ensuring workloads remain reliable, observable, and performant from day one.

This is a highly hands-on systems role. You'll partner closely with distributed systems, runtime, compiler, networking, and hardware engineers to build the infrastructure foundation that powers the next generation of AI workloads.

What Success Looks Like

In your first 12–18 months, you will help:

  • Design, deploy, and operate large-scale CPU, GPU, and accelerator clusters powering production AI inference.

  • Build provisioning and lifecycle management systems that automate deployment, upgrades, validation, and fleet operations.

  • Improve cluster scheduling, resource utilization, isolation, and capacity management across heterogeneous hardware.

  • Build highly observable infrastructure that enables rapid debugging, incident response, and operational excellence.

  • Partner with distributed systems, runtime, compiler, networking, and hardware engineers to bring new accelerator platforms into production.

  • Influence the architecture of the infrastructure platform that will power the next generation of AI workloads.

You may be a good fit if

  • Experience in infrastructure, cluster engineering, platform engineering, SRE, HPC, or distributed systems.

  • Deep Linux systems experience, including debugging performance, networking, storage, processes, and kernel-level issues.

  • Experience operating Kubernetes, Slurm, Nomad, or similar orchestration and scheduling systems.

  • Strong automation skills using tools such as Terraform, Ansible, Helm, Python, Go, or equivalent.

  • Experience with GPU or accelerator infrastructure, including drivers, firmware, CUDA/ROCm stacks, or hardware validation.

  • Familiarity with high-performance networking such as InfiniBand, RoCE, high-speed Ethernet, or datacenter fabrics.

  • Strong operational judgment: you know how to build systems that are observable, recoverable, and boring in production.

  • Comfort working in a fast-moving startup environment with high ownership and ambiguity.

  • Bachelor's degree in a relevant field, or an equivalent combination of education, training, and professional experience.

Strong candidates may also have

  • Experience building or operating AI inference, training, HPC, or neocloud infrastructure.

  • Experience with bare-metal provisioning, PXE/iPXE, image pipelines, BIOS/firmware management, or rack bring-up.

  • Experience with multi-tenant cluster isolation, quota systems, fair scheduling, or usage accounting.

  • Experience debugging distributed workload performance across compute, memory, network, and storage bottlenecks.

  • Experience building observability platforms using technologies such as Prometheus, OpenTelemetry, Grafana, or similar tooling.

  • Familiarity with heterogeneous hardware environments across NVIDIA, AMD, Intel, ARM, or emerging accelerators.

Why join now?

Gimlet is at the very beginning of its journey, and that's what makes this moment special. Most AI infrastructure companies are focused on deploying more compute. We are focused on making increasingly diverse compute work together, and that ambition touches every part of how we build and run this company.

As an early member of the team, you will have significant ownership over your work, partner directly with a small group of highly capable people, and help shape not just what we build, but how we scale the company.

We value people who are excited to work across domains, take ownership of meaningful problems, and help define what Gimlet becomes over the next several years.

Agency Policy: Gimlet Labs does not accept unsolicited resumes from recruitment agencies or search firms. Any unsolicited resumes submitted without a signed agreement will be considered the property of Gimlet Labs, and no fees will be paid.

Skills

  • Linux
  • Kubernetes
  • Python
  • Go

Never be applicant #200 again

Every job here is indexed straight from company career pages — often hours after it opens, before it reaches the big boards. Create a free account and get your best matches in a twice-daily digest.

  • Your best matches, twice a day
  • No duplicates, no ghost jobs, no recruiter spam
  • Every job free to browse — pay only when you apply
Get my matched jobs

Free account — no card required

93 117 live jobs · 17 641 companies tracked · 167 added today

Similar jobs

Member of Technical Staff - Infrastructure — Gimlet · Real Job Offers