Member of Technical Staff, GPU / ML Systems
Boost your chances before you apply.
SkyPilot
San Mateo, CA, US
Summary
Own the GPU and ML-systems layer, focusing on accelerator scheduling, utilization, and health monitoring across clouds and Kubernetes. Build optimizations for large-scale pre-training, high-throughput inference, and deepen integrations with frameworks like vLLM and PyTorch.
Job Description
About SkyPilot
SkyPilot accelerates the world's most ambitious AI teams. SkyPilot turns fragmented AI compute across clusters into one optimized, highly available and easy-to-use pool: a single AI supercomputer.
SkyPilot (10k+ GitHub stars, 18M+ downloads) manages the GPU fleets of 100s of companies, from Fortune 500s to top AI-natives like Abridge, Applied Compute, Mistral, Unconventional AI, H Company, and Nubank, with usage growing exponentially. Born in the UC Berkeley lab behind Spark and Databricks, our growing team includes top-tier talent from Databricks, Google, Berkeley, MIT, CMU, and Cornell. To date, SkyPilot has raised over $20M in seed funding from top investors (incl. Lux, Coatue, Amplify) and operators (incl. Ali Ghodsi, Jeff Dean, Guillermo Rauch, Amjad Masad, Clem Delangue, Aaron Levie).
The role
SkyPilot exists because GPUs are scarce, expensive, and scattered — and the workloads that need them (pre-training, post-training, RL, high-throughput inference) push hardware to its limits. We're looking for an engineer to own the GPU and ML-systems layer that frontier AI teams run on: accelerator scheduling and utilization, the serving path, and the integrations that make SkyPilot the fastest, most cost-efficient place to run demanding AI workloads. A few points of GPU utilization here can save a team millions in compute and days on every training run.
What you'll do
- Own GPU scheduling, utilization and health: how SkyPilot discovers, packs, and binpacks accelerator capacity across clouds and Kubernetes, with real-time GPU health monitoring and automatic failure recovery.
- Build optimizations for training and serving: Enable large scale pre-training with node hot-swapping, design storage systems for fast model checkpointing, container migration, inference autoscaling and multi-cluster serving, preemption handling, and sandboxes for training, RL rollouts, and evals.
- Make the AI stack run great out of the box: deepen integrations with vLLM, PyTorch, Slime, and the frameworks teams use for pre-training and high-throughput inference.
What we're looking for
- Hands-on experience with GPU or accelerator systems and with ML training or inference infrastructure.
- Strongly preferred: Familiarity with the modern ML ecosystem (e.g. vLLM, PyTorch, CUDA, verl/slime) and workload-orchestration frameworks (e.g. Kueue, KAI, KServe).
- You've done real ML-systems performance work - tell us about a bottleneck you hunted down (a stalled data pipeline, GPUs idling on a scheduling gap, communication you overlapped with compute) and what you measured before and after.
- Strong Python, and comfort reaching into systems-level and GPU-adjacent details.
- You care about squeezing most from the compute available to you
- Experience operating large-scale training or high-throughput inference in production
What we offer
- Competitive compensation and equity
- Comprehensive medical, dental, vision coverage for you and your dependents
- The chance to work with some of the best minds in cloud, distributed, and AI systems — with significant autonomy and ownership.
- A front-row seat at the latest open-source infra startup from Berkeley (lineage: Databricks, Anyscale).
- Gourmet lunch & dinner for the team to do their best work
Location: San Mateo, CA. Remote will be considered for exceptional candidates.
Run, manage, and scale AI workloads on any AI infrastructure (Kubernetes, 17+ clouds, or on-prem).
SkyPilot gives AI teams a simple interface to run jobs on any infra. Infra teams get a unified control plane to manage any AI compute — with advanced scheduling, scaling, and orchestration.
Company size
1 employee
Industry
Technology, Information and Internet
Org type
Privately Held
Headquarters
San Francisco
Run, manage, and scale AI workloads on any AI infrastructure (Kubernetes, 17+ clouds, or on-prem).
SkyPilot gives AI teams a simple interface to run jobs on any infra. Infra teams get a unified control plane to manage any AI compute — with advanced scheduling, scaling, and orchestration.
Company size
1 employee
Industry
Technology, Information and Internet
Org type
Privately Held
Headquarters
San Francisco