Member of the Technical Staff - Platform
Boost your chances before you apply.
Andromeda Cluster
San Francisco, CA, US
Summary
You will build and operate the control plane for a large-scale GPU fleet, managing the automated lifecycle of clusters from bare metal to customer-ready. This includes maintaining Kubernetes and Postgres systems while scaling infrastructure to support thousands of nodes.
Job Description
Member of the Technical Staff, Platform
Location: North America Remote / San Francisco, CA · Full-Time
About Andromeda
Andromeda is a market and infrastructure platform to buy, sell, and operate compute.
We believe demand for compute will grow exponentially. So fast that a handful of vertically integrated providers won't be able to scale across operations, capital, supply chains, and politics to serve it. The result is a massive wave of fragmentation, with AI factories of every shape and size coming to market to fill this demand. Our job is to enable all of that fragmented compute to flow through one platform, delivering reliable capacity to model builders, research labs, and inference providers when they need it. We believe every spare electron should be made productive for AI and we're building the platform that makes that possible.
We sit at the center of three forces:
-
Companies that need reliable, high-performance compute fast
-
A fragmented global supply of GPUs across hyperscalers, neoclouds, and independent data centers
-
Capital, risk, and operational complexity that most teams are not equipped to manage
When we succeed, trillions of dollars of compute will flow through Andromeda. Builders get capacity when they need it. Providers get a reliable way to monetize, operate, and finance infrastructure at scale. Capital gets an easy way to deploy, hedge, and underwrite.
In five years, Andromeda won't just participate in the AI infrastructure market. We will shape it.
The Role
We are looking for engineers to build and operate the control plane that runs our fleet. Your responsibilities will include:
-
The automated systems that take a cluster from bare machines to customer-ready
-
Machine lifecycle between tenants: join, wipe, verify, rejoin
-
Operating Kubernetes and Postgres across the fleet
-
Contributing to our custom Kubernetes operators
-
Scaling clusters from tens of nodes to thousands
-
Participating in on-call rotations
Requirements
-
Impressive technical work you can go deep on, with impact in the world. That can take three years or twenty.
-
2+ years of on-call experience for critical production services
-
Deep Kubernetes experience
-
Strong Linux fundamentals: kernel, cgroups, containers, networking, storage
-
Experience operating databases, monitoring, CI/CD, and cloud infrastructure at scale
-
Familiarity with fleet management and capacity planning
Prior experience writing operators, or with GPUs, HPC scheduling, or bare-metal hardware is nice to have, but not required.
Why You’ll Love It Here
-
High-growth environment: Get in early at a company at the center of the AI infrastructure boom
-
Competitive compensation: + meaningful equity
-
Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage, 401(k), and unlimited PTO
Andromeda Cluster is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.
Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers.
We began with a single managed cluster — but it filled almost instantly. Since then, we’ve been quietly building the systems, network, and orchestration layer that makes the world’s AI infrastructure more accessible.
Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it’s needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth.
Our long-term vision is to build the liquidity layer for global AI compute. We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering.
Company size
11-50 employees
Industry
Technology, Information and Internet
Org type
Privately Held
Headquarters
San Francisco, California
Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers.
We began with a single managed cluster — but it filled almost instantly. Since then, we’ve been quietly building the systems, network, and orchestration layer that makes the world’s AI infrastructure more accessible.
Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it’s needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth.
Our long-term vision is to build the liquidity layer for global AI compute. We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering.
Company size
11-50 employees
Industry
Technology, Information and Internet
Org type
Privately Held
Headquarters
San Francisco, California