Senior Platform Engineer
Boost your chances before you apply.
STN Inc
San Francisco, CA, US
Summary
Build and operate the multi-tenant orchestration and scheduling layer to transform raw GPU infrastructure into a cloud service. Design customer-facing APIs, CLIs, and automation for node provisioning and image management.
Job Description
Senior Platform Engineer
Platform and software · shared across customers
Reports to: Director, Platform Engineering (or Chief Architect)
Location: Remote (US) or Pleasanton, CA (hybrid)
Department: Cloud Platform Engineering / GPU Platform Engineering
Position summary
The Senior Platform Engineer builds and operates the multi-tenant orchestration, scheduling, and customer-facing platform layer that turns raw GPU infrastructure into a usable cloud service. This role is the software backbone of GPU One (GPUaaS).
Key responsibilities
-
Design and build the orchestration layer (Kubernetes, Slurm, Run:ai, or comparable)
-
Manage multi-tenant isolation including namespaces, networking, storage, and quotas
-
Build customer-facing platform APIs, CLIs, web portals, and SDKs
-
Implement and operate image management, GPU operator, and node provisioning automation
-
Drive infrastructure-as-code and automation across the platform stack
-
Partner with SRE on platform reliability, SLO definition, and observability
-
Support TAM and Support engineers on customer-impacting platform issues
-
Maintain customer environment templates, configuration management, and rollout tooling
-
Participate in architecture review, design discussions, and technical roadmap
-
Drive continuous platform improvement and reduce operational toil
Required qualifications
-
6+ years in platform engineering, SRE, or cloud engineering at scale
-
Deep Kubernetes expertise including CRDs, operators, and multi-tenant patterns
-
Strong programming skills in Go, Python, or both
-
Experience operating GPU clusters or AI infrastructure at production scale
-
Bachelor's degree in computer science or equivalent experience
Preferred qualifications
-
Experience with NVIDIA GPU Operator, MIG, MPS, and NCCL operator patterns
-
Familiarity with Slurm operator, Run:ai, KubeRay, or comparable AI orchestration
-
Service mesh experience (Istio, Linkerd) and multi-cluster networking
-
Open source contributions in the cloud-native or AI infrastructure ecosystem
At STN, we don’t just deliver technology, we build the foundation that modern organizations run on. From enterprise IT to AI infrastructure, we design, operate, and support systems that are reliable, secure, and built for performance. But what sets us apart isn’t just our stack it’s how we show up.
We don’t believe in one-size-fits-all. We don’t drop in hardware and disappear. We work side by side with our customers to understand what they actually need and build solutions that fit, flex, and scale as they grow.Whether you're running business-critical systems, deploying AI models, or training large-scale workloads on NVIDIA GPUs, we’re here to make sure your infrastructure isn’t just running, it’s working for you.
Our team brings deep technical expertise, hands-on support, and a people-first mindset to everything we do. Because we believe technology should unlock potential, not create more problems. STN exists to help teams thrive in complex environments with custom engineering, real partnership, and a clear plan forward.
Founded
2016
Company size
11-50 employees
Industry
IT Services and IT Consulting
Org type
Privately Held
Headquarters
Pleasanton, California
At STN, we don’t just deliver technology, we build the foundation that modern organizations run on. From enterprise IT to AI infrastructure, we design, operate, and support systems that are reliable, secure, and built for performance. But what sets us apart isn’t just our stack it’s how we show up.
We don’t believe in one-size-fits-all. We don’t drop in hardware and disappear. We work side by side with our customers to understand what they actually need and build solutions that fit, flex, and scale as they grow.Whether you're running business-critical systems, deploying AI models, or training large-scale workloads on NVIDIA GPUs, we’re here to make sure your infrastructure isn’t just running, it’s working for you.
Our team brings deep technical expertise, hands-on support, and a people-first mindset to everything we do. Because we believe technology should unlock potential, not create more problems. STN exists to help teams thrive in complex environments with custom engineering, real partnership, and a clear plan forward.
Founded
2016
Company size
11-50 employees
Industry
IT Services and IT Consulting
Org type
Privately Held
Headquarters
Pleasanton, California