Staff Software Engineer, AI Inference Gateway
Boost your chances before you apply.
Salad
US
Summary
You will own the end-to-end AI Gateway system, including request routing, streaming, and horizontal scaling across the infrastructure. Additionally, you will build and tune fleet-efficiency algorithms to optimize throughput and performance for inference nodes.
Job Description
Staff Software Engineer, AI Inference Gateway
Salad Technologies · Remote · Full-time
The role
SaladCloud runs inference on thousands of geodistributed workstation and consumer GPUs — NVIDIA and AMD — that nobody else can use. Sitting in front of that fleet is a unique AI Gateway written in Rust, built on Pingora and WireGuard, that serves real-time and batch inference and embeddings to customers as a single, reliable API. You'll own it.
This is an infrastructure role with real autonomy. You'll design and ship the systems that decide how requests are fulfilled, when to add or drop capacity, and how to bill for every token — and you'll be the person who knows whether a new model, quantization, or GPU is actually worth running.
What you'll do
- Own the AI Gateway system end to end: request routing, streaming, batch/async job handling, and horizontal scaling across multiple gateway servers
- Build and tune the fleet-efficiency algorithms — autoscaling nodes, scoring performance, and evicting underperformers so the network converges on peak throughput per dollar
- Operate and scale the underlying inference nodes running vLLM, llama.cpp, and similar servers
- Design and maintain OpenTelemetry-based observability across the gateway and the fleet; pay-per-token and subscription billing already rides on this telemetry, so you'll keep that integration accurate
- Spend roughly 25% of your time benchmarking new models, quantizations, and GPU hardware, and work with the product and marketing teams to make the call on what goes into production
- Collaborate on cross-system design and integrations with the rest of Salad engineering; explain trade-offs clearly to both technical and non-technical teammates
The team
You'll join a small team — five engineers plus the CTO, who still writes code, and two deeply technical product managers who ship their own fixes rather than queue them for you. You'll report directly to the CTO. Everyone on the engineering team has 10+ years of experience and has built systems that already run at massive scale — with minimal incidents in a typical year, because we ship things that are production-worthy the first time.
You'll be the first hire with inference-gateway expertise and you'll be asked to own the whole surface. You won't do it alone: your teammates are eager to learn the domain, will pair on design and integrations, and share after-hours support. The bar is high; so is the support.
We've listed a lot below. If you're strong in two of the first three — Rust, LLM inference servers, distributed systems — apply, and expect to learn the rest fast with people who'll help.
What we're looking for
- Strong production Rust experience, ideally on high-throughput networked services
- Hands-on experience running LLM inference servers (vLLM, llama.cpp, TGI, TensorRT-LLM, or similar) — you know what a KV cache is and why it fills up
- Distributed systems fundamentals: load balancing, backpressure, failure handling, unreliable nodes
- Comfort operating what you build: metrics, tracing, on-call instincts
- Willing to be on-call for a system you helped make quiet — we treat after-hours pages as a design bug to fix, not a lifestyle
- Clear written and verbal communication; you can run a design without hand-holding and bring others along
Nice to have
- Pingora, Tokio, or proxy/gateway internals
- Experience with heterogeneous or consumer-grade GPU fleets, CUDA, or ROCm
- Quantization formats (GGUF, AWQ, GPTQ, FP8) and their performance characteristics
Why Salad
You'll work on a hard, unusual problem — coaxing datacenter-grade reliability out of hardware in people's homes — with the freedom to make major decisions and see them hit production quickly.
We're scaling rapidly: demand for inference on SaladCloud currently exceeds the capacity we can bring online, and this role exists to keep up with it. There is no shortage of work and no prospect of it ramping down.
The platform handles workload security and isolation, so you can focus on performance. We don't log or train on our customers' prompts or data.
Benefits
- Unlimited PTO
- 75% of health insurance premiums covered for you and your dependents
- Dental and vision coverage
- 401(k) plan
- Stock options
- Company-provided computer
- $500 WFH budget
- Fully remote, with flexible hours
Compensation
$180,000-$220,000/year
400 million consumer-grade GPUs in the world lie unused most of the day. Meanwhile, AI/ML companies are fighting for hard-to-find, expensive AI-focused GPUs. Salad connects the two.
SaladCloud is the world's most affordable GPU cloud for AI/ML inference at scale. Our recipe? We built the web’s most trusted computesharing network on a principle of mutual exchange.
Salad Cloud 1.0 provides easy, on-demand access to 10k+ consumer-grade GPUs on the world's largest distributed cloud at the lowest cost, supporting any framework/application without extra development work.
The Salad desktop application securely applies latent consumer compute resources like GPU VRAM, CPU L3 cache, and bandwidth to enterprise-scale container workloads, and rewards our users—the mighty Salad Chefs—with meaningful rewards, like video games, gift cards, streaming subscriptions, PayPal credits, and more epic loot. Salad Chefs redeem over 3,000,000 digital and real-world items from the Salad Storefront every year!
Learn more about the SaladCloud at https://salad.com.
Founded
2018
Company size
11-50 employees
Industry
Software Development
Org type
Privately Held
Headquarters
Salt Lake City, Utah
400 million consumer-grade GPUs in the world lie unused most of the day. Meanwhile, AI/ML companies are fighting for hard-to-find, expensive AI-focused GPUs. Salad connects the two.
SaladCloud is the world's most affordable GPU cloud for AI/ML inference at scale. Our recipe? We built the web’s most trusted computesharing network on a principle of mutual exchange.
Salad Cloud 1.0 provides easy, on-demand access to 10k+ consumer-grade GPUs on the world's largest distributed cloud at the lowest cost, supporting any framework/application without extra development work.
The Salad desktop application securely applies latent consumer compute resources like GPU VRAM, CPU L3 cache, and bandwidth to enterprise-scale container workloads, and rewards our users—the mighty Salad Chefs—with meaningful rewards, like video games, gift cards, streaming subscriptions, PayPal credits, and more epic loot. Salad Chefs redeem over 3,000,000 digital and real-world items from the Salad Storefront every year!
Learn more about the SaladCloud at https://salad.com.
Founded
2018
Company size
11-50 employees
Industry
Software Development
Org type
Privately Held
Headquarters
Salt Lake City, Utah