Don't see your role? Reach out anyway.
Boost your chances before you apply.
uRun
San Francisco, CA, US
Summary
Build and architect the Universal Runtime (uRun) layer to enable real-time, stateful inference for AI. Lead technical direction or execute high-velocity engineering to solve ambiguous infrastructure problems.
Job Description
Don't see your role? Reach out anyway.
Our roadmap moves faster than our job board. Some of our best hires were never attached to an open req, they just made it obvious we'd be foolish not to. If you scan our openings and think "none of these, but I'd be one of the best things to happen to this company," this is the post for you.
What we're building
uRun, Universal Runtime, is the layer that makes real-time, stateful inference possible. The bottleneck in interactive AI isn't the models, it's the runtime underneath them. We're an infrastructure company, and we build the layer model labs, builders, and research teams ship on top of.
Who we want to hear from
A technical leader (CTO, VP/Head of Engineering, founding level) who has architected large-scale production systems and built the teams around them. Distributed systems, GPU-heavy workloads, and low-latency infra are home turf here.
A cracked engineer whose shipped work says more than any resume. Inference, systems, performance, serving infra, or raw full-stack velocity, pick your weapon.
Someone exceptional in a function we haven't opened yet. GTM, design, devrel, ops. If you've operated at the top of your field, make the case.
What's consistent across everyone we hire: a track record of owning ambiguous problems end-to-end, a high bar you hold yourself to, and the judgment to set direction with limited scaffolding.
What you'll get
Competitive salary and real equity, you're early, and the grant reflects that. Full health, dental, and vision. 401(k). FSA/HSA. PTO we trust you to manage. Top-tier AI tooling (Claude, Codex, Kimi, whatever moves you faster). MacBook Pro and AirPods on us.
Real-time stateful inference wasn't possible. We built it.
uRun is the infrastructure layer for interactive generative AI - real-time, steerable, and production-grade from day one. Our stateful inference runtime keeps sessions alive so every interaction builds on the last, eliminating latency between prompt and output entirely.
Founded
2025
Company size
2-10 employees
Industry
Software Development
Org type
Privately Held
Headquarters
San Francisco Bay Area, CA
Real-time stateful inference wasn't possible. We built it.
uRun is the infrastructure layer for interactive generative AI - real-time, steerable, and production-grade from day one. Our stateful inference runtime keeps sessions alive so every interaction builds on the last, eliminating latency between prompt and output entirely.
Founded
2025
Company size
2-10 employees
Industry
Software Development
Org type
Privately Held
Headquarters
San Francisco Bay Area, CA