• Skip to primary navigation
  • Skip to main content
  • Skip to footer

Side Hustles

Side Hustles

Side Hustles For All

  • Best Side Hustles
    • Woman sitting on a pile of coins and working on a laptop surrounded by icons representing different side hustle ideas

      31 Best Side Hustles to Earn Extra Money in 2026

    • Bicycle courier delivering food for their side hustle.

      What Is a Side Hustle?

    • Remote worker sitting at his desk making money from home

      18 Ways to Make Money from Home (Online and Offline Jobs)

    • By Category
      • Arts & Crafts
      • Business Services
      • Caregiving
      • Creative Services
      • Digital Freelance Services
      • View All
    • By Lifestyle
      • I’m introverted
      • I’m a man
      • I’m a woman
      • I’m a stay-at-home mom
      • I’m unique
      • View All
    • By Profession
      • Artists & Creatives
      • Musicians
      • Nurses
      • Physicians
      • Teachers
      • View All
    • By Age Group
      • College Students
      • Teens
      • Age 50+
      • Seniors
      • View All
    • By Skills & Interests
      • Get Paid to Lose Weight
      • Get Paid to Play Games
      • Get Paid to Read
      • Get Paid to Sleep
      • Get Paid to Travel
      • View All
  • Best Gig Apps
    • Freelance worker popping out of a phone screen and considering gig apps on the App Store and Google Play

      Top 6 Gig Apps to Make Real Cash in 2026

    • Smartphone surrounded by the icons of different money-making apps

      Top 10 Best Money-Making Apps to Try in 2026

    • two teenagers using job apps on a phone and laptop

      19 Job Apps for Teens to Find Jobs and Make Money

    • By Gig Type
      • Cashback
      • Data Entry
      • Delivery
      • Games
      • Product Testing
      • View All
    • By Payment Method
      • Bingo Games that Pay to Cash App
      • Games that Pay Real Money
      • Games that Pay to Cash App
      • Games that Pay via PayPal
      • Surveys that Pay to Cash App
      • View All
    • By Benefits
      • $20 Signup Bonuses
      • $25 Signup Bonuses
      • $50 Signup Bonuses
      • Best Signup Bonuses
      • Instant Signup Bonuses
      • View All
    • By Skills & Interests
      • Driving
      • Losing Weight
      • Playing Games
      • Product Testing
      • Watching Ads
      • View All
  • Job Hunting
    • Freelance worker browsing a job post on a freelance job board.

      23 Job Boards You Can Use to Find Remote Work

    • Freelance writer sitting at her laptop working on a project

      15 Best Remote Jobs That Require No Paid Work Experience

    • Teenager sitting at laptop working an online job

      14 Online Jobs for Teens (With No Experience)

    • Freelancing
      • Freelance Writing Sites
      • Freelance Writing Job Boards
      • Freelance Writing Platforms
      • View More
    • Gig & Shift Work
      • Gig Work Apps
      • On-Demand Work Apps
      • Shift Work Apps
      • View More
    • GPT (Get Paid To)
      • Microtasking
      • Product Testing
      • Survey Taking
      • View More
    • Remote Working
      • Best Remote Job Boards
      • Top 15 Remote Jobs
      • View More
  • Job Board
    • Work Schedule
      • Part-Time Jobs
      • Per-Diem Jobs
      • Go Search
    • Work Environment
      • Hybrid Jobs
      • Remote Jobs
      • Go Search
    • Employment Type
      • Contractor Jobs
      • Internship Jobs
      • Temporary Jobs
      • Go Search
    • Job Title
      • Accounting Jobs
      • Data Entry Jobs
      • Nursing Jobs
      • Online Teaching Jobs
      • Software Engineer Jobs
      • Go Search
    • State
      • California Jobs
      • Florida Jobs
      • New York Jobs
      • Pennsylvania Jobs
      • Texas Jobs
      • Go Search
    • City
      • Chicago, IL
      • Houston, TX
      • Los Angeles, CA
      • New York City, NY
      • Phoenix, AZ
      • Go Search

Home Flexible Job Board Founding Engineer - ML Performance

$250,000–395,000/yr 125d ago

Founding Engineer - ML Performance

Boost your chances before you apply.

  • ✨ Apply 10x Faster Free

    It takes 30+ tailored applications to land jobs like this one. We'll help you get that done in 1 hour.

    No Credit Card Required

  • Proceed to Application Go directly to the company's job page to apply.
Logo

uRun

CA

Full-time Permanent Remote

✨ Apply 10x Faster

Analyze your resume for missing keywords, then one-click optimize it. Don't be anything less than a 100% match candidate.

Free

No Credit Card Required

Summary

Develop custom CUDA kernels and optimize model inference to achieve sub-50ms latency and 10-100x performance gains. Own the end-to-end inference pipeline, focusing on GPU utilization, memory bandwidth, and distributed memory optimizations.

Job Description

The problem we saw

Most AI infrastructure is built for batch: send a query, wait, get a response, reset. Powerful, but transactional. AI is becoming interactive — sessions that hold state, models that stay alive between turns, generation that responds as it runs — and the infrastructure to deliver that at scale doesn't really exist yet.

The bottleneck isn't the models anymore. It's the infrastructure underneath them.

What we're building to fix it

uRun is the inference cloud for interactive AI: the compute layer that makes real-time, stateful inference possible at scale. We came out of stealth in April 2026, are backed by top-tier investors, and are founded by Keegan McCallum, who scaled inference infrastructure for some of the most demanding generative AI workloads in production.

We're an infrastructure company. We build the layer that model labs, builders, and research teams ship on top of.

Where you come in

Performance is uRun's core differentiator. We're not chasing incremental gains — we're building infrastructure that runs 10–100x faster than the status quo. As our ML Performance Engineer, you will be the person who makes that true.

This is a founding technical hire. You will write custom CUDA kernels, push GPU utilization to its limits, and own inference latency end-to-end across the stack. You will work directly with the founding team on the hardest performance problems in production AI infrastructure — and your fingerprints will be on everything we ship.

What you'll actually be doing day-to-day

  • Write custom CUDA kernels that unlock performance headroom unavailable through off-the-shelf frameworks

  • Optimize model inference end-to-end, targeting sub-50ms latency across our inference platform

  • Drive 10x performance improvements across the stack: memory bandwidth, kernel fusion, operator scheduling, and beyond

  • Implement zero-copy distributed memory optimizations across multi-GPU and multi-node environments

  • Own GPU utilization and memory management, squeezing every available FLOP out of the hardware we run

  • Profile, benchmark, and instrument the full inference pipeline to find and eliminate bottlenecks systematically

  • Set the performance engineering bar for the team: define what fast looks like and build the tooling to measure it

What skills you need for the journey

  • Deep, hands-on CUDA expertise: you have written custom kernels in production, not just called into cuBLAS

  • Strong background in model inference and post-training optimization at scale

  • Fluency in GPU memory hierarchy, warp scheduling, kernel fusion, and hardware-aware algorithm design

  • Experience profiling and benchmarking complex inference pipelines: you know where the time goes and how to get it back

  • Able to operate at the frontier with minimal guidance — you identify the problem, design the approach, and ship the fix

Things that will give you an edge

  • Public work in GPU optimization or inference efficiency — open source contributions, a published paper, or a side project that shows your depth (vLLM, Flash-Attention, TensorRT-LLM, PyTorch, or equivalent)

  • Experience with hardware-aware optimization frameworks: CuTe, Triton, TileLang, or similar

  • Familiarity with distributed memory and communication primitives: NCCL, InfiniBand, NVLink, RoCE

  • Contributions to or deep familiarity with PyTorch Distributed, Ray core, or similar systems

  • Experience optimizing for video generation or other high-throughput, latency-sensitive generative workloads

  • Prior work at an inference-focused company or research lab pushing the boundary of what GPU hardware can do

What you'll get in return

Competitive salary and meaningful equity in an early-stage AI infrastructure company. The band above is our target; for an exceptional candidate we'll go higher. Equity is real — you're early, and the grant reflects that.

  • Health, dental, and vision — full coverage

  • 401(k) — company-supported retirement savings

  • FSA/HSA — flexible spending accounts for healthcare costs

  • Paid time off — we trust you to manage your time

  • Top-tier tooling — access to the best AI tools available: Claude, Codex, Kimi, and whatever else helps you move faster

  • MacBook Pro and AirPods — the hardware you need, on us

How we work (and what that feels like day-to-day)

We build the stage, not the show. We're an infrastructure company, a developer-tools company, and a production partner for model labs, and focus is a deliberate choice we've made and hold to.

Day-to-day, that means a small team, a high bar, and real ownership. You won't wait for permission or inherit a backlog of someone else's decisions, in a founding security role, the function is what you make it.

It also means ambiguity: priorities shift, not everything is documented, and you'll often be the person who decides what "secure enough, for now" means. That suits some people and not others, and we'd rather you know that before you apply.

  • Watch our launch party video

  • Read the manifesto

  • Follow us on LinkedIn

  • Follow us on X

About the company

uRun

Real-time stateful inference wasn't possible. We built it.

uRun is the infrastructure layer for interactive generative AI - real-time, steerable, and production-grade from day one. Our stateful inference runtime keeps sessions alive so every interaction builds on the last, eliminating latency between prompt and output entirely.

Founded

2025

Company size

2-10 employees

Industry

Software Development

Org type

Privately Held

Headquarters

San Francisco Bay Area, CA

Apply Now

About the company

uRun

Real-time stateful inference wasn't possible. We built it.

uRun is the infrastructure layer for interactive generative AI - real-time, steerable, and production-grade from day one. Our stateful inference runtime keeps sessions alive so every interaction builds on the last, eliminating latency between prompt and output entirely.

Founded

2025

Company size

2-10 employees

Industry

Software Development

Org type

Privately Held

Headquarters

San Francisco Bay Area, CA

Footer

sidehustles.com
Facebook Twitter Instagram LinkedIn Reddit TikTok YouTube

Show Me The Money

  • Side Hustle Basics
  • Side Hustle Job Board (Remote & Part-Time Jobs)
  • Gig App Reviews
  • Job Hunting
  • Manage Your Money
  • The Gig Apple: News & Events

Company

  • About Us
  • Contact Us
  • Become a Contributor
  • Advertising & Sponsorships
  • Partner With Us
  • Editorial Guidelines

Side Hustles © All rights reserved

  • Privacy Policy
  • Terms of Service

Thanks for using our free job board

Your review would mean a lot to us.

If you love that we're just giving away remote jobs for free with no paywall, please spread the word. (You will need to create an account on Trustpilot, for which we'll be eternally grateful.) Good luck out there!

Leave a Review Not yet. Send me to the job post.