• Skip to primary navigation
  • Skip to main content
  • Skip to footer

Side Hustles

Side Hustles

Side Hustles For All

  • Best Side Hustles
    • Woman sitting on a pile of coins and working on a laptop surrounded by icons representing different side hustle ideas

      31 Best Side Hustles to Earn Extra Money in 2026

    • Bicycle courier delivering food for their side hustle.

      What Is a Side Hustle?

    • Remote worker sitting at his desk making money from home

      18 Ways to Make Money from Home (Online and Offline Jobs)

    • By Category
      • Arts & Crafts
      • Business Services
      • Caregiving
      • Creative Services
      • Digital Freelance Services
      • View All
    • By Lifestyle
      • I’m introverted
      • I’m a man
      • I’m a woman
      • I’m a stay-at-home mom
      • I’m unique
      • View All
    • By Profession
      • Artists & Creatives
      • Musicians
      • Nurses
      • Physicians
      • Teachers
      • View All
    • By Age Group
      • College Students
      • Teens
      • Age 50+
      • Seniors
      • View All
    • By Skills & Interests
      • Get Paid to Lose Weight
      • Get Paid to Play Games
      • Get Paid to Read
      • Get Paid to Sleep
      • Get Paid to Travel
      • View All
  • Best Gig Apps
    • Freelance worker popping out of a phone screen and considering gig apps on the App Store and Google Play

      Top 6 Gig Apps to Make Real Cash in 2026

    • Smartphone surrounded by the icons of different money-making apps

      Top 10 Best Money-Making Apps to Try in 2026

    • two teenagers using job apps on a phone and laptop

      19 Job Apps for Teens to Find Jobs and Make Money

    • By Gig Type
      • Cashback
      • Data Entry
      • Delivery
      • Games
      • Product Testing
      • View All
    • By Payment Method
      • Bingo Games that Pay to Cash App
      • Games that Pay Real Money
      • Games that Pay to Cash App
      • Games that Pay via PayPal
      • Surveys that Pay to Cash App
      • View All
    • By Benefits
      • $20 Signup Bonuses
      • $25 Signup Bonuses
      • $50 Signup Bonuses
      • Best Signup Bonuses
      • Instant Signup Bonuses
      • View All
    • By Skills & Interests
      • Driving
      • Losing Weight
      • Playing Games
      • Product Testing
      • Watching Ads
      • View All
  • Job Hunting
    • Freelance worker browsing a job post on a freelance job board.

      23 Job Boards You Can Use to Find Remote Work

    • Freelance writer sitting at her laptop working on a project

      15 Best Remote Jobs That Require No Paid Work Experience

    • Teenager sitting at laptop working an online job

      14 Online Jobs for Teens (With No Experience)

    • Freelancing
      • Freelance Writing Sites
      • Freelance Writing Job Boards
      • Freelance Writing Platforms
      • View More
    • Gig & Shift Work
      • Gig Work Apps
      • On-Demand Work Apps
      • Shift Work Apps
      • View More
    • GPT (Get Paid To)
      • Microtasking
      • Product Testing
      • Survey Taking
      • View More
    • Remote Working
      • Best Remote Job Boards
      • Top 15 Remote Jobs
      • View More
  • Job Board
    • Work Schedule
      • Part-Time Jobs
      • Per-Diem Jobs
      • Go Search
    • Work Environment
      • Hybrid Jobs
      • Remote Jobs
      • Go Search
    • Employment Type
      • Contractor Jobs
      • Internship Jobs
      • Temporary Jobs
      • Go Search
    • Job Title
      • Accounting Jobs
      • Data Entry Jobs
      • Nursing Jobs
      • Online Teaching Jobs
      • Software Engineer Jobs
      • Go Search
    • State
      • California Jobs
      • Florida Jobs
      • New York Jobs
      • Pennsylvania Jobs
      • Texas Jobs
      • Go Search
    • City
      • Chicago, IL
      • Houston, TX
      • Los Angeles, CA
      • New York City, NY
      • Phoenix, AZ
      • Go Search

Home Flexible Job Board Senior Site Reliability Engineer - AI Infrastructure

Salary Unstated 169d ago

Senior Site Reliability Engineer - AI Infrastructure

Boost your chances before you apply.

  • ✨ Apply 10x Faster Free

    It takes 30+ tailored applications to land jobs like this one. We'll help you get that done in 1 hour.

    No Credit Card Required

  • Proceed to Application Go directly to the company's job page to apply.
Logo

Andromeda Cluster

San Francisco, CA, US

Full-time Permanent Remote

✨ Apply 10x Faster

Analyze your resume for missing keywords, then one-click optimize it. Don't be anything less than a 100% match candidate.

Free

No Credit Card Required

Summary

You will design, operate, and debug large-scale GPU infrastructure for distributed AI training and inference. You will also serve as the primary technical partner for customers, ensuring the reliability and performance of high-speed interconnects and compute clusters.

Job Description

Senior Site Reliability Engineer

Location: Global Remote / San Francisco · Full-Time

About Andromeda

Andromeda is a market and infrastructure platform to buy, sell, and operate compute.

We believe demand for compute will grow exponentially. So fast that a handful of vertically integrated providers won't be able to scale across operations, capital, supply chains, and politics to serve it. The result is a massive wave of fragmentation, with AI factories of every shape and size coming to market to fill this demand. Our job is to enable all of that fragmented compute to flow through one platform, delivering reliable capacity to model builders, research labs, and inference providers when they need it. We believe every spare electron should be made productive for AI and we're building the platform that makes that possible.

We sit at the center of three forces:

  • Companies that need reliable, high-performance compute fast

  • A fragmented global supply of GPUs across hyperscalers, neoclouds, and independent data centers

  • Capital, risk, and operational complexity that most teams are not equipped to manage

When we succeed, trillions of dollars of compute will flow through Andromeda. Builders get capacity when they need it. Providers get a reliable way to monetize, operate, and finance infrastructure at scale. Capital gets an easy way to deploy, hedge, and underwrite.

In five years, Andromeda won't just participate in the AI infrastructure market. We will shape it.

The Role

This is not a generalist SRE role.

You will design, operate, and debug large-scale GPU infrastructure used for distributed training and inference, working directly with customers pushing the limits of modern AI systems.

We’re looking for engineers who have personally run GPU clusters in production, understand the failure modes of distributed training, and can reason about performance from network fabric → kernel → framework.

What You’ll Own

  • GPU Cluster Architecture: Design and evolve multi-provider, multi-region GPU compute clusters optimized for large-scale training. Make topology-aware scheduling, networking, and storage decisions that directly impact training throughput and cost efficiency.

  • Customer Technical Partnership: Serve as the primary technical point of contact for customers running large-scale training workloads. Onboard, troubleshoot, and optimize, often in real time.

  • Reliability & Performance Engineering: Define SLOs and error budgets that account for the unique failure modes of GPU infrastructure (ECC errors, NVLink degradation, NCCL timeouts). Own capacity planning across heterogeneous GPU fleets optimized for training throughput.

  • Networking & Fabric Health: Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) that underpin distributed training. Diagnose and resolve fabric-level issues that degrade collective operations.

  • Observability: Build deep visibility into GPU utilization, memory pressure, interconnect throughput, training job performance, and hardware health. Go well beyond standard infrastructure metrics.

  • Automation & Tooling: Build production-grade automation for cluster provisioning, GPU health checks, job scheduling, self-healing, and firmware/driver lifecycle management.

  • Incident Leadership: Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks. Drive blameless postmortems and systemic fixes.

What We’re Looking For

  • GPU Systems Expertise: Deep, hands-on experience operating large-scale GPU clusters (NVIDIA A100/H100/B200 or equivalent). You understand GPU memory hierarchies, ECC behavior, thermal throttling, and hardware failure modes from direct experience not documentation.

  • High-Performance Networking: Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training. You can diagnose why an all-reduce is slow, identify a degraded link in a fat-tree topology, and reason about congestion control at scale.

  • Distributed Training & ML Frameworks: Working knowledge of how large training jobs actually run — NCCL, CUDA, PyTorch distributed, DeepSpeed, Megatron, FSDP, or similar. You don't need to write the models, but you need to understand what's happening at the systems level when a 1,000-GPU training run stalls.

  • Linux & Systems Internals: Expert-level Linux knowledge: kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, performance profiling at the syscall and hardware level.

  • Kubernetes & Orchestration: Strong experience running Kubernetes in production with GPU workloads, including device plugins, topology-aware scheduling, multi-cluster federation, and custom operators. Experience with Slurm or other HPC schedulers is equally valued.

  • Automation & Software Engineering: Strong engineering skills in Python, Go, or Bash. You build production-grade tools and services, not just scripts. Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent).

  • Observability & Monitoring: Hands-on experience building monitoring and alerting for GPU infrastructure, not just Prometheus/Grafana basics, but GPU-specific telemetry (DCGM, nvidia-smi, fabric manager metrics) integrated into actionable dashboards.

  • Incident Management: Proven track record leading incident response for complex distributed systems where the failure could be in hardware, firmware, networking, drivers, orchestration, or application code and you need to narrow it down fast.

Strong Candidates May Have

  • Distributed Storage: Experience with high-performance parallel file systems (VAST, Weka, Lustre, GPFS) and the checkpoint I/O and data-loading bottlenecks that come with large training runs.

  • Training Optimization: Experience profiling and optimizing distributed training performance: identifying stragglers, tuning collective communication strategies, improving MFU (Model FLOPs Utilization), and reducing idle GPU time across large runs.

  • Cluster Buildout & Hardware: Experience involved in physical cluster design - rack layout, power/cooling constraints, network topology design, and hardware validation/burn-in at scale.

  • Team Leadership: Experience leading or mentoring a team of infrastructure engineers. We're growing and need people who raise the bar for everyone around them.

Why You’ll Love It Here

This is a high-impact, senior builder’s role. You’ll have significant ownership and autonomy to shape how our systems run at a foundational level, working directly with customers and providers while architecting the infrastructure backbone for reliable, scalable AI compute. You’ll influence technical direction and help define what world-class AI infrastructure operations look like.

Andromeda Cluster is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

About the company

Andromeda Cluster

Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers.

We began with a single managed cluster — but it filled almost instantly. Since then, we’ve been quietly building the systems, network, and orchestration layer that makes the world’s AI infrastructure more accessible.

Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it’s needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth.

Our long-term vision is to build the liquidity layer for global AI compute. We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering.

Company size

11-50 employees

Industry

Technology, Information and Internet

Org type

Privately Held

Headquarters

San Francisco, California

Apply Now

About the company

Andromeda Cluster

Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers.

We began with a single managed cluster — but it filled almost instantly. Since then, we’ve been quietly building the systems, network, and orchestration layer that makes the world’s AI infrastructure more accessible.

Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it’s needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth.

Our long-term vision is to build the liquidity layer for global AI compute. We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering.

Company size

11-50 employees

Industry

Technology, Information and Internet

Org type

Privately Held

Headquarters

San Francisco, California

Footer

sidehustles.com
Facebook Twitter Instagram LinkedIn Reddit TikTok YouTube

Show Me The Money

  • Side Hustle Basics
  • Side Hustle Job Board (Remote & Part-Time Jobs)
  • Gig App Reviews
  • Job Hunting
  • Manage Your Money
  • The Gig Apple: News & Events

Company

  • About Us
  • Contact Us
  • Become a Contributor
  • Advertising & Sponsorships
  • Partner With Us
  • Editorial Guidelines

Side Hustles © All rights reserved

  • Privacy Policy
  • Terms of Service

Thanks for using our free job board

Your review would mean a lot to us.

If you love that we're just giving away remote jobs for free with no paywall, please spread the word. (You will need to create an account on Trustpilot, for which we'll be eternally grateful.) Good luck out there!

Leave a Review Not yet. Send me to the job post.