• Skip to primary navigation
  • Skip to main content
  • Skip to footer

Side Hustles

Side Hustles

Side Hustles For All

  • Best Side Hustles
    • Woman sitting on a pile of coins and working on a laptop surrounded by icons representing different side hustle ideas

      31 Best Side Hustles to Earn Extra Money in 2026

    • Bicycle courier delivering food for their side hustle.

      What Is a Side Hustle?

    • Remote worker sitting at his desk making money from home

      18 Ways to Make Money from Home (Online and Offline Jobs)

    • By Category
      • Arts & Crafts
      • Business Services
      • Caregiving
      • Creative Services
      • Digital Freelance Services
      • View All
    • By Lifestyle
      • I’m introverted
      • I’m a man
      • I’m a woman
      • I’m a stay-at-home mom
      • I’m unique
      • View All
    • By Profession
      • Artists & Creatives
      • Musicians
      • Nurses
      • Physicians
      • Teachers
      • View All
    • By Age Group
      • College Students
      • Teens
      • Age 50+
      • Seniors
      • View All
    • By Skills & Interests
      • Get Paid to Lose Weight
      • Get Paid to Play Games
      • Get Paid to Read
      • Get Paid to Sleep
      • Get Paid to Travel
      • View All
  • Best Gig Apps
    • Freelance worker popping out of a phone screen and considering gig apps on the App Store and Google Play

      Top 6 Gig Apps to Make Real Cash in 2026

    • Smartphone surrounded by the icons of different money-making apps

      Top 10 Best Money-Making Apps to Try in 2026

    • two teenagers using job apps on a phone and laptop

      19 Job Apps for Teens to Find Jobs and Make Money

    • By Gig Type
      • Cashback
      • Data Entry
      • Delivery
      • Games
      • Product Testing
      • View All
    • By Payment Method
      • Bingo Games that Pay to Cash App
      • Games that Pay Real Money
      • Games that Pay to Cash App
      • Games that Pay via PayPal
      • Surveys that Pay to Cash App
      • View All
    • By Benefits
      • $20 Signup Bonuses
      • $25 Signup Bonuses
      • $50 Signup Bonuses
      • Best Signup Bonuses
      • Instant Signup Bonuses
      • View All
    • By Skills & Interests
      • Driving
      • Losing Weight
      • Playing Games
      • Product Testing
      • Watching Ads
      • View All
  • Job Hunting
    • Freelance worker browsing a job post on a freelance job board.

      23 Job Boards You Can Use to Find Remote Work

    • Freelance writer sitting at her laptop working on a project

      15 Best Remote Jobs That Require No Paid Work Experience

    • Teenager sitting at laptop working an online job

      14 Online Jobs for Teens (With No Experience)

    • Freelancing
      • Freelance Writing Sites
      • Freelance Writing Job Boards
      • Freelance Writing Platforms
      • View More
    • Gig & Shift Work
      • Gig Work Apps
      • On-Demand Work Apps
      • Shift Work Apps
      • View More
    • GPT (Get Paid To)
      • Microtasking
      • Product Testing
      • Survey Taking
      • View More
    • Remote Working
      • Best Remote Job Boards
      • Top 15 Remote Jobs
      • View More
  • Job Board
    • Work Schedule
      • Part-Time Jobs
      • Per-Diem Jobs
      • Go Search
    • Work Environment
      • Hybrid Jobs
      • Remote Jobs
      • Go Search
    • Employment Type
      • Contractor Jobs
      • Internship Jobs
      • Temporary Jobs
      • Go Search
    • Job Title
      • Accounting Jobs
      • Data Entry Jobs
      • Nursing Jobs
      • Online Teaching Jobs
      • Software Engineer Jobs
      • Go Search
    • State
      • California Jobs
      • Florida Jobs
      • New York Jobs
      • Pennsylvania Jobs
      • Texas Jobs
      • Go Search
    • City
      • Chicago, IL
      • Houston, TX
      • Los Angeles, CA
      • New York City, NY
      • Phoenix, AZ
      • Go Search

Home Flexible Job Board Machine Learning Engineer, Speech - Joint Audio-Video Modeling

$200,000–220,000/yr 56d ago

Machine Learning Engineer, Speech - Joint Audio-Video Modeling

Boost your chances before you apply.

  • ✨ Apply 10x Faster Free

    It takes 30+ tailored applications to land jobs like this one. We'll help you get that done in 1 hour.

    No Credit Card Required

  • Proceed to Application Go directly to the company's job page to apply.
Logo

Cantina

US

Full-time Permanent Remote

✨ Apply 10x Faster

Analyze your resume for missing keywords, then one-click optimize it. Don't be anything less than a 100% match candidate.

Free

No Credit Card Required

Summary

You will design and implement state-of-the-art speech and audio generation systems, focusing on joint audio-video modeling and production inference. You will also drive the end-to-end model development lifecycle, including data curation, experimental design, and performance optimization.

Job Description

About Cantina:

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!

About the Role:

We're looking for a Research / ML Engineer to join our Speech Team to build state-of-the-art speech and audio generation systems end-to-end from data specs through production inference with a focus on joint audio-video modeling.

You'll own the audio side of multimodal generation: the representations (audio VAEs, neural codecs), the generative backbone (diffusion / flow-matching transformers), and the conditioning and alignment machinery that makes characters speak, sing, and emote in sync with what's on screen. That includes voice cloning and multi-speaker conditioning inside joint AV models, cinematic dialogue with music and sound design, and adjacent speech tasks (controllable TTS, voice conversion) that feed the same stack.

You'll drive the model ↔ data ↔ eval flywheel, partnering closely with research, video, data, and infra to ship fast, reliable, and cost-aware models. In this role you'll work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems.

You will thrive in this role if you:

  • See research and engineering as two sides of the same coin and enjoy owning work end-to-end.

  • Are excited to work across modalities and collaborate closely with a video generation team rather than staying inside audio.

  • Are results-oriented, flexible, and willing to pick up whatever moves the needle.

  • Like collaborating closely with infra, data, and product to ship measurable improvements.

  • Enjoy designing experiments, listening tests, and metrics that correlate with user-perceived quality.

  • Are eager to learn every day, and to find and solve unique large-scale problems.

What You’ll Do:

  • Audio Representations: Design, train, and improve the audio VAEs, neural codecs, and vocoders our generative models sit on top of latent design, reconstruction and perceptual objectives, compression-vs-fidelity tradeoffs.

  • Model Building: Architect, implement, pre-train, fine-tune, and post-train/alignment (e.g., GRPO/DPO) diffusion and flow-matching transformers for large-scale audio and video generation.

  • Joint Audio-Video Modeling: Design the audio conditioning and cross-modal alignment inside joint AV models, audio latents alongside video latents, reference-audio and multi-speaker conditioning, multi shot generation audio/video modeling.

  • Experimental Design: Design, run, and analyze scientific experiments to advance our understanding of the models.

  • Data Ownership: Define data requirements and collaborate on acquisition, curation, AV-sync and quality filtering, annotation quality, and synthetic data strategies for paired audio-video and speech corpora.

  • Rigorous Evaluation: Design automated objective/subjective evaluations audio fidelity and intelligibility metrics, AV-sync, listening and viewing tests, robustness & bias checks, and red-team studies.

  • Inference Efficiency: Drive distillation, step-count reduction, quantization, and kernel/memory optimization to meet interactive latency and cost targets.

  • Pipeline Delivery: Harden the training → evaluation → inference pipeline; profile latency, memory, and cost; and meet production SLAs with robust monitoring and rollback.

  • GPU Scaling: Partner with infrastructure to run distributed training/inference on cloud fleets and productionize models with reliability and observability.

  • Project Leadership: Independently lead small research projects while collaborating on larger team initiatives, including cross-team work with video generation.

  • Tool Development: Develop and improve dev tooling to enhance team productivity.

  • Safety & Responsibility: Contribute to safety/consent guardrails, watermarking, and misuse/abuse mitigation for responsible voice and likeness technology.

What You’ll Bring:

  • Exceptional research/development experience with large-scale audio models (>8B parameters, >500k hours of data).

  • Deep hands-on experience with diffusion and/or flow-matching transformers, including practical knowledge of samplers, schedules, conditioning mechanisms, and distillation.

  • Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders latent/tokenizer design, reconstruction and perceptual objectives, adversarial training.

  • Strong experience with multi-node, multi-GPU distributed training (FSDP/DeepSpeed or equivalent).

  • Strong software engineering skills with a proven track record of building complex systems.

  • Strong with PyTorch and performance work (profiling, CUDA/Triton/C++ as needed) and writing reliable production-quality code.

  • Shipped large-scale speech/audio or multimodal generative models to production.

  • Background in working with large-scale ML data, and the ability to iterate on data and triangulate quality using both subjective and objective signals.

  • Experience with voice cloning, speech control/steerability, or expressive speech generation.

  • Notable publications and/or open-source contributions in speech/audio/ML.

  • Strongly preferred:

    • Experience with multimodal audio-video modeling: joint AV generation of multi-shot, multi-speaker scenes with dialogue, music, and sound design generated jointly with video, and the cross-modal alignment that keeps them in sync.

    • Experience with video generation: video diffusion/flow-matching transformers, video VAEs, conditioned and multi-shot generation, building data pipelines for video models.

    • Streaming or real-time generation, causal distillation (e.g., Self Forcing / Self Forcing++).

Compensation:

The anticipated annual base salary range for this role is between $200,000-$220,000 (€170,000-€190,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits for U.S.-based roles:

  • Competitive salary and generous company equity

  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina

  • 42 days of paid time off, including:

    • 15 PTO days

    • 10 sick days

    • 15 company holidays

    • 2 floating holidays

  • Generous parental leave & fertility support

  • 401(k) retirement savings plan

  • Lifestyle spending account – $500/month to use however you’d like

  • Complimentary lunch and snacks for in-office employees

  • One Medical membership, and more!

About the company

Cantina

Cantina is an innovation consultancy headquartered in Boston. For more than 15 years, we’ve partnered with visionary business leaders to challenge the status quo, reimagine products & services for a customer-centric age, and drive new growth and efficiencies for our client’s organizations.

What we’ve learned throughout our history is that innovation is not magic. It’s the result of disciplined craft and methodology that bridges the gap between design thinking and design doing. By applying a creative, multidisciplinary approach, Cantina helps our clients create and sustain innovation, transform business-as-usual, and improve the lives of customers and employees.

Founded

2007

Company size

51-200 employees

Industry

Design

Org type

Privately Held

Headquarters

Boston, MA

Apply Now

About the company

Cantina

Cantina is an innovation consultancy headquartered in Boston. For more than 15 years, we’ve partnered with visionary business leaders to challenge the status quo, reimagine products & services for a customer-centric age, and drive new growth and efficiencies for our client’s organizations.

What we’ve learned throughout our history is that innovation is not magic. It’s the result of disciplined craft and methodology that bridges the gap between design thinking and design doing. By applying a creative, multidisciplinary approach, Cantina helps our clients create and sustain innovation, transform business-as-usual, and improve the lives of customers and employees.

Founded

2007

Company size

51-200 employees

Industry

Design

Org type

Privately Held

Headquarters

Boston, MA

Footer

sidehustles.com
Facebook Twitter Instagram LinkedIn Reddit TikTok YouTube

Show Me The Money

  • Side Hustle Basics
  • Side Hustle Job Board (Remote & Part-Time Jobs)
  • Gig App Reviews
  • Job Hunting
  • Manage Your Money
  • The Gig Apple: News & Events

Company

  • About Us
  • Contact Us
  • Become a Contributor
  • Advertising & Sponsorships
  • Partner With Us
  • Editorial Guidelines

Side Hustles © All rights reserved

  • Privacy Policy
  • Terms of Service

Thanks for using our free job board

Your review would mean a lot to us.

If you love that we're just giving away remote jobs for free with no paywall, please spread the word. (You will need to create an account on Trustpilot, for which we'll be eternally grateful.) Good luck out there!

Leave a Review Not yet. Send me to the job post.