General Clinician (MD/DO) - HLS
Boost your chances before you apply.
Weekday AI
US
Summary
Physicians will apply clinical judgment to develop grading criteria, evaluate clinical AI dialogues, and perform structured annotation for AI systems. This role involves defining standards of care and clinical reasoning without direct patient care or live diagnosis responsibilities.
Job Description
This role is for one of our clients
Compensation: $150 per hour
We are hiring residency-trained physicians across specialties for non-clinical work developing and evaluating clinical AI systems. You will apply your clinical judgment to grading-criteria development, dialogue evaluation, and structured annotation work that determines how these systems are measured.
This is a non-clinical role — no direct patient care, and no responsibility for live diagnosis.
This is a shared expert pool. After onboarding you may be matched to any of several concurrent clinical workstreams based on your specialty, availability, and interest. You are not committing to a single project, and you may move between streams as priorities shift.
What you may work on
Work varies by workstream and may include:
- Grading criteria development — taking a clinical question and breaking the ideal answer into discrete, checkable criteria, so a model response can be graded consistently rather than impressionistically.
- Clinical dialogue evaluation — reviewing multi-turn clinical conversations and judging them for accuracy, safety, completeness, appropriate hedging, and whether escalation advice was correct.
- Clinical reasoning annotation — recording how you would work through a case, including the differential you considered and rejected, not only the conclusion.
- Output review — flagging hallucinated findings, dangerous omissions, unsupported certainty, and advice that is technically correct but clinically unsafe.
- Guideline authoring — defining edge cases and standards of care for your specialty so annotation stays consistent across a large group of clinicians.
- Difficult-case writing — constructing clinical questions that probe the limits of current model reasoning.
Task length varies by stream, from roughly 45 minutes for a dialogue evaluation up to an hour or more for grading-criteria authoring. You will get a specific throughput target for whichever stream you are matched to.
Required qualifications
- MD or DO with a completed residency in any specialty
- Active, unrestricted medical license in your country of practice
- 2+ years post-residency clinical experience, practising or previously practising
- Comfort writing structured clinical rationale that a non-specialist reviewer can follow
- Written and spoken English fluency
- Minimum 20 hours per week, with the ability to concentrate hours when a stream is time-boxed
Preferred qualifications
- Board certification in your specialty
- U.S. licensure, and familiarity with U.S. standards of care and clinical guidelines
- Primary care, internal medicine, emergency medicine, or hospitalist background, where breadth of presentation matters most
- Prior clinical annotation, AI evaluation, medical education, or question-writing experience
- Grading-criteria design, resident assessment, or clinical guideline development experience
- Published research or sustained technical writing (please link a sample)
Why this work
Most clinical AI failures are not exotic — they are ordinary questions answered with misplaced confidence. Catching that requires someone who has actually carried clinical responsibility. The standard these systems get held to is written by physicians, and here you would be writing it.
We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.
Finding a great startup job should not feel like a second job.
Weekday AI brings the most exciting roles from premium YC and VC-backed startups into one place. Our team hand-picks every listing, so you skip the agency spam and the dead links and go straight to companies building things that matter.
Why people use Weekday AI?
Curated roles from premium YC and VC-backed startups
Fresh jobs every week, picked by humans
AI that helps you apply in minutes
Free to join
Your next role should come from a company you are proud to work at. We help you find it.
Browse live roles at jobs.weekday.works
Founded
2022
Company size
11-50 employees
Industry
Technology, Information and Internet
Org type
Privately Held
Finding a great startup job should not feel like a second job.
Weekday AI brings the most exciting roles from premium YC and VC-backed startups into one place. Our team hand-picks every listing, so you skip the agency spam and the dead links and go straight to companies building things that matter.
Why people use Weekday AI?
Curated roles from premium YC and VC-backed startups
Fresh jobs every week, picked by humans
AI that helps you apply in minutes
Free to join
Your next role should come from a company you are proud to work at. We help you find it.
Browse live roles at jobs.weekday.works
Founded
2022
Company size
11-50 employees
Industry
Technology, Information and Internet
Org type
Privately Held