ML Challenge Task Auditor

Posted 1 September 2026 Hourly Remote Bi-weekly Payout English
Mercor
Apply on → Mercor
$70 – $90 per hour

Evaluate the quality, correctness, and methodological rigor of applied machine-learning tasks used to train and evaluate a frontier AI lab’s models. You’ll assess experiment design, model-selection reasoning, and evaluation methodology — and provide clear, rubric-based written feedback.

Basic Qualifications
• 3+ years hands-on applied/experimental ML (experiment design, model selection, hyperparameter tuning, evaluation methodology)
• Strong grasp of data-quality rigor: leakage detection, metric gaming, and train/test/CV hygiene
• Proficiency with standard ML frameworks (PyTorch, TensorFlow, scikit-learn, XGBoost)
• Ability to critique ML claims against evidence and reproduce results

Preferred Qualifications
• Competition / benchmark experience (e.g., Kaggle)
• Graduate research or publication record in applied ML
• Prior task-grading or peer-review experience

Note: this role evaluates applied/experimental ML rigor — it is not an LLM-application-building or MLOps role.

Compensation

  • Pay: $70 – $90/hour
  • Type: Hourly contract
  • Location: Remote — United States

3 slots remaining.

Getting Started

New to Remote Gig Work?

No fluff, no theory. The First Month Playbook walks you through profile setup, landing your first client, and building a workflow that actually sticks.

Read the Playbook
New to Remote Gig Work?
Featured Platform

Apply to Mercor

Mercor matches you with AI and tech companies looking for remote talent. One application, multiple opportunities. Affiliate link — we may earn a commission.

Apply Now on Mercor
Apply to Mercor

Browse remote jobs by platform