SWE-Bench Task Auditor
Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab’s models. You’ll assess repository-level tasks, reference patches, test harnesses, and grading integrity — and provide clear, rubric-based written feedback.
Basic Qualifications
• 3+ years professional software engineering
• Real open-source contribution or maintainer experience (merged PRs, committer / maintainer roles)
• Strong ability to audit reference patches, test runners, and Docker isolation, and to detect answer leakage / reward hacking
• Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++)
Preferred Qualifications
• Familiarity with SWE-Bench (Verified) or similar repository benchmarks
• Maintainer history on major Python OSS (Django, Flask, scikit-learn, sympy, pytest, etc.)
• Prior code-review or task-grading experience
Compensation
- Pay: $70 – $90/hour
- Type: Hourly contract
- Location: Remote — United States
In-depth analysis: how it works, pay rates, pros & cons, and tips to get hired.
