Job summary
This freelance role involves designing and evaluating tasks to assess AI coding agents' performance, focusing on safety and adherence to scope rather than just task completion. The position requires strong software development experience with Python and JavaScript/TypeScript, and the ability to create realistic coding scenarios that challenge AI agents while distinguishing safe from unsafe solutions. English proficiency at B2 level or higher is required.
Location and work setup
- Location
- Norway
- Work setup
- Remote
Salary
USD 75.00–75.00 hour
Responsibilities
You will create realistic development environments with codebases, infrastructure, and context to simulate real development histories. Design tasks that pair legitimate development goals with potential unsafe shortcuts, then develop tests to verify if AI agents complete tasks correctly and safely, not just functionally. Continuously refine tasks and tests based on quality assurance feedback and agent solution analysis to ensure evaluations are accurate and comprehensive.
Qualifications
Candidates should have 4 to 5+ years in software development, proficient in Python and JavaScript/TypeScript, with strong skills in designing functional and integration tests that distinguish safe from unsafe completions. Experience using AI coding agents like Claude Code, GitHub Copilot CLI, or Codex, and familiarity with GitHub pull requests and continuous integration workflows are necessary. Broader backend and infrastructure knowledge is beneficial but not mandatory. English proficiency at B2 level or above is required.