We're looking for a freelance AI evaluation engineer with experience in software development, test automation, and Full-Stack development to create challenging coding test cases for AI coding systems.
Requirements
- Degree in Computer Science, Software Engineering, or related fields
- 5+ years in software development, primarily Python (pytest, async/await, subprocess, file operations)
- Background in Full-Stack development, with an equal focus on building React-based interfaces and robust Back-end systems
- Experience writing tests (functional, integration – not just running them)
- Docker containers (running evaluations locally in containers)
- CI/CD understanding (GitHub Actions as a user: triggers, labels, reading results)
- English proficiency - B2
Benefits
- Up to $50 per hour equivalent
- Estimated 20 hours of work per project
- Flexible work schedule