04 sep
|
DevFixr
|
Argentina
- Weekly Hours: Minimum 15 – 25 hours per week (Versátil & Asynchronous)
- Start: Immediately. The client has sample work that needs turning around this week, with a substantially larger programme expected to follow.
- Engagement: Contract, paid hourly. No 40-hour minimum — work arrives in sprints, and you take on
what suits you. Please find the job details below:
About the Role
We are seeking exceptional Software Engineers to join our team as AI Benchmark Task Authors . In this role, you will not simply write standard feature code; you will design, engineer, and author complex, production-grade software evaluation environments used to benchmark frontier Artificial Intelligence models (such as GPT-5, Claude, and Gemini).
You will build realistic multi-module code repositories, author ambiguous real-world issue specifications, and write hidden automated test suites designed to stress-test model reasoning, edge-case handling, and architectural correctness.
Key Responsibilities
- Repository Fixture Engineering: Design and build self-contained, dependency-light, multi-module software codebases (1,000–2,000 lines of code) in languages like Python, TypeScript, C++, Java, or Go.
- Specification Authoring: Write clear issue descriptions framing problems or feature requirements from an end-user perspective—without revealing the exact code paths or files that need modification.
- Hidden Test Suite Authoring:
Construct comprehensive hidden unit and integration test suites (pytest, Jest, etc.) that evaluate exact functional correctness, boundary conditions, edge cases, and refusal honesty.
- Gold & Partial Solution Authoring: Implement 100% correct "Gold" reference solutions alongside plausible, imperfect "partial credit" implementations to ensure test suites measure a gradient of model understanding.
- Quality & Contamination Assurance: Ensure all authored fixtures are clean, leak-free, and deterministic with zero network or wall-clock dependencies.
Key Requirements & Qualifications
- Core Backend Proficiency: Strong hands-on development experience in at least one primary backend language (Python, Node.js/TypeScript, Java, C#, C++, or Go).
- Rigorous Testing Mindset: Deep experience writing unit, integration, and edge-case test suites. You know how to break code and catch silent failures, race conditions, or lossy state conversions.
- Architectural Literacy: Ability to understand emergent system behaviors, data pipeline transformations, and complex algorithms.
- Autonomous & Asynchronous: Ability to work independently without hand-holding, document assumptions clearly, and manage project deliverables self-sufficiently.
- Fluency in English: Strong written and verbal communication skills for writing technical task specifications and collaborating asynchronously.
📌 AI Benchmark Task Authors (Argentina)
🏢 DevFixr
📍 Argentina