LILT is seeking experienced software engineers to design, build, and validate benchmarks for multilingual models. This remote, freelance role focuses on creating high-quality tasks in the candidate’s native language, evaluating agents, and writing deterministic verifier scripts.
You’ll work across data, prompts, and evaluation rubrics, collaborating with a global linguistics and ML team. Idóneo candidates have 5+ years of software engineering experience, strong Python and shell skills, and a deep
📌 Remote Native-Language AI Benchmark Engineer (Argentina)
🏢 Lilt
📍 Argentina
Postulate a este anuncio
Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.