Anyone AI Labs is seeking a Research Scientist to lead evaluations and benchmarking for large language models. You will design frontier-grade evaluation methods and benchmarks across reasoning, coding, agents, tool use, and multimodal tasks, grounded in expert-verified truth and validated across multiple models.
You will own the study lifecycle from framing questions to publishing results, with opportunities to contribute to public benchmarks and papers at top venues, while collaborating with
📌 LLM Evaluation Scientist: Frontier Benchmarking (Buenos Aires)
🏢 Anyone AI
📍 Buenos Aires
Postulate a este anuncio
Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.