01 ago
|
Anyone AI
|
Buenos Aires
01 ago
Anyone AI
Buenos Aires
Anyone AI Labs is seeking a Research Scientist to lead evaluations and benchmarking for large language models.
You will design frontier-grade evaluation methods and benchmarks across reasoning, coding, agents, tool use, and multimodal tasks, grounded in expert-verified truth and validated across multiple models.You will own the study lifecycle from framing questions to publishing results, with opportunities to contribute to public benchmarks and papers at top venues, while collaborating with
#J-*****-Ljbffr
📌 Llm Evaluation Scientist: Frontier Benchmarking (Buenos Aires)
🏢 Anyone AI
📍 Buenos Aires