Senior Data Scientist – AI/LLM Evaluation

Square One Resources REMOTE 2026-08-28
  • Proven experience as a Senior Data Scientist or in a similar role.
  • Strong data science, statistical analysis, and experimentation skills.
  • Hands-on experience with AI/LLM evaluation.
  • Strong understanding of LLMs and agentic workflows.
  • Experience designing, validating, and optimizing evaluation metrics for AI/ML models.
  • Strong analytical skills and experience investigating model errors, including false positives and false negatives.
  • Experience working with ground-truth datasets and defining quality criteria.
  • Ability to independently analyze experimental results and translate findings into actionable recommendations.
  • Strong problem-solving and analytical mindset.
Nice to Have
  • Basic knowledge of 3D, geometry, or rendering.
  • Experience with LLM-as-a-Judge / LLM-based evaluation systems.
  • Experience with human-in-the-loop evaluation methodologies.
  • Previous experience evaluating AI agents or agentic systems.
  • Familiarity with advanced approaches to evaluating LLM-generated content and outputs.
We are looking for a Senior Data Scientist to join an international project focused on Artificial Intelligence, Large Language Models (LLMs), and AI evaluation.
The project focuses on developing and improving methodologies for evaluating the quality and performance of AI models and agentic workflows. The role involves designing evaluation metrics, building ground-truth datasets, analyzing model performance, and developing advanced approaches to automated and human-feedback-based evaluation.
,[Design and validate evaluation metrics and ground-truth datasets., Analyze evaluation quality, including accuracy, false positives, and false negatives., Optimize metric normalization and scoring methodologies., Develop and improve LLM-based judges for automated evaluation., Design and implement human-feedback-based evaluation approaches., Explore and develop new evaluation metrics, including:, complexity,, prompt adherence,, output quality and correctness., Analyze experiment results and identify opportunities to improve model performance and evaluation methodologies., Collaborate with Data Science, AI/ML, and Engineering teams to develop evaluation solutions for AI models and agentic workflows.] Requirements: Data science, Statistical analysis, Experimantation skills, AI, LLM, Agentic workflow, 3D, Geometry, Rendering, LLM-as-a-Judge, LLM-based evaluation Additionally: Private healthcare, Sport subscription.