AI Quality Assurance & Evaluation Specialist
About Gramian
Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.
About the Role
We are seeking an AI Quality & Evaluation Specialist to validate task quality, review AI agent performance, and audit grading logic. You will analyze execution traces, tool calls, reference solutions, and evaluation criteria to identify task defects, grading errors, and unjustified model failures. The ideal candidate combines technical fluency with strong analytical judgment and the ability to provide clear, evidence-based feedback.
CONTRACT: Contractor (Hourly)
COMMITMENT: 40 hours per week
LOCATION: Remote — Latin America (LATAM)
Responsibilities
- Validate task instructions, source materials, reference solutions, and evaluation criteria for consistency and completeness.
- Review AI agent execution traces, tool calls, and generated deliverables to assess whether outcomes are justified.
- Audit grading logic to identify brittle checks, incorrect expected answers, and unsupported rubric criteria.
- Identify cases where valid alternative solutions are unfai