Content Evaluator – Bilingual (French and English)- Flexible Hours
Rate not listed
- Listed pay
- Not listed
- Field
- Languages
- Languages
- French, English
- Where
- United States
- Type
- Contract, remote
- Posted on Innodata
- First seen here
- 9 October 2026
Summary
We are seeking detail-oriented evaluators to conduct human quality evaluations for an enterprise AI customer support product on Instagram, WhatsApp, and Messenger. You will evaluate and benchmark...
From the Innodata listing
Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.
Role Overview
We are seeking detail-oriented evaluators to conduct human quality evaluations for an enterprise AI customer support product on Instagram, WhatsApp, and Messenger. You will evaluate and benchmark AI model responses against complex evaluation rubrics using provided business knowledge bases.
What You’ll Own:
- Model Evaluation: Review and score AI-generated customer interactions across Foundational, Experiential, and Operational dimensions (e.g., Action Fidelity, Faithfulness, Hallucination, Compliance, Tone, and Handoff).
- Intent & Fact Verification: Benchmark both informational (R1) and transactional (R2) customer queries against authoritative business sources (FAQs, product catalogs, SOPs) within the task UI.
- Quality Assurance: Participate in dual-review processes and daily calibration audits to ensure inter-rater agreement and establish ground-truth performance targets.
- Performance Targets: Deliver precise evaluation
You’ll Thrive in This Role If You Have:
- Customer Service Background: Prior experience in customer service, call centers, retail, or handling customer communications via email, chat, or phone (highly prioritized).
- English Proficiency: Exceptional written English skills with a strong command of tone, brand voice, grammar, and nuance.
- French Proficiency: Exceptional written French skills with a strong command of tone, brand voice, grammar, and nuance.
- Analytical Precision: Ability to strictly follow multi-tier evaluation guidelines, complex logic trees, and technical rubrics without deviation.
- Tech Adaptability: Comfort using dedicated web-based tools and labeling interfaces.
Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at https://consumer.ftc.gov/articles/job-scams.
If you believe you’ve been targeted by a recruitment scam, please report it to Innodata at verifyjoboffer@innodata.com and consider reporting it to the FTC at ReportFraud.ftc.gov.
Text above is the platform's own listing, shown as published. Check the details on Innodata before applying.
Prepare for the assessment
Practice questions and what each platform says about its screening for this kind of role:
Terms in this listing
- Benchmark
- A fixed set of test tasks used to compare models or track progress over time.
- Rubric
- A written list of criteria and scores used to judge a response, for example accuracy, instruction following and tone, each with clear pass or fail descriptions.
- Model evaluation
- Testing how well a model performs on a set of tasks, often by people scoring answers against a rubric.
- Hallucination
- When a model states something false or made up as if it were true, such as a fake citation or an invented fact.
- Calibration
- Exercises where raters do the same tasks and compare results with the expected answers, so everyone applies the guidelines the same way.
Related roles
AI Voice Evaluation Specialist
$20-27/hr as listed
- Languages
- English
- United States
Content Evaluator – Bilingual (Vietnamese and English)- Flexible Hours
Up to $20/hr as listed
- Languages
- Vietnamese
- United States
Hebrew Language Transcription Expert
$100/hr as listed
- Languages
- Hebrew
- Worldwide
Korean Language Expert
$65-98/hr as listed
- Languages
- Korean
- 58 countries
Japanese Language Expert
$65-98/hr as listed
- Languages
- Japanese
- 58 countries