AI Evaluators: Assessing A Shopping Assistant
$50/hr as listed
- Listed pay
- $50/hr as listed
- Field
- Generalist
- Languages
- English
- Where
- United States
- Type
- Contract, remote
- Posted on Terac
- First seen here
- 9 October 2026
Summary
We're hiring AI evaluators to assess the accuracy and helpfulness of a new digital shopping assistant. This project focuses on understanding how well the system handles real-world e-commerce...
From the Terac listing
We're hiring AI evaluators to assess the accuracy and helpfulness of a new digital shopping assistant. This project focuses on understanding how well the system handles real-world e-commerce queries and where it falls short in its logic. Your analysis will directly feed into improving the underlying model and its response quality.
How It Works
You will review real interaction traces between users and the shopping assistant within our custom platform. As you analyze these conversations, you will pinpoint specific failures, logical errors, or unhelpful product recommendations. From there, you will create structured rubrics and verifiers to consistently judge future response quality. This is an ongoing remote engagement requiring 20+ hours per week.
Who This Is For
This opportunity is ideal for quality assurance specialists, AI data evaluators, and e-commerce professionals with a strong eye for detail. We welcome applicants with prior experience in prompt engineering, complex data annotation, or software testing. You should be comfortable analyzing text interactions deeply and building structured evaluation frameworks from scratch.
What You'll Do
- Review real user interaction traces with an AI shopping assistant
- Identify logical failures, inaccuracies, or poor recommendations in the text
- Create structured rubrics and verifiers to judge response quality
- Commit to 20+ hours per week of evaluation work on our internal platform
Who Should Apply
- Experience in data evaluation, quality assurance, or AI training
- Strong analytical skills with the ability to spot subtle errors in text
- Familiarity with e-commerce search and digital shopping experiences
- Ability to commit to a sustained workload of 20+ hours per week
Compensation
$50 per hour
Ready to participate?
About Terac
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Learn more at terac.com or on YouTube at @jointerac.
Text above is the platform's own listing, shown as published. Check the details on Terac before applying.
Prepare for the assessment
Practice questions and what each platform says about its screening for this kind of role:
Terms in this listing
- Rubric
- A written list of criteria and scores used to judge a response, for example accuracy, instruction following and tone, each with clear pass or fail descriptions.
- Prompt
- The input given to a model: a question, an instruction or a conversation so far.
- Annotation
- Adding labels or notes to data such as text, images, audio or video so a model can learn from it.
How applying on Terac works
You create an expert account on Terac and then apply to individual opportunities. Terac's documentation says anyone with an account can take part in its referral scheme without having completed opportunities themselves (referral policy). A sitemap sample also showed that closed listings can stay online and read "This listing has closed", so a stale page is not a live opening (Terac sitemap).
Related roles
Brand Evaluators: 15-20-Minute Image Evaluation Task
$85/hr as listed
- Generalist
- English
- United States
General Participants: Paid Video Recording Study
$2/hr as listed
- Generalist
- English
- Brazil
IT and Business Managers: Paid Interview on Technology Operations
$150/task as listed
- Generalist
- English
- United States
Mid-to-Senior Policy Professionals: Survey on Current Issues
$95/task as listed
- Generalist
- English
- United States
Luxury Fashion Consumers: Ready-To-Wear Shopping Experiences
$105/task as listed
- Generalist
- English
- United States