AI Training Roles

AI training glossary

AI training jobs use a lot of shorthand. These are plain definitions of the terms you will meet in postings, guidelines and screenings, written for people new to the work.

Adversarial prompt
A prompt written to expose a model's weak spots, such as tricky reasoning, misleading premises or edge cases.
Related: Red teaming, Stump the model
AI interview
A screening interview run by an AI system instead of a person, usually spoken or on video, with questions about your background and skills.
Related: Assessment
Annotation
Adding labels or notes to data such as text, images, audio or video so a model can learn from it.
Related: Labeling, Bounding box, Ground truth
Assessment
Any screening step a platform uses before giving you work: a quiz, a skills test, a language test, a writing sample or an AI interview.
Related: Qualification test, AI interview
Benchmark
A fixed set of test tasks used to compare models or track progress over time.
Related: Model evaluation, Golden set
Bounding box
A rectangle drawn around an object in an image or video frame to mark where it is.
Related: Annotation, Segmentation
Calibration
Exercises where raters do the same tasks and compare results with the expected answers, so everyone applies the guidelines the same way.
Related: Golden set, Inter-rater agreement
Chain of thought
The step-by-step reasoning a model writes before its final answer. Some tasks ask you to check each step, not only the result.
Related: Process supervision
Data collection
Projects where you create new data, such as photos, videos, recordings or written samples, instead of judging existing data.
Related: Speech data, Consent form
Demonstration data
Example answers written by people to show the model what an ideal response looks like. It is the main input for SFT.
Related: SFT, Golden response
Domain expert
Someone hired for professional knowledge in a field, such as law, medicine or engineering, to write or judge expert-level tasks.
Related: Stump the model, Rubric
Edge case
An unusual item the guidelines do not clearly cover. Projects usually ask you to flag it rather than guess.
Related: Guidelines
Fact-checking
Checking each claim in a model's answer against reliable sources and marking what is wrong or unsupported.
Related: Hallucination, Grounding
Golden response
A model answer written or rewritten by an expert to be the best possible reply to a prompt.
Related: Demonstration data, Rewrite
Golden set
A set of tasks with answers already agreed by experts, mixed into your queue to measure accuracy and consistency.
Related: Quality check, Calibration
Ground truth
The correct answer or label that model output is compared against.
Related: Golden set, Annotation
Grounding
Whether an answer is supported by a given source, such as a document, search result or retrieved passage.
Related: Hallucination, RAG
Guidelines
The project's written instructions that define how to do a task and how to judge answers. On most projects they take precedence over personal preference.
Related: Rubric, Edge case
Hallucination
When a model states something false or made up as if it were true, such as a fake citation or an invented fact.
Related: Fact-checking, Grounding
Harm categories
The named types of harmful content a project tracks, such as hate speech, dangerous instructions or medical misinformation.
Related: Safety policy, Red teaming
Hourly rate
Pay set per hour worked, usually tracked by a timer or time log on the platform. A listed rate is not a promise of hours.
Related: Pay per task
Instruction following
Whether a model did exactly what was asked, including constraints such as length, format, language or things to avoid.
Related: Rubric, System prompt
Inter-rater agreement
How often different raters give the same judgement on the same item. Low agreement usually means unclear guidelines or inconsistent work.
Related: Calibration, Rubric
Jailbreak
A prompt designed to get a model to ignore its safety rules, for example by role-play or by hiding the real request.
Related: Red teaming, Adversarial prompt
Labeling
Assigning a category or tag to an item, for example marking a review as positive or negative. Often used as a synonym for annotation.
Related: Annotation, Taxonomy
Likert scale
A scale of ordered options, such as strongly disagree to strongly agree, used for ratings.
Related: Rating scale
Locale
A specific language and region combination, such as Spanish for Mexico or Portuguese for Brazil. Many language roles are tied to one locale.
Related: Localization
Localization
Adapting content to a specific language and culture, including units, dates, idioms and references, not just translating words.
Related: Locale, Post-editing
Model evaluation
Testing how well a model performs on a set of tasks, often by people scoring answers against a rubric.
Related: Rubric, Benchmark
Native proficiency
Command of a language at the level of someone who grew up speaking it, which many language roles require.
Related: Locale
Pairwise comparison
A preference task with exactly two answers, where you pick the better one or mark them as equal, sometimes on a scale.
Related: Preference ranking, Likert scale
Pay per task
Pay set per completed item rather than per hour, so actual earnings depend on your speed and on task availability.
Related: Hourly rate
Post-editing
Correcting machine-translated text so it reads accurately and naturally.
Related: Localization, Rewrite
Preference ranking
A task where you see two or more model answers to the same prompt and order them from best to worst, usually with a short reason.
Related: Pairwise comparison, RLHF, Rationale
Process supervision
Judging each step of a model's reasoning rather than only the final answer, common in maths, science and coding work.
Related: Chain of thought
Project pause
When a project stops sending tasks, sometimes for days or permanently. Common in AI training work and a reason pay is not guaranteed.
Related: Task queue
Prompt
The input given to a model: a question, an instruction or a conversation so far.
Related: Prompt writing, System prompt
Prompt writing
Creating realistic or challenging prompts for a model, often with a target answer or rubric, to build training or test data.
Related: Prompt, Stump the model
Qualification test
A test you must pass before you can work on a project. It usually checks that you understood the guidelines on a set of sample tasks.
Related: Assessment, Guidelines
Quality check
Any review of your work by a reviewer, an automated test or a golden set. Many projects use quality scores to decide who keeps getting tasks.
Related: Golden set, Reviewer
RAG (Retrieval-augmented generation)
A setup where the model first retrieves documents and then writes an answer based on them. Raters often check whether the answer actually follows the sources.
Related: Grounding
Rating scale
A fixed scale such as 1 to 5 or 1 to 7 used to score answers on one or more criteria.
Related: Likert scale, Rubric
Rationale
The short written explanation you give for a rating or ranking. Reviewers often weigh it as much as the rating itself.
Related: Preference ranking, Rubric
Red teaming
Deliberately trying to make a model produce harmful, false or policy-breaking output, so the weaknesses can be found and fixed.
Related: Jailbreak, Adversarial prompt, Safety policy
Relevance
How well a result or answer matches what the user was looking for.
Related: Search quality rating, User intent
Reviewer
A more experienced worker who checks other people's tasks, gives feedback and scores quality.
Related: Quality check
Reward model
A separate model trained on human preference judgements that predicts how a person would score an answer. It is used to steer the main model during RLHF.
Related: RLHF, Preference ranking
Rewrite
A task where you fix a model's answer so it is correct, complete and follows the guidelines, instead of only scoring it.
Related: Golden response
RLHF (Reinforcement learning from human feedback)
A way of improving a model by having people compare or score its answers, training a reward model on those judgements, then tuning the model to produce answers people prefer.
Related: Reward model, Preference ranking, SFT
Rubric
A written list of criteria and scores used to judge a response, for example accuracy, instruction following and tone, each with clear pass or fail descriptions.
Related: Guidelines, Rating scale, Inter-rater agreement
Safety policy
The rules that define what a model must refuse or handle carefully, such as self-harm, violence or illegal activity.
Related: Red teaming, Harm categories
Search quality rating
Judging how useful and relevant search results, ads or map results are for a given query, following long rater guidelines.
Related: Relevance, User intent
Segmentation
Marking the exact outline of objects in an image, pixel by pixel, rather than with a box.
Related: Bounding box, Annotation
SFT (Supervised fine-tuning)
Training a model on examples of good prompts and ideal answers written or checked by people, so it learns the expected format and quality.
Related: RLHF, Demonstration data
Speech data
Recordings of people speaking, collected to train or test speech recognition and voice models. Projects usually set rules for setup, noise and script reading.
Related: Transcription, Data collection
Stump the model
A task where you write a hard question the model gets wrong and record the correct answer, often used in expert STEM and coding projects.
Related: Adversarial prompt, Ground truth
System prompt
Hidden instructions given to a model before the user's message, setting its role, tone or rules.
Related: Prompt, Instruction following
Talent pool
A list of approved people a platform can invite when a matching project opens. Joining one does not mean work is available yet.
Related: Project pause
Task queue
The list of tasks available to you on a project. It can empty out without warning when a project pauses or ends.
Related: Project pause
Taxonomy
The fixed list of categories a project uses for labeling, often with definitions and examples for each.
Related: Labeling, Guidelines
Timestamping
Marking the start and end times of words, sentences or speakers in a recording.
Related: Transcription
Transcription
Writing down exactly what is said in an audio or video recording, following the project's rules for spelling, fillers, noise and speaker changes.
Related: Timestamping, Speech data
User intent
What a person actually wants when they type a query or prompt, which may differ from the literal words.
Related: Relevance

See the terms in use: interview prep by role and the open roles.