AI training glossary
AI training jobs use a lot of shorthand. These are plain definitions of the terms you will meet in postings, guidelines and screenings, written for people new to the work.
- Adversarial prompt
- A prompt written to expose a model's weak spots, such as tricky reasoning, misleading premises or edge cases.
- Related: Red teaming, Stump the model
- AI interview
- A screening interview run by an AI system instead of a person, usually spoken or on video, with questions about your background and skills.
- Related: Assessment
- Annotation
- Adding labels or notes to data such as text, images, audio or video so a model can learn from it.
- Related: Labeling, Bounding box, Ground truth
- Assessment
- Any screening step a platform uses before giving you work: a quiz, a skills test, a language test, a writing sample or an AI interview.
- Related: Qualification test, AI interview
- Benchmark
- A fixed set of test tasks used to compare models or track progress over time.
- Related: Model evaluation, Golden set
- Bounding box
- A rectangle drawn around an object in an image or video frame to mark where it is.
- Related: Annotation, Segmentation
- Calibration
- Exercises where raters do the same tasks and compare results with the expected answers, so everyone applies the guidelines the same way.
- Related: Golden set, Inter-rater agreement
- Chain of thought
- The step-by-step reasoning a model writes before its final answer. Some tasks ask you to check each step, not only the result.
- Related: Process supervision
- Consent form
- A document you sign agreeing to how your recordings, images or other data will be used, for example to train or synthesize voices.
- Related: Data collection
- Data collection
- Projects where you create new data, such as photos, videos, recordings or written samples, instead of judging existing data.
- Related: Speech data, Consent form
- Demonstration data
- Example answers written by people to show the model what an ideal response looks like. It is the main input for SFT.
- Related: SFT, Golden response
- Domain expert
- Someone hired for professional knowledge in a field, such as law, medicine or engineering, to write or judge expert-level tasks.
- Related: Stump the model, Rubric
- Edge case
- An unusual item the guidelines do not clearly cover. Projects usually ask you to flag it rather than guess.
- Related: Guidelines
- Fact-checking
- Checking each claim in a model's answer against reliable sources and marking what is wrong or unsupported.
- Related: Hallucination, Grounding
- Golden response
- A model answer written or rewritten by an expert to be the best possible reply to a prompt.
- Related: Demonstration data, Rewrite
- Golden set
- A set of tasks with answers already agreed by experts, mixed into your queue to measure accuracy and consistency.
- Related: Quality check, Calibration
- Ground truth
- The correct answer or label that model output is compared against.
- Related: Golden set, Annotation
- Grounding
- Whether an answer is supported by a given source, such as a document, search result or retrieved passage.
- Related: Hallucination, RAG
- Guidelines
- The project's written instructions that define how to do a task and how to judge answers. On most projects they take precedence over personal preference.
- Related: Rubric, Edge case
- Hallucination
- When a model states something false or made up as if it were true, such as a fake citation or an invented fact.
- Related: Fact-checking, Grounding
- Harm categories
- The named types of harmful content a project tracks, such as hate speech, dangerous instructions or medical misinformation.
- Related: Safety policy, Red teaming
- Hourly rate
- Pay set per hour worked, usually tracked by a timer or time log on the platform. A listed rate is not a promise of hours.
- Related: Pay per task
- Instruction following
- Whether a model did exactly what was asked, including constraints such as length, format, language or things to avoid.
- Related: Rubric, System prompt
- Inter-rater agreement
- How often different raters give the same judgement on the same item. Low agreement usually means unclear guidelines or inconsistent work.
- Related: Calibration, Rubric
- Jailbreak
- A prompt designed to get a model to ignore its safety rules, for example by role-play or by hiding the real request.
- Related: Red teaming, Adversarial prompt
- Labeling
- Assigning a category or tag to an item, for example marking a review as positive or negative. Often used as a synonym for annotation.
- Related: Annotation, Taxonomy
- Likert scale
- A scale of ordered options, such as strongly disagree to strongly agree, used for ratings.
- Related: Rating scale
- Locale
- A specific language and region combination, such as Spanish for Mexico or Portuguese for Brazil. Many language roles are tied to one locale.
- Related: Localization
- Localization
- Adapting content to a specific language and culture, including units, dates, idioms and references, not just translating words.
- Related: Locale, Post-editing
- Model evaluation
- Testing how well a model performs on a set of tasks, often by people scoring answers against a rubric.
- Related: Rubric, Benchmark
- Native proficiency
- Command of a language at the level of someone who grew up speaking it, which many language roles require.
- Related: Locale
- Pairwise comparison
- A preference task with exactly two answers, where you pick the better one or mark them as equal, sometimes on a scale.
- Related: Preference ranking, Likert scale
- Pay per task
- Pay set per completed item rather than per hour, so actual earnings depend on your speed and on task availability.
- Related: Hourly rate
- Post-editing
- Correcting machine-translated text so it reads accurately and naturally.
- Related: Localization, Rewrite
- Preference ranking
- A task where you see two or more model answers to the same prompt and order them from best to worst, usually with a short reason.
- Related: Pairwise comparison, RLHF, Rationale
- Process supervision
- Judging each step of a model's reasoning rather than only the final answer, common in maths, science and coding work.
- Related: Chain of thought
- Project pause
- When a project stops sending tasks, sometimes for days or permanently. Common in AI training work and a reason pay is not guaranteed.
- Related: Task queue
- Prompt
- The input given to a model: a question, an instruction or a conversation so far.
- Related: Prompt writing, System prompt
- Prompt writing
- Creating realistic or challenging prompts for a model, often with a target answer or rubric, to build training or test data.
- Related: Prompt, Stump the model
- Qualification test
- A test you must pass before you can work on a project. It usually checks that you understood the guidelines on a set of sample tasks.
- Related: Assessment, Guidelines
- Quality check
- Any review of your work by a reviewer, an automated test or a golden set. Many projects use quality scores to decide who keeps getting tasks.
- Related: Golden set, Reviewer
- RAG (Retrieval-augmented generation)
- A setup where the model first retrieves documents and then writes an answer based on them. Raters often check whether the answer actually follows the sources.
- Related: Grounding
- Rating scale
- A fixed scale such as 1 to 5 or 1 to 7 used to score answers on one or more criteria.
- Related: Likert scale, Rubric
- Rationale
- The short written explanation you give for a rating or ranking. Reviewers often weigh it as much as the rating itself.
- Related: Preference ranking, Rubric
- Red teaming
- Deliberately trying to make a model produce harmful, false or policy-breaking output, so the weaknesses can be found and fixed.
- Related: Jailbreak, Adversarial prompt, Safety policy
- Relevance
- How well a result or answer matches what the user was looking for.
- Related: Search quality rating, User intent
- Reviewer
- A more experienced worker who checks other people's tasks, gives feedback and scores quality.
- Related: Quality check
- Reward model
- A separate model trained on human preference judgements that predicts how a person would score an answer. It is used to steer the main model during RLHF.
- Related: RLHF, Preference ranking
- Rewrite
- A task where you fix a model's answer so it is correct, complete and follows the guidelines, instead of only scoring it.
- Related: Golden response
- RLHF (Reinforcement learning from human feedback)
- A way of improving a model by having people compare or score its answers, training a reward model on those judgements, then tuning the model to produce answers people prefer.
- Related: Reward model, Preference ranking, SFT
- Rubric
- A written list of criteria and scores used to judge a response, for example accuracy, instruction following and tone, each with clear pass or fail descriptions.
- Related: Guidelines, Rating scale, Inter-rater agreement
- Safety policy
- The rules that define what a model must refuse or handle carefully, such as self-harm, violence or illegal activity.
- Related: Red teaming, Harm categories
- Search quality rating
- Judging how useful and relevant search results, ads or map results are for a given query, following long rater guidelines.
- Related: Relevance, User intent
- Segmentation
- Marking the exact outline of objects in an image, pixel by pixel, rather than with a box.
- Related: Bounding box, Annotation
- SFT (Supervised fine-tuning)
- Training a model on examples of good prompts and ideal answers written or checked by people, so it learns the expected format and quality.
- Related: RLHF, Demonstration data
- Speech data
- Recordings of people speaking, collected to train or test speech recognition and voice models. Projects usually set rules for setup, noise and script reading.
- Related: Transcription, Data collection
- Stump the model
- A task where you write a hard question the model gets wrong and record the correct answer, often used in expert STEM and coding projects.
- Related: Adversarial prompt, Ground truth
- System prompt
- Hidden instructions given to a model before the user's message, setting its role, tone or rules.
- Related: Prompt, Instruction following
- Talent pool
- A list of approved people a platform can invite when a matching project opens. Joining one does not mean work is available yet.
- Related: Project pause
- Task queue
- The list of tasks available to you on a project. It can empty out without warning when a project pauses or ends.
- Related: Project pause
- Taxonomy
- The fixed list of categories a project uses for labeling, often with definitions and examples for each.
- Related: Labeling, Guidelines
- Timestamping
- Marking the start and end times of words, sentences or speakers in a recording.
- Related: Transcription
- Transcription
- Writing down exactly what is said in an audio or video recording, following the project's rules for spelling, fillers, noise and speaker changes.
- Related: Timestamping, Speech data
- User intent
- What a person actually wants when they type a query or prompt, which may differ from the literal words.
- Related: Relevance
See the terms in use: interview prep by role and the open roles.