Search⌘ K
AI Features

Preparing a Dataset

Explore techniques for preparing datasets that ensure clear prompt and target boundaries, prevent validation leakage, and produce stable fine-tuned AI models. Understand data normalization, duplicate grouping, and masking strategies to achieve accurate training outcomes and improve model generalization.

A fine-tuning run only optimizes against the tokens we mark as targets, so unclear boundaries between prompt and response produce ambiguous supervision. Here is one example that looks readable to a human but is hard for a training loop to separate into input versus target.

Instruction: Grade the student answer.
Rubric: 2 points if it mentions photosynthesis and sunlight.
Student: Plants make food using water.
Answer: 1/2. Missing sunlight and photosynthesis.
Unstructured prompt lacking explicit boundaries between input instructions and target completion tokens

Here is the same example with explicit role delimiters and a clear target, so the loss can be applied only to the assistant completion.

<|user|>
Grade the student answer using the rubric.
Rubric: 2 points if it mentions photosynthesis and sunlight.
Student: Plants make food using water.
<|assistant|>
Score: 1/2
Reason: Missing sunlight and photosynthesis.
Explicit user and assistant tags isolate loss calculation strictly to the target completion

Nothing crashed in the first format. The bug is that the model can be rewarded for copying prompt text, because the target boundary is not encoded.

What the dataset is supposed to teach

With boundary control in place, the next decision is what behavior this dataset should encode. For this chapter, treat fine-tuning data as a way to teach stable behavior such as output format, tone, refusal policy, or decision rules, not fast-changing facts that will go stale. A practical schema keeps each example inspectable and makes it obvious what the target tokens are. Two common patterns are an instruction style record and a messages style record.{“instruction”: “...”, “input”: “...”, “output”: “...”} where the training code concatenates instruction and  ...