Annotations
Highlight a response and explain what should change. Comments and explicit user corrections can train reward models and guide skill generation.
What an annotation is
An annotation belongs to one trace and, optionally, one step in it. It carries a rating, a comment, or both. The UI creates comments; ratings are available through the API. A trace can have any number of annotations.
| Parameter | Type | Description |
|---|---|---|
| step_id | string | null | null annotates the whole trace. Set it to one of the trace's step ids to annotate that step. |
| rating | 1 | -1 | null | 1 is +, −1 is −, null is a comment with no rating. Not allowed on system steps. |
| comment | string | null | Free text. Actionable feedback can produce training pairs and is also sent to GEPA. |
| quote | object | null | {text, prefix, suffix}: the highlighted span inside the step's content. Requires step_id, and text must appear in that step. |
| label_error | boolean | true when someone has flagged this annotation as wrong. Flagged annotations are excluded from training and GEPA. |
| label_error_note | string | null | Why the label is wrong. |
An annotation needs a rating, a comment, or both. A request that breaks any of the rules in the table gets a 422 that says which one.
API ratings
Rate the whole trace when the outcome is what matters ("resolved the ticket"). Rate a step when one decision is the problem ("called issue_refund before lookup_order"). The two are trained as separate granularities and never paired against each other; see training pairs below.
System steps
You can comment on a system step but not rate it. The reward model never sees system steps: GEPA puts the skill it writes into the system message, and a reward model that read the system message could be satisfied by the skill's text rather than by what the agent does. Comments on a system prompt still reach GEPA as feedback, which is where they're useful.
Quoted spans
Select text inside a step in the app and the comment is anchored to it. The anchor is the selected text plus a short prefix and suffix from around it, the same scheme Stash uses for comments on pages. The surrounding context tells two identical phrases in the same step apart.
{"text": "I've issued a full refund", "prefix": "Sure! ", "suffix": " to your card"}The quote's text must appear in the step's content, or the request fails with a 422. The quote shows reviewers what the comment is about. A step-level rating applies to the whole step either way: training renders the step, not the quoted span.
Flagging label errors
Some labels are wrong. Instead of deleting the annotation, flag it with label_error: true and a note. The annotation stays visible in the trace, and training and GEPA skip it.
curl -s -X PATCH "$STASH_URL/api/v1/rm/annotations/<annotation_id>" \
-H "Authorization: Bearer $STASH_API_KEY" \
-H "Content-Type: application/json" \
-d '{"label_error": true, "label_error_note": "Policy allows refunds on damaged items"}'On each trace summary, label_error_count counts flagged annotations, and positive_count / negative_count leave them out. These count explicit API ratings; feedback-derived pairs are not included.
Annotating over the API
Get step ids from GET /traces/{trace_id}, then post one annotation per call. Omit step_id to annotate the whole trace. The response is the stored annotation, with its id, author_id, author_name, and created_at.
curl -s "$STASH_URL/api/v1/rm/traces/<trace_id>/annotations" \
-H "Authorization: Bearer $STASH_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"step_id": "<step_id>",
"comment": "Promised a refund without checking the policy",
"quote": {"text": "I'"'"'ve issued a full refund", "prefix": "Sure! ", "suffix": " to your card"}
}'This is also how to load labels you already have: import the traces, then post their existing ratings as annotations.
From annotations to training pairs
Manual comments are optional. Automatic judgments provide training supervision; the trained reward model then supplies per-action credit in the trace viewer. Those learned scores stay separate from annotations. See Action credit for scoring new traces and interpreting the heatmap.
The reward model trains on preference pairs: one text that should score higher (chosen) and one that should score lower (rejected). Only the model's selected traces supply pairs. There are two sources: feedback-derived alternatives and explicit API ratings.
Automatic assessments, comments and user corrections
The worker assesses assistant responses and tool calls, including tool choice and arguments, and identifies explicit, actionable feedback about them. Tool calls with no accompanying text are included. Results from tools provide context; they are not assistant decisions or user feedback.
For tool calls, generated alternatives change the tool or its arguments and are independently reviewed against the context available before the call. Later results are excluded from the comparison. Choose View learning on a reward model to see these assessments and their step numbers. These findings do not automatically create comments.
Each pair compares an original response or tool call with a generated alternative in the same context. It records the preference, source ID, evidence quote, and rationale, distinguishing AI judgments from interpretations of user feedback. These are not direct human preference votes.
Silence, a new question, or a tool error is not a preference. Ambiguous feedback is skipped. The cited source must exist, a user correction must follow the response, and a step comment must refer to that response or call. Tool results and system steps cannot be revision targets. Malformed extraction fails the job. Flagged comments are excluded.
Both responses share the original conversation prefix. System steps and later messages, including the correction, are excluded from the training text. The exact pairs and evidence are persisted with the model. Training combines API-rating pairs followed by feedback pairs, caps them at max_pairs (default 4000), and requires at least two usable pairs.
Explicit API ratings
Ratings supplied through the API produce pairs as follows:
1. Drop flagged annotations
Every annotation with label_error = true is removed before anything else.
2. Collapse ratings per target
A target is a whole trace or a single step. Its ratings are summed. A positive sum makes it chosen, a negative sum makes it rejected, and zero skips it. Opposite ratings cancel out.
| target | ratings | sum | result |
|---|---|---|---|
| trace A | +1, +1 | +2 | chosen |
| trace B | −1 | −1 | rejected |
| trace C | +1, −1 | 0 | skipped |
| trace D, step 4 | −1 | −1 | rejected |
| trace E, step 2 | +1, +1, −1 | +1 | chosen |
3. Render each target to text
A trace renders all of its steps except system steps. A step renders the trace's steps up to and including itself, again without system steps, so the model sees the context the agent had. Each step is <role>: <content>, steps are joined with blank lines, and tool calls are written as assistant → <tool_name>(<tool_input json>). The trace on Trace format renders as:
user: I want a refund for order 1182
assistant → lookup_order({"order_id": "1182"})
tool: {"status": "delivered"}
assistant: Your order was delivered on May 3.An assistant step with both text and a tool call renders both lines, text first:
assistant: Let me check that order.
assistant → lookup_order({"order_id": "1182"})4. Pair within the same granularity
Every chosen target is paired with every rejected target of the same granularity: traces with traces, steps with steps. In the table above that gives two pairs: (A, B) and (E step 2, D step 4). The pairs are shuffled with seed 0 and capped at max_pairs (default 4000), so the same labels always produce the same training set.
Exporting annotations
GET /api/v1/rm/export/annotations returns one annotation per line. Flagged annotations are included, with label_error: true, so the export is a complete record.
{"trace_id": "…", "trace_external_id": "…", "step_index": 4, "rating": -1,
"comment": "Promised a refund without checking the policy",
"quote": {"text": "I've issued a full refund", "prefix": "Sure! ", "suffix": " to your card"},
"label_error": false, "label_error_note": null, "author": "Henry Dowling",
"created_at": "2026-09-29T03:12:00+00:00"}| Parameter | Type | Description |
|---|---|---|
| trace_id | string | Stash's id for the trace. |
| trace_external_id | string | null | The id you imported the trace with. |
| step_index | integer | null | Position of the step in the trace, or null for a trace-level annotation. |
| rating | 1 | -1 | null | +1, −1, or no rating. |
| comment | string | null | The reviewer's comment. |
| quote | object | null | {text, prefix, suffix}. |
| label_error | boolean | Whether the annotation is flagged as wrong. |
| label_error_note | string | null | Why it was flagged. |
| author | string | The name of the user who wrote it. |
| created_at | string | ISO 8601. |