API reference
Everything in the reward model app is available over REST under /api/v1/rm.
Authentication
Send a Stash API key as a bearer token. Create one on the API Keys page. Every endpoint reads and writes only the caller's own traces, annotations, models, and runs. If you run your own Stash, set STASH_URL to your backend instead.
export STASH_URL=https://api.joinstash.ai
export STASH_API_KEY="<your key>"
curl -s "$STASH_URL/api/v1/rm/formats" -H "Authorization: Bearer $STASH_API_KEY"Examples on this page use those two variables. Request and response bodies are JSON unless noted.
Reward-model APIs require an account enrolled in the experiment. Accounts that existed at rollout remain disabled and receive 404 from every endpoint here. New accounts are enabled by default; eligibility is per account, including new accounts in existing organizations.
Errors
Errors return {"detail": "…"}. An id that belongs to another user is a 404, the same as one that doesn't exist. A request Stash can't act on, such as an unparseable import, a rating on a system step, or a rejected SQL query, is a 422 whose detail says why.
| method | path | summary |
|---|---|---|
| GET | /formats | List import formats |
| POST | /traces/import | Import traces |
| GET | /traces | List traces |
| GET | /traces/{trace_id} | Get a trace with steps, annotations, scores |
| DELETE | /traces/{trace_id} | Delete a trace |
| POST | /traces/{trace_id}/annotations | Annotate a trace or step |
| PATCH | /annotations/{annotation_id} | Update an annotation or flag a label error |
| DELETE | /annotations/{annotation_id} | Delete an annotation |
| GET | /export/traces | Export traces as JSONL |
| GET | /export/annotations | Export annotations as JSONL |
| GET | /export/pairs | Export training pairs as JSONL |
| POST | /reward-models | Train a reward model |
| GET | /reward-models | List reward models |
| GET | /reward-models/{id} | Get a reward model |
| GET | /reward-models/{id}/weights | Download the trained weights |
| POST | /gepa-runs | Start a GEPA run |
| GET | /gepa-runs | List GEPA runs |
| GET | /gepa-runs/{id} | Get a GEPA run |
| GET | /gepa-runs/{id}/skill | Download the best skill as SKILL.md |
| POST | /query | Run read-only SQL |
Formats
GET/api/v1/rm/formatsList the import formats this server accepts.[{"name": "stash", "description": "…"}, {"name": "openai_chat", "description": "…"}, …]Traces
POST/api/v1/rm/traces/importImport one payload of traces.| Parameter | Type | Description |
|---|---|---|
| format* | string | A format name from /formats, or auto. |
| data* | string | The file contents, as one string. |
jq -Rs '{format: "auto", data: .}' traces.jsonl \
| curl -s "$STASH_URL/api/v1/rm/traces/import" \
-H "Authorization: Bearer $STASH_API_KEY" \
-H "Content-Type: application/json" \
--data @-{"format": "openai_chat", "imported": 128, "trace_ids": ["…", …]}format in the response is the format that was used, which tells you what auto detected. The import is all or nothing: any error is a 422 and stores no traces. That includes auto matching no format, a payload that doesn't parse as its format, a trace with only system steps, and a trace with no title and no user step. Re-importing a trace with the same id replaces its steps and deletes the annotations on them. See Trace format.
GET/api/v1/rm/tracesList traces, paginated.| Parameter | Type | Description |
|---|---|---|
| limit | integer | Page size, 1 to 500. Default 50. |
| offset | integer | Rows to skip. Default 0. |
curl -s "$STASH_URL/api/v1/rm/traces?limit=50&offset=0" -H "Authorization: Bearer $STASH_API_KEY"Returns {"traces": [TraceSummary], "total": int}, newest first.
GET/api/v1/rm/traces/{trace_id}A trace with its steps, annotations, and reward model scores.curl -s "$STASH_URL/api/v1/rm/traces/<trace_id>" -H "Authorization: Bearer $STASH_API_KEY"Returns a TraceDetail. Step ids for annotating come from this response.
POST/api/v1/rm/traces/{trace_id}/scoreScore assistant actions with a saved trained model.Send {} for the current shared evaluator, or{"reward_model_id": "<model_id>"} for a personal or previously released shared model. Returns 202 with a scoring job (id, trace_id, reward_model_id, status, error and timestamps). Poll trace detail's scoring_runs for its status and action_scores for results. Repeated requests reuse an active job. Unowned traces or unreleased models owned by others return 404; models without action training, unfinished models and traces without assistant actions return 422. This invokes the trained checkpoint, without LLM labeling or retraining.
GET/api/v1/rm/evaluatorCurrent shared release, or null before bootstrap.Returns {"default": {"id": "…", "name": "…", "revision": 1, "updated_at": "…"}}. Shared weights and pooled training data are not exposed.
PATCH/api/v1/rm/traces/{trace_id}/training-contributionSet permission to use this trace in shared training.Send {"allowed": true} to opt in, or false to revoke future contribution. Scoring does not require contribution permission. Revocation does not unlearn released models.
DELETE/api/v1/rm/traces/{trace_id}Delete a trace and its annotations. Returns 204.curl -s -X DELETE "$STASH_URL/api/v1/rm/traces/<trace_id>" -H "Authorization: Bearer $STASH_API_KEY"Annotations
POST/api/v1/rm/traces/{trace_id}/annotationsAnnotate a whole trace, or one step.| Parameter | Type | Description |
|---|---|---|
| step_id | string | A step of this trace. Omit to annotate the whole trace. |
| rating | 1 | -1 | + or −. Not allowed on a system step. |
| comment | string | Free text. An annotation needs a rating, a comment, or both. |
| quote | object | {text, prefix, suffix}: the highlighted span. Requires step_id, and text must appear in the step's content. |
curl -s "$STASH_URL/api/v1/rm/traces/<trace_id>/annotations" \
-H "Authorization: Bearer $STASH_API_KEY" \
-H "Content-Type: application/json" \
-d '{"step_id": "<step_id>", "rating": -1, "comment": "Skipped the policy check"}'Returns the Annotation. Breaking any rule in the table is a 422.
PATCH/api/v1/rm/annotations/{annotation_id}Change a rating or comment, or flag a label error.Only the fields you send change. The step and quote are fixed once created.
| Parameter | Type | Description |
|---|---|---|
| rating | 1 | -1 | null | New rating. |
| comment | string | null | New comment. |
| label_error | boolean | true excludes the annotation from training and GEPA. |
| label_error_note | string | Why the label is wrong. |
curl -s -X PATCH "$STASH_URL/api/v1/rm/annotations/<annotation_id>" \
-H "Authorization: Bearer $STASH_API_KEY" \
-H "Content-Type: application/json" \
-d '{"label_error": true, "label_error_note": "Refund was within policy"}'DELETE/api/v1/rm/annotations/{annotation_id}Delete an annotation. Returns 204.curl -s -X DELETE "$STASH_URL/api/v1/rm/annotations/<annotation_id>" -H "Authorization: Bearer $STASH_API_KEY"Export
All three return JSONL (application/x-ndjson), one object per line, for everything you own.
GET/api/v1/rm/export/tracesTraces in the Stash Trace Format. Re-importable as-is.GET/api/v1/rm/export/pairsThe training pairs your current labels produce, capped at 4000: {"chosen": str, "rejected": str, "trace_ids": [str], "granularity": str, "action_type": str}.curl -s "$STASH_URL/api/v1/rm/export/traces" -H "Authorization: Bearer $STASH_API_KEY" > traces.jsonl
curl -s "$STASH_URL/api/v1/rm/export/annotations" -H "Authorization: Bearer $STASH_API_KEY" > annotations.jsonl
curl -s "$STASH_URL/api/v1/rm/export/pairs" -H "Authorization: Bearer $STASH_API_KEY" > pairs.jsonlThe pairs export contains pairs built from explicit API ratings. It does not run feedback extraction or export a model's saved feedback-derived training evidence.
Reward models
POST/api/v1/rm/reward-modelsQueue a training job.| Parameter | Type | Description |
|---|---|---|
| name* | string | Display name. |
| trace_ids* | string[] | At least one trace UUID owned by you. Only these traces supply training pairs. |
| base_model | string | Hugging Face model id. Default Qwen/Qwen3-0.6B. |
| epochs | integer | Training epochs. Default 1. |
| max_pairs | integer | Cap on training pairs. Default 4000. |
curl -s "$STASH_URL/api/v1/rm/reward-models" \
-H "Authorization: Bearer $STASH_API_KEY" \
-H "Content-Type: application/json" \
-d '{"name": "refund-policy", "trace_ids": ["<trace_id>"]}'Returns a RewardModel with status queued. Feedback extraction and the minimum of two usable pairs are checked in the worker. The deployment sets RM_COMPUTE; sending compute or any unknown request field returns 422. An unowned trace ID returns 404.
GET/api/v1/rm/reward-modelsList your reward models, newest first.GET/api/v1/rm/reward-models/{id}Get one reward model. Poll this for status and metrics.curl -s "$STASH_URL/api/v1/rm/reward-models/<id>" -H "Authorization: Bearer $STASH_API_KEY"GET/api/v1/rm/reward-models/{id}/weightsGet a private checkpoint download URL.Returns {"url": "https://…"} after checking ownership and succeeded status; otherwise 404. The URL expires after five minutes. Download from that URL without forwarding your Stash API key. The gzip archive contains a reward-model/ directory with model weights, tokenizer files, reward_stats.json, and action_reward_stats.json when trained on action comparisons.
set -euo pipefail
reward_weights_url="$(curl -fsS "$STASH_URL/api/v1/rm/reward-models/<id>/weights" -H "Authorization: Bearer $STASH_API_KEY" | jq -er '.url')"
curl -fL "$reward_weights_url" --output reward-model.tar.gz
tar xzf reward-model.tar.gzGEPA runs
POST/api/v1/rm/gepa-runsQueue a GEPA run that writes a skill.| Parameter | Type | Description |
|---|---|---|
| reward_model_id* | string | One of your reward models with status succeeded, used as the metric. Another user's is a 404; an unfinished one is a 422. |
| task_model | string | LiteLLM task model. Default anthropic/claude-haiku-4-5. |
| task_api_base | string | OpenAI-compatible base URL (vLLM, SGLang, …) for the task model. |
| reflection_model | string | Model that writes skill bodies. Default anthropic/claude-sonnet-5. |
| max_metric_calls | integer | Budget of example evaluations. Default 40. |
curl -s "$STASH_URL/api/v1/rm/gepa-runs" \
-H "Authorization: Bearer $STASH_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"reward_model_id": "<reward_model_id>",
"task_model": "openai/Qwen/Qwen3-8B",
"task_api_base": "http://gpu-box.internal:8000/v1",
"reflection_model": "anthropic/claude-sonnet-5"
}'Returns a GepaRun with status queued. Only reward_model_id is required. The worker generates the skill name and description; they are null until success.
GET/api/v1/rm/gepa-runsList your GEPA runs, newest first.GET/api/v1/rm/gepa-runs/{id}Get one run, including its result once it succeeds.GET/api/v1/rm/gepa-runs/{id}/skillThe best skill as a text/markdown attachment named SKILL.md.curl -s "$STASH_URL/api/v1/rm/gepa-runs/<id>/skill" -H "Authorization: Bearer $STASH_API_KEY" -o SKILL.mdA 404 until the run has succeeded.
SQL query
POST/api/v1/rm/queryRun one read-only SELECT over your data.Each request gets a fresh in-memory DuckDB holding only your rows. Send exactly one SELECT (or WITH … SELECT), checked with DuckDB's own parser. File and network access are disabled. Results are capped at 1000 rows, and truncated is true when there were more. A query that runs longer than 10 seconds is stopped with a 422, query exceeded 10s; DuckDB errors come back as a 422 with DuckDB's message.
curl -s "$STASH_URL/api/v1/rm/query" \
-H "Authorization: Bearer $STASH_API_KEY" \
-H "Content-Type: application/json" \
-d '{"sql": "SELECT source_format, count(*) AS n FROM traces GROUP BY 1"}'{"columns": ["source_format", "n"],
"rows": [["openai_chat", 128], ["otel", 40]],
"truncated": false}Tables
| table | columns |
|---|---|
| traces | id, external_id, title, source_format, metadata (JSON), created_at |
| steps | id, trace_id, idx, role, content, tool_name, tool_input (JSON), tool_call_id, metadata (JSON) |
| annotations | id, trace_id, step_id, step_index, rating, comment, quote (JSON), label_error, label_error_note, created_at |
| scores | reward_model_id, reward_model_name, trace_id, score, created_at |
| action_scores | reward_model_id, reward_model_name, trace_id, step_id, step_index, score, credit, created_at |
SELECT * FROM steps LIMIT 0 returns a table's column names without any rows.
Example: negative rate per tool
Which tools are your reviewers rating down most often?
SELECT s.tool_name,
count(*) AS rated_steps,
avg(CASE WHEN a.rating = -1 THEN 1 ELSE 0 END) AS negative_rate
FROM annotations a
JOIN steps s ON s.id = a.step_id
WHERE a.rating IS NOT NULL
AND NOT a.label_error
AND s.tool_name IS NOT NULL
GROUP BY s.tool_name
ORDER BY negative_rate DESCExample: worst-scored traces nobody has reviewed
SELECT t.title, sc.score
FROM scores sc
JOIN traces t ON t.id = sc.trace_id
WHERE sc.reward_model_name = 'refund-policy'
AND t.id NOT IN (SELECT trace_id FROM annotations)
ORDER BY sc.score ASC
LIMIT 20Example: where reviewers disagree
SELECT trace_id, step_index,
count(*) FILTER (WHERE rating = 1) AS plus,
count(*) FILTER (WHERE rating = -1) AS minus
FROM annotations
WHERE NOT label_error
GROUP BY trace_id, step_index
HAVING plus > 0 AND minus > 0jq -Rs '{sql: .}' query.sql.Objects
TraceSummary
| Parameter | Type | Description |
|---|---|---|
| id | string | Stash's trace id. |
| external_id | string | null | The id you imported it with. |
| title | string | Title, or the first 80 characters of the first user step. |
| source_format | string | The format it was imported from. |
| step_count | integer | Number of steps. |
| positive_count | integer | + ratings on the trace and its steps, not counting flagged ones. API-rating counts only; feedback-derived pairs are not counted here. |
| negative_count | integer | − ratings on the trace and its steps, not counting flagged ones. |
| comment_count | integer | Annotations with a comment. |
| label_error_count | integer | Annotations flagged as label errors. |
| latest_score | object | null | {reward_model_id, reward_model_name, score} from the most recently finished reward model that scored this trace. |
| action_credit | object | null | {mean, count, revision}: average action reward and coverage for the current shared evaluator; not an outcome score. |
| shared_training_allowed | boolean | Owner permission to contribute this trace and feedback to shared training. Defaults false. |
| created_at | string | ISO 8601. |
TraceDetail
Every TraceSummary field, plus:
| Parameter | Type | Description |
|---|---|---|
| metadata | object | The trace's metadata. |
| steps | [Step] | In order. |
| annotations | [Annotation] | Oldest first. |
| scores | array | [{reward_model_id, reward_model_name, score, created_at}], one per reward model that scored this trace, most recent model first. |
| action_scores | array | [{reward_model_id, reward_model_name, step_id, score, credit, created_at}]. Raw learned rewards and relative credit in [-1, 1] for assistant actions only. |
| scoring_runs | array | Latest job per model: id, trace_id, reward_model_id, status, error, created_at, started_at, finished_at. Status is queued, running, succeeded or failed. |
| default_evaluator | object | null | {id, name, revision, updated_at} for the released Stash evaluator. |
| automatic_scoring | object | null | {attempts, error} while automatic scoring is pending. Automatic attempts are capped at three. |
| training_collection | object | null | {status, error} for an opted-in trace's contribution extraction. |
Step
| Parameter | Type | Description |
|---|---|---|
| id | string | Step id. Use it as step_id when annotating. |
| index | integer | Position in the trace, from 0. |
| role | string | system, user, assistant, or tool. |
| content | string | The step's text. |
| tool_name | string | null | Tool called, or that produced this result. |
| tool_input | object | null | Tool call arguments. |
| tool_call_id | string | null | Links a call and its result. |
| metadata | object | null | Per-step metadata. |
Annotation
| Parameter | Type | Description |
|---|---|---|
| id | string | Annotation id. |
| trace_id | string | The trace it belongs to. |
| step_id | string | null | The step, or null for the whole trace. |
| rating | 1 | -1 | null | +, −, or none. |
| comment | string | null | Free text. |
| quote | object | null | {text, prefix, suffix}. |
| label_error | boolean | Flagged as wrong; excluded from training and GEPA. |
| label_error_note | string | null | Why it was flagged. |
| author_id | string | User who wrote it. |
| author_name | string | That user's name. |
| created_at | string | ISO 8601. |
RewardModel
| Parameter | Type | Description |
|---|---|---|
| id | string | Reward model id. |
| name | string | Display name. |
| base_model | string | Hugging Face model id it was trained from. |
| compute | string | Deployment-selected runtime: local or modal. Read-only. |
| trace_count | integer | Number of selected traces. |
| trace_ids | string[] | Selected trace UUIDs; included by GET /reward-models/{id}. |
| epochs | integer | Requested epochs. |
| max_pairs | integer | Requested cap on pairs. |
| status | string | queued, running, succeeded, or failed. |
| num_pairs | integer | null | Usable pairs from selected traces, including feedback-derived pairs and the held-out split. |
| metrics | object | null | Training counts, trace-held-out accuracy (nullable), action_scoring_version, action/tool evaluation counts and accuracy, loss, epochs, device and seconds. See Training for the full schema. |
| error | string | null | Failure message when status is failed. |
| created_at | string | ISO 8601. |
| started_at | string | null | ISO 8601. |
| finished_at | string | null | ISO 8601. |
GepaRun
| Parameter | Type | Description |
|---|---|---|
| id | string | Run id. |
| reward_model_id | string | The reward model used as the metric. |
| skill_name | string | null | Generated name; null until success. |
| skill_description | string | null | Generated description; null until success. |
| task_model | string | LiteLLM model string. |
| task_api_base | string | null | OpenAI-compatible base URL, if set. |
| reflection_model | string | LiteLLM model string. |
| max_metric_calls | integer | Evaluation budget. |
| status | string | queued, running, succeeded, or failed. |
| seed_skill | string | null | The starting SKILL.md, with the description as its body. |
| seed_score | number | null | The seed skill's mean score, for comparison. |
| best_skill | string | null | The winning SKILL.md, frontmatter included. |
| best_score | number | null | Its mean calibrated score over all examples, 0 to 1; 0.5 is the calibration mean of chosen/rejected texts. |
| candidates | array | null | Every skill tried, as a full SKILL.md: [{skill, score}]. |
| error | string | null | Failure message when status is failed. |
| created_at | string | ISO 8601. |
| started_at | string | null | ISO 8601. |
| finished_at | string | null | ISO 8601. |