stashDocumentation

API reference

Everything in the reward model app is available over REST under /api/v1/rm.

Authentication

Send a Stash API key as a bearer token. Create one on the API Keys page. Every endpoint reads and writes only the caller's own traces, annotations, models, and runs. If you run your own Stash, set STASH_URL to your backend instead.

export STASH_URL=https://api.joinstash.ai
export STASH_API_KEY="<your key>"

curl -s "$STASH_URL/api/v1/rm/formats" -H "Authorization: Bearer $STASH_API_KEY"

Examples on this page use those two variables. Request and response bodies are JSON unless noted.

Reward-model APIs require an account enrolled in the experiment. Accounts that existed at rollout remain disabled and receive 404 from every endpoint here. New accounts are enabled by default; eligibility is per account, including new accounts in existing organizations.

Errors

Errors return {"detail": "…"}. An id that belongs to another user is a 404, the same as one that doesn't exist. A request Stash can't act on, such as an unparseable import, a rating on a system step, or a rejected SQL query, is a 422 whose detail says why.

methodpathsummary
GET/formatsList import formats
POST/traces/importImport traces
GET/tracesList traces
GET/traces/{trace_id}Get a trace with steps, annotations, scores
DELETE/traces/{trace_id}Delete a trace
POST/traces/{trace_id}/annotationsAnnotate a trace or step
PATCH/annotations/{annotation_id}Update an annotation or flag a label error
DELETE/annotations/{annotation_id}Delete an annotation
GET/export/tracesExport traces as JSONL
GET/export/annotationsExport annotations as JSONL
GET/export/pairsExport training pairs as JSONL
POST/reward-modelsTrain a reward model
GET/reward-modelsList reward models
GET/reward-models/{id}Get a reward model
GET/reward-models/{id}/weightsDownload the trained weights
POST/gepa-runsStart a GEPA run
GET/gepa-runsList GEPA runs
GET/gepa-runs/{id}Get a GEPA run
GET/gepa-runs/{id}/skillDownload the best skill as SKILL.md
POST/queryRun read-only SQL

Formats

GET/api/v1/rm/formatsList the import formats this server accepts.
[{"name": "stash", "description": "…"}, {"name": "openai_chat", "description": "…"}, …]

Traces

POST/api/v1/rm/traces/importImport one payload of traces.
ParameterTypeDescription
format*stringA format name from /formats, or auto.
data*stringThe file contents, as one string.
jq -Rs '{format: "auto", data: .}' traces.jsonl \
  | curl -s "$STASH_URL/api/v1/rm/traces/import" \
      -H "Authorization: Bearer $STASH_API_KEY" \
      -H "Content-Type: application/json" \
      --data @-
{"format": "openai_chat", "imported": 128, "trace_ids": ["…", …]}

format in the response is the format that was used, which tells you what auto detected. The import is all or nothing: any error is a 422 and stores no traces. That includes auto matching no format, a payload that doesn't parse as its format, a trace with only system steps, and a trace with no title and no user step. Re-importing a trace with the same id replaces its steps and deletes the annotations on them. See Trace format.

GET/api/v1/rm/tracesList traces, paginated.
ParameterTypeDescription
limitintegerPage size, 1 to 500. Default 50.
offsetintegerRows to skip. Default 0.
curl -s "$STASH_URL/api/v1/rm/traces?limit=50&offset=0" -H "Authorization: Bearer $STASH_API_KEY"

Returns {"traces": [TraceSummary], "total": int}, newest first.

GET/api/v1/rm/traces/{trace_id}A trace with its steps, annotations, and reward model scores.
curl -s "$STASH_URL/api/v1/rm/traces/<trace_id>" -H "Authorization: Bearer $STASH_API_KEY"

Returns a TraceDetail. Step ids for annotating come from this response.

POST/api/v1/rm/traces/{trace_id}/scoreScore assistant actions with a saved trained model.

Send {} for the current shared evaluator, or{"reward_model_id": "<model_id>"} for a personal or previously released shared model. Returns 202 with a scoring job (id, trace_id, reward_model_id, status, error and timestamps). Poll trace detail's scoring_runs for its status and action_scores for results. Repeated requests reuse an active job. Unowned traces or unreleased models owned by others return 404; models without action training, unfinished models and traces without assistant actions return 422. This invokes the trained checkpoint, without LLM labeling or retraining.

GET/api/v1/rm/evaluatorCurrent shared release, or null before bootstrap.

Returns {"default": {"id": "…", "name": "…", "revision": 1, "updated_at": "…"}}. Shared weights and pooled training data are not exposed.

PATCH/api/v1/rm/traces/{trace_id}/training-contributionSet permission to use this trace in shared training.

Send {"allowed": true} to opt in, or false to revoke future contribution. Scoring does not require contribution permission. Revocation does not unlearn released models.

DELETE/api/v1/rm/traces/{trace_id}Delete a trace and its annotations. Returns 204.
curl -s -X DELETE "$STASH_URL/api/v1/rm/traces/<trace_id>" -H "Authorization: Bearer $STASH_API_KEY"

Annotations

POST/api/v1/rm/traces/{trace_id}/annotationsAnnotate a whole trace, or one step.
ParameterTypeDescription
step_idstringA step of this trace. Omit to annotate the whole trace.
rating1 | -1+ or −. Not allowed on a system step.
commentstringFree text. An annotation needs a rating, a comment, or both.
quoteobject{text, prefix, suffix}: the highlighted span. Requires step_id, and text must appear in the step's content.
curl -s "$STASH_URL/api/v1/rm/traces/<trace_id>/annotations" \
  -H "Authorization: Bearer $STASH_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"step_id": "<step_id>", "rating": -1, "comment": "Skipped the policy check"}'

Returns the Annotation. Breaking any rule in the table is a 422.

PATCH/api/v1/rm/annotations/{annotation_id}Change a rating or comment, or flag a label error.

Only the fields you send change. The step and quote are fixed once created.

ParameterTypeDescription
rating1 | -1 | nullNew rating.
commentstring | nullNew comment.
label_errorbooleantrue excludes the annotation from training and GEPA.
label_error_notestringWhy the label is wrong.
curl -s -X PATCH "$STASH_URL/api/v1/rm/annotations/<annotation_id>" \
  -H "Authorization: Bearer $STASH_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"label_error": true, "label_error_note": "Refund was within policy"}'
DELETE/api/v1/rm/annotations/{annotation_id}Delete an annotation. Returns 204.
curl -s -X DELETE "$STASH_URL/api/v1/rm/annotations/<annotation_id>" -H "Authorization: Bearer $STASH_API_KEY"

Export

All three return JSONL (application/x-ndjson), one object per line, for everything you own.

GET/api/v1/rm/export/tracesTraces in the Stash Trace Format. Re-importable as-is.
GET/api/v1/rm/export/annotationsOne annotation per line; fields in Annotations.
GET/api/v1/rm/export/pairsThe training pairs your current labels produce, capped at 4000: {"chosen": str, "rejected": str, "trace_ids": [str], "granularity": str, "action_type": str}.
curl -s "$STASH_URL/api/v1/rm/export/traces"      -H "Authorization: Bearer $STASH_API_KEY" > traces.jsonl
curl -s "$STASH_URL/api/v1/rm/export/annotations" -H "Authorization: Bearer $STASH_API_KEY" > annotations.jsonl
curl -s "$STASH_URL/api/v1/rm/export/pairs"       -H "Authorization: Bearer $STASH_API_KEY" > pairs.jsonl

The pairs export contains pairs built from explicit API ratings. It does not run feedback extraction or export a model's saved feedback-derived training evidence.

Reward models

POST/api/v1/rm/reward-modelsQueue a training job.
ParameterTypeDescription
name*stringDisplay name.
trace_ids*string[]At least one trace UUID owned by you. Only these traces supply training pairs.
base_modelstringHugging Face model id. Default Qwen/Qwen3-0.6B.
epochsintegerTraining epochs. Default 1.
max_pairsintegerCap on training pairs. Default 4000.
curl -s "$STASH_URL/api/v1/rm/reward-models" \
  -H "Authorization: Bearer $STASH_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"name": "refund-policy", "trace_ids": ["<trace_id>"]}'

Returns a RewardModel with status queued. Feedback extraction and the minimum of two usable pairs are checked in the worker. The deployment sets RM_COMPUTE; sending compute or any unknown request field returns 422. An unowned trace ID returns 404.

GET/api/v1/rm/reward-modelsList your reward models, newest first.
GET/api/v1/rm/reward-models/{id}Get one reward model. Poll this for status and metrics.
curl -s "$STASH_URL/api/v1/rm/reward-models/<id>" -H "Authorization: Bearer $STASH_API_KEY"
GET/api/v1/rm/reward-models/{id}/weightsGet a private checkpoint download URL.

Returns {"url": "https://…"} after checking ownership and succeeded status; otherwise 404. The URL expires after five minutes. Download from that URL without forwarding your Stash API key. The gzip archive contains a reward-model/ directory with model weights, tokenizer files, reward_stats.json, and action_reward_stats.json when trained on action comparisons.

set -euo pipefail
reward_weights_url="$(curl -fsS "$STASH_URL/api/v1/rm/reward-models/<id>/weights" -H "Authorization: Bearer $STASH_API_KEY" | jq -er '.url')"
curl -fL "$reward_weights_url" --output reward-model.tar.gz
tar xzf reward-model.tar.gz

GEPA runs

POST/api/v1/rm/gepa-runsQueue a GEPA run that writes a skill.
ParameterTypeDescription
reward_model_id*stringOne of your reward models with status succeeded, used as the metric. Another user's is a 404; an unfinished one is a 422.
task_modelstringLiteLLM task model. Default anthropic/claude-haiku-4-5.
task_api_basestringOpenAI-compatible base URL (vLLM, SGLang, …) for the task model.
reflection_modelstringModel that writes skill bodies. Default anthropic/claude-sonnet-5.
max_metric_callsintegerBudget of example evaluations. Default 40.
curl -s "$STASH_URL/api/v1/rm/gepa-runs" \
  -H "Authorization: Bearer $STASH_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "reward_model_id": "<reward_model_id>",
    "task_model": "openai/Qwen/Qwen3-8B",
    "task_api_base": "http://gpu-box.internal:8000/v1",
    "reflection_model": "anthropic/claude-sonnet-5"
  }'

Returns a GepaRun with status queued. Only reward_model_id is required. The worker generates the skill name and description; they are null until success.

GET/api/v1/rm/gepa-runsList your GEPA runs, newest first.
GET/api/v1/rm/gepa-runs/{id}Get one run, including its result once it succeeds.
GET/api/v1/rm/gepa-runs/{id}/skillThe best skill as a text/markdown attachment named SKILL.md.
curl -s "$STASH_URL/api/v1/rm/gepa-runs/<id>/skill" -H "Authorization: Bearer $STASH_API_KEY" -o SKILL.md

A 404 until the run has succeeded.

SQL query

POST/api/v1/rm/queryRun one read-only SELECT over your data.

Each request gets a fresh in-memory DuckDB holding only your rows. Send exactly one SELECT (or WITH … SELECT), checked with DuckDB's own parser. File and network access are disabled. Results are capped at 1000 rows, and truncated is true when there were more. A query that runs longer than 10 seconds is stopped with a 422, query exceeded 10s; DuckDB errors come back as a 422 with DuckDB's message.

curl -s "$STASH_URL/api/v1/rm/query" \
  -H "Authorization: Bearer $STASH_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"sql": "SELECT source_format, count(*) AS n FROM traces GROUP BY 1"}'
{"columns": ["source_format", "n"],
 "rows": [["openai_chat", 128], ["otel", 40]],
 "truncated": false}

Tables

tablecolumns
tracesid, external_id, title, source_format, metadata (JSON), created_at
stepsid, trace_id, idx, role, content, tool_name, tool_input (JSON), tool_call_id, metadata (JSON)
annotationsid, trace_id, step_id, step_index, rating, comment, quote (JSON), label_error, label_error_note, created_at
scoresreward_model_id, reward_model_name, trace_id, score, created_at
action_scoresreward_model_id, reward_model_name, trace_id, step_id, step_index, score, credit, created_at

SELECT * FROM steps LIMIT 0 returns a table's column names without any rows.

Example: negative rate per tool

Which tools are your reviewers rating down most often?

SELECT s.tool_name,
       count(*)                                        AS rated_steps,
       avg(CASE WHEN a.rating = -1 THEN 1 ELSE 0 END)  AS negative_rate
FROM annotations a
JOIN steps s ON s.id = a.step_id
WHERE a.rating IS NOT NULL
  AND NOT a.label_error
  AND s.tool_name IS NOT NULL
GROUP BY s.tool_name
ORDER BY negative_rate DESC

Example: worst-scored traces nobody has reviewed

SELECT t.title, sc.score
FROM scores sc
JOIN traces t ON t.id = sc.trace_id
WHERE sc.reward_model_name = 'refund-policy'
  AND t.id NOT IN (SELECT trace_id FROM annotations)
ORDER BY sc.score ASC
LIMIT 20

Example: where reviewers disagree

SELECT trace_id, step_index,
       count(*) FILTER (WHERE rating = 1)  AS plus,
       count(*) FILTER (WHERE rating = -1) AS minus
FROM annotations
WHERE NOT label_error
GROUP BY trace_id, step_index
HAVING plus > 0 AND minus > 0
To send multi-line SQL with curl, save it to a file and build the body with jq -Rs '{sql: .}' query.sql.

Objects

TraceSummary

ParameterTypeDescription
idstringStash's trace id.
external_idstring | nullThe id you imported it with.
titlestringTitle, or the first 80 characters of the first user step.
source_formatstringThe format it was imported from.
step_countintegerNumber of steps.
positive_countinteger+ ratings on the trace and its steps, not counting flagged ones. API-rating counts only; feedback-derived pairs are not counted here.
negative_countinteger− ratings on the trace and its steps, not counting flagged ones.
comment_countintegerAnnotations with a comment.
label_error_countintegerAnnotations flagged as label errors.
latest_scoreobject | null{reward_model_id, reward_model_name, score} from the most recently finished reward model that scored this trace.
action_creditobject | null{mean, count, revision}: average action reward and coverage for the current shared evaluator; not an outcome score.
shared_training_allowedbooleanOwner permission to contribute this trace and feedback to shared training. Defaults false.
created_atstringISO 8601.

TraceDetail

Every TraceSummary field, plus:

ParameterTypeDescription
metadataobjectThe trace's metadata.
steps[Step]In order.
annotations[Annotation]Oldest first.
scoresarray[{reward_model_id, reward_model_name, score, created_at}], one per reward model that scored this trace, most recent model first.
action_scoresarray[{reward_model_id, reward_model_name, step_id, score, credit, created_at}]. Raw learned rewards and relative credit in [-1, 1] for assistant actions only.
scoring_runsarrayLatest job per model: id, trace_id, reward_model_id, status, error, created_at, started_at, finished_at. Status is queued, running, succeeded or failed.
default_evaluatorobject | null{id, name, revision, updated_at} for the released Stash evaluator.
automatic_scoringobject | null{attempts, error} while automatic scoring is pending. Automatic attempts are capped at three.
training_collectionobject | null{status, error} for an opted-in trace's contribution extraction.

Step

ParameterTypeDescription
idstringStep id. Use it as step_id when annotating.
indexintegerPosition in the trace, from 0.
rolestringsystem, user, assistant, or tool.
contentstringThe step's text.
tool_namestring | nullTool called, or that produced this result.
tool_inputobject | nullTool call arguments.
tool_call_idstring | nullLinks a call and its result.
metadataobject | nullPer-step metadata.

Annotation

ParameterTypeDescription
idstringAnnotation id.
trace_idstringThe trace it belongs to.
step_idstring | nullThe step, or null for the whole trace.
rating1 | -1 | null+, −, or none.
commentstring | nullFree text.
quoteobject | null{text, prefix, suffix}.
label_errorbooleanFlagged as wrong; excluded from training and GEPA.
label_error_notestring | nullWhy it was flagged.
author_idstringUser who wrote it.
author_namestringThat user's name.
created_atstringISO 8601.

RewardModel

ParameterTypeDescription
idstringReward model id.
namestringDisplay name.
base_modelstringHugging Face model id it was trained from.
computestringDeployment-selected runtime: local or modal. Read-only.
trace_countintegerNumber of selected traces.
trace_idsstring[]Selected trace UUIDs; included by GET /reward-models/{id}.
epochsintegerRequested epochs.
max_pairsintegerRequested cap on pairs.
statusstringqueued, running, succeeded, or failed.
num_pairsinteger | nullUsable pairs from selected traces, including feedback-derived pairs and the held-out split.
metricsobject | nullTraining counts, trace-held-out accuracy (nullable), action_scoring_version, action/tool evaluation counts and accuracy, loss, epochs, device and seconds. See Training for the full schema.
errorstring | nullFailure message when status is failed.
created_atstringISO 8601.
started_atstring | nullISO 8601.
finished_atstring | nullISO 8601.

GepaRun

ParameterTypeDescription
idstringRun id.
reward_model_idstringThe reward model used as the metric.
skill_namestring | nullGenerated name; null until success.
skill_descriptionstring | nullGenerated description; null until success.
task_modelstringLiteLLM model string.
task_api_basestring | nullOpenAI-compatible base URL, if set.
reflection_modelstringLiteLLM model string.
max_metric_callsintegerEvaluation budget.
statusstringqueued, running, succeeded, or failed.
seed_skillstring | nullThe starting SKILL.md, with the description as its body.
seed_scorenumber | nullThe seed skill's mean score, for comparison.
best_skillstring | nullThe winning SKILL.md, frontmatter included.
best_scorenumber | nullIts mean calibrated score over all examples, 0 to 1; 0.5 is the calibration mean of chosen/rejected texts.
candidatesarray | nullEvery skill tried, as a full SKILL.md: [{skill, score}].
errorstring | nullFailure message when status is failed.
created_atstringISO 8601.
started_atstring | nullISO 8601.
finished_atstring | nullISO 8601.