Training
A shared action evaluator that improves from permitted feedback, with personal models for custom rubrics.
The shared evaluator
The trace viewer uses Stash's released evaluator by default. It automatically scores imported and updated traces, including tool calls, without asking you to train a model or write comments. Shared inputs include system instructions and optional task/tool context inmetadata.evaluation_context. Only include information available before the action. Inputs currently retain the last 1024 tokens; early context may be truncated on long traces.
Scoring does not contribute your data to training. Opt in per trace to contribute corrections, step ratings and independently reviewed comparisons. Turning contribution off excludes them from future candidates, but cannot remove what previously released weights have learned. Human-reviewed or verified-outcome comparisons form a separate benchmark. Candidates must beat the incumbent overall without regressions on sufficiently populated tool, domain or agent slices. Predictions from the evaluator are never reused as their own training labels.
The model
A reward model reads a rendered trace and returns one number. Stash loads your base model as AutoModelForSequenceClassification with num_labels=1, so the output is a single score r(x), and trains it with the Bradley–Terry loss:
loss = −log σ( r(chosen) − r(rejected) )The loss only cares about the gap between the two scores: it pushes the chosen text above the rejected one. A score has no fixed scale on its own. Compare scores from the same model, never across models.
Personal models use automatic assessment of assistant actions (including tool calls), actionable comments, later user corrections, and API ratings on the selected traces. Manual comments are optional. Personal models omit system steps; Annotations describes exactly how pairs are built and rendered.
Start a personal training job
In the app, select traces and choose Create new reward model. There is no separate Auto mode or rating control in the UI. The experiment is enabled for new accounts; accounts that existed at rollout receive 404 from reward-model APIs.
curl -s "$STASH_URL/api/v1/rm/reward-models" \
-H "Authorization: Bearer $STASH_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "refund-policy-v2",
"trace_ids": ["<trace_id>"],
"epochs": 2
}'| Parameter | Type | Description |
|---|---|---|
| name* | string | Display name. |
| trace_ids* | string[] | At least one trace UUID owned by you. Only these traces supply training pairs. |
| base_model | string | Hugging Face model id to fine-tune. Default Qwen/Qwen3-0.6B. |
| epochs | integer | Passes over the training pairs. Default 1. |
| max_pairs | integer | Cap on combined API-rating and feedback-derived pairs. Default 4000. |
The worker extracts supported preferences and saves each generated alternative with its source ID, verbatim evidence quote, and rationale. These are model interpretations of feedback, not direct human votes. Ambiguous feedback is skipped; invalid extraction fails the job. At least two usable pairs are required. num_pairs includes held-out pairs.
The deployment selects RM_COMPUTE; sending compute or any unknown request field returns 422. Extraction and the minimum pair count are checked after queueing, before downloading the base model.
Status
A reward model moves through queued → running → succeeded or failed. Metrics and scores are stored in the same step that marks it succeeded, so a succeeded model has scores, metrics, and a private stored checkpoint. On failure, error holds the reason: the last 2000 characters of worker.log when the worker failed.
Training settings
| setting | value |
|---|---|
| max length | 1024 tokens. Longer texts lose their beginning, so the end of the conversation, the part being judged, is kept. |
| learning rate | 1e-5, AdamW |
| batch size | 4 pairs |
| held-out split | 10% of source traces, at least 1 when possible. Cross-partition pairs are excluded. If no independent split is usable, evaluation is unavailable. Training needs at least 2 pairs. |
| shuffling | Seed 0, for both the split and the batch order. |
These are fixed in the worker; the request sets only the fields above.
Choosing a base model
The base model must support AutoModelForSequenceClassification and fit the deployment's memory and time limits. Start with the default, Qwen/Qwen3-0.6B, on a worker with sufficient memory, such as the hosted Modal GPU. Move to a larger base model when held-out accuracy stops improving as you add labels.
Compute
| compute | runs on |
|---|---|
| local | The machine running the Stash worker: MPS on Apple silicon, CUDA when a GPU is present, otherwise CPU. |
| modal | A Modal A10G GPU, with a 20-minute limit per training or GEPA invocation. Same training code; only the device changes. |
With modal, the GPU process receives the inputs, uploads the checkpoint directly to private S3, and returns only JSON results and scores. It receives storage and model-provider credentials, but no database, queue, or integration credentials. The ML image is built from rm_worker/requirements.txt.
The device used is recorded in metrics.device.
Metrics
The split is by source trace: neighboring actions stay together, and pairs spanning training and evaluation traces are excluded. Held-out pairs are used neither to update the model nor to set its display scale. Accuracy measures agreement with the preference supervision; it does not establish task success. When no usable independent split exists, accuracy is null.
| metrics field | meaning |
|---|---|
| train_pairs | Pairs the model trained on. |
| eval_pairs | Held-out pairs. |
| eval_accuracy | Fraction of held-out pairs where r(chosen) > r(rejected). 0.5 is chance. |
| action_scoring_version | 1 when trained on action comparisons; null otherwise. Older models lack this field. |
| action_train_pairs | Action comparisons used to train the scorer. |
| action_eval_pairs / action_eval_accuracy | Number of held-out action comparisons and their pairwise accuracy; accuracy is null if none. |
| tool_eval_pairs / tool_eval_accuracy | Number of held-out tool-call comparisons and their pairwise accuracy; accuracy is null if none. |
| excluded_cross_trace_pairs | Comparisons excluded because their source traces span both partitions. |
| final_loss | Mean Bradley–Terry loss over the last epoch's batches. |
| epochs | Epochs run. |
| device | mps, cuda, or cpu. |
| seconds | Worker training time, including scoring and checkpoint upload. |
eval_accuracy is too noisy to establish quality. Check behavior on separate traces before relying on scores.Scores
Action credit
The shared evaluator automatically scores each assistant response and tool call in imported traces. Open a trace to see badges and a heatmap under Action credit, including collapsed tool calls. User messages, system messages, and tool results are observations and have no action score. Each action is scored using its preceding context, without later tool results or replies.
Score again requests a fresh run of the saved checkpoint without labeling calls or retraining. Advanced options let you select a personal model and chooseScore actions. Older personal models need new training with action comparisons. Reimporting clears old action scores and queues the current shared evaluator again. The trace list shows mean action credit and the number of actions scored, separately from personal whole-trace scores.
credit = tanh((raw_reward − action_training_mean) / (2 × action_training_std))Credit ranges from −1 to +1. Higher means the model prefers the action relative to its training reference; zero is the reference midpoint. It is not a correctness probability or an exact causal contribution, and action credits do not add up to the whole-trace score. The learned scores are separate from human comments and AI-generated training judgments. The API returns them in action_scores, and the SQL endpoint exposes an action_scores table.
Whole-trace reward
After training, the worker scores every trace you own, including ones nobody annotated. A score is the model's raw reward: any real number, higher is better. Scores appear on each trace in the app, in GET /traces/{trace_id} (the latest one also in latest_score on the trace list), and in the scores table of the SQL endpoint. Sorting unannotated traces by score is a fast way to find what to review next.
GEPA uses the mean and standard deviation of rewards for both chosen and rejected texts in the training pairs only. Its score falls between 0 and 1, with 0.5 at that calibration mean, not a probability of correctness; see the calibrated score.
Downloading the weights
Choose Download weights on a succeeded model. The owner-authorized API returns {"url": "https://…"} with a five-minute signed URL, not archive bytes. Download from that URL without forwarding your Stash API key. Unowned or unfinished models return 404. Checkpoints live in private S3 storage, independent of the training filesystem.
set -euo pipefail
reward_weights_url="$(curl -fsS "$STASH_URL/api/v1/rm/reward-models/<id>/weights" -H "Authorization: Bearer $STASH_API_KEY" | jq -er '.url')"
curl -fL "$reward_weights_url" --output reward-model.tar.gz
tar xzf reward-model.tar.gzreward-model/
config.json
model.safetensors
tokenizer.json
tokenizer_config.json
chat_template.jinja
reward_stats.json
action_reward_stats.json # when trained on action comparisonsWith worker dependencies installed, run this from the repository root after extracting the archive. The loader applies the same truncation as training:
from rm_worker.scoring import RewardModel
model = RewardModel("reward-model")
rewards = model.score([rendered_trace])Render input as described in Annotations. GEPA downloads this same checkpoint into a fresh temporary workspace.
Self-hosting the worker
A dedicated Celery worker consumes the reward exchange/queue with concurrency 1. The API and ingestion workers do not load torch. For local execution, install the ML environment from the repository root:
uv venv -p 3.12 rm_worker/.venv
uv pip install --python rm_worker/.venv/bin/python -r rm_worker/requirements.txtSet these in the root .env, loaded by ./start.sh:
RM_COMPUTE=local
RM_WORKER_PYTHON=/absolute/path/to/stash/rm_worker/.venv/bin/python
RM_ARTIFACT_DIR=/absolute/path/to/rm-job-scratch
ANTHROPIC_API_KEY=<provider-key>
S3_ENDPOINT=<storage-origin>
S3_BUCKET=<private-bucket>
S3_ACCESS_KEY=<access-key>
S3_SECRET_KEY=<secret-key>
S3_REGION=<region>S3 is required even for local training. The API and worker need the same storage settings. RM_ARTIFACT_DIR is temporary scratch space; the API does not need that directory. ANTHROPIC_API_KEY supports feedback extraction and default GEPA models. Start the local stack with ./start.sh, which includes the reward worker.
For hosted Modal execution, set RM_COMPUTE=modal on the API and reward worker. The backend image includes the runner and sets RM_WORKER_PYTHON=/usr/local/bin/python and RM_ARTIFACT_DIR=/tmp/stash-rm. The reward worker also needs MODAL_TOKEN_ID, MODAL_TOKEN_SECRET, and the same DATABASE_URL and REDIS_URL as the API. Its command is:
celery -A backend.celery_app worker --loglevel=info --concurrency=1 -Q rewardThe worker README and job directory contract describe subprocess inputs, outputs, and failure handling.