Implementation
The backend stores traces, annotations, and training evidence in Postgres. A dedicated reward worker extracts preferences through a model-provider API, then dispatches training and skill generation to rm_worker. GPU libraries run in a local ML environment or inside Modal; the API never imports torch.
import trace file ─ adapter ─▶ traces, steps
annotate annotations ─────────▶ annotations
train selected feedback ─ pairs.jsonl ─ rm_worker ─▶ private S3 checkpoint, scores
skill gepa_examples.jsonl ─ rm_worker ─▶ SKILL.md
query your rows ─▶ in-memory DuckDB ─▶ resultThis data is separate from Stash sessions: traces imported here don't appear in your session history, and sessions aren't imported here.
Import
Each input format has an adapter that converts a file into the Stash Trace Format: a list of steps with a role of system, user, assistant, or tool. A tool call becomes an assistant step with tool_name and tool_input; its result becomes a tool step with the same tool_call_id. Formats that log one entry per LLM call (OpenTelemetry, Langfuse, LangSmith) are joined into one conversation by dropping the history each call re-sends. An import is parsed completely before anything is stored, so it either stores every trace or none. Details: Trace format.
Annotations to preference pairs
Training uses only selected traces. Comments and explicit later user corrections can produce an original/revised response pair with a shared context and cited evidence. The model-generated revision and preference are interpretations of feedback, not direct human votes. Ambiguous feedback is skipped; invalid extraction fails the job. Explicit API ratings also produce pairs: Stash drops annotations flagged as label errors, sums the ratings on each trace and each step, and calls a target chosen if its sum is positive and rejected if negative. Each target is rendered as <role>: <content> lines without system steps; a step target includes the steps before it. Every chosen target is paired with every rejected target of the same kind (trace with trace, step with step), shuffled with seed 0, and capped at 4000 pairs. Details: Annotations.
Training
A training job runs on the dedicated Celery reward exchange/queue. The backend writes the pairs and the traces to score into a job directory under RM_ARTIFACT_DIR, runs rm_worker with RM_WORKER_PYTHON, and reads the results back. The worker fine-tunes the base model (Qwen/Qwen3-0.6B by default) as a single-output classifier with the Bradley–Terry loss −log σ(r(chosen) − r(rejected)), holding out 10% of pairs to report accuracy. It runs on MPS, CUDA, or CPU on the worker machine, or on a Modal A10G. Details: Training.
Scoring
After training, the worker scores every trace the owner has, and those raw rewards are what the app and the API show, attributed to their model. Separately, it scores both chosen and rejected texts from the training and held-out pairs and saves their mean and standard deviation in model/reward_stats.json. GEPA uses them to calibrate each reward to sigmoid((reward − mean) / std), so 0.5 corresponds to the calibration mean and a good reply still has room to score higher. Details: The calibrated score.
Skill writing (GEPA)
Each selected trace with replayable input becomes an example: its own system prompt, plus the non-system steps before the agent's first reply. For each candidate skill, the worker calls your task model with the system prompt and the skill in the system message, scores the reply with the reward model, and sends the score, unflagged comments, and recorded feedback evidence to a reflection model, which writes the next skill body. The reward model never sees the system message, so a skill can only raise its score by changing what the agent says. The result is the highest-scoring skill. Details: Skills (GEPA).
Checkpoint storage
Training uploads a private S3 checkpoint before succeeding. The API authorizes a five-minute download URL; GEPA downloads that checkpoint into a fresh temporary workspace. The API never needs access to the training worker's filesystem. Modal returns only JSON results and scores, and receives no database or queue credentials.
REST and SQL
Everything is under /api/v1/rm with bearer auth, and every row belongs to one user. Accounts outside the new-account experiment receive 404 from every reward-model endpoint. The SQL endpoint loads only the caller's traces, steps, annotations, and scores into a new in-memory DuckDB per query, with file and network access turned off, a 1000-row cap, and a 10-second limit. Details: API reference.