stashDocumentation

Skills (GEPA)

Turn your reviewers' comments into a skill your agent loads, scored by your reward model.

What you get

A GEPA run produces a skill: a SKILL.md file with YAML frontmatter (name and description) followed by Markdown instructions. Your agent loads it into its context. Your agent's own system prompt is never rewritten.

---
name: refund-policy
description: Use when a customer asks for a refund, return, or exchange.
---

Before promising a refund, look the order up with lookup_order and check
days_since_delivery. Refunds are allowed within 30 days of delivery.
Past 30 days, say so plainly and offer store credit instead.

The reflection model generates the name and description from the selected traces and their feedback. The description says when to use the skill and becomes its seed body; GEPA then evolves the body. In the app, View skill opens an existing run or starts one.

What GEPA is

GEPA is an optimizer for text, from Agrawal et al., 2025, "GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning" (arXiv 2507.19457). It improves text in a loop:

  1. 01Run the agent with the current text on some examples.
  2. 02Collect a score and written feedback for each result.
  3. 03Ask a reflection model to read the feedback and propose better text.
  4. 04Keep a Pareto front of candidates: any text that is best on at least one example survives.

Written feedback is what separates it from optimizers that only see a number. A comment like "promised a refund without checking the policy" tells the reflection model what to change, not only that something went wrong. In Stash, the text GEPA evolves is the skill's body.

How Stash wires it up

GEPA needsStash provides
examplesThe traces selected for the reward model, whether or not they have annotations. Each keeps its own system prompt (its system steps joined, or none) and its input: the non-system steps before the first assistant step.
candidateOne text field, the skill body. The seed body is the generated skill_description as a single line.
task modelCalled once per example: a system message holding the example's own system prompt and the rendered skill, then the example's input.
metricYour reward model's reward for the conversation, reply included, calibrated to a score between 0 and 1. See below.
feedbackThe score, unflagged comments, and recorded evidence from feedback-derived training pairs.

The task model's system message is the example's own system prompt, a blank line, then the skill. A trace with no system steps gets the skill block alone.

<the example's own system prompt, when it has one>

<skill name="refund-policy">
---
name: refund-policy
description: Use when a customer asks for a refund, return, or exchange.
---

<candidate skill body>
</skill>

The conversation the reward model scores is rendered like training text (see Annotations), without the system message. That is why the reward model never sees system steps: the skill is in the system message, and if the reward model read it, GEPA could raise its score by writing what the reward model likes into the skill without changing what the agent does.

The comments come from the original trace. They describe what went wrong last time, which is what the reflection model needs to write the next version. It is told that it is writing the body of a SKILL.md and must return only the body; the name and description stay fixed for the whole run.

The calibrated score

GEPA's score for one example is

score = sigmoid( (reward − mean) / std )

where mean and std are the mean and standard deviation of the reward model's scores over both chosen and rejected texts in the training pairs only, saved in reward_stats.json inside the checkpoint. A score of 0.5 corresponds to that calibration mean; higher is better. It is not a probability of correctness. Raw rewards are the wrong scale for this: a confident reward model gives a decent reply a raw sigmoid(reward) of about 0.99, which leaves GEPA no room to tell a better skill from the seed.

GEPA downloads the private checkpoint into a fresh temporary workspace. It uses the same runtime recorded on the reward model; it does not need the original training filesystem.

A trace with no input before its first assistant step has nothing to replay, so it is skipped. If no examples are left, the run fails.

Start a run

You need one of your reward models with status succeeded. Another user's model is a 404; one that hasn't finished training is a 422.

curl -s "$STASH_URL/api/v1/rm/gepa-runs" \
  -H "Authorization: Bearer $STASH_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "reward_model_id": "<reward_model_id>",
    "task_model": "openai/gpt-4.1-mini",
    "reflection_model": "anthropic/claude-sonnet-5"
  }'
ParameterTypeDescription
reward_model_id*stringA trained reward model. It is the metric.
task_modelstringLiteLLM task model. Default anthropic/claude-haiku-4-5.
task_api_basestringBase URL for an OpenAI-compatible server. Use with an openai/<name> task_model.
reflection_modelstringModel that writes skill bodies. Default anthropic/claude-sonnet-5.
max_metric_callsintegerBudget: how many example evaluations GEPA may run, each one a task model call plus a reward model score. Default 40.

Use the model your agent actually runs on as task_model, so the skill is tuned for it. Use a strong model as reflection_model; it generates the skill identity and writes every candidate body.

Both models are called through LiteLLM, which reads the standard provider keys from the worker's environment. The Modal runner forwards ANTHROPIC_API_KEY and OPENAI_API_KEY; other providers need explicit runner configuration. A vLLM or SGLang server set as task_api_base needs no key unless you started it with one.

Your own model as the task model

Any OpenAI-compatible server works, including vLLM and SGLang. Serve your model, then pass openai/<served model name> as task_model and the server's /v1 URL as task_api_base.

vllm serve Qwen/Qwen3-8B --port 8000

# in the run request:
"task_model": "openai/Qwen/Qwen3-8B",
"task_api_base": "http://gpu-box.internal:8000/v1"
The task model is called from the local worker or Modal GPU process, according to the model's runtime. The task_api_base URL must be reachable from there.

Reading the results

curl -s "$STASH_URL/api/v1/rm/gepa-runs/<id>" \
  -H "Authorization: Bearer $STASH_API_KEY" \
  | jq '{status, seed_score, best_score}'
fieldmeaning
statusqueued, running, succeeded, or failed.
seed_skillThe starting SKILL.md, with the description as its body.
seed_scoreThe seed skill's mean score over all examples.
best_skillThe full SKILL.md with the highest mean score.
best_scoreIts mean score over all examples.
candidatesEvery skill tried, as a full SKILL.md, with its mean score: [{skill, score}].
errorWhy the run failed, when status is failed.

Scores are means of the calibrated score, between 0 and 1, where 0.5 is the calibration mean. Compare best_score to seed_score: both come from the same reward model on the same examples, so the difference is what the skill bought you. Selected traces are often few, so every example is used both to propose skills and to rank them. Treat the gain as an in-sample number, not a held-out one.

Install the skill

Download the best skill as SKILL.md. Until the run succeeds, this is a 404.

mkdir -p .claude/skills/refund-policy
curl -s "$STASH_URL/api/v1/rm/gepa-runs/<id>/skill" \
  -H "Authorization: Bearer $STASH_API_KEY" \
  -o .claude/skills/refund-policy/SKILL.md

That path is where Claude Code looks for project skills. For another agent, put the file wherever it loads skills or instructions from.

When a run fails

If the task model rejects one example's request, that example scores 0 and the error goes to the reflection model as feedback, so the run continues. Authentication, rate-limit, and connection errors from the task model fail the run, and so does any error from the reflection model. A run where the reflection model was never called also fails, usually because max_metric_calls ran out while evaluating the seed skill. error on the run holds the end of the worker's log.

Before you ship the skill

GEPA optimizes whatever your reward model rewards, including its mistakes. Read best_skill and a few of the candidates before deploying one. If the winning skill games something your reviewers wouldn't approve of, run the agent with it, import the resulting traces, annotate them, and retrain. Each round of labels closes a gap the last model left open.