Trace format
One JSONL shape for import, export, and every adapter. Bring any of eight formats; Stash converts them to this.
Stash Trace Format
One trace per line. A trace is an ordered list of steps. The same shape is what GET /api/v1/rm/export/traces returns, so an export re-imports cleanly with format: "stash". Pretty-printed here; in a file, each trace is one line.
{
"id": "refund-1182",
"title": "Refund request for order 1182",
"metadata": {"agent": "support-bot", "model": "qwen3-8b"},
"steps": [
{"role": "system", "content": "You are a support agent."},
{"role": "user", "content": "I want a refund for order 1182"},
{"role": "assistant", "content": "", "tool_name": "lookup_order",
"tool_input": {"order_id": "1182"}, "tool_call_id": "call_1"},
{"role": "tool", "content": "{\"status\": \"delivered\"}", "tool_name": "lookup_order",
"tool_call_id": "call_1"},
{"role": "assistant", "content": "Your order was delivered on May 3."}
]
}Trace fields
| Parameter | Type | Description |
|---|---|---|
| steps* | array | The conversation, in order. At least one step, and at least one that isn't a system step. |
| id | string | Your external id. Re-importing a trace with the same id replaces its steps, and the annotations on the old steps are deleted with them. |
| title | string | Display title. Defaults to the first 80 characters of the first user step. A trace with no title and no user step fails import. |
| metadata | object | Free-form key/value data about the trace, such as agent name or model. |
Step fields
| Parameter | Type | Description |
|---|---|---|
| role* | string | One of system, user, assistant, tool. |
| content* | string | Always a string. Structured tool output is serialized JSON. An assistant step that only calls a tool has empty content. |
| tool_name | string | On an assistant step: the tool it calls. On a tool step: the tool that produced the result. |
| tool_input | object | The arguments of the tool call. Must be an object. |
| tool_call_id | string | Links a tool step to the assistant step that called it. |
| metadata | object | Free-form, per step. Adapters set {"thinking": true} on model reasoning and {"is_error": true} on failed tool results. |
Every adapter below writes a tool call and its result as that pair of steps, and fills in the tool step's tool_name from the matching call.
Adapters put an assistant message's text and each of its tool calls in separate steps. In the Stash format you can also put both on one assistant step; the reward model then sees the text line followed by the tool call line.
System steps
System steps are stored and shown, but the reward model never sees them. GEPA puts the skill it writes into the system message, so a reward model that read the system message could be satisfied by the skill's text instead of by what the agent does. You can comment on a system step, but not rate it, and a trace with only system steps fails import.
Supported input formats
Pass format on import as one of these names, or auto. One import call takes one payload: the contents of one file.
| format | input | one trace per | trace id |
|---|---|---|---|
| stash | Stash Trace Format JSONL | line | id |
| openai_chat | JSONL of {"messages": [...]} (fine-tuning format), or a JSON array of messages | line | none |
| anthropic_messages | JSONL of {"system": ..., "messages": [...]} | line | none |
| otel | OTLP/JSON with resourceSpans and/or resourceLogs: one export request, or Collector file-exporter JSONL | traceId | traceId |
| langfuse | The GET /api/public/traces/{traceId} response, one object or a JSON array | trace | id |
| langsmith | LangSmith run export JSONL | trace_id | trace_id |
| claude_code | Claude Code session transcript (~/.claude/projects/**/*.jsonl) | file | sessionId |
| codex | Codex CLI rollout (~/.codex/sessions/**/rollout-*.jsonl) | file | session id |
The trace id column is what becomes the trace's id. For every format except openai_chat and anthropic_messages, re-importing the same file replaces the traces it created instead of adding copies.
Auto-detection
With format: "auto", Stash checks the payload's shape against each format in the order of the table above and uses the first match. The response's format field tells you which one it picked. When nothing matches, the import fails with a 422:
could not detect the trace format; tried: stash, openai_chat, anthropic_messages, otel, langfuse, langsmith, claude_code, codexThe whole payload is parsed and checked before anything is stored, so a failed import stores nothing. Parse errors name the format, for example could not parse as otel: …. Pass an explicit format when you want a payload to fail rather than be read as something else.
Agent loops with many LLM calls
otel, langfuse, and langsmith record one entry per LLM call, and each call re-sends the conversation so far. Stash sorts a trace's calls by start time and joins them: when the steps collected so far are a prefix of the next call's messages, only the new messages are appended. A call that doesn't continue the conversation is appended whole. Model reasoning is left out of that comparison, since most APIs don't re-send it.
Only LLM calls are read. Tool spans, tool observations, and tool or chain runs are skipped, because every tool call and its result already appear in the next LLM call's messages.
Format examples
A minimal input for each format, and how it maps onto steps. Each example here imports as-is; JSONL examples are pretty-printed, so put each record on one line (jq -c . does this).
openai_chat
{"messages": [
{"role": "system", "content": "You are a weather bot."},
{"role": "user", "content": "Weather in Paris?"},
{"role": "assistant", "content": null, "tool_calls": [
{"id": "call_w1", "type": "function",
"function": {"name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}}]},
{"role": "tool", "tool_call_id": "call_w1", "content": "18C, cloudy"},
{"role": "assistant", "content": "It's 18C and cloudy in Paris."}
]}system and developer messages become system steps. Each entry in tool_calls becomes an assistant step with tool_name from function.name, tool_input parsed from function.arguments, and tool_call_id from id. Each role: "tool" message becomes a tool step. When content is a list of parts, each part becomes its own step. A bare JSON array of messages is read as one trace.
anthropic_messages
{"system": "You are a coding assistant.",
"messages": [
{"role": "user", "content": "List the files."},
{"role": "assistant", "content": [
{"type": "thinking", "thinking": "I should call ls.", "signature": "sig"},
{"type": "text", "text": "Let me look."},
{"type": "tool_use", "id": "toolu_1", "name": "bash", "input": {"command": "ls"}}]},
{"role": "user", "content": [
{"type": "tool_result", "tool_use_id": "toolu_1",
"content": [{"type": "text", "text": "a.py\nb.py"}]}]},
{"role": "assistant", "content": [{"type": "text", "text": "There are two files: a.py and b.py."}]}
]}system becomes the first step. Content blocks become steps in order: text keeps its role, tool_use becomes an assistant tool call step, tool_result becomes a tool step even though Anthropic sends it inside a user message, and thinking becomes an assistant step with {"thinking": true} in its metadata. Empty and redacted thinking is dropped. Other blocks, such as images, become a marker like [image].
otel
Three conventions are read. Any of them can appear in the same export.
{"resourceSpans": [{"scopeSpans": [{"spans": [{
"traceId": "5b8efff798038103d269b633813fc60c",
"spanId": "a4",
"name": "ChatCompletion",
"startTimeUnixNano": "4000",
"attributes": [
{"key": "openinference.span.kind", "value": {"stringValue": "LLM"}},
{"key": "llm.input_messages.0.message.role", "value": {"stringValue": "system"}},
{"key": "llm.input_messages.0.message.content", "value": {"stringValue": "You are a travel agent."}},
{"key": "llm.input_messages.1.message.role", "value": {"stringValue": "user"}},
{"key": "llm.input_messages.1.message.content", "value": {"stringValue": "Find a flight to Rome."}},
{"key": "llm.input_messages.2.message.role", "value": {"stringValue": "assistant"}},
{"key": "llm.input_messages.2.message.tool_calls.0.tool_call.id", "value": {"stringValue": "call_f1"}},
{"key": "llm.input_messages.2.message.tool_calls.0.tool_call.function.name", "value": {"stringValue": "search_flights"}},
{"key": "llm.input_messages.2.message.tool_calls.0.tool_call.function.arguments", "value": {"stringValue": "{\"to\": \"FCO\"}"}},
{"key": "llm.input_messages.3.message.role", "value": {"stringValue": "tool"}},
{"key": "llm.input_messages.3.message.tool_call_id", "value": {"stringValue": "call_f1"}},
{"key": "llm.input_messages.3.message.content", "value": {"stringValue": "AZ 609, 09:40, 180 EUR"}},
{"key": "llm.output_messages.0.message.role", "value": {"stringValue": "assistant"}},
{"key": "llm.output_messages.0.message.content", "value": {"stringValue": "AZ 609 leaves at 09:40 for 180 EUR."}}
]
}]}]}]}| convention | read from |
|---|---|
| OpenInference | llm.input_messages.N.message.* and llm.output_messages.N.message.*, including message.tool_calls and message.contents |
| OTel GenAI | gen_ai.system_instructions, gen_ai.input.messages, gen_ai.output.messages on spans or on a gen_ai.client.inference.operation.details log event, plus the deprecated per-message log events (gen_ai.user.message, gen_ai.choice, …) |
| OpenLLMetry | gen_ai.prompt.N.* and gen_ai.completion.N.* |
The GenAI message attributes may be structured OTLP values or JSON strings. Spans and log records are grouped by traceId, and each group's LLM calls are joined into one conversation as described above.
langfuse
{
"id": "lf-trace-1",
"name": "support-chat",
"observations": [
{"type": "GENERATION", "startTime": "2026-09-01T10:00:01.000Z",
"input": {"messages": [
{"role": "system", "content": "You are a support agent."},
{"role": "user", "content": "Where is my package?"}]},
"output": {"role": "assistant", "content": null, "tool_calls": [
{"id": "call_t1", "type": "function",
"function": {"name": "track", "arguments": "{\"id\": \"PKG9\"}"}}]}},
{"type": "TOOL", "startTime": "2026-09-01T10:00:02.000Z",
"input": {"id": "PKG9"}, "output": "in transit, ETA Friday"},
{"type": "GENERATION", "startTime": "2026-09-01T10:00:03.000Z",
"input": [
{"role": "system", "content": "You are a support agent."},
{"role": "user", "content": "Where is my package?"},
{"role": "assistant", "content": null, "tool_calls": [
{"id": "call_t1", "type": "function",
"function": {"name": "track", "arguments": "{\"id\": \"PKG9\"}"}}]},
{"role": "tool", "tool_call_id": "call_t1", "content": "in transit, ETA Friday"}],
"output": "It arrives Friday."}
]
}Send the body of Langfuse's GET /api/public/traces/{traceId}, or a JSON array of them. Each trace keeps its Langfuse id and uses its name as the title. Its GENERATION observations are read in startTime order: the input is a message list, bare or under messages, and the output is an assistant message or plain text. The TOOL observation above is skipped; its call and result are already in the second generation's input.
langsmith
{"id": "run-root", "trace_id": "trace-1", "run_type": "chain",
"start_time": "2026-09-02T09:00:00.000000",
"inputs": {"input": "What is 12*7?"}, "outputs": {"output": "84"}}
{"id": "run-llm-1", "trace_id": "trace-1", "run_type": "llm",
"start_time": "2026-09-02T09:00:01.000000",
"inputs": {"messages": [
{"role": "system", "content": "You are a calculator."},
{"role": "user", "content": "What is 12*7?"}]},
"outputs": {"choices": [{"message": {"role": "assistant", "content": "",
"tool_calls": [{"id": "call_m1", "type": "function",
"function": {"name": "multiply", "arguments": "{\"a\": 12, \"b\": 7}"}}]}}]}}
{"id": "run-llm-2", "trace_id": "trace-1", "run_type": "llm",
"start_time": "2026-09-02T09:00:02.000000",
"inputs": {"messages": [
{"role": "system", "content": "You are a calculator."},
{"role": "user", "content": "What is 12*7?"},
{"role": "assistant", "content": "", "tool_calls": [{"id": "call_m1", "type": "function",
"function": {"name": "multiply", "arguments": "{\"a\": 12, \"b\": 7}"}}]},
{"role": "tool", "tool_call_id": "call_m1", "content": "84"}]},
"outputs": {"choices": [{"message": {"role": "assistant", "content": "12*7 = 84"}}]}}One run per line, grouped by trace_id. Only llm runs are read, in start_time order. Messages can be role/content objects, LangChain-serialized messages ({"lc": 1, "id": [..., "AIMessage"], "kwargs": {...}}), or {"type": "ai", ...} objects. Outputs can be LangChain generations, OpenAI choices, a messages list, or a single message.
claude_code
{"type": "user", "sessionId": "3f1c2a8e-0000-4000-8000-00000000abcd", "cwd": "/work/demo",
"gitBranch": "main", "message": {"role": "user", "content": "Fix the failing test in calc.py"}}
{"type": "assistant", "sessionId": "3f1c2a8e-0000-4000-8000-00000000abcd",
"message": {"role": "assistant", "content": [
{"type": "tool_use", "id": "toolu_r1", "name": "Read", "input": {"file_path": "/work/demo/calc.py"}}]}}
{"type": "user", "sessionId": "3f1c2a8e-0000-4000-8000-00000000abcd",
"message": {"role": "user", "content": [
{"type": "tool_result", "tool_use_id": "toolu_r1", "content": "def add(a, b):\n return a - b"}]}}
{"type": "assistant", "sessionId": "3f1c2a8e-0000-4000-8000-00000000abcd",
"message": {"role": "assistant", "content": [{"type": "text", "text": "add() subtracts. Fixing it."}]}}
{"type": "ai-title", "sessionId": "3f1c2a8e-0000-4000-8000-00000000abcd", "aiTitle": "Fix add() in calc.py"}One session file is one trace, keyed by sessionId and titled with the last ai-title line. Only user and assistant lines are read, and lines marked isMeta, isSidechain, isCompactSummary, or isApiErrorMessage are skipped. Content blocks map the same way as anthropic_messages. cwd and gitBranch go into the trace's metadata. To import every session in a project, send one request per file:
for f in ~/.claude/projects/-Users-me-myrepo/*.jsonl; do
jq -Rs '{format: "claude_code", data: .}' "$f" \
| curl -s "$STASH_URL/api/v1/rm/traces/import" \
-H "Authorization: Bearer $STASH_API_KEY" \
-H "Content-Type: application/json" --data @-
donecodex
{"type": "session_meta", "payload": {"id": "01a0ffff-0000-7000-8000-000000000001",
"cwd": "/work/app", "cli_version": "0.150.0"}}
{"type": "response_item", "payload": {"type": "message", "role": "developer",
"content": [{"type": "input_text", "text": "You are Codex, a coding agent."}]}}
{"type": "response_item", "payload": {"type": "message", "role": "user",
"content": [{"type": "input_text", "text": "How many lines in main.py?"}]}}
{"type": "response_item", "payload": {"type": "reasoning",
"summary": [{"type": "summary_text", "text": "Count lines with wc."}], "encrypted_content": "gAAAA"}}
{"type": "response_item", "payload": {"type": "function_call", "name": "shell",
"arguments": "{\"command\": [\"wc\", \"-l\", \"main.py\"]}", "call_id": "call_s1"}}
{"type": "response_item", "payload": {"type": "function_call_output",
"call_id": "call_s1", "output": "42 main.py"}}
{"type": "response_item", "payload": {"type": "message", "role": "assistant",
"content": [{"type": "output_text", "text": "main.py has 42 lines."}]}}One rollout file is one trace, keyed by the session_meta id. Only response_item lines are read. developer messages become system steps, function_call and custom_tool_call become assistant tool call steps, their outputs become tool steps joined on call_id, and a reasoning item's readable summary becomes a thinking step. cwd and cli_version go into the trace's metadata.