Decision Recording and Models

A train's decisions run without Trax.Effect; Trax.Core needs only an IDecider. Trax.Effect adds three things around them: an adapter for typed decision models, a record of every decision against the run that made it, and re-queued runs that take the tracks the original took.

services.AddTrax(trax => trax.AddEffects(effects => effects
    .UsePostgres(connectionString)
    .AddDecisionRecording()
    .AddNimbleDecider(o => o.Endpoint = new Uri("http://localhost:8000/v1/systemone"))));

Nimble

Nimble is Bespoke Labs' typed decision model, and the one Trax builds around: its 9B weights are open (published under Apache 2.0, a fine-tune of Qwen3.5-9B), so a train's decisions can run on hardware you own, with no data leaving it and no per-request price.

AddNimbleDecider talks to a Nimble server you run, started from Nimble's own serving code (nimble/serving/server.py in bespokelabsai/nimble), which answers POST /v1/systemone. There is no default endpoint. Nimble's documentation describes no production hosted API, and the public demo it links to is unauthenticated and runs on one GPU, so it is not somewhere to send a train's state and Trax does not point at it. A host that registers Nimble without an Endpoint fails to start.

SettingDefaultWhy
Endpointnone, requiredThe full URL of POST /v1/systemone on your server
Modelbespokelabs/Bespoke-Nimble-9BThe checkpoint id the server answers to. nimble-latest is refused.
ApiKeynoneSent as a bearer token only when set. The server checks one only when it is started with OPENJEV_API_KEY.
MaxConcurrentRequests4One server container runs four evaluations at once and turns a fifth away with a 529. Raise it when the server scales to more containers.
MaxOptions26The most options or levels the server accepts on one question
Questions per request64The server's limit; a Decide asking more is refused at startup
AttemptTimeout30 secondsA self-hosted server can be slow on its first request after loading

The request cannot pin a model revision: the server serves whichever revision it was deployed with. Pin the revision where you deploy the server, and re-check your confidence bars when you change it. Token and body limits (8,192 prompt tokens a question, a 2 MiB body) cannot be checked before sending, so a request over them is answered 413 or 422 and fails as Permanent.

On Bespoke's benchmarks Nimble is slightly less accurate than Jev (90.1% against 93.2% on their held-out set); put it in front of a larger model with a cascade when that matters, and check its confidence bars against your own labelled cases either way.

Other typed decision models

Trax.Effect.Decisions.SystemOne answers a train's questions through the System One request format, which Jev introduced and Nimble's server accepts: one POST carrying the model, the state, and the questions, each with its question key as its id, answered with one typed answer per question and the model's name. That name is the one the request asked for, echoed back, not a version the server confirms, so pin the version where the model is deployed. AddNimbleDecider is this adapter with Nimble's model name and limits filled in; AddSystemOneDecider is the same adapter for Jev or another server that accepts the format, so moving between them is a change of endpoint and model name:

effects.AddSystemOneDecider(o =>
{
    o.Endpoint = new Uri("https://api.typesafe.ai/v1/systemone");
    o.Model = "jev-1.13.0";
    o.ApiKey = configuration["Jev:ApiKey"];
});
Trax questionSent asCriteria sentRead back
Choicechoiceeach option's name and descriptionchoice, confidence (or, without one, the chosen option's probability), probabilities
Scorescoreeach level's description, lowest firstscore, confidence (or, without one, the nearest level's probability), probabilities keyed by every level index from 0
Yes/nonoulwhat yes and no mean (default Yes and No)noul, the probability of yes

The state is sent as a string when it is one, and otherwise as JSON with camel-cased names, so a question's instructions can name a field. Send a state that holds only what the questions need: these models answer less accurately with detail that does not bear on the question.

AddNimbleDecider and AddSystemOneDecider refuse, at startup, settings that would make decisions hard to trust or leak data:

  • There must be an endpoint, an http or https URL. Nimble has no default to fall back on.
  • The model must be pinned to a version (jev-1.13.0, not jev-latest or a bare jev; Nimble's bespokelabs/Bespoke-Nimble-9B, not nimble-latest). A confidence bar tuned against one version does not carry over to the next. AllowFloatingModel turns this off for AddSystemOneDecider.
  • The endpoint must be HTTPS unless it is a loopback address, where a self-hosted model usually runs. The request carries the train's state and the API key.

The API key is optional; a blank one, such as an unset configuration value, sends no Authorization header.

A throttled, unavailable or slow model, a connection that is refused, reset or times out, or a 200 whose body is not a System One response (including one with no answers object) is retried with a doubling, jittered wait that stops growing at MaxRetryDelay (30 seconds by default). A Retry-After is honoured up to MaxRetryDelay; when the model asks for longer, or the retries run out, the failure is classified Transient.

These are classified Permanent and not retried, because only a change to the request or the configuration cures them:

  • a request the model refuses: bad criteria, a bad key, a 501 or 505
  • a redirect. The adapter does not follow redirects; the endpoint is the one configured.
  • an endpoint that cannot be reached as configured: its name does not resolve, its TLS handshake fails, or a proxy refuses the credentials
  • a request that cannot be sent as it is: a state written as JSON null, or one that cannot be serialized at all. These depend on the value, so they are refused when the request is about to be sent.

What the model refuses whatever the state is refused earlier, at startup. SystemOneDecider implements IVetsQuestions, so the startup check refuses a declaration with more questions in one Decide than MaxQuestions, a question with blank instructions, fewer than two or more than MaxOptions options or levels, an option named twice, a kind of question the format cannot ask, or a state type JSON writes as a bare number, true or false, and the host does not start with it.

A failure's message gives the status, a few words on what it means, and the provider's request id when the response carries one in its x-typesafe-request-id header. It never quotes the response body, because a server's validation error can echo the request, and the request carries the train's state, while the message is stored as the run's failure reason. Cancelling the run cancels the request. An answer the adapter cannot read is left out, and the train fails on the unanswered question, Transient, rather than acting on a guess.

Putting the model in front of a larger one

Without a name, each method registers the decider every train asks, as both IDecider and SystemOneDecider, and may be called once; a second unnamed registration fails at startup instead of replacing the first. With a name, it registers a keyed SystemOneDecider and nothing else, so several models can sit side by side. Compose them yourself, and register the result as the IDecider:

services.AddTrax(trax => trax.AddEffects(effects => effects
    .AddNimbleDecider("nimble", o => o.Endpoint = new Uri("http://localhost:8000/v1/systemone"))
    .AddSystemOneDecider("jev", o =>
    {
        o.Endpoint = new Uri("https://api.typesafe.ai/v1/systemone");
        o.Model = "jev-1.13.0";
        o.ApiKey = configuration["Jev:ApiKey"];
    })));
 
services.AddSingleton<IDecider>(sp => new CascadingDecider(
    sp.GetRequiredKeyedService<SystemOneDecider>("nimble"),
    sp.GetRequiredKeyedService<SystemOneDecider>("jev"),
    escalateBelow: 0.8));

A keyed decider is not something DecidedBy<TDecider>() can name, because that resolves by type. To give one step a different decider, register a type of your own for it.

See Escalating what the fast decider is unsure of.

Recording decisions

AddDecisionRecording() records, for every question a run asks, the question as asked, the answer it acted on, the model and decider that gave it, any shadow's answer and whether it agreed, and every track taken on it. Each decision is logged and written to trax.decision against the run's metadata as it is made, before the train acts on it, through a short-lived data context of its own rather than the run's. A run that fails, times out or is killed after deciding still shows what it decided, and so does a step that failed on a decider's answer: a missing or unfit live answer is recorded with the reason in refused. Each track a routing step takes is added to that question's row when it is taken; more than one step can route on one decision (a Decide followed by two Switch steps on the same choice), so the row keeps every routing in order rather than only the last.

Recording is required, not best effort. A decision that cannot be written fails its step before any track is taken, classified Transient (or Permanent when the store refuses the value), because a re-queue of that run would otherwise have nothing to replay. A shadow's answer can never cost the live record: a number JSON cannot hold is written as the string "NaN", "Infinity" or "-Infinity", and an answer of a type the journal has no stored form for is recorded with the reason in the shadow's error. A live answer of such a type fails its step, Permanent, since it could never be read back. So does a decision the run's own train reports under an external id other than the run's (the train changed its ExternalId while running), rather than being acted on without a record. Decisions are matched to the run through its async flow; when code in the run lost that flow (it suppressed ExecutionContext flow, say), the journal looks the run up by its external id among the runs of that train in progress on this host, and records against it when exactly one matches. Otherwise the decision is logged only. Calling AddDecisionRecording() more than once registers it once.

ColumnHolds
metadata_idThe run. Rows are deleted with it.
question_keyThe question key: the type's name without its namespace, or the Key set on [Asks]
occurrenceWhich asking of the question this was in the run, from 0
fingerprintThe fingerprint of the asking the answer was given to, 64 lowercase hex characters. A replay hands it back, and an answer whose fingerprint differs from the question as it is asked now is not replayed.
kindchoice, score or yes_no
questionThe question's instructions and criteria (jsonb)
answerThe answer acted on (jsonb), with replay_refused when an earlier run's answer was not replayed and the decider was asked afresh. On a refused row, the answer the run would not act on, or null when the decider gave none.
refusedWhy the run would not act on the decider's answer, or null for an answer it acted on. Every row has an answer or a refused (a check constraint holds it). A refused row is never replayed.
modelThe model that answered, as the decider names it, or null for a decider that is not a model. For a System One model it is the name the request asked for, echoed back.
deciderThe decider's type, or null for a replayed answer
replayedWhether the answer came from an earlier run
shadowsEach shadow's answer, whether it agreed, and why it gave none (jsonb)
routesEvery track a routing step took on this decision, in order, as a jsonb array of {"track": ..., "fallback_reason": ...}; fallback_reason says why the decision was not followed, and is null when it was. Null when nothing routed on it.
decided_atWhen it was answered

The same migration adds decisions_recorded to trax.metadata. It is set on a run's first write when the host records decisions, before any junction, so a replay can tell a run that reached no questions from one whose decisions were never recorded.

It needs a data provider, and is a compile error before one. The table ships in the core migration set (Postgres 054, Sqlite 019) and is read through IDataContext.RecordedDecisions.

-- How often each support track was taken this week, and how often it was overruled.
-- question_key is the type's name without its namespace, unless [Asks] sets a Key.
SELECT r ->> 'track' AS track, count(*) AS routings, count(r ->> 'fallback_reason') AS overruled
FROM trax.decision d
JOIN trax.metadata m ON m.id = d.metadata_id
CROSS JOIN LATERAL jsonb_array_elements(d.routes) AS r
WHERE d.question_key = 'TicketTrack' AND m.start_time > now() - interval '7 days'
GROUP BY 1;

That is the data a confidence bar should be tuned from: label a few hundred recorded decisions with what should have happened, and pick the bar where the cost of a wrong track meets the cost of sending work to the fallback.

Re-queued runs replay their decisions

Re-queueing an execution, with the dashboard's Re-queue button or the requeueExecution mutation, queues a run that replays the original's recorded decisions. Both go through IOperationsService.RequeueExecutionAsync, which sets the new run's ReplayDecisionsOf to the original's id when the original has decisions to replay: it recorded a decision it acted on, or was itself queued to replay another run. A run of a train that never decides is re-queued as an ordinary enqueue. Each question the new run asks is answered from what was recorded for the same question and asking, without calling a decider or its shadows, and recorded with replayed set. A re-queue repeats a run, usually because something after a decision failed, and asking a model again could take a different track. Trax.Docs/adr/0041 records why.

The requeue is the only way to set the link. Queueing a train through queueTrain or QueueTrainAsync never replays, and neither does a dead-letter retry or a manifest's scheduled run.

A requeue of a requeue

A replay follows replay_decisions_of back through every run it repeats. For each question and occurrence the nearest run's recorded answer wins, since that is what the nearest run acted on, whether it replayed it or was answered afresh. A question that run never reached falls back to the run it replayed, and so on. So a requeue of a requeue that failed before reaching a question still takes the track the first run took there. The chain is followed at most 32 runs back.

All of this is loaded once, before the run's first junction, so answering a question never waits on the database. Metadata cleanup does not delete a run while a Queued work queue entry or any other run names it in replay_decisions_of, so a chain stays whole while anything links to it; when the linking run expires too, it is deleted first and the run it links to in a later batch or sweep.

The replay meetsWhat happens
A question no run in the chain reached (they failed earlier, or the chain changed since)Asked afresh
A recorded answer whose fingerprint differs from the question as it is asked now, or that no longer fits it (an option renamed or removed, a scale with fewer levels, another kind of question)Asked afresh, with the reason stored as replay_refused
A recorded choice of a member the switch has no track forReplayed; it takes the Otherwise track again, as it did the first time
A refused row (the step failed on the decider's answer)Skipped: never replayed, so the question is asked afresh
A host that does not call AddDecisionRecording()The run fails before its first junction, Permanent
A run in the chain that no longer exists, or is a run of another trainThe run fails before its first junction, Permanent
A run in the chain that ran without recording its decisions (decisions_recorded false, and it replayed nothing itself)The run fails before its first junction, Permanent: what it decided cannot be known
A chain that leads back on itself, or goes back more than 32 runsThe run fails before its first junction, Permanent
A recorded answer that cannot be readThe run fails before its first junction, Permanent
A database failure while loading the chainThe run fails before its first junction, Transient

A run that cannot honour its replay fails instead of asking afresh, because it was queued to repeat the original. A run in the chain that recorded its decisions but reached no questions is not a failure: there is nothing of its own to repeat, and the replay goes on to the run before it.

A host that does not record

A requeue is linked only when the run has decisions to replay, so a replay reaches a host without AddDecisionRecording() when one host recorded the run and another runs the requeue: one worker of a fleet left without the call. Every host that runs trains (AddScheduler, AddTraxJobRunner, and AddTraxWorker through it) checks for this at startup. When IDecisionReplay is not registered and some registered train's chain, or a track in it, asks a decider, the host refuses to start: an InvalidOperationException lists those trains and says to call AddDecisionRecording(). The check runs as the host starts, before any worker claims work.

Trax fails closed here rather than warning. The replay on such a host would already fail rather than ask afresh, but only when a requeue happened to land on it, and a warning on one worker of a fleet is easy to miss. Refusing to start puts the gap in front of whoever deploys the host, before it takes any work.

SDK Reference

AddDecisionRecording | AddNimbleDecider | AddSystemOneDecider | IOperationsService | TrainExecution