Scheduling Sample

samples/Scheduling in Trax.Samples is one ASP.NET process that schedules trains, runs them on the built-in local workers, retries the ones that fail and dead-letters the ones that keep failing. The Trax GraphQL operations surface and the dashboard show and steer it. It shows the scheduling surface and nothing else, and its E2E suite proves every behaviour on this page against the running host.

What it proves

ManifestDeclared withWhat it shows
refresh-exchange-ratesSchedule(..., Every.Seconds(5))An interval manifest
reprice-catalogThenInclude(...)A dependent: runs after each successful refresh
alert-rate-spikeInclude(..., o => o.Dormant())A dormant dependent: runs only when the refresh activates it, with the input the refresh gives it
send-daily-digestSchedule(..., Cron.Daily(hour: 7))A cron manifest, evaluated in UTC
send-launch-announcementScheduleOnce(..., TimeSpan.FromSeconds(10))A one-off that disables itself after its first success
import-supplier-feedSchedule(..., Every.Seconds(15), o => o.MaxRetries(2))A failing train: two retries with backoff, a dead letter, and a requeue that runs it again

The supplier the last train calls is down for its first three calls after startup. So the first run fails, the retry 2 seconds later fails, the retry 4 seconds after that fails, and the manifest is dead-lettered. Requeue the dead letter and the run succeeds, because the outage is over.

Run it

You need the .NET 10 SDK and Docker. From the Trax.Samples folder:

docker compose up -d database
dotnet run --project samples/Scheduling/Trax.Samples.Scheduling.Host

The host listens on http://localhost:5230 and starts in Development (its launchSettings.json says so). It reads the TraxDatabase connection string from appsettings.json, Host=localhost;Port=5432;Database=trax_scheduling;Username=trax;Password=trax123, a database docker compose creates (see Running the samples for a volume made before it). Trax creates its trax schema on first start.

URLWhat
http://localhost:5230/traxThe dashboard. Development only, no sign-in
http://localhost:5230/trax/graphqlThe GraphQL endpoint. operations needs the header X-Api-Key: operator-key-do-not-use-in-production

Try it

Within about ten seconds of starting, the dashboard's Data > Dead Letters page lists import-supplier-feed, and Data > Metadata shows its three failed runs, each with the junction that threw (DownloadFeedJunction), the exception type and the stack trace.

The same over GraphQL:

KEY='X-Api-Key: operator-key-do-not-use-in-production'
URL=http://localhost:5230/trax/graphql
 
curl -s $URL -H 'Content-Type: application/json' -H "$KEY" -d '{"query":
  "{ operations { deadLetters { deadLetters(status: AWAITING_INTERVENTION) { items { id manifestName reason } } } } }"}'
{"data":{"operations":{"deadLetters":{"deadLetters":{"items":[{"id":1,
  "manifestName":"Trax.Samples.Scheduling.Trains.ImportSupplierFeed.IImportSupplierFeedTrain",
  "reason":"Max retries exceeded: (3) failures > (2) max retries"}]}}}}}

Requeue it, then read it back. Use the id the query returned; on a fresh database it is 1:

curl -s $URL -H 'Content-Type: application/json' -H "$KEY" -d '{"query":
  "mutation { operations { deadLetters { requeueDeadLetter(id: 1) { success workQueueId message } } } }"}'
# {"data":{"operations":{"deadLetters":{"requeueDeadLetter":{"success":true,"workQueueId":14,"message":"Dead letter requeued"}}}}}
 
curl -s $URL -H 'Content-Type: application/json' -H "$KEY" -d '{"query":
  "{ operations { deadLetters { deadLetter(id: 1) { status retryMetadataId } } } }"}'
# {"data":{"operations":{"deadLetters":{"deadLetter":{"status":"RETRIED","retryMetadataId":81}}}}}

retryMetadataId is the run the requeue started; it completed. Run the daily digest now, outside its schedule:

curl -s $URL -H 'Content-Type: application/json' -H "$KEY" -d '{"query":
  "mutation { operations { triggerManifest(externalId: \"send-daily-digest\") { success message } } }"}'
# {"data":{"operations":{"triggerManifest":{"success":true,"message":"Manifest triggered"}}}}

Leave the key off and every operations field answers {"errors":[{"message":"Not authorized.","extensions":{"code":"TRAX_AUTHORIZATION"}}]}.

How it is built

Two projects, following the trains library pattern:

samples/Scheduling/
  Trax.Samples.Scheduling/            class library: trains, ManifestNames, the two fake services
  Trax.Samples.Scheduling.Host/       Microsoft.NET.Sdk.Web: Program.cs, appsettings.json, DemoKeys

The library references Trax.Effect, Trax.Mediator and Trax.Scheduler (the last for IDormantDependentContext). The host adds Trax.Effect.Data.Postgres, Trax.Effect.Provider.Parameter, Trax.Api.Auth.ApiKey, Trax.Api.GraphQL and Trax.Dashboard.

Program.cs

The whole file, less its banner comment:

using Trax.Api.Auth.ApiKey;
using Trax.Api.GraphQL.Extensions;
using Trax.Dashboard.Extensions;
using Trax.Effect.Data.Postgres.Extensions;
using Trax.Effect.Extensions;
using Trax.Effect.Provider.Parameter.Extensions;
using Trax.Mediator.Extensions;
using Trax.Samples.Scheduling;
using Trax.Samples.Scheduling.Host;
using Trax.Samples.Scheduling.Services;
using Trax.Samples.Scheduling.Trains.AlertRateSpike;
using Trax.Samples.Scheduling.Trains.ImportSupplierFeed;
using Trax.Samples.Scheduling.Trains.RefreshExchangeRates;
using Trax.Samples.Scheduling.Trains.RepriceCatalog;
using Trax.Samples.Scheduling.Trains.SendDailyDigest;
using Trax.Samples.Scheduling.Trains.SendLaunchAnnouncement;
using Trax.Scheduler.Extensions;
using Trax.Scheduler.Services.Scheduling;
 
var builder = WebApplication.CreateBuilder(args);
 
var connectionString =
    builder.Configuration.GetConnectionString("TraxDatabase")
    ?? throw new InvalidOperationException("Connection string 'TraxDatabase' not found.");
 
builder.Services.AddSingleton<ExchangeRateFeed>();
builder.Services.AddSingleton<SupplierFeed>();
 
builder.Services.AddTrax(trax =>
    trax.AddEffects(effects => effects.UsePostgres(connectionString).SaveTrainParameters())
        .AddMediator(typeof(ManifestNames).Assembly)
        .AddScheduler(scheduler =>
            scheduler
                // Demo speed. The production defaults are in the comments.
                .ManifestManagerPollingInterval(TimeSpan.FromSeconds(1)) // default 5 s
                .JobDispatcherPollingInterval(TimeSpan.FromSeconds(1)) // default 2 s
                .DefaultRetryDelay(TimeSpan.FromSeconds(2)) // default 5 min
                .RetryBackoffMultiplier(2.0) // default 2.0
                .MaxRetryDelay(TimeSpan.FromSeconds(30)) // default 1 h
                .ConfigureLocalWorkers(workers =>
                    workers.PollingInterval = TimeSpan.FromMilliseconds(250) // default 1 s
                )
                .AddMetadataCleanup(cleanup =>
                {
                    cleanup.RetentionPeriod = TimeSpan.FromMinutes(30);
                    cleanup.AddTrainType<IRefreshExchangeRatesTrain>(TimeSpan.FromDays(1));
                    cleanup.AddTrainType<IRepriceCatalogTrain>(TimeSpan.FromDays(1));
                })
                .Schedule<IRefreshExchangeRatesTrain>(
                    ManifestNames.RefreshExchangeRates,
                    new RefreshExchangeRatesInput { BaseCurrency = "USD" },
                    Every.Seconds(5)
                )
                .ThenInclude<IRepriceCatalogTrain>(
                    ManifestNames.RepriceCatalog,
                    new RepriceCatalogInput { Catalog = "storefront" }
                )
                .Include<IAlertRateSpikeTrain>(
                    ManifestNames.AlertRateSpike,
                    new AlertRateSpikeInput(),
                    options => options.Dormant()
                )
                .Schedule<ISendDailyDigestTrain>(
                    ManifestNames.SendDailyDigest,
                    new SendDailyDigestInput { Audience = "subscribers" },
                    Cron.Daily(hour: 7)
                )
                .ScheduleOnce<ISendLaunchAnnouncementTrain>(
                    ManifestNames.SendLaunchAnnouncement,
                    new SendLaunchAnnouncementInput(),
                    TimeSpan.FromSeconds(10)
                )
                .Schedule<IImportSupplierFeedTrain>(
                    ManifestNames.ImportSupplierFeed,
                    new ImportSupplierFeedInput { Supplier = "acme" },
                    Every.Seconds(15),
                    options => options.MaxRetries(2)
                )
        )
);
 
// NO WARRANTY: a demo key, registered only in Development.
if (builder.Environment.IsDevelopment())
    builder.Services.AddTraxApiKeyAuth(keys =>
        keys.Add(DemoKeys.OperatorKey, id: "operator", DemoKeys.OperatorRole)
    );
builder.Services.AddAuthentication();
builder.Services.AddAuthorization();
 
builder.Services.AddTraxGraphQL(graphql =>
    graphql
        .ExposeOperationQueries()
        .ExposeOperationMutations()
        .GateOperations(roles: DemoKeys.OperatorRole)
);
 
if (builder.Environment.IsDevelopment())
    builder.AddTraxDashboard(options => options.AllowAnonymousDashboard());
 
var app = builder.Build();
 
app.UseAuthentication();
app.UseAuthorization();
 
if (app.Environment.IsDevelopment())
    app.UseTraxDashboard();
 
app.UseTraxGraphQL();
 
app.Run();
 
namespace Trax.Samples.Scheduling.Host
{
    public partial class Program;
}

DemoKeys holds two constants: OperatorRole = "Operator" and OperatorKey = "operator-key-do-not-use-in-production".

Registration order

CallMust comeWhy
AddTrax(...) with AddEffects, then AddMediator, then AddSchedulerFirstThe builder types enforce the order inside it at compile time: AddScheduler exists only on what AddMediator returns
AddTraxApiKeyAuth, AddAuthentication, AddAuthorizationAnywhereThe subscription interceptor and GateOperations read the registered schemes once the container is built
AddTraxGraphQL(...)After AddTraxIt throws AddTrax() must be called before AddTraxGraphQL() otherwise
AddTraxDashboard(...)After AddTraxIt throws AddTraxDashboard() requires AddTrax() to be called first otherwise, and lists only the trains registered before it
UseAuthentication, UseAuthorizationBefore UseTraxDashboard and UseTraxGraphQLSo the dashboard and the operations gate see the caller

The schedules

Interval. Every.Seconds(5) counts from the manifest's last success, so consecutive runs start at least five seconds apart, plus up to a polling interval or two. An interval manifest that has never run is due on the first poll after the host starts.

Dependent. ThenInclude makes reprice-catalog depend on the manifest declared just before it. It has no schedule: the scheduler queues it once its parent has succeeded since its own latest run started. Disable the parent and the dependent stops too.

Dormant dependent. Include parents from the root Schedule rather than from the previous call, so alert-rate-spike also depends on refresh-exchange-rates. .Dormant() means the parent's success never queues it. FlagSpikeJunction, in the parent's chain, activates it with the reading as its input:

public class FlagSpikeJunction(IDormantDependentContext dormants, ILogger<FlagSpikeJunction> logger)
    : Junction<(RefreshExchangeRatesInput Input, RateReading Reading), RefreshExchangeRatesOutput>
{
    public override async Task<RefreshExchangeRatesOutput> Run(
        (RefreshExchangeRatesInput Input, RateReading Reading) input
    )
    {
        var (settings, reading) = input;
        var spike = Math.Abs(reading.ChangePercent) > settings.SpikeThresholdPercent;
 
        if (spike)
            await dormants.ActivateAsync<IAlertRateSpikeTrain, AlertRateSpikeInput, Unit>(
                ManifestNames.AlertRateSpike,
                new AlertRateSpikeInput
                {
                    BaseCurrency = reading.BaseCurrency,
                    ChangePercent = reading.ChangePercent,
                }
            );
 
        return new RefreshExchangeRatesOutput(reading.BaseCurrency, reading.ChangePercent, spike);
    }
}

The context knows which manifest's run it is in, so it can only activate dormant dependents declared under that manifest. Run the same train on the train bus, outside the scheduler, and ActivateAsync logs a warning and queues nothing.

Cron. Cron.Daily(hour: 7) is stored as 0 7 * * * and evaluated in UTC. A new cron manifest does not run at startup: seeding records its next occurrence in NextScheduledRun. Triggering it by hand counts as a run; after that success the next run is the next occurrence after it.

One-off. ScheduleOnce stores ScheduledAt as the seed time plus the delay. Once that passes, the manifest runs; after its first success it sets IsEnabled = false and never runs on its own again, even if re-enabled. Every start re-seeds it, which moves a one-off that has not run yet to the new start plus the delay, so a host that restarts more often than the delay never reaches it.

Retries and the dead letter. MaxRetries(2) allows the first run and two retries. The retry delay is DefaultRetryDelay * RetryBackoffMultiplier ^ (failures - 1), capped at MaxRetryDelay:

AttemptWaits after the previous failureResult
1Fails: HttpRequestException, Supplier 'acme' answered 503 Service Unavailable.
2at least 2 sFails
3at least 4 sFails; the next ManifestManager cycle writes a dead letter with reason Max retries exceeded: (3) failures > (2) max retries
No fourth attempt: a manifest with a dead letter awaiting intervention is skipped

A retry is not a separate timer. It is the manifest's next due run, held back by the backoff. An interval or cron manifest that fails when it was due stays due, so it is retried as soon as the backoff passes. A run started off-schedule, by a trigger, that fails is retried only once the manifest is next due by its schedule.

Requeueing the dead letter (the dashboard, the requeueDeadLetter mutation or ITraxScheduler.RequeueDeadLetterAsync) queues a run with the manifest's own input, marks the dead letter Retried and, once the run is dispatched, sets the dead letter's RetryMetadataId to it. Resolving a dead letter either way resets the failure count, so the requeued run starts a fresh budget of retries. Acknowledging resolves it without running anything, and a later requeue of the same dead letter is refused (success: false).

Demo speed

SettingSampleDefaultEffect here
ManifestManagerPollingInterval1 s5 sHow soon a due manifest, a retry or a dead letter is noticed
JobDispatcherPollingInterval1 s2 sHow soon a queued entry is handed to a worker
DefaultRetryDelay2 s5 minThe first retry's wait
RetryBackoffMultiplier2.02.0Doubles each further wait
MaxRetryDelay30 s1 hCaps the wait
Local worker PollingInterval250 ms1 sHow soon a worker claims a dispatched job

Keep the defaults in production. One second is the floor for both polling intervals.

The operations surface and who may use it

ExposeOperationQueries() and ExposeOperationMutations() add the operations namespace; GateOperations(roles: "Operator") requires that role on all of it. The demo key carries the role and is registered only in Development. Started in Production the host still starts, registers no key at all, and refuses every operations call with TRAX_AUTHORIZATION; the dashboard is not mapped, so /trax is a 404. Register real keys with AddHashed or a resolver (AddTraxApiKeyAuth) before serving it.

What fails at startup

Each of these refuses to start the host with an InvalidOperationException; the messages are the ones the sample printed when the line was changed.

Change to the sampleMessage starts
Drop .GateOperations(...)ExposeOperationMutations() exposes scheduler-control mutations (...) but the GraphQL endpoint is not gated and the namespace carries no gate of its own
.GateOperations() with no policy or rolesGateOperations() needs a policy or roles: GateOperations(policy: "...") or GateOperations(roles: "...")
AddTraxDashboard() with no postureUseTraxDashboard() needs to know who may use the dashboard
Register the demo key outside DevelopmentAddTraxApiKeyAuth() registered a key containing 'do-not-use-in-production', which marks a published demo key, and the environment is 'Production'
cleanup.DeleteBatchSize = 20_000The scheduler configuration has values it cannot run with. AddMetadataCleanup: DeleteBatchSize must be between 1 and 10000.

AddScheduler checks every interval, delay and count against its range when the host builds; the ranges are in AddScheduler.

Testing it

tests/Trax.Samples.Scheduling.E2E starts the real host with WebApplicationFactory against its own database, scheduling_e2e_tests (port overridable with TRAX_TEST_PG_PORT), and asserts on the Trax tables and the GraphQL endpoint:

  • Drop the trax schema before the host starts, so every run sees a first start: the one-off still due, the supplier down. The host recreates the schema.
  • Keep the scheduler running. Before acting, read the newest metadata id, then wait for runs of one manifest newer than it. Waits poll the database until a condition holds or a budget runs out, never a fixed sleep.
  • Mark the assembly [NonParallelizable]: the tests share one host and steer the same manifests.
  • Run a metadata cleanup sweep on demand by resolving IMetadataCleanupTrain from the host's services and calling Run(new MetadataCleanupRequest()).
  • Start a second factory with UseEnvironment("Production") to prove the demo key and the dashboard are gone.

SDK Reference

AddScheduler | Schedule | ThenInclude / Include | IDormantDependentContext | ScheduleOnce / TriggerAsync | Every / Cron | AddMetadataCleanup | ITraxScheduler | AddTraxGraphQL | AddTraxApiKeyAuth | AddTraxDashboard

Feature pages: Scheduling, Dependent Trains, Delayed / One-Off Jobs, Dead Letters & Cleanup, Operations queries.