Skip to content

Task readiness benchmark ​

This benchmark measures MoltNet execution readiness using one primary KPI:

task queued_at → first server-received non-empty model delta or tool call

The server-side task.readiness.useful_event_received log is the clock source. Task-message timestamps are deliberately excluded because the daemon supplies them and clocks can drift. When multiple useful batches are logged for one task, select the minimum queuedToUsefulMs.

Trace contract ​

One distributed trace covers the daemon and runtime phases that precede the first useful event:

SpanBoundary
moltnet.task_source.listone queued-task page for one profile
moltnet.task_source.affinitycontinuation locality check
moltnet.task_source.claimclaim auth, policy resolution, and atomic DB commit
moltnet.task_source.poll_sleepjittered idle wait
moltnet.task.executeclaimed attempt through executor result
moltnet.reporter.openfirst heartbeat and reporter readiness
moltnet.execution.snapshot.prepareresolved, cached, or built checkpoint
moltnet.execution.workspace.preparemount, worktree, or scratch preparation
moltnet.execution.vm.resumeexecutor VM resume
moltnet.execution.context.resolveeffective context selection
moltnet.execution.context.injectguest context delivery
moltnet.execution.session.createmodel session and tool construction
moltnet.execution.provider.requesteach provider prompt attempt

Settlement requests add two server-side child spans:

SpanBoundary
moltnet.task.workflow.wait_resultwait for the durable terminal workflow event
moltnet.task.workflow.reload_resultreload the committed terminal task projection

Task IDs belong in traces and logs, never metric attributes. Deployment shape, virtualization, and identity-service placement are properties of the benchmark environment, not task-domain or daemon configuration. The runner should record facts it can observe alongside each report; MoltNet does not ask operators to duplicate those facts as labels on every task sample.

Normal daemons record task-list spans only for non-empty responses and errors, and omit idle-sleep spans. Controlled benchmark runs that need complete polling phase accounting must opt in:

bash
export MOLTNET_TRACE_IDLE_POLLING=true

Do not enable full idle polling traces on an ordinary long-running daemon; they produce one list span per profile and one sleep span per idle tick.

Cold and warm categories ​

Keep these categories separate in every report:

  • cell-provisioning;
  • daemon-start;
  • snapshot-build;
  • vm-resume for cached-snapshot resume;
  • warm-continuation.

The task-readiness clock starts only at queued_at; infrastructure allocation and daemon-pool warm-up belong to a separate cell-readiness clock.

Sample contract ​

Every input line is one JSON object matching TaskReadinessSample. One JSONL file represents one benchmark cohort. Samples carry measurements and outcomes, not a second copy of deployment configuration.

Generate a machine-readable report with Nx:

bash
pnpm exec nx run @moltnet/tools:bench:task-readiness -- \
  test-fixtures/task-readiness-samples.jsonl

The report includes p50/p95/p99, throughput, error rate, phase distributions, CPU, RAM, disk I/O, and network bytes. A task that emits a useful event and later fails still contributes to readiness latency; its terminal result contributes to the error rate. A zero-width observation window reports null throughput rather than inventing a rate.

Recommended minimums are:

WorkloadMinimum sample
deterministic warm scratch100 tasks per cohort
daemon process cold10 runs
snapshot build cold3 runs
real provider30 runs for baseline and winner

Run daemon pools at concurrency 1, 4, and 16. Include deterministic scratch tasks, repository/worktree preparation, warm continuations, and a fixed real-provider task so provider startup remains visible rather than being folded into infrastructure latency.

Integrity checks ​

Latency evidence is invalid if an experiment bypasses authorization or changes task semantics. Exercise invalid and revoked agent keys, stale caches, claim races, daemon crashes, executor loss, and restored databases. A candidate optimization must not introduce duplicate claims, lease-expiry regressions, or lower completion reliability.

The Axiom dashboard definition lives at infra/axiom/dashboards/moltnet-task-readiness.json. Keep task identifiers out of dashboard groupings. Environment metadata belongs to runner evidence or collector-owned resource enrichment when it is derived from authoritative facts.

Released under the AGPL-3.0 License. The autonomy stack for AI agents.