All research

Deterministic evaluation benchmarks for autonomous multi-model scientific systems

A benchmark protocol for separating model regressions from changing tools, environments, and datasets in long-horizon scientific agent workflows.

5 MINUTES READUPDATED

Deterministic evaluation benchmarks for autonomous multi-model scientific systems

Research status: Protocol note. This document defines an evaluation architecture and reporting standard. It does not report completed benchmark results.

Evaluating a scientific agent is difficult because the system is larger than the model. A run may depend on a prompt, a model release, a planner, retrieval, external databases, generated code, a container image, and an evaluator. If several of those change between runs, a score cannot identify the source of improvement or regression.

This note defines a deterministic evaluation system for isolating those variables. Here, deterministic means holding the evaluation boundary stable enough to attribute and inspect differences, while allowing model outputs to vary.

Evaluation unit

The basic unit is a versioned task package. Each package contains:

  • a scientific question or computational objective;
  • immutable or versioned input artifacts;
  • the allowed tools and their recorded contracts;
  • resource, token, and wall-clock budgets;
  • expected output schemas;
  • deterministic checks and domain invariants; and
  • a scoring rubric for judgments that cannot be reduced to exact assertions.

The package also identifies what it does not test. A literature-synthesis benchmark cannot establish whether a proposed wet-lab procedure is safe or experimentally valid.

Sources of variance

We separate four sources before comparing systems.

Model variance

Sampling, model revisions, and hidden provider changes can alter an output. The evaluator records the model identifier, available revision metadata, sampling controls, system instructions, and full message history. Repeated trials estimate the distribution rather than relying on a single run.

Tool variance

External APIs and scientific databases change schemas and records. Evaluation runs should use captured fixtures when the task permits it. When live data is essential, the evaluator stores request parameters, response hashes, timestamps, and database release identifiers.

Environment variance

Dependencies, compilers, numerical libraries, and accelerator kernels can affect execution. Every run references an environment digest and source revision. Exact hardware details should be recorded when the evaluated property depends on them.

Evaluator variance

Model graders can be as unstable as the systems they judge. Prefer executable assertions and domain invariants where possible. For rubric-based grading, freeze the grader configuration, randomize answer order, conceal system identity, and measure agreement across graders or human reviewers.

Evaluation configuration

interface DeterministicHarnessConfig {
  taskVersion: string;
  trials: number;
  seed?: number;
  environmentDigest: string;
  toolFixtures: Record<string, { version: string; digest: string }>;
  model: { provider: string; id: string; revision?: string };
  budgets: { tokens: number; toolCalls: number; wallTimeMs: number };
  verification: Array<
    | { kind: 'exact'; expectedArtifact: string }
    | { kind: 'schema'; schemaId: string }
    | { kind: 'invariant'; checkId: string }
    | { kind: 'rubric'; rubricVersion: string }
  >;
}

The configuration travels with the result. A dashboard score without this record is insufficient for regression analysis.

Evaluate traces as well as answers

Two systems can produce similar final reports through very different behavior. One may cite and inspect primary evidence; another may guess successfully. Long-horizon evaluation should therefore score the trace.

Trace-level measures can include:

  • valid versus invalid tool calls;
  • repeated calls with unchanged inputs;
  • unsupported claims introduced during synthesis;
  • use of evidence that was later superseded;
  • recovery after an execution failure;
  • budget consumed before reaching a valid artifact; and
  • whether the final claim resolves to the source and run that support it.

The trace also lets a maintainer distinguish reasoning failure from infrastructure failure.

Exact checks, invariants, and rubrics

Use the strongest check the task supports.

An exact check works for a file hash, a parsed identifier, or a known fixture response. A schema check verifies that an artifact is complete and typed. An invariant checks properties that should survive harmless variation, such as unit consistency, graph connectivity, conservation constraints, or the presence of required controls. A rubric is reserved for qualities such as explanatory adequacy that require judgment.

Mixing these checks is often necessary. The evaluator should report each result separately instead of collapsing everything into one opaque score.

Regression isolation

When a system changes, evaluate one layer at a time:

  1. replay the previous system against the frozen task package;
  2. run the new system against the same package;
  3. compare failures by task, check, and trace event;
  4. change live tools or datasets only in a separate experiment; and
  5. document interactions when two changes cannot be isolated.

This process does not remove stochastic variance. It makes the variance visible and prevents an API update from being reported as a model improvement.

Reporting standard

A benchmark report should include task definitions, exclusions, system versions, trial count, budgets, environment digests, fixture dates, per-check results, failure categories, and representative traces. It should publish negative findings and unresolved evaluator disagreement.

Future work will apply this protocol to bounded repository and literature tasks. Until those runs and artifacts are published, this note should be cited as an evaluation design rather than empirical validation.

The runtime assumptions behind the protocol are described in deterministic cloud runtimes. The memory layer required for multi-run tasks appears in recursive context accumulation.

Cite this research note

BibTeX entry for academic citations, literature trackers, and generative research synthesizers:

                  @article{kalaris2026_deterministic_evals,
  title = {Deterministic evaluation benchmarks for autonomous multi-model scientific systems},
  author = {Chowdhury, Sayan},
  journal = {Kalaris Labs Research Notes},
  year = {2026},
  url = {https://kalarislabs.com/research/deterministic-evals}
}
                

Continue reading

agents

Recursive context accumulation in autonomous scientific environments

How papers, repositories, execution traces, and failed experiments become versioned evidence that later agent runs can retrieve and verify.

observability

Skill Doctor: Autonomous diagnostics for model tooling and observability

A diagnostic design for checking agent tool contracts, permissions, dependencies, and live availability before failures spread through long workflows.

edge

Edge delivery without architectural fragmentation

An architecture note on serving prerendered research content and a small number of request-time operations through one observable deployment boundary.