Skip to content

Python API

Autobench exposes the same runtime through a fluent builder, typed specification models, and lower-level extension seams. Use the highest-level surface that can express the benchmark clearly.

Surface Selection

Surface Use it when
Benchmark Application code composes a benchmark dynamically
BenchmarkSpec You need the complete typed configuration surface
YAML + load_benchmark_spec Humans or agents author portable benchmark definitions
Runtime/evaluation functions You are building an adapter, service, or custom runner

All three authoring paths execute through run_benchmark_spec().

Fluent Builder

from autobench import (
    Benchmark,
    Case,
    Direction,
    ExactScorer,
    FactorValue,
    ObservationRole,
    PassFailScorer,
    Semantic,
    Variant,
)

benchmark = (
    Benchmark("builder-demo")
    .description("Compare current and candidate behavior.")
    .correlation(
        group_id="routing-proposal-42",
        attempt=1,
        phase="validation",
        labels={"owner": "evaluation"},
    )
    .dataset(
        [
            Case(
                id="refund",
                input={"message": "Refund order 42"},
                expected={"route": "billing"},
            )
        ],
        dataset_id="routing-regressions",
        version="v3",
    )
    .variants(
        [
            Variant(
                id="current",
                factors=[FactorValue(name="routing_profile", value="v3")],
            ),
            {
                "id": "candidate",
                "factors": {
                    "routing_profile": {
                        "value": "v4",
                        "optimize": True,
                    }
                },
            },
        ]
    )
    .task("my_app.benchmarks:run_case")
    .scoring(
        [
            ExactScorer(
                name="route",
                actual="output.route",
                expected="case.expected.route",
                semantic_type=Semantic.QUALITY_CORRECTNESS,
                direction=Direction.MAXIMIZE,
                role=ObservationRole.OBJECTIVE,
            ),
            PassFailScorer(
                name="success",
                path="output.success",
                semantic_type=Semantic.RESULT_SUCCESS,
                role=ObservationRole.CONSTRAINT,
            ),
        ]
    )
)

result = benchmark.run(experiment_id="routing-candidate-42", concurrency_limit=4)

Attach live lifecycle observers without changing the benchmark definition:

from autobench import ProgressEvent


def observe(event: ProgressEvent) -> None:
    print(event.sequence, event.kind, event.run_id, event.run_status)


result = benchmark.run(progress_handlers=(observe,))

Builder Methods

Method Configures
description(value) Benchmark description
correlation(...) Static invocation group, attempt, phase, ancestry hints, and scalar labels
capture(policy) ABP and asset capture policy
dataset(...) Inline cases or a typed dataset source
variants(items) Typed variants or normalized dictionaries
task(target, kind="python") Task target
scoring(items) Built-in or Python scorer specs
derive(items) Per-run derivers
instrument(*items) Typed built-ins or runtime custom instrumentors
instrument_all(...) Compatible built-in discovery
to_spec() Canonical BenchmarkSpec
run(...) / run_async(...) Sync or async execution

Post-derivation, policies, report views, and custom semantic registries currently live on the full BenchmarkSpec. Extend the compiled spec rather than inventing builder-only state:

import asyncio

from autobench import PolicySpec, run_benchmark_spec

spec = benchmark.to_spec().model_copy(
    update={
        "policies": [
            PolicySpec(
                name="quality-floor",
                metric=Semantic.QUALITY_CORRECTNESS,
                must_greater_equal=0.9,
            )
        ]
    }
)
result = asyncio.run(run_benchmark_spec(spec, concurrency_limit=4))

Task Contract

from autobench import Case, RunContext


def run_case(ctx: RunContext, case: Case) -> Result:
    ...

ctx is always first and case is always second. A task may be sync or async and may return any serializable result. A Pydantic model is useful because scorers can resolve output fields reliably.

The runtime resolves module:function targets relative to the benchmark file before falling back to normal Python import paths.

RunContext

RunContext owns evidence for one case x variant run:

Method Purpose
factor(name) Read a configured factor value
span(...) Time and nest an operation
metric(...) / metrics(...) Record numeric, boolean, or structured metrics
factor_observation(...) Record a factor discovered at runtime
event(...) Record a discrete event
diagnostic(...) Record non-objective evidence
outcome(...) Record semantic success
check(...) Record a correctness constraint and reason
record_measurement(...) Record summaries plus optional raw samples
artifact(...) Attach a payload
error(...) Preserve a structured error
attach_tracked_asset(...) Bind an explicit tracked asset version
await checkpoint(name) Commit all currently available evidence to durable staging

Evidence emitted before an exception remains in the failed run.

phase reports whether the run is resolving, executing, scoring, deriving, post-processing, or finalizing. Checkpoints preserve that phase automatically. Application tasks normally only call checkpoint():

async def run_case(ctx: RunContext, case: Case) -> Result:
    draft = await build_draft(case.input)
    ctx.artifact("draft", draft)
    await ctx.checkpoint("draft-built")
    return await validate_draft(draft)

The method requires an active Recorder; without one it raises RuntimeError. It is cancellation aware: if the caller is cancelled while persistence is active, Autobench gives the atomic write bounded time to finish and then propagates cancellation. Names beginning with autobench. belong to framework lifecycle checkpoints and are rejected for application calls.

Load And Run YAML

import asyncio
from pathlib import Path

from autobench import (
    ExecutionCorrelation,
    load_benchmark_spec,
    run_benchmark_path,
    run_benchmark_spec,
)

path = Path("benchmarks/routing.yaml")
spec = load_benchmark_spec(path)

sync_result = run_benchmark_path(
    path,
    experiment_id="routing-42",
    concurrency_limit=4,
)

async_result = asyncio.run(
    run_benchmark_spec(
        spec,
        experiment_id="routing-43",
        concurrency_limit=4,
        correlation=ExecutionCorrelation(attempt=2, labels={"review": "holdout"}),
        progress_handlers=(observe,),
    )
)

Import ExecutionCorrelation for invocation-level overrides. Explicit fields merge with spec.execution.correlation; omitted fields remain unchanged and label maps merge by key. The resolved value is immutable and identical on ExperimentResult, every RunResult, durable records, and replayed results. parent_experiment_id and resumed_from_experiment_id are grouping metadata only. Replay ancestry remains RecordLineage / parent_run_id, and Autobench does not infer workflow resume from either correlation field.

run_benchmark_path(), run_benchmark_spec(), Benchmark.run(), and Benchmark.run_async() share the same progress_handlers, progress_error_policy, and progress_error_handler contract. Bare handlers are strict by default; see Tasks and Runtime for ordering, terminal status, and backpressure guarantees.

Loading resolves dataset, pricing, task, and Python scorer references relative to the YAML file.

Record And Replay

from pathlib import Path

from autobench import (
    collect_benchmark_source_files,
    record_experiment,
    replay_experiment,
)

record_dir = Path("runs/routing-42")
record = record_experiment(
    async_result,
    record_dir,
    source_files=list(collect_benchmark_source_files(path)),
    path_root=Path.cwd(),
    durability="atomic",
)
replayed = replay_experiment(record_dir)

record_experiment() refuses to overwrite a non-empty experiment directory. It assembles the record in a temporary sibling, validates its integrity manifest, and publishes it atomically. Referenced tracked-asset histories and large trace artifacts are persisted automatically. Use durability="synced" when supported POSIX file and directory fsync calls are also required.

For durable execution, pass a recorder to the pipeline instead of waiting for the full result:

from autobench import FileRecorder

durable_result = asyncio.run(
    run_benchmark_spec(
        spec,
        experiment_id="routing-durable-44",
        concurrency_limit=4,
        recorder=FileRecorder(
            Path("runs/routing-durable-44"),
            source_files=collect_benchmark_source_files(path),
            path_root=Path.cwd(),
            durability="atomic",
        ),
    )
)

FileRecorder commits complete run snapshots during execution and atomically publishes the final directory after post-processing. Its sibling .routing-durable-44.staging directory survives an interrupted process. Use inspect_staging, recover_staging, and finalize_staging to recover it; use archive_staging or discard_staging for explicit lifecycle decisions.

Cooperative task cancellation, KeyboardInterrupt, and supported CLI SIGTERM handling commit a terminal partial checkpoint before propagating. A hard process kill can preserve only the last explicit checkpoint that had already returned; it cannot execute a final write.

from autobench import finalize_staging, inspect_staging

staging = Path("runs/.routing-durable-44.staging")
inspection = inspect_staging(staging)
if inspection.recoverable and inspection.missing_run_ids:
    partial_record = finalize_staging(
        staging,
        Path("runs/routing-durable-44-partial"),
        allow_partial=True,
    )

Recorder and RecordSession are the typed extension contracts for another persistence backend. The pipeline owns open, stage, finish or abort, and close; application code should not open a live session just to run a normal benchmark. ExperimentStart, ExecutionSnapshot, and PartialRunSnapshot are frozen transfer models. Recovery never imports the task or optional SDKs.

A custom session exposes immediate file and stream storage explicitly through artifact_sink: ArtifactSink | None. The session may return itself, as FileRecordSession does, or delegate to a separate local, remote, or composite sink. Returning None is valid, but a task that calls ctx.artifact_file() or ctx.artifact_stream() then receives ArtifactSinkRequiredError before Autobench consumes the source. The lifecycle session does not need to proxy the complete sink protocol merely to delegate storage.

ArtifactSink provides synchronous file/stream methods, prepare_file_async(), and prepare_stream_async(). Async implementations own in-flight transfers after caller cancellation and must settle them before the associated session closes or publishes its final record.

Load one exact record when building an audit or optimizer adapter:

from autobench import load_experiment_record, load_run_record

experiment = load_experiment_record(record_dir)
run = load_run_record(record_dir / experiment.run_paths[0], root_dir=record_dir)

Reports And Exports

from pathlib import Path

from autobench import (
    build_report,
    compare_variants,
    export_markdown_report,
    export_runs_csv,
    export_summary_yaml,
    load_experiment_record,
    write_markdown_report,
)

record = load_experiment_record(record_dir)
report = build_report(
    replayed,
    experiment_record=record,
    experiment_root=record_dir,
)
comparison = compare_variants(
    replayed,
    baseline="current",
    candidate="candidate",
)

export_summary_yaml(replayed, Path("analysis/summary.yaml"))
export_runs_csv(replayed, Path("analysis/runs.csv"))
export_markdown_report(replayed, Path("analysis/report.md"))
publication = write_markdown_report(
    report,
    Path("analysis/report-bundle"),
    layout="bundle",
    immutable_root=record_dir,
)

build_leaderboard, build_case_matrix, build_metric_distribution, and build_run_metric_rows expose individual projections. write_markdown_report() returns profile, selected layout, output paths, byte counts, and SHA-256 hashes. See Markdown Reports for configuration and safety boundaries.

For a case-level benchmark verdict, return a mapping or Pydantic model containing hard_pass, score, metrics, and feedback from the task. build_report() keeps that quality outcome separate from task execution status and projects it into KPIs, case tables, and purposeful inline SVG. Technical run, trace, asset, artifact, hash, and provenance detail belongs to audit.

Group or select several invocation results without changing their records:

from autobench import ExecutionCorrelation, build_grouped_reports, filter_experiments

validation = filter_experiments(
    results,
    correlation=ExecutionCorrelation(group_id="routing-proposal-42", phase="validation"),
)
groups = build_grouped_reports(results)

Filters match only explicitly supplied fields. Grouped reports retain each experiment report and summarize the attempts and phases present under each group_id.

Optional OTLP Export

from pathlib import Path

from autobench import OTLPSettings, export_record_otlp

delivery = export_record_otlp(
    Path("runs/routing-42"),
    settings=OTLPSettings(
        endpoint="https://collector.example/v1/traces",
        service_name="routing-benchmark",
    ),
)

Install autobench[otlp] on the exporting process. export_record_otlp() loads immutable record models; export_otlp() accepts already-loaded ExperimentRecord and RunRecord values. Both return OTLPExportResult and raise OTLPExportError without modifying evidence. Vendor settings remain separate from BenchmarkSpec. See OTLP Export.

Native Instrumentation

from autobench import Benchmark

benchmark = Benchmark("agent").instrument_all(
    exclude={"httpx"},
    strict=False,
    assets={
        "representations": ["definition", "effective"],
        "include": ["prompt", "tool", "output_schema"],
    },
)

Unavailable integrations become diagnostic observations. strict=True instead requires every selected integration to be compatible.

Use typed settings for explicit control:

from autobench import HTTPXCaptureSettings, HTTPXInstrumentation, OpenAIInstrumentation

benchmark.instrument(
    OpenAIInstrumentation(),
    HTTPXInstrumentation(
        capture=HTTPXCaptureSettings(
            path="hash",
            response_headers=("x-request-id",),
        )
    ),
)

Explicit settings override automatic discovery, including enabled=False. A custom runtime Instrumentor can also be passed to instrument() and remains Python-only.

Explicit Tracking

from autobench import track

SYSTEM_PROMPT = track.prompt(
    name="support_system",
    source="prompts/support.md",
)


@track.tool
def lookup_order(order_id: str) -> dict[str, str]:
    """Return the current order status."""
    ...

track.prompt, track.tool, track.type, track.dataclass, and track.asset register exact versions. track.write_assets(path) writes DSL-shaped manifests plus one content.sqlite3 registry. load_asset_content(...) resolves an exact historical snapshot and load_asset_diff(...) resolves the corresponding readable diff. Native discovery can attach unadorned SDK-visible components to runs. Experiment recording uses the same contract at artifacts/asset-content.sqlite3.

Production And Generated Cases

from pathlib import Path

from autobench import (
    CaseGeneratorInput,
    SamplingPolicy,
    generate_dataset_sync,
    generated_batch_from_cases,
    samples_to_cases,
    write_generation_result,
)

review_cases = samples_to_cases(production_samples, policy=SamplingPolicy(max_samples=50))
generated = generated_batch_from_cases(
    synthetic_cases,
    generator_asset_version="prompt.generator@v4",
    model_provider="openrouter",
    model_name="openai/gpt-5.6-luna",
)

result = generate_dataset_sync(
    generate_cases,
    CaseGeneratorInput(seed=17, settings={"count": 20}),
    generator_id="generation:generate_cases",
    dataset_id="generated-routing",
    version="v1",
)
write_generation_result(result, Path("datasets/generated-routing.yaml"))

Production helpers normalize reviewed samples. Generated-dataset APIs own the typed preparation, hashing, review projection, manifest, and safe publication boundary, while the application owns the actual generator/provider logic. Generation finishes before normal benchmark planning. See Generated Datasets.

Extension Rules

  • Put subject execution in a task.
  • Put domain judgment in a Python scorer.
  • Use a deriver for same-run computations and a post-deriver for matched runs.
  • Use a policy for acceptance boundaries.
  • Use an instrumentor for a stable SDK boundary.
  • Use source maps and extractors for external field normalization.
  • Use metric packs for reusable domain defaults.
  • Never mutate recorded evidence; create a derived record or a new experiment.

See API Reference for generated signatures and model fields.