Skip to content

Autobench

Define what better means once. Autobench evaluates every version by the same rules, compares the results, and keeps the full history ready for inspection.

Applications change: a team may switch models, revise a prompt, replace a tool, tune an algorithm, or ship a new configuration. To decide whether the change is actually better, they commonly write a benchmark script. That script runs representative inputs, checks the outputs, records values such as correctness, latency, token usage, or cost, and compares one version with another.

The script works, but every project tends to build this machinery again. Results use incompatible formats, measurement and scoring logic become mixed with application code, and an old result is often impossible to inspect without rerunning the original program.

We built Autobench to solve this problem. You describe the inputs to test, the variants to compare, the application task, and the meaning of success in YAML or Python. Autobench then runs the full matrix, collects measurements and traces, evaluates each result, records the application assets that affected it, and stores an immutable experiment record. That same record can be replayed, reported, compared, or exported later without calling the application again.

uv add autobench
autobench validate autobench.yaml
autobench run autobench.yaml --record runs/latest

After the run, inspect the recorded experiment without executing the application again:

autobench replay runs/latest
autobench report runs/latest
autobench compare runs/latest --baseline current --candidate proposed
autobench export runs/latest --format csv --path analysis/runs.csv

The Framework Loop

BenchmarkSpec
  Dataset[Case] x Variant[Factor]
    -> task(ctx, case)
    -> observations + ABP trace + artifacts + asset versions
    -> scorers + per-run derivation
  -> cross-run derivation + policies
  -> immutable RunRecords
  -> replay + Rich reports + comparison + exports + optimization feedback

The task is the only application-specific part. Autobench owns the repeated infrastructure around it: matrix planning, context propagation, instrumentation, scoring, derivation, persistence, reporting, and replay.

Why Semantic Evidence Matters

Raw names such as prompt_tokens, input_tokens, accuracy, and answer_quality are local conventions. Autobench observations can also declare stable meaning:

llm.tokens.input
llm.tokens.output
llm.model.name
quality.correctness
time.latency
money.cost
agent.tool.argument.correctness

That semantic layer lets reports, pricing derivation, policy checks, and optimization systems use evidence from different applications without guessing what every local metric name means.

What You Can Benchmark

Autobench is optimized for AI systems but does not require one:

System Cases Variants Evidence
LLM application prompts and expected answers model, prompt, temperature quality, tokens, latency, cost
Agent user goals and expected actions instructions, tools, model action selection, arguments, sequence, completion
Search or retrieval queries and relevant items index, reranker, limits recall, precision, latency
Service/API requests and expected responses release, configuration correctness, errors, throughput, SLA
Algorithm input fixtures implementation correctness, repeated timings, speedup
Data pipeline source batches parser or policy coverage, validity, loss, runtime

See Use Cases for complete patterns.

Core Capabilities

Area Included
Definition Human-readable YAML DSL, typed Python builder, JSON Schema completion
Data Inline/file/glob datasets, defaults, attachments, generated and production cases
Execution Sync/async tasks, deterministic matrices, bounded concurrency, failure isolation
Evidence Semantic observations, checks, events, artifacts, measurements, ABP traces
Evaluation Built-in and custom scorers, expected actions, policies, metric packs
Derivation Token cost, tiered pricing, paired baselines, comparison verdicts
Instrumentation Manual spans, method instrumentation, Pydantic AI, pydantic-gepa, OpenAI, Agents, HTTPX
Lineage Explicit and automatic prompt/tool/schema/capability/agent asset versioning
Persistence Immutable YAML records, source hashes, environment metadata, portable artifacts
Analysis Replay, Rich reports, leaderboards, matrices, distributions, comparisons, exports

Choose A Starting Point

Goal Read
Install the right extras Installation
Run a complete benchmark First Benchmark
Find a pattern for your system Use Cases
Understand ownership and data flow Architecture
Author the full DSL YAML Spec
Compose benchmarks in Python Python API
Instrument an existing SDK application Native Instrumentation
Record an optimizer run and candidate lineage Pydantic-GEPA Instrumentation
Collect prompt/tool/schema lineage automatically Automatic Asset Discovery
Inspect all shipped features Capability Map

Project Boundaries

Autobench records and evaluates evidence. It does not own your application, make causal claims from confounded runs, keep provider pricing permanently current, or choose an optimization algorithm. Those boundaries keep the core usable for arbitrary systems while allowing pydantic-gepa, autoptimize, or another consumer to build on stable experiment records.