Skip to content

Example Projects

The repository examples use the public Autobench runtime. They are ordered by the amount of framework surface they demonstrate, not by whether the subject is AI-based.

Offline Release Matrix

These examples are credential-free and run in make examples:

Example Subject Main features
minimal text transformation inline cases, variants, exact score, matrix, comparison
basic support routing file dataset, spans, checks, artifacts, Rich reports
mid response generation semantic usage, pricing, cost, policies, distributions
advanced search implementations repeated samples, noise, paired speedup
abp_manual ticket router manual span plus method instrumentation
abp_concurrent async workers task-local trace context and concurrent runs
automatic_assets Pydantic AI and custom SDK automatic behavioral asset lineage
generated_dataset support-routing case preparation typed generator, request YAML, review state, frozen dataset and provenance manifest
otlp_export immutable ABP record offline OTLP hierarchy mapping through an injected exporter
pydantic_gepa Optimize Anything Pipeline optimizer lifecycle, engine branches, evaluation budgets, candidate lineage, and component asset versions

Run all offline examples:

make examples

Minimal: Learn The Matrix

autobench validate examples/minimal/autobench.yaml
autobench run examples/minimal/autobench.yaml --record /tmp/autobench-minimal
autobench replay /tmp/autobench-minimal

Read examples/minimal/autobench.yaml together with minimal_benchmark.py. This is the shortest complete case x variant -> task -> score -> record -> report implementation.

Generated Dataset: Prepare Before Planning

cd examples/generated_dataset
autobench dataset generate generator:generate_routing_cases \
  --request request.yaml \
  --output generated-cases.yaml \
  --id routing-generated \
  --version v1

This example is deterministic and credential-free. It writes a normal dataset and a separate generation manifest, showing the boundary between data preparation and benchmark execution.

Basic: Application Evidence

autobench run examples/basic/autobench.yaml --record /tmp/autobench-basic
autobench report /tmp/autobench-basic \
  --format markdown --layout bundle --output /tmp/autobench-basic-report

The task validates typed input, reads a factor, opens a workflow span, stores its output as an artifact, and lets declarative scorers evaluate correctness and handling. The candidate fixes a known routing failure, so the case matrix and comparison contain a visible behavioral delta. Its YAML also publishes reports/benchmark.md inside the record before the manifest is sealed; the second command demonstrates a post-hoc bundle generated from replayed evidence.

Mid: Quality, Cost, And Constraints

autobench run examples/mid/autobench.yaml --record /tmp/autobench-mid

This example records input/output tokens and latency, resolves a local model pricing table, derives money.cost, checks success and cost policies, and configures leaderboard and distribution views. It is the best starting point for an LLM benchmark that already has a task implementation.

Advanced: Measurement And Paired Baselines

autobench run examples/advanced/autobench.yaml --record /tmp/autobench-advanced

The task uses measure_callable() and ctx.record_measurement() instead of custom timing loops. The post-deriver matches runs by case and computes candidate speedup against the baseline while correctness remains a constraint.

Pydantic AI: Live Layered Instrumentation

uv sync --extra instrumentation
export OPENROUTER_API_KEY=...
export OPENROUTER_MODEL=openrouter:openai/gpt-5.6-luna
uv run python examples/pydantic_ai/openrouter_instrument_all.py \
  --record /tmp/autobench-openrouter

The program makes a real OpenRouter request through Pydantic AI, uses a tool, streams structured output, and calls Benchmark.instrument_all(). The task has no manual metrics, spans, or tracking decorators. Autobench collects layered Pydantic AI, OpenAI client, and HTTPX evidence plus prompt, tool, output-schema, and agent versions.

Inspect it afterward:

autobench instrumentation trace /tmp/autobench-openrouter
autobench replay /tmp/autobench-openrouter

examples/pydantic_ai/agent_benchmark.py is provider-neutral and accepts any configured Pydantic AI model identifier through PYDANTIC_AI_MODEL.

Pydantic-GEPA: Optimizer Evidence

uv run autobench run examples/pydantic_gepa/autobench.yaml \
  --record /tmp/autobench-pydantic-gepa
uv run autobench report /tmp/autobench-pydantic-gepa

The directory contains four credential-free benchmarks: standard GEPA, Optimize Anything Omni, multi-component prompt/tool/output-schema optimization, and staged checkpoint/resume. Native instrumentation records engine contenders, selection evidence, resource-specific budgets, candidate lineage, effective component versions, and a replayable typed projection.

for name in standard autobench multi_component resume; do
  uv run autobench run "examples/pydantic_gepa/$name.yaml" \
    --record "/tmp/autobench-pydantic-gepa-$name"
done

An optional live example layers Pydantic AI, OpenAI-compatible OpenRouter, and HTTPX evidence under the optimizer evaluation without duplicate accounting:

export OPENROUTER_API_KEY=...
uv run python examples/pydantic_gepa/live_pydantic_ai.py \
  --record /tmp/autobench-pydantic-gepa-live

It uses openrouter:openai/gpt-5.6-luna and is intentionally excluded from offline CI. See Pydantic-GEPA Instrumentation.

Automatic Asset Discovery

uv run python examples/automatic_assets/pydantic_ai_discovery.py \
  --record /tmp/autobench-pydantic-assets

uv run python examples/automatic_assets/custom_sdk_discovery.py \
  --record /tmp/autobench-custom-assets

Both are offline. The first uses a real Pydantic AI Agent, AbstractCapability, tool, and Pydantic output model with no explicit tracking. The second adds prompt, tools, and output-schema extraction to an arbitrary method with InstrumentAssetSpec.

ABP Manual And Concurrent

autobench run examples/abp_manual/autobench.yaml --record /tmp/abp-manual
autobench run examples/abp_concurrent/autobench.yaml \
  --concurrency 2 \
  --record /tmp/abp-concurrent

Use the manual example to learn RunContext.span() and instrument_method(). Use the concurrent example to inspect sibling span parentage and task-local context under async execution.

OpenAI Streaming

uv sync --extra instrumentation
uv run python examples/abp_openai/run_openai_streaming.py

This uses the official OpenAI client and a real streaming parser over an offline HTTPX mock transport. It demonstrates first-chunk and stream-completion evidence without network access.

OpenAI Agents

uv sync --extra openai-agents
uv run python examples/abp_openai_agents/run_openai_agents.py

The example sends real OpenAI Agents workflow/function/custom trace events through the Autobench trace processor. It requires no model request.

Replay And Extraction

uv run python examples/abp_replay/replay_and_extract.py /tmp/recorded-experiment

The script loads records without provider SDKs and creates extraction-derived records with explicit parent lineage.

Offline OTLP Export

uv run autobench run examples/abp_manual/autobench.yaml --record /tmp/abp-manual
uv run python examples/otlp_export/export_record.py /tmp/abp-manual

The example maps a real recorded experiment to OTel SDK spans through an injected in-memory exporter, so it verifies hierarchy and delivery without a collector or network request. Production delivery uses autobench telemetry export; see OTLP Export.

CodeMode: Migrating A Real Benchmark Runner

export OPENROUTER_API_KEY=...
uv run python examples/codemode/run_benchmark.py --only parse_cron \
  --record /tmp/autobench-codemode

The CodeMode example replaces a bespoke benchmark script with cases, model-pair factors, a task, semantic coverage/success/latency evidence, generated-spec artifacts, recording, and reports. Its task still owns Vowel CodeMode calls; Autobench remains generic. The external CodeMode runtime and network credentials are required.

What To Copy

Copy the pattern, not generated run directories:

  • task signature and typed input/output from minimal or basic;
  • pricing and policies from mid;
  • measurement and paired comparison from advanced;
  • automatic SDK setup from pydantic_ai;
  • optimizer evidence from pydantic_gepa;
  • custom instrumentation from automatic_assets;
  • replay processing from abp_replay.

For combinations not represented by one project, use Use Cases and the Capability Map.