Example Projects¶
The repository examples use the public Autobench runtime. They are ordered by the amount of framework surface they demonstrate, not by whether the subject is AI-based.
Offline Release Matrix¶
These examples are credential-free and run in make examples:
| Example | Subject | Main features |
|---|---|---|
minimal |
text transformation | inline cases, variants, exact score, matrix, comparison |
basic |
support routing | file dataset, spans, checks, artifacts, Rich reports |
mid |
response generation | semantic usage, pricing, cost, policies, distributions |
advanced |
search implementations | repeated samples, noise, paired speedup |
abp_manual |
ticket router | manual span plus method instrumentation |
abp_concurrent |
async workers | task-local trace context and concurrent runs |
automatic_assets |
Pydantic AI and custom SDK | automatic behavioral asset lineage |
generated_dataset |
support-routing case preparation | typed generator, request YAML, review state, frozen dataset and provenance manifest |
otlp_export |
immutable ABP record | offline OTLP hierarchy mapping through an injected exporter |
pydantic_gepa |
Optimize Anything Pipeline | optimizer lifecycle, engine branches, evaluation budgets, candidate lineage, and component asset versions |
Run all offline examples:
make examples
Minimal: Learn The Matrix¶
autobench validate examples/minimal/autobench.yaml
autobench run examples/minimal/autobench.yaml --record /tmp/autobench-minimal
autobench replay /tmp/autobench-minimal
Read examples/minimal/autobench.yaml together with minimal_benchmark.py. This is the shortest
complete case x variant -> task -> score -> record -> report implementation.
Generated Dataset: Prepare Before Planning¶
cd examples/generated_dataset
autobench dataset generate generator:generate_routing_cases \
--request request.yaml \
--output generated-cases.yaml \
--id routing-generated \
--version v1
This example is deterministic and credential-free. It writes a normal dataset and a separate generation manifest, showing the boundary between data preparation and benchmark execution.
Basic: Application Evidence¶
autobench run examples/basic/autobench.yaml --record /tmp/autobench-basic
autobench report /tmp/autobench-basic \
--format markdown --layout bundle --output /tmp/autobench-basic-report
The task validates typed input, reads a factor, opens a workflow span, stores its output as an
artifact, and lets declarative scorers evaluate correctness and handling. The candidate fixes a
known routing failure, so the case matrix and comparison contain a visible behavioral delta.
Its YAML also publishes reports/benchmark.md inside the record before the manifest is sealed; the
second command demonstrates a post-hoc bundle generated from replayed evidence.
Mid: Quality, Cost, And Constraints¶
autobench run examples/mid/autobench.yaml --record /tmp/autobench-mid
This example records input/output tokens and latency, resolves a local model pricing table, derives
money.cost, checks success and cost policies, and configures leaderboard and distribution views.
It is the best starting point for an LLM benchmark that already has a task implementation.
Advanced: Measurement And Paired Baselines¶
autobench run examples/advanced/autobench.yaml --record /tmp/autobench-advanced
The task uses measure_callable() and ctx.record_measurement() instead of custom timing loops.
The post-deriver matches runs by case and computes candidate speedup against the baseline while
correctness remains a constraint.
Pydantic AI: Live Layered Instrumentation¶
uv sync --extra instrumentation
export OPENROUTER_API_KEY=...
export OPENROUTER_MODEL=openrouter:openai/gpt-5.6-luna
uv run python examples/pydantic_ai/openrouter_instrument_all.py \
--record /tmp/autobench-openrouter
The program makes a real OpenRouter request through Pydantic AI, uses a tool, streams structured
output, and calls Benchmark.instrument_all(). The task has no manual metrics, spans, or tracking
decorators. Autobench collects layered Pydantic AI, OpenAI client, and HTTPX evidence plus prompt,
tool, output-schema, and agent versions.
Inspect it afterward:
autobench instrumentation trace /tmp/autobench-openrouter
autobench replay /tmp/autobench-openrouter
examples/pydantic_ai/agent_benchmark.py is provider-neutral and accepts any configured Pydantic AI
model identifier through PYDANTIC_AI_MODEL.
Pydantic-GEPA: Optimizer Evidence¶
uv run autobench run examples/pydantic_gepa/autobench.yaml \
--record /tmp/autobench-pydantic-gepa
uv run autobench report /tmp/autobench-pydantic-gepa
The directory contains four credential-free benchmarks: standard GEPA, Optimize Anything Omni, multi-component prompt/tool/output-schema optimization, and staged checkpoint/resume. Native instrumentation records engine contenders, selection evidence, resource-specific budgets, candidate lineage, effective component versions, and a replayable typed projection.
for name in standard autobench multi_component resume; do
uv run autobench run "examples/pydantic_gepa/$name.yaml" \
--record "/tmp/autobench-pydantic-gepa-$name"
done
An optional live example layers Pydantic AI, OpenAI-compatible OpenRouter, and HTTPX evidence under the optimizer evaluation without duplicate accounting:
export OPENROUTER_API_KEY=...
uv run python examples/pydantic_gepa/live_pydantic_ai.py \
--record /tmp/autobench-pydantic-gepa-live
It uses openrouter:openai/gpt-5.6-luna and is intentionally excluded from offline CI. See
Pydantic-GEPA Instrumentation.
Automatic Asset Discovery¶
uv run python examples/automatic_assets/pydantic_ai_discovery.py \
--record /tmp/autobench-pydantic-assets
uv run python examples/automatic_assets/custom_sdk_discovery.py \
--record /tmp/autobench-custom-assets
Both are offline. The first uses a real Pydantic AI Agent, AbstractCapability, tool, and Pydantic
output model with no explicit tracking. The second adds prompt, tools, and output-schema extraction
to an arbitrary method with InstrumentAssetSpec.
ABP Manual And Concurrent¶
autobench run examples/abp_manual/autobench.yaml --record /tmp/abp-manual
autobench run examples/abp_concurrent/autobench.yaml \
--concurrency 2 \
--record /tmp/abp-concurrent
Use the manual example to learn RunContext.span() and instrument_method(). Use the concurrent
example to inspect sibling span parentage and task-local context under async execution.
OpenAI Streaming¶
uv sync --extra instrumentation
uv run python examples/abp_openai/run_openai_streaming.py
This uses the official OpenAI client and a real streaming parser over an offline HTTPX mock transport. It demonstrates first-chunk and stream-completion evidence without network access.
OpenAI Agents¶
uv sync --extra openai-agents
uv run python examples/abp_openai_agents/run_openai_agents.py
The example sends real OpenAI Agents workflow/function/custom trace events through the Autobench trace processor. It requires no model request.
Replay And Extraction¶
uv run python examples/abp_replay/replay_and_extract.py /tmp/recorded-experiment
The script loads records without provider SDKs and creates extraction-derived records with explicit parent lineage.
Offline OTLP Export¶
uv run autobench run examples/abp_manual/autobench.yaml --record /tmp/abp-manual
uv run python examples/otlp_export/export_record.py /tmp/abp-manual
The example maps a real recorded experiment to OTel SDK spans through an injected in-memory
exporter, so it verifies hierarchy and delivery without a collector or network request. Production
delivery uses autobench telemetry export; see OTLP Export.
CodeMode: Migrating A Real Benchmark Runner¶
export OPENROUTER_API_KEY=...
uv run python examples/codemode/run_benchmark.py --only parse_cron \
--record /tmp/autobench-codemode
The CodeMode example replaces a bespoke benchmark script with cases, model-pair factors, a task, semantic coverage/success/latency evidence, generated-spec artifacts, recording, and reports. Its task still owns Vowel CodeMode calls; Autobench remains generic. The external CodeMode runtime and network credentials are required.
What To Copy¶
Copy the pattern, not generated run directories:
- task signature and typed input/output from
minimalorbasic; - pricing and policies from
mid; - measurement and paired comparison from
advanced; - automatic SDK setup from
pydantic_ai; - optimizer evidence from
pydantic_gepa; - custom instrumentation from
automatic_assets; - replay processing from
abp_replay.
For combinations not represented by one project, use Use Cases and the Capability Map.