Troubleshooting¶
Start with the narrowest command that can identify the failing layer.
Task Module Cannot Be Imported¶
Could not import task module 'benchmarks.tasks'
Task targets must use module:function, not a file path:
run:
python: benchmark_task:run
Autobench first uses normal Python imports, then searches relative to the benchmark spec. Common fixes:
- place
benchmark_task.pynext toautobench.yamland usebenchmark_task:run; - for a package, ensure package directories have the expected Python import structure;
- do not include
.pyin the target; - run
autobench validate path/to/autobench.yamlfrom any directory to test resolution.
The function must accept (ctx, case) in that order.
YAML Validates In One Editor But Not In Autobench¶
The schema directive improves editor completion; Autobench's installed Pydantic models remain the runtime authority. Match the schema version to the installed package:
python -c "import autobench; print(autobench.__version__)"
# yaml-language-server: $schema=./schemas/0.3.0/benchmark_schema.json
Run autobench validate and use its file/line diagnostics. Unknown scorer, policy,
instrumentation, and capture fields are rejected intentionally.
Dataset File Is Not Found¶
file:// references and glob patterns resolve relative to the benchmark YAML, not the shell's
current directory:
dataset:
source: file://datasets/cases.yaml
For a glob:
dataset:
source: file://datasets/cases/*.yaml
An unmatched glob is an error. Dataset files can contain a dataset DSL document, a case list, or a single case mapping.
Record Directory Already Exists¶
Autobench records are immutable. record_experiment() and autobench run --record refuse to
overwrite an existing experiment:
Experiment record already exists
Use a new directory or remove/archive the old directory explicitly outside Autobench. Do not merge unrelated experiments by copying run files together.
A Metric Is Missing From Reports¶
Reports query semantic types, not only local names. Check:
- the task/scorer/deriver emitted the observation;
semantic_typematches the report metric;- the value is numeric or boolean for the selected aggregate;
- the selected span/query is not filtering it out;
- source precedence did not intentionally select a score or derived observation instead.
Inspect the per-run YAML or use Python:
from autobench import ObservationQuery
query = ObservationQuery(observations=run.task_result.observations)
matches = query.exact("money.cost")
Missing cost is not converted to zero. Ensure token, model, provider, and pricing inputs are all available to the token-cost deriver.
Paired Baseline Does Not Produce A Value¶
The baseline and candidate must match on the configured key, normally case_id, and both must have
the source metric. Check:
baseline_variantexactly matches a variant ID;- both runs emit the same semantic metric;
- the metric unit is compatible;
match_onidentifies a unique counterpart;- the configured missing policy is appropriate.
Use a case matrix for the source metric before debugging the formula.
instrument_all() Records Skipped Integrations¶
This is normal when optional SDKs are not installed. Automatic discovery records
instrumentation.skipped diagnostics and continues by default.
autobench instrumentation doctor
Install the relevant extra, remove the integration from exclude, or use strict=True when absence
must fail the benchmark.
Explicit enabled: false wins over automatic discovery. A custom runtime instrumentor with the
same ID also prevents a duplicate built-in installation.
No Automatic Assets Appear¶
Automatic discovery only observes values that cross a supported instrumented SDK boundary while a benchmark run is active. Check:
- the corresponding instrumentor is compatible and installed;
- asset discovery is enabled;
includecontains the expected family;representationsincludesdefinitionoreffectiveas needed;- the SDK call occurs inside the task;
- capture policy does not reduce the asset below the expected content level.
Run the offline examples/automatic_assets/ programs to separate environment issues from
application behavior.
Pydantic-GEPA Optimization Evidence Is Missing¶
The observer records only while an Autobench run context is active. Check:
autobench[pydantic-gepa]is installed on Python 3.11-3.13;autobench instrumentation doctorreports event contract version1as compatible;- the YAML contains
instrumentation.pydantic_gepaorinstrument_all()selected it; - the optimization call occurs inside
task(ctx, case); - the instrumentor is not disabled or suppressed;
- the record contains the
autobench.pydantic_gepa/v1extension.
Use detail: evaluations or detail: full when per-candidate and per-case spans are expected.
summary intentionally omits those spans but still records budgets, selections, candidates, and
the durable projection. The complete workflow is in
Pydantic-GEPA Instrumentation.
Duplicate Or Conflicting Instrumentation¶
Autobench prevents unsafe double patching. Do not install two instrumentors with the same ID or instrument the same owner/method with incompatible specs. Prefer one of:
- automatic discovery only;
- explicit typed settings only;
- a custom runtime instrumentor that owns the same ID.
InstrumentationConflictError and patch diagnostics identify the owner and method involved.
Trace Is Partial¶
A partial trace can be valid evidence. It may result from cancellation, an interrupted stream, a task exception, or unmatched start/end signals. Inspect:
autobench instrumentation trace runs/example
ABP materialization keeps completed spans and diagnostics instead of dropping the trace. Accounting extractors avoid double counting aggregate and leaf usage even when evidence is incomplete.
Captured Content Is Missing Or Hashed¶
Runtime evidence is metadata-first, while behavioral asset definitions are full by default so
historical candidates remain reconstructable. Both are controlled by the benchmark
CapturePolicy, path rules, semantic overrides, and SDK-specific HTTP settings.
To avoid retaining asset bodies, use a preset that changes both defaults or set the asset fallback explicitly:
from autobench import CaptureLevel, CapturePolicy
policy = CapturePolicy.hashed(
semantic_overrides={"output_schema": CaptureLevel.FULL},
)
The experiment-local bodies and readable diffs are in artifacts/asset-content.sqlite3;
assets/*.yaml contains typed references and changed paths rather than copies. Secret names,
denied paths, truncation limits, and binary rules still apply.
Replay Needs An Optional SDK¶
It should not. replay, report, compare, export, and instrumentation trace are designed to
load records without benchmark or provider imports. If replay fails, verify that:
experiment.yamland every path inruns.pathsexist;- trace artifact paths remain inside the experiment directory;
- referenced artifacts were copied with the records;
- the record version is supported.
Do not solve a missing artifact by re-executing the benchmark implicitly.
Python Type Errors In Tasks¶
Autobench keeps case input and factor values generic because applications define their schemas. Validate at the task boundary:
from pydantic import BaseModel, TypeAdapter
request = Request.model_validate(case.input)
mode = TypeAdapter(Mode).validate_python(ctx.factor("mode"))
This gives application-specific errors without weakening Autobench's public types.
Get A Reproducible Diagnostic Bundle¶
For a bug report, include:
autobench --help
autobench instrumentation doctor
autobench validate path/to/autobench.yaml
Also include the package version, Python version, failing record directory when it contains no sensitive data, and the smallest benchmark/task that reproduces the problem. Review capture policy before sharing records.