multivon-eval bootstrap proposes an answer for you to review. Hand it a one-paragraph product description and a few real traces, and it emits a runnable EvalSuite with metrics suggested from your trace shape, provisional thresholds from observed scores, and adversarial seed cases targeting the most likely failure mode.
The whole flow
Cost and latency depend on trace count, evaluators, model pricing, and retries. Local Ollama avoids hosted API charges but still uses your compute. The reported bootstrap cost is an estimate, not a complete provider bill.
As of 0.9.4 the emitted
eval_suite.py is genuinely runnable end-to-end: python eval_suite.py --runs 1 executes the full suite without further edits. Two things to know before that first run. It makes real judge calls (each LLM evaluator can make multiple judge calls per case), and its main() exits 1 when the pass rate is under 50% — which is expected against the placeholder stub_model. The all-fail run is the signal to wire in your real model. 0.9.4 also fixed a stale suite.run(cases=...) kwarg and report.print_summary() call that previously broke the generated file; if you hit those on a 0.9.3-and-earlier install, upgrade.
Scaled, gated generation (0.12.0)
--n-seed-cases now goes to 500. Generation runs in batches of ≤30, with
later batches steered away from already-accepted inputs, and every case
passes gates before acceptance: well-formed (structural), duplicate
(NFC-normalized identity or token-Jaccard ≥ 0.85, across batches), and, with
--validate-cases --baseline-model-file model.py, the hardness band from
auto.validate_adversarial_cases (judge-priced, so opt-in).
No silent caps. The CLI summary and DISCOVERY_REPORT.md print the full
accounting — generated 500, accepted 431 — dropped 38 duplicates, 12 malformed, 19 outside hardness band [0.5, 1.0] — and a skipped hardness gate
says so explicitly. Each accepted case’s metadata["generation"] records its
batch, the gates it passed, and its hardness score. --budget-usd (default
2.00) checks the estimated seed-generation cost before the first
LLM call. It does not cap proposal, trace-scoring, validation, or model-under-test
costs and is not a hard ceiling on your total bill.
What the inputs look like
product.md — free-form markdown, under 5000 words
# Product, # Users, # Inputs / Outputs, # Known risks. The LLM uses the description to ground its metric recommendations in your domain.
traces.jsonl — newline-delimited JSON, max 10 000 rows
Only input is required. Optional keys are detected and used to infer product shape:
If you don’t have output traces yet (pre-launch), the tool still proposes evaluators based on the product description alone — threshold calibration is the only step that’s skipped.
load_traces accepts your existing dump shape. As of 0.9.4, the loader auto-aliases field names from LangSmith (query → input, answer → output, retrieved_context → context), LangFuse (prompt / completion), and Phoenix (input / output). It prints a one-line summary to stderr — loaded 287/300 traces · renamed 814 fields · skipped 13 (missing input) — so you can see exactly what was renamed and what was dropped. The silent-skip behavior from earlier releases is gone.How the recommendation works
Three layers, in order:-
Heuristic anchor.
auto_evaluators()inspects the trace shape and picks a deterministic starting set — Faithfulness + Hallucination for RAG, ToolCallAccuracy for agents, etc. This runs in microseconds, no LLM call, and provides the safety net. -
LLM proposal. A single call to Claude Haiku (configurable via
--judge-provider/--judge-model) reads your product description + trace summary + a sample of traces, then proposes 4–6 evaluators with per-metric rationale and threshold suggestions. The LLM is constrained to an enumerated allow-list of evaluator names — it cannot invent metrics, and any proposal outside the list is silently dropped. - Merge + threshold suggestions. LLM picks take priority for the same evaluator name (more contextual); heuristic picks fill in tiers the LLM missed (e.g. NotEmpty as a guardrail). Thresholds are then re-tuned by running each evaluator over a sample of your traces (50 for deterministic evaluators, 15 for LLM-judge evaluators) and setting the threshold to the 25th percentile of observed scores.
PII handling — no surprises
Every trace is locally scanned for high-confidence secrets and PII before any data leaves your machine: AWS keys, OpenAI / Anthropic / GitHub tokens, JWTs, private-key PEM headers, US SSNs, emails, Luhn-valid credit cards, and more. Three policies:
If PII was detected and redacted, the report’s
## Trace evidence section surfaces the counts (email=12, ssn=2) and the bootstrap pipeline auto-adds PIIEvaluator as a guardrail in the generated suite.
The discovery report
DISCOVERY_REPORT.md records inferred product shape, suggested metrics, rejected proposals, and reasons. Use it for team review alongside real failure examples. The generated report does not establish launch readiness.
Drift detection comes free — --repo (0.10.0)
As of 0.10.0,
bootstrap also sets up prompt-drift staleness detection,
at no extra cost or latency:
--repo PATH(default.) tells bootstrap which app repo to scan for prompt call sites. It writesprompt_baseline.jsonat that root — the committed snapshot thatmultivon-eval stalenesslater diffs against.- Every generated case is stamped with
metadata._provenance:authored_by="bootstrap", the repo SHA, andtargets=[]. The empty targets list means “authored against this repo state”, and nothing more. Bootstrap generates cases from your product description and traces; it knows nothing about call sites, so case→site bindings are never fabricated. Bind cases explicitly later withmultivon-eval staleness stamp. - The completion checklist gains one line confirming both:
suite.lock (the lockfile’s cases hash excludes metadata by
design). From that point on, a zero-arg multivon-eval staleness . in the
repo reports which prompts changed since the suite was bootstrapped.
What’s NOT in the bootstrap (and why)
- No model-side wrapper. The bootstrap emits a suite that calls your
model_fn; you bring the wiring to your real model. No vendor lock-in. - There’s no “monitor production for me” loop. The bootstrap is a setup-time tool: once you have the suite, you run it with
python eval_suite.pyor wire it into CI via--fail-threshold. - No template marketplace, either. Picking your suite from a 20-template menu is the old approach; the bootstrap picks for you and tells you why.
- No dashboard. Local-first by design. Use
multivon-eval view report.jsonif you want a browsable HTML report.
When to use it (and when not)
Use it when:- You’re starting a new LLM feature and don’t yet have an eval suite.
- You inherited an LLM product with no evals and need a credible starting point.
- You’re switching evaluation tools and want a clean baseline.
- You’re preparing for a launch and want a forwardable “what we eval and why” document for your team.
- You have a mature, hand-tuned eval suite and just want to add one more evaluator. Use
add_evaluators()directly. - You only need a deterministic check (length, regex, schema). The bootstrap is overkill — write the assertion in two lines.
Configuration reference
Local judge — run bootstrap fully offline
The whole pipeline routes throughmake_judge_call as of 0.9.4 / 0.9.6 — including the adversarial seed-case generator (auto.py). That means --judge-provider ollama (or litellm) runs end-to-end, not just at argparse level.
OLLAMA_HOSTenv var controls the daemon address (default127.0.0.1:11434).--judge-base-url http://localhost:8000/v1overrides for vLLM, LM Studio, or a remote Ollama.- Runtime depends on local hardware and model size. Local execution avoids hosted API charges.
- Cases generated under a local judge carry
metadata['judge_used'] = "ollama:qwen2.5:14b"andmetadata['prompt_version']for downstream replay.
Programmatic API
If you’d rather invoke the bootstrap from Python (e.g. inside a Jupyter notebook or CI pipeline), the callable is exported asmultivon_eval.bootstrap (implemented in multivon_eval.discover):
See also
- Intelligent eval primitives — full reference for the three primitives the bootstrap composes (
auto_evaluators,generate_adversarial_cases,validate_adversarial_cases). Use these directly when you want fine-grained control or want to compose them into your own pipeline. - Generate eval cases — the higher-level
generate_from_file/generate_from_texthelpers, when you want generation without targeting a specific failure mode. - Prompt-drift staleness — the drift report that consumes the
prompt_baseline.jsonand provenance stamps bootstrap writes. - Quickstart — the manual path: write
EvalCaseobjects directly. - /eval-bootstrap Claude Code skill — the auto-invoking wrapper around this CLI for Claude Code users. Install with
multivon-eval install-skills.

