Skip to main content
Writing assertions is the easy part of LLM evaluation. Knowing what to measure for your specific product is the hard part, and multivon-eval bootstrap proposes an answer for you to review. Hand it a one-paragraph product description and a few real traces, and it emits a runnable EvalSuite with metrics suggested from your trace shape, provisional thresholds from observed scores, and adversarial seed cases targeting the most likely failure mode.

The whole flow

You get five artifacts: Cost and latency depend on trace count, evaluators, model pricing, and retries. Local Ollama avoids hosted API charges but still uses your compute. The reported bootstrap cost is an estimate, not a complete provider bill. As of 0.9.4 the emitted eval_suite.py is genuinely runnable end-to-end: python eval_suite.py --runs 1 executes the full suite without further edits. Two things to know before that first run. It makes real judge calls (each LLM evaluator can make multiple judge calls per case), and its main() exits 1 when the pass rate is under 50% — which is expected against the placeholder stub_model. The all-fail run is the signal to wire in your real model. 0.9.4 also fixed a stale suite.run(cases=...) kwarg and report.print_summary() call that previously broke the generated file; if you hit those on a 0.9.3-and-earlier install, upgrade.

Scaled, gated generation (0.12.0)

--n-seed-cases now goes to 500. Generation runs in batches of ≤30, with later batches steered away from already-accepted inputs, and every case passes gates before acceptance: well-formed (structural), duplicate (NFC-normalized identity or token-Jaccard ≥ 0.85, across batches), and, with --validate-cases --baseline-model-file model.py, the hardness band from auto.validate_adversarial_cases (judge-priced, so opt-in). No silent caps. The CLI summary and DISCOVERY_REPORT.md print the full accounting — generated 500, accepted 431 — dropped 38 duplicates, 12 malformed, 19 outside hardness band [0.5, 1.0] — and a skipped hardness gate says so explicitly. Each accepted case’s metadata["generation"] records its batch, the gates it passed, and its hardness score. --budget-usd (default 2.00) checks the estimated seed-generation cost before the first LLM call. It does not cap proposal, trace-scoring, validation, or model-under-test costs and is not a hard ceiling on your total bill.

What the inputs look like

product.md — free-form markdown, under 5000 words

Suggested sections (none are required): # Product, # Users, # Inputs / Outputs, # Known risks. The LLM uses the description to ground its metric recommendations in your domain.

traces.jsonl — newline-delimited JSON, max 10 000 rows

Only input is required. Optional keys are detected and used to infer product shape:
If you don’t have output traces yet (pre-launch), the tool still proposes evaluators based on the product description alone — threshold calibration is the only step that’s skipped.
load_traces accepts your existing dump shape. As of 0.9.4, the loader auto-aliases field names from LangSmith (query → input, answer → output, retrieved_context → context), LangFuse (prompt / completion), and Phoenix (input / output). It prints a one-line summary to stderr — loaded 287/300 traces · renamed 814 fields · skipped 13 (missing input) — so you can see exactly what was renamed and what was dropped. The silent-skip behavior from earlier releases is gone.

How the recommendation works

Three layers, in order:
  1. Heuristic anchor. auto_evaluators() inspects the trace shape and picks a deterministic starting set — Faithfulness + Hallucination for RAG, ToolCallAccuracy for agents, etc. This runs in microseconds, no LLM call, and provides the safety net.
  2. LLM proposal. A single call to Claude Haiku (configurable via --judge-provider / --judge-model) reads your product description + trace summary + a sample of traces, then proposes 4–6 evaluators with per-metric rationale and threshold suggestions. The LLM is constrained to an enumerated allow-list of evaluator names — it cannot invent metrics, and any proposal outside the list is silently dropped.
  3. Merge + threshold suggestions. LLM picks take priority for the same evaluator name (more contextual); heuristic picks fill in tiers the LLM missed (e.g. NotEmpty as a guardrail). Thresholds are then re-tuned by running each evaluator over a sample of your traces (50 for deterministic evaluators, 15 for LLM-judge evaluators) and setting the threshold to the 25th percentile of observed scores.
Generation may use multiple batches and trace-side scoring may use many calls. A p25 threshold describes the current score distribution: it does not learn which outputs humans consider acceptable. Do not use these suggestions as a production acceptance policy until you have reviewed labeled examples, tuned on a development split, and checked a separate held-out split.

PII handling — no surprises

Every trace is locally scanned for high-confidence secrets and PII before any data leaves your machine: AWS keys, OpenAI / Anthropic / GitHub tokens, JWTs, private-key PEM headers, US SSNs, emails, Luhn-valid credit cards, and more. Three policies: If PII was detected and redacted, the report’s ## Trace evidence section surfaces the counts (email=12, ssn=2) and the bootstrap pipeline auto-adds PIIEvaluator as a guardrail in the generated suite.

The discovery report

DISCOVERY_REPORT.md records inferred product shape, suggested metrics, rejected proposals, and reasons. Use it for team review alongside real failure examples. The generated report does not establish launch readiness.

Drift detection comes free — --repo (0.10.0)

As of 0.10.0, bootstrap also sets up prompt-drift staleness detection, at no extra cost or latency:
  • --repo PATH (default .) tells bootstrap which app repo to scan for prompt call sites. It writes prompt_baseline.json at that root — the committed snapshot that multivon-eval staleness later diffs against.
  • Every generated case is stamped with metadata._provenance: authored_by="bootstrap", the repo SHA, and targets=[]. The empty targets list means “authored against this repo state”, and nothing more. Bootstrap generates cases from your product description and traces; it knows nothing about call sites, so case→site bindings are never fabricated. Bind cases explicitly later with multivon-eval staleness stamp.
  • The completion checklist gains one line confirming both:
Stamping flows through the existing case-serialization path and never perturbs suite.lock (the lockfile’s cases hash excludes metadata by design). From that point on, a zero-arg multivon-eval staleness . in the repo reports which prompts changed since the suite was bootstrapped.

What’s NOT in the bootstrap (and why)

  • No model-side wrapper. The bootstrap emits a suite that calls your model_fn; you bring the wiring to your real model. No vendor lock-in.
  • There’s no “monitor production for me” loop. The bootstrap is a setup-time tool: once you have the suite, you run it with python eval_suite.py or wire it into CI via --fail-threshold.
  • No template marketplace, either. Picking your suite from a 20-template menu is the old approach; the bootstrap picks for you and tells you why.
  • No dashboard. Local-first by design. Use multivon-eval view report.json if you want a browsable HTML report.

When to use it (and when not)

Use it when:
  • You’re starting a new LLM feature and don’t yet have an eval suite.
  • You inherited an LLM product with no evals and need a credible starting point.
  • You’re switching evaluation tools and want a clean baseline.
  • You’re preparing for a launch and want a forwardable “what we eval and why” document for your team.
Don’t use it when:
  • You have a mature, hand-tuned eval suite and just want to add one more evaluator. Use add_evaluators() directly.
  • You only need a deterministic check (length, regex, schema). The bootstrap is overkill — write the assertion in two lines.

Configuration reference

Local judge — run bootstrap fully offline

The whole pipeline routes through make_judge_call as of 0.9.4 / 0.9.6 — including the adversarial seed-case generator (auto.py). That means --judge-provider ollama (or litellm) runs end-to-end, not just at argparse level.
  • OLLAMA_HOST env var controls the daemon address (default 127.0.0.1:11434).
  • --judge-base-url http://localhost:8000/v1 overrides for vLLM, LM Studio, or a remote Ollama.
  • Runtime depends on local hardware and model size. Local execution avoids hosted API charges.
  • Cases generated under a local judge carry metadata['judge_used'] = "ollama:qwen2.5:14b" and metadata['prompt_version'] for downstream replay.

Programmatic API

If you’d rather invoke the bootstrap from Python (e.g. inside a Jupyter notebook or CI pipeline), the callable is exported as multivon_eval.bootstrap (implemented in multivon_eval.discover):

See also

  • Intelligent eval primitives — full reference for the three primitives the bootstrap composes (auto_evaluators, generate_adversarial_cases, validate_adversarial_cases). Use these directly when you want fine-grained control or want to compose them into your own pipeline.
  • Generate eval cases — the higher-level generate_from_file / generate_from_text helpers, when you want generation without targeting a specific failure mode.
  • Prompt-drift staleness — the drift report that consumes the prompt_baseline.json and provenance stamps bootstrap writes.
  • Quickstart — the manual path: write EvalCase objects directly.
  • /eval-bootstrap Claude Code skill — the auto-invoking wrapper around this CLI for Claude Code users. Install with multivon-eval install-skills.