Datasets
Evaluation Regression Comparer
Compare evaluation runs by test ID with your own metric thresholds, cost units and pass/fail rules.
Processed on our CPU server.Submitted input is processed for this run and is not stored.
1024 KiB limitInput
Result
A useful result starts here.
Paste your input or load an example,
then run the tool.
Supported formats & limitations
- JSON is an array of records; JSONL is one object per line. Every record requires a unique non-empty string id. Invalid records reject the comparison.
- For boolean pass/fail use {"metric":"pass"}. For a numeric field declare metric, unit, direction (higher/lower), threshold and optional non-negative minDelta. Each supplied metric must declare the matching unit.
- Crossing the pass threshold takes precedence; otherwise numeric changes larger than minDelta classify improvements/regressions. Exact or tolerance ties remain ties. Missing metrics remain unknown. Numeric metrics and deltas use finite IEEE-754 float64 precision.
- Optional cost requires a non-negative number and costUnit; matching records must use identical cost units. Optional errors is a non-negative integer. Missing values produce no delta; deltas are after minus before.
- Maximum 2,000 records per set, 128 KiB per JSONL line, depth 48 and 24,000 JSON nodes. No judge, statistical significance, causal inference or inferred prices. Reports cap findings at 200 with truncation declared.
Execution budget: 5s; output limit: 4096 KiB. Browser worker startup has a separate 3s allowance.