OPEN-JEV / EVALUATION

Benchmark results.
Scope included.

This index consolidates completed measurements from the published reports. Each table keeps its own metric and denominator. The Open-Jev columns refer to the released 2B/9B LoRA adapters, scalar decision heads and saved calibration. No new 27B result is available.

External decision benchmarks

Correct decisions and accuracy. JF100 rotations have a separate denominator.
Evaluation / metricReleased Open-Jev 2BReleased Open-Jev 9BJev 1.13.0GPT-5.6 Luna (none)GPT-6 Astra (low)
JevBench public · correct / 231150/231 · 64.94%179/231 · 77.49%200/231 · 86.58%206/231 · 89.18%231/231 · 100.00%
JevBench original · correct / 7256/7265/7271/7269/7272/72
JevBench easy · correct / 4848/4848/4848/4848/4848/48
JevBench hard · correct / 11146/11166/11181/11189/111111/111
JF100 · correct / 300 rotationsPendingPending232/300227/300300/300

JevBench covers all 231 available public tasks out of 534 total: 72 original, 48 easy and 111 hard. The other 303 private/judge tasks were unavailable. All five streams completed and passed independent replay audits. Native and GPT adapters use different candidate orders on 119 of 139 Choice tasks. Open-Jev and Jev return native probabilities; GPT returns verbalized probability vectors in constrained JSON, not token logprobs. Jev had one vector normalized under the upstream rounding policy (230/231 strict-valid); the other four streams had 231/231 strict-valid vectors.

JF100 has 100 external items in three option rotations: 300 correlated decisions, not 300 independent problems. Jev’s 232/300 is categorical-only accounting and retains two probability-mass flags. The earlier Open-Jev pilot checkpoints are different from the released models; their archived scores do not fill the pending released-model cells.

Retrieval on TREC-DL

Primary strict nDCG@10. Higher is better; failed queries remain zero in the full denominator.
Evaluation / metricReleased Open-Jev 2BReleased Open-Jev 9BJev 1.13.0GPT-5.6 Luna (none)GPT-6 Astra (low)
TREC-DL19 · 43 queriesPendingPending0.2758360.7299110.736610
TREC-DL20 · 54 queriesPendingPending0.1906670.7020820.714484

TREC uses the downloaded BM25 top 100, nine adaptive 20-passage windows per query and full official qrels with direct-grade gains. Each hosted model completed 97 queries / 873 requests. Strict failures contribute zero over the full 43/54 query denominators. Jev had 108 mass-validation failures affecting 66 queries: 15/43 DL19 and 16/54 DL20 queries were strict-complete. Luna and Astra had no transport or strict-validation failures.

Jev’s predeclared supplementary actual-scalar nDCG@10 is 0.728218 on DL19 and 0.715734 on DL20. These scalar scores do not validate the probability vectors and do not replace the strict scores. Downloaded BM25 is 0.505831 / 0.479637. Later windows adapt to each provider’s rankings; Jev expected scores and GPT integer grades have different resolution. This is an independently authored protocol, not an exact reproduction of the community demo.

Usage-based standard-list estimates for the complete TREC run: Luna $0.799273; Astra $39.62622. These are estimates, not invoices. Jev cost is unknown; Open-Jev compute is unmetered, not free.

TREC protocol, summaries & independent audits ↗

Prepared-data reference matches

Hard-target reference matches. Each row has its own sample, task definition and denominator.
Evaluation / metricReleased Open-Jev 2BReleased Open-Jev 9BJev 1.13.0GPT-5.6 Luna (none)GPT-6 Astra (low)
Release-v2 test · hard reference matches9,515/10,046 · 94.71%9,799/10,046 · 97.54%Not evaluatedNot evaluatedNot evaluated
Release-v2 OOD · hard reference matches13,287/15,446 · 86.02%14,205/15,446 · 91.97%Not evaluatedNot evaluatedNot evaluated
Same 76 hard coverage cases65/7672/7666/7660/7671/76
Broader coverage · 140 hard casesPartial: 64 pendingPartial: 64 pending117/140109/140135/140
IR controls · 165 hard decisionsPendingPending165/165160/165165/165
FizzBuzz · 300 typed decisionsPendingPending299/300300/300300/300
Mailroom · 921 labeled decisionsPendingPending908/921900/921913/921

Release-v2 evaluates every 26,452 test/OOD row per model: 10,532 test and 15,920 OOD rows. Hard accuracy uses only 10,046 test and 15,446 OOD one-hot targets; the 960 soft-target rows are excluded from these fractions. These are reference matches in prepared task data, not end-to-end workflow success or game win rates. There was no full-data base-model comparison.

The same-76 row is the matched coverage comparison: historical 2B/9B predictions were reused only after exact row and checkpoint checks, adding no new latency. Their additional 64 hard cases remain pending in the broader 140-case suite. A later label audit found equivalent platformer actions and omitted ViZDoom policy constants. The separately labeled post-hoc six-case exclusion gives 2B 60/70, 9B 67/70, Jev 64/70, Luna 57/70 and Astra 69/70; it does not replace the original row.

The IR pilot contains six original query instances and is separate from TREC. FizzBuzz tests integers 1–100 with three typed questions each. Mailroom uses 87 requests, shared families and 921 labeled decisions; 36 inapplicable questions have no gold. These finite controls and correlated views are not representative population or production-success estimates. In these suites GPT emits categorical decisions, not probability vectors.

Latency on repeated workloads

P50 / P95 full-response wall time in milliseconds, 20 timed requests per workload.
Evaluation / metricReleased Open-Jev 2BReleased Open-Jev 9BJev 1.13.0GPT-5.6 Luna (none)GPT-6 Astra (low)
Customer service · 8 questions85.03 / 133.91 msNot measured295.26 / 330.37 ms918.13 / 1443.13 ms1938.39 / 2375.71 ms
1,024 state tokens · 32 candidates1015.90 / 1369.72 msNot measured301.37 / 361.21 ms690.12 / 787.92 ms1388.07 / 1741.09 ms

The repeated latency study uses 11 fixed workloads, 20 measured requests and three warmups per workload and provider, concurrency one and no retries. The table shows two workloads; all 11 remain in the linked full report. Open-Jev uses a warm local H100 and loopback HTTP with prefix cache off. Hosted providers use fresh HTTPS connections, including Internet/TLS/routing/scheduling. These are client-observed deployment times, not matched-hardware model speedups, throughput or energy-efficiency measurements. Response validity does not establish equal task quality.

The larger candidate workload is an unfavorable case for uncached Open-Jev 2B. Experimental prefix caching exceeded the probability tolerance on 9/11 workloads despite all 440 paired selected decisions matching; it remains off by default.

Separate JevBench timing diagnostics

One observation per public JevBench task; separate from the repeated-workload study.
ModelP50P95
Released Open-Jev 2B138.0 ms205.8 ms
Released Open-Jev 9B189.2 ms839.3 ms
Jev 1.13.0291.3 ms353.7 ms
GPT-5.6 Luna953.8 ms1307.5 ms
GPT-6 Astra2206.4 ms3581.6 ms

JevBench timing is a separate diagnostic: one full-response observation per heterogeneous task, no warmups or retries. It must not be pooled with the repeated-workload latency study. Local model loading is outside the timer; full response completion and validation are inside.

Models, evidence and pending work

The 2B and 9B packages are released LoRA adapters, scalar heads and calibration parameters; pinned base weights and the Open-Jev loader are required. Exact model and package hashes are recorded in the linked reports. Different GPT reasoning settings are not equal budgets.

Released Open-Jev JF100, TREC and the remaining provider-control suites are pending. New 27B training and v3 data preparation have no audited final quality result. The prepared 107,922-row broader held-out registry and the fixed 1,280-row / 840-group v2–v3 panel are evaluation plans, not completed results. Historical pilots remain archived under their original identities.