1,350 / 1,350
Matched decision pairs
Direct readout and native one-token generation produced identical labels and candidate logits in every matched pair.
No latency advantage over native one-token generation was established.
Open research · Local language models
Formerly Mini Jev · Renamed 29 September 2026
An inspectable local interface for typed decisions from frozen language models
Give a language model a state, explicit criteria, and a set of candidates. Inspect the typed decision, the candidate distribution, and the request that produced it.
Updated 29 September 2026 · Research manuscript · Not peer reviewed
The linked PDF is the current LogitTrail conference-preparation revision; NAACL submission is still pending. The earlier preprint submitted to arXiv under the title “Mini Jev” is awaiting moderation. An arXiv link will appear after announcement. Repository and page URLs retain mini-jev so existing links continue to work.
The interface
true candidate within a false/true pair.The released native system uses a frozen model with no trained decision head. The public API and browser expose candidate probabilities and typed values; raw candidate token IDs and logits remain in internal engine/research traces. Probabilities are normalized over the supplied candidates; concentration is not the probability of being correct.
See the system
A stable decision can still be wrong. This walkthrough explains that observation using retained experiment evidence. Candidate concentration is not a calibrated probability of correctness.
This earlier recording and its poster retain the former Mini Jev name. Its final caption names the earlier source branch; later research integrations are not shown. English examples illustrate the interface, while published task-quality evaluations use Japanese data. This page hosts recordings; run the application locally for live inference.
What we measured
The contribution is a runnable local workbench and an auditable evaluation of its behavior. Candidate scoring and model-inspection interfaces have prior art, discussed in the paper.
1,350 / 1,350
Direct readout and native one-token generation produced identical labels and candidate logits in every matched pair.
No latency advantage over native one-token generation was established.
26,050
One matched output-mode study, followed by two studies of presentation sensitivity and averaging across three checkpoints.
Requests are not independent questions. Reduced order sensitivity can coexist with incorrect or nearly constant predictions.
56 / 56
CPU-only reanalysis reproduced the derived outputs byte for byte from retained records.
This is analysis replay. Native build and smoke checks used the same Mac and verified caches, not an independent replication.
Task-quality evidence comes from Japanese public-data subsets on one Apple Silicon Mac. Public-data contamination is unknown. The separate 2,400-item local regression suite is AI-authored. No human usability study or second-machine replication has been completed.
Probability averaging requires multiple model calls and does not consistently improve accuracy. The model comparisons also differ in runtime and precision, so differences cannot be attributed to model size alone.
LogitTrail, formerly Mini Jev, was inspired by TypeSafe's Jev. It does not reproduce Jev's architecture, training, calibration, speed, or protocol. The separate r-ms/mini-jev project is closely related prior work; LogitTrail is the UPHASH project formerly published under the same name. The paper situates candidate scoring and inspection alongside prior work including Ma et al., Hanna and Feng, LMQL, ChainForge, LLM Comparator, Langfuse, and Phoenix. Closely related scoring work includes Zhuang et al., Beyond Yes and No and Wang et al., Single-Token Expected-Value Scoring; LogitTrail does not claim to introduce candidate-logit normalization or expected-value scoring.
AI assistance in implementation, analysis, and writing is disclosed in the manuscript. The author is responsible for its content. As of 29 September 2026, the earlier preprint submitted to arXiv under the Mini Jev title is awaiting moderation. The continuing conference manuscript has not been submitted or accepted; neither version is peer reviewed.
The frozen source/evidence ZIP covers the application and three paper studies. The manuscript, video, and framework diagnostics are provided separately above. Framework diagnostics are not general performance or usability benchmarks.
Run it yourself
The browser workbench connects to a resident local model through the LogitTrail API. Inspect each candidate and copy the actual request and response.
Installation instructions ↗Model weights and native binaries are not bundled. Lower-memory machines have not been validated. For CPU-only analysis replay, use the reproduction instructions above.
Reference
This citation identifies the current LogitTrail manuscript. The earlier arXiv submission retains its Mini Jev title pending moderation; it has no public identifier yet.
Download BibTeX ↓@misc{oshio2026logittrail,
title = {LogitTrail: An Inspectable Local Interface
for Typed Decisions from Frozen Language Models},
author = {Oshio, Yuki},
year = {2026},
note = {Research manuscript; not peer reviewed},
url = {https://uphash-network.github.io/mini-jev/}
}