LogitTrail

Open research · Local language models

LogitTrail

Formerly Mini Jev · Renamed 29 September 2026

An inspectable local interface for typed decisions from frozen language models

Yuki Oshio UPHASH Inc.

Give a language model a state, explicit criteria, and a set of candidates. Inspect the typed decision, the candidate distribution, and the request that produced it.

Updated 29 September 2026 · Research manuscript · Not peer reviewed

The linked PDF is the current LogitTrail conference-preparation revision; NAACL submission is still pending. The earlier preprint submitted to arXiv under the title “Mini Jev” is awaiting moderation. An arXiv link will appear after announcement. Repository and page URLs retain mini-jev so existing links continue to work.

The interface

One model. Three kinds of decision.

InputState + criteriaExplicit candidate meanings
Frozen modelCandidate logitsRead at the next-token position
Typed outputChoice · Noul · ScoreInspect probabilities and JSON
CChoice
The most likely candidate, returned as its semantic label.
NNoul
The probability assigned to the true candidate within a false/true pair.
SScore
The expected zero-based stage index, assuming equally spaced ordinal stages.

The released native system uses a frozen model with no trained decision head. The public API and browser expose candidate probabilities and typed values; raw candidate token IDs and logits remain in internal engine/research traces. Probabilities are normalized over the supplied candidates; concentration is not the probability of being correct.

See the system

From message to inspectable decision.

Interface guide ↗
60-second evidence walkthrough · Updated branding; retained experiment results and an archival Mini Jev interface excerpt, not a new live run · Open short video directly · Transcript, sources & CC BY-SA license

A stable decision can still be wrong. This walkthrough explains that observation using retained experiment evidence. Candidate concentration is not a calibrated probability of correctness.

Watch the full interface recording · 2 min 25 sec
Recorded under the former Mini Jev name · Interaction with the local model · Captions embedded in the video · Open full video directly

This earlier recording and its poster retain the former Mini Jev name. Its final caption names the earlier source branch; later research integrations are not shown. English examples illustrate the interface, while published task-quality evaluations use Japanese data. This page hosts recordings; run the application locally for live inference.

What we measured

A system you can inspect. Evidence you can check.

The contribution is a runnable local workbench and an auditable evaluation of its behavior. Candidate scoring and model-inspection interfaces have prior art, discussed in the paper.

1,350 / 1,350

Matched decision pairs

Direct readout and native one-token generation produced identical labels and candidate logits in every matched pair.

No latency advantage over native one-token generation was established.

26,050

Measured research requests

One matched output-mode study, followed by two studies of presentation sensitivity and averaging across three checkpoints.

Requests are not independent questions. Reduced order sensitivity can coexist with incorrect or nearly constant predictions.

56 / 56

Analysis outputs reproduced

CPU-only reanalysis reproduced the derived outputs byte for byte from retained records.

This is analysis replay. Native build and smoke checks used the same Mac and verified caches, not an independent replication.

Scope, limitations, and prior work

Task-quality evidence comes from Japanese public-data subsets on one Apple Silicon Mac. Public-data contamination is unknown. The separate 2,400-item local regression suite is AI-authored. No human usability study or second-machine replication has been completed.

Probability averaging requires multiple model calls and does not consistently improve accuracy. The model comparisons also differ in runtime and precision, so differences cannot be attributed to model size alone.

LogitTrail, formerly Mini Jev, was inspired by TypeSafe's Jev. It does not reproduce Jev's architecture, training, calibration, speed, or protocol. The separate r-ms/mini-jev project is closely related prior work; LogitTrail is the UPHASH project formerly published under the same name. The paper situates candidate scoring and inspection alongside prior work including Ma et al., Hanna and Feng, LMQL, ChainForge, LLM Comparator, Langfuse, and Phoenix. Closely related scoring work includes Zhuang et al., Beyond Yes and No and Wang et al., Single-Token Expected-Value Scoring; LogitTrail does not claim to introduce candidate-logit normalization or expected-value scoring.

AI assistance in implementation, analysis, and writing is disclosed in the manuscript. The author is responsible for its content. As of 29 September 2026, the earlier preprint submitted to arXiv under the Mini Jev title is awaiting moderation. The continuing conference manuscript has not been submitted or accepted; neither version is peer reviewed.

The frozen source/evidence ZIP covers the application and three paper studies. The manuscript, video, and framework diagnostics are provided separately above. Framework diagnostics are not general performance or usability benchmarks.

Run it yourself

Keep inference on your Mac.

The browser workbench connects to a resident local model through the LogitTrail API. Inspect each candidate and copy the actual request and response.

Installation instructions ↗

Native setup

Hardware
Apple M5 Pro · 64 GB unified memory
System
macOS 26.4 (tested)
Required tools
Python 3.10+, Xcode command-line tools, CMake, Git
Model
Qwen3.6-35B-A3B Q4_K_M · ~20.4 GB download

Model weights and native binaries are not bundled. Lower-memory machines have not been validated. For CPU-only analysis replay, use the reproduction instructions above.

Reference

Cite the manuscript

This citation identifies the current LogitTrail manuscript. The earlier arXiv submission retains its Mini Jev title pending moderation; it has no public identifier yet.

Download BibTeX ↓
@misc{oshio2026logittrail,
  title = {LogitTrail: An Inspectable Local Interface
    for Typed Decisions from Frozen Language Models},
  author = {Oshio, Yuki},
  year = {2026},
  note = {Research manuscript; not peer reviewed},
  url = {https://uphash-network.github.io/mini-jev/}
}