Skip to content
MrJev

JevBERT

A local server that speaks Jev's /v1/systemone shape from a BERT encoder, with a numbered account of every request it refuses that Jev might accept.

View on GitHub →

Hands-on review

A Jev-shaped server whose own compatibility document says not to call it compatible yet. The official SDK drove it with only a base-URL change.

Good for

  • A local `/v1/systemone` the official Python SDK drives with a base-URL change
  • A written account of where a Jev-shaped server diverges, and what is unverified
  • A benchmark harness that refuses to spend money unless you say so twice

Watch out for

  • The backend is a public zero-shot NLI classifier, uncalibrated and unevaluated
  • Python 3.13 only, and `uv sync` pulls a CUDA 12.8 torch by default
  • Nothing here has been checked against the real Jev API, as it says itself

Tested Sep 22, 2026 at f0994cd4c91e · python:3.13-slim in Docker on 4 CPU cores, no GPU; its 1294 tests, the server driven by typesafe-sdk 0.7.0, and its published rejection table walked row by row

How we reviewed this: we ran its suite in python:3.13-slim on four CPU cores, fetched the pinned model and tampered with a byte of it, started the server and drove it with the official typesafe-sdk 0.7.0, then sent every row of its published rejection table as a raw body. We made no Jev calls — and neither, by its own account, has the author.

What it is, in its own words

It is: a local HTTP server that accepts the same JSON shape as Jev’s POST /v1/systemone and returns the same shape. It works from the official Python SDK typesafe-sdk==0.7.0 by swapping only the endpoint and the API key.

It is not: a trained JevBERT. The contents are a public zero-shot NLI classifier (MoritzLaurer/bge-m3-zeroshot-v2.0) plugged into the A0 shape.

And on the first screen of compat/differences.md, in bold:

These “differences” have not been checked against the real Jev API. There is no guarantee that real Jev behaves as this document says, and none that it does not. This is a stage at which it must not be labelled “Jev-compatible”.

We have reviewed twenty-odd projects this month that claim compatibility on the strength of one curl. This is the first that wrote down what it had not checked and then refused itself the word.

The drop-in claim holds

The official SDK, pointed at 127.0.0.1:8765 with a key and nothing else. Three question types at once, in Japanese, restored into the SDK’s typed answers:

model            : jevbert-poc-nli-ja-en-0.2.0
refund_requested : noul=0.6689
department       : choice='billing' confidence=0.8045
urgency          : score=1.9057 confidence=0.7084
  legend         : {0: '対応期限の指定がない', 1: '数日以内の対応を求めている', 2: '当日中または直ちに対応することを求めている'}
usage            : input_tokens=476 output_tokens=0

Its own demo then checks nine structural invariants against the wire body — that IDs round-trip, that distributions sum to 1 within 1e-6, that the choice is the argmax with ties broken by the smallest key, that the score is the expected value of its distribution, that a Noul carries no confidence, that no unspecified fields appear. All nine passed.

GET /v1/models works, which is worth saying because it is the endpoint that has broken every other API-shaped project we have tested: OpenThai-SystemOne does not implement it and models.list() returns 404, and Von returned the OpenAI shape where the SDK wanted another. Here it returns three names, and the description attached to each of them reads:

Not a trained JevBERT model. Probabilities are uncalibrated and no quality gate has been run, so the answers must not be presented as evaluated judgements.

That warning travels with the model list, into any client that enumerates. It is the most unusual thing in this repository.

The rejection table is real

compat/differences.md opens with a table of every request that Jev might accept and this will refuse, each with a status, an error code, a reason and a workaround. We sent them:

what we sent documented what we got
model: "jev-1.13.0" (R11) 422 model_not_found 422 model_not_found
a Choice with one candidate (R4) 422 validation_error 422 validation_error
a Score with one level (R5) 422 validation_error 422 validation_error
an unknown field inside a question (R7) 422 validation_error 422 validation_error
a duplicate key in the JSON (R8) 400 invalid_json 400 invalid_json
NaN in the body (R8) 400 invalid_json 400 invalid_json
a bare number as state (R9) 422 validation_error 422 validation_error
a null nested inside a state object (R9) allowed 200
an unknown top-level field (R7) ignored 200
a wrong API key 401 401 unauthorized

R11 is the one to notice. Accepting jev-1.13.0 as a wildcard, the way several servers do, “would make a caller believe it had run that version of Jev”. Refusing it is a three-line decision that most of this directory got wrong.

The 422 body carries both its own error object and a detail array in Jev’s schema shape — and the document says plainly that the 401, 429 and 5xx shapes could not be determined from public material and are not claimed to match.

The weights are pinned, and checked

The manifest records the source repository, its revision, and a SHA-256 per file. The server never downloads weights itself; fetch-model does, and verifies. We flipped one bit inside config.json and started it:

bundle jevbert-poc-nli-ja-en-0.2.0 failed to load: error_class=ValueError
INFO:     Shutting down

Restored the byte, and it started. The failure message is terser than it needs to be — it does not say which file or that a checksum failed — but the gate is real.

The manifest is equally plain about what it does not know: calibration.state: uncalibrated, validated.quality: unevaluated, validated.max_choice_options: unknown.

The benchmark harness will not spend your money

docs/BENCHMARK.md says the benchmark is an environment only, and that the experiment against real Jev has never been run, because it costs money. The harness enforces that rather than trusting the author to remember:

what we did what happened
a target on a non-loopback host with no billable: key refused: unmarked_remote is billable and 8 requests would be sent (up to 24 HTTP attempts with retries); pass --allow-billable to send them
the same target with an explicit billable: false sent
plain http:// to a non-loopback host rejected when the config loads: “base_url must use https for a host other than this machine; the API key would otherwise travel in clear text”

Anything that is not 127.0.0.1, localhost or ::1 is billable unless you say otherwise, the refusal counts the worst case including retries, and a failure the other side answered is not retried because it may already have been billed. We run this directory under the same rule by hand; this is the first project we have seen that encodes it.

Engineering

1294 tests pass, 1 skipped, in about two minutes on four cores — unit, property (hypothesis), contract, SDK-over-a-real-socket and benchmark suites. For an eighteen-star repository that is an unusual amount of test.

One is flaky. test_an_understated_length_cannot_smuggle_the_rest_of_the_body failed once in twenty identical runs, and the failure is in the test’s own teardown rather than the server: it guards the write against OSError with the comment “The refusal arrived before we finished; the response says so”, then calls connection.shutdown(SHUT_WR) outside that guard and gets OSError: [Errno 107] Transport endpoint is not connected. The behaviour being asserted — that an early refusal closes the connection — is what breaks the assertion. We have reported it.

Two things to budget for. requires-python is >=3.13,<3.14, so this is a 3.13-only project. And uv sync resolves torch from the CUDA 12.8 index by an explicit source pin — the comment says it is for an RTX 5090 — so a CPU-only machine downloads a GPU wheel it will not use. It works; it is 2.5 GB.

Verdict

The backend is a stand-in and the project says so in four places, so there is nothing here to benchmark and we did not try. What there is, is a worked answer to a question the rest of this category keeps skipping: what does “Jev-compatible” actually commit you to, and which parts of it can you not verify without the real thing in front of you?

Read compat/differences.md before you build a Jev-shaped server of your own. R1 to R13 is the list of places yours will diverge, written by someone who went looking.

For a server on an open model that does claim to be usable today, see OpenThai-SystemOne; for the same question of API convergence across projects, the open models category.

See how it compares with other tools in Best Jev tools, tested hands-on.

Review updated Sep 22, 2026. Numbers quoted from the project are its author's own; we don't publish our own measurements of Jev.

More in Open Models & Reproductions

Laya

★ 27k▲ 12k

NandhaKishorM/laya

Non-autoregressive decision engine over 100+ languages: three checkpoints and a router that detects the script and dispatches per request. Its benchmarks end with a limits section naming the datasets it does not generalise to and the headline figure that came from a training split.

PythonReviewed

kev

★ 7.5k▲ 6.2k

jaredpalmer/kev

Jev-style decision models from 0.5B to 8B, built as LoRA adapters on Qwen and served behind a Jev-compatible /v1/systemone API.

PythonReviewed

SemIf

★ 4.5k▲ 1.9k

TheoLeeCJ/SemIf-OpenJev

Jev-style decisions from a frozen 4B model on a single RTX 3090, with a browser demo. Formerly OpenJev.

PythonReviewed

Get new Jev projects every week

New Jev releases, pricing changes, and the best new projects, once a week. No spam; unsubscribe anytime.

Powered by Buttondown. See our privacy policy.