Skip to content
MrJev

semgrep (uehaj)

Grep by meaning: Jev scores each line against a description in any language, with AND, OR, and NOT. A single dependency-free Node file.

View on GitHub →

Hands-on review

One file, no dependencies, and a grep that matches a proposition rather than a pattern. Every claim in its README reproduced on our run.

Good for

  • Searching logs or tickets for a situation you can describe but not regex
  • Mixed-language corpora: the query and the text need not share a language
  • Ad-hoc questions over files you have right now, with no index to build

Watch out for

  • Every line of every file you search is sent to TypeSafe, and each query pays again
  • One request per 30 lines, eight in flight; a big directory is a big bill
  • It is a filter, not an explanation: you get matching lines, not reasons

Tested Sep 20, 2026 at ba6ef50f85d0 · Node 24 in Docker with real Jev calls, against the project's own test corpora

How we reviewed this: we ran it in Docker on Node 24 with a real Jev key, against the project’s own test corpora — including the contrast set it uses to argue it is not vector search — and ran its self-check suite.

What it does

semgrep, by Junji UEHARA, is grep where the pattern is a meaning:

./semgrep -n -e "customer is angry or frustrated" tickets.txt

For every non-blank line it asks Jev one Noul question — does this line match the meaning “…”? — and keeps the lines whose probability clears a threshold. Meanings combine with AND (-a), OR (repeated -e) and NOT (-v), which works because each meaning yields an independent probability rather than a distance.

It is one file, semgrep.mjs, with zero dependencies, needing Node 20.12+ and a key. It also ships as a Claude Code skill that falls back to npx @uehaj/semgrep.

Its claims, on our run

The README’s three headline claims all reproduced, with a real key against its own corpora.

1. It judges a proposition, not a topic. All six lines below are about a refund; only two are a customer asking for one. Printed with probabilities and the threshold disabled:

1:返金してほしい。商品が壊れていた                          [0.98]
2:返金処理が完了しましたのでご確認ください                   [0.10]
3:当社の返金ポリシーは購入後30日以内です                     [0.10]
4:The manager denied the refund request yesterday          [0.19]
5:I demand a full refund immediately                       [0.95]
6:Refunds are processed within 5 business days             [0.08]

That gap is the argument for this tool over embedding search, and it held: “refund completed”, “refund policy” and “manager denied the refund” all sit near the floor.

2. The query’s language doesn’t matter. One Japanese meaning, 顧客が返金を求めている, over a file of refund requests in six languages, returned all six — French, Russian, German, Spanish, Chinese and Korean — and left the thank-you notes alone.

3. AND and NOT are ordinary booleans. “Someone is angry” AND “the customer, not the staff, is the one acting” kept the customer’s line and dropped the one where a support agent hung up angrily — two lines whose embeddings would be nearly identical.

Its self-check suite (tests/check.sh), which asserts OR, AND, NOT, the -l/-c/-r/-C output shapes and several error paths against a live API, passed.

We’re not publishing speed or cost figures for the service, and we’d treat the README’s own as the author’s.

What it sends

This is the thing to understand before pointing it at a directory.

  • Every non-blank line of every file you search goes to TypeSafe. Lines are batched 30 per request (capped at 20,000 characters), each line truncated to 2,000 characters, with up to eight requests in flight. The file name is not sent; the lines are.
  • Each query pays for the corpus again. There is no index and no cache. That is the trade the README makes explicitly: nothing to build, nothing stale — and for repeated queries over a large fixed corpus, a vector index is cheaper.
  • Your key comes from the environment, ./.env, or ~/.config/semgrep/.env. Nothing else leaves; there is no telemetry and no second provider.
  • Retries are bounded and sensible: up to six attempts with exponential backoff on 429, 529 and 5xx, and on connection errors.

The practical rule: this is a tool for logs and tickets you would be comfortable pasting into a hosted API, run on a directory you have deliberately chosen. -r over a repository will happily send the repository.

Where it fits

Against grep, it finds what you meant when you can’t write the pattern. Against vector search, it answers a different question — whether a statement holds for a line, not how close the line is to a topic — at the cost of paying per query instead of per index. Against handing the file to an LLM, you get a calibrated probability per line and a threshold you can set once, rather than prose you must parse and trust.

What it doesn’t give you: any explanation. A line either clears the threshold or doesn’t. For triage that’s the right shape; for “why did this match”, it isn’t the tool.

The README’s own caveat is worth repeating: TypeSafe documents English as the most accurate language, and the author notes Japanese meanings wobble slightly more near the threshold. When a query sits on the line, phrase the meaning in English.

Maintenance

23 commits through September 19, MIT licensed (the file is there, though GitHub’s detector currently reports NOASSERTION), published as @uehaj/semgrep 0.2.2, with a changelog, a release script, a Japanese README, a self-check suite and a judged evaluation in tests/. No open issues. Both READMEs are unusually well argued — the section on why this is not vector search is the clearest statement of that distinction we’ve read in this ecosystem.

Verdict

A small, sharp tool that does exactly what it says. Every claim we tested reproduced, the boolean combinations work the way the argument for them says they should, and the whole thing is one dependency-free file you can read in a sitting.

Use it on logs and tickets, not on your source tree, and remember there is no index: the cost scales with corpus times queries, not with corpus.

For a different take on search, see neo4jev, which uses Jev to walk a graph rather than filter lines. Our Best Jev Tools roundup compares what each tool sends.

See how it compares with other tools in Best Jev tools, tested hands-on.

Review updated Sep 20, 2026. Numbers quoted from the project are its author's own; we don't publish our own measurements of Jev.

More in Command-Line Tools

jevgrep

★ 866

dzhng/jevgrep

Finds code by what it does rather than what it matches: a CLI for coding agents that asks which files and which regions are relevant to a description.

TypeScript

JevRev

★ 434

Alex314618-create/JevRev

Puts an LLM and Jev in one workflow, with sift, loop and long modes, a TUI, and a CLI that can be pointed at any /v1/systemone host. English and Chinese.

TypeScript

webctl

★ 145▲ 57

dorkitude/webctl

Search CLI for agents: results from up to three backends are scored by Jev against your query and an explicit --goal, and only the relevant ones reach the agent's context. Its benchmark excludes an arm it could not observe.

GoReviewed

Get new Jev projects every week

New Jev releases, pricing changes, and the best new projects, once a week. No spam; unsubscribe anytime.

Powered by Buttondown. See our privacy policy.