Skip to content
MrJev

webctl

Search CLI for agents: results from up to three backends are scored by Jev against your query and an explicit --goal, and only the relevant ones reach the agent's context. Its benchmark excludes an arm it could not observe.

View on GitHub →

Hands-on review

A search CLI that scores results with Jev so your agent reads less, with the most careful benchmark write-up in this directory.

Good for

  • A benchmark with controlled arms, blind grading and recorded violations
  • Searching with no key at all, through keyless provider endpoints
  • Saying plainly where it does not help: 'for a short fact it adds nothing'

Watch out for

  • Chunk selection and the summary are separate costs; budget for both
  • Two days old when first reviewed: pin a version and read the changelog
  • Chunk text is sent to Jev twice — once in state, once in the question

Tested Sep 23, 2026 at e9bc54aacc96 · golang:1.26 in Docker running its own `go vet ./... && go test ./...`, plus a probe of ours run against both the reviewed commit and the fix

How we reviewed this: we ran its own CI — go vet ./... && go test ./... — reproduced the one failure three times, and wrote a minimal Go program to find its cause. We made no Jev calls and ran no live searches.

What it is

webctl "San Francisco giants MLB score recent games" \
    --goal "The user asked for the final score of last night's Giants game"

A search CLI for coding agents. It queries up to three backends, hands every result to Jev with the query and the goal, keeps the high-scoring subset, dedupes, and prints that. The agent reads a handful of relevant snippets instead of twenty-five results. With --scrape --filter-chunks it goes further: fetch the page, chunk it, score the chunks, return only the relevant ones.

The pitch is token cost, and the author is unusually plain about when it does not apply: for a long PDF or a Reddit thread it saves “an enormous number of tokens compared to reading the page; for a short fact it adds nothing”.

The benchmark is the best thing here

benchmarks/README.md describes the sort of comparison most projects in this directory gesture at and do not build:

  • Four arms over 30 cases, differing in one thing — which tools the agent may use. The webctl arm has Bash restricted to webctl and WebSearch/WebFetch disallowed; the native arm has Bash disallowed.
  • Violations are recorded as a first-class metric: “a native search in a webctl arm, or a shell command in a native arm”.
  • Grading is by a third model, blind to arm and in shuffled order, against expected facts or a rubric.
  • Recency cases, which have no fixed answer, are graded by cross-answer agreement and sourcing instead.

And the part that decided how we read the rest:

codex-terra-* (not default) … Left out of the results because Codex’s native search is server-side and what it puts into context cannot be observed.

An arm was built, run, and then excluded from the headline because the measurement could not be trusted. We have read a lot of benchmark write-ups this month and that is a rare instinct. (The numbers themselves are the author’s, on the author’s cases; we have not reproduced them and do not republish them.)

The timeout that did not stop the clock, and now does

When we first reviewed this, go vet ./... was clean and every package passed except one, three runs out of three, with the same line showing on main in CI:

--- FAIL: TestCommandBackendFailures (5.00s)
    summarize_test.go:56: timeout did not cut the command short

The cause is a Go trap worth knowing. exec.CommandContext killed sh on time, but Stdout/Stderr were bytes.Buffer rather than files, so os/exec copied through pipes in goroutines and Wait would not return while the forked sleep still held the write end. We reported it and suggested cmd.WaitDelay.

The author went further than the suggestion, in ab99fe7 (v0.1.7): the command now runs in its own process group and the timeout kills the group, with WaitDelay as a backstop — so a wrapper script’s child is actually stopped rather than left running with a closed pipe.

We wrote our own probe rather than trust the badge: a summarize command of sleep 30 & sleep 30, a two-second timeout, and an assertion on the wall clock. Same probe, both commits:

e9bc54a  (what we reviewed)   elapsed=30s   err=summarize command timed out after 2s   FAIL
ab99fe7  (v0.1.7)             elapsed=2s    err=summarize command timed out after 2s   PASS

The error message was always right; before the fix it arrived twenty-eight seconds after the timeout it named. CI on main and on the v0.1.7 tag is green, and so is the full suite here: go vet clean, every package ok.

What goes to Jev

Search results — titles, URLs, snippets — along with your query and your --goal, which is a sentence about what the user actually wants. With --scrape --filter-chunks, the parsed page text goes too.

One detail worth knowing, from a comment in chunks.go:

// The chunk text and its context are sent twice: once in state,
// once in the question.
cost := 2*(estimateTokens(ch.Text)+estimateTokens(ch.Before)) + perChunkOverhead

The batch planner budgets for that doubling rather than hiding it, which is the right way to handle a cost you have decided to pay. But it means a scraped page reaches TypeSafe at roughly twice its token count, and the saving is measured against what would otherwise enter your agent’s context, not against what leaves your machine.

Searching works with no key at all, through keyless Exa, Parallel, Keenable, You.com and Firecrawl endpoints and then DuckDuckGo, with a documented cooldown ladder — and the README says straight out to “expect thin results without a key”.

Engineering

MIT, Go 1.26, installable by Homebrew, apt, npm or go install. Thirteen packages, all but one green; the test suites cover the provider layer, chunking, dedupe, scraping, config and key handling.

Sixty-eight stars and two days old when we tested it, which explains the CI state more than it excuses it.

Verdict

Worth installing if your agent searches the web often and you pay for its context. The --goal flag is the good idea: scoring results against what the user actually wants, rather than against the query string, is what a typed decision model is for.

Read benchmarks/README.md even if you never install the tool — as a piece of measurement design it is better than the thing it measures. And note the response to the one defect we found: reported in the morning, fixed the same evening, and fixed a level deeper than the fix we proposed.

For chunk-level filtering of a different corpus, see jev-pruner and jev-sift; for bulk scoring where the packing itself is the subject, jev-ultralightspeed.

See how it compares with other tools in Best Jev tools, tested hands-on.

Review updated Sep 23, 2026. Numbers quoted from the project are its author's own; we don't publish our own measurements of Jev.

More in Command-Line Tools

jevgrep

★ 866

dzhng/jevgrep

Finds code by what it does rather than what it matches: a CLI for coding agents that asks which files and which regions are relevant to a description.

TypeScript

JevRev

★ 434

Alex314618-create/JevRev

Puts an LLM and Jev in one workflow, with sift, loop and long modes, a TUI, and a CLI that can be pointed at any /v1/systemone host. English and Chinese.

TypeScript

semgrep (uehaj)

★ 145▲ 26

uehaj/jev-semgrep

Grep by meaning: Jev scores each line against a description in any language, with AND, OR, and NOT. A single dependency-free Node file.

JavaScriptReviewed

Get new Jev projects every week

New Jev releases, pricing changes, and the best new projects, once a week. No spam; unsubscribe anytime.

Powered by Buttondown. See our privacy policy.