← All writing
· 4 min readjevaiengineeringpathfinderagentssecurity

Jev in Pathfinder: letting the model judge without letting it drive

Pathfinder now asks Jev, TypeSafe's structured-decision model, two narrow questions about evidence. Every answer goes through deterministic code before anything happens, and nothing happens without me.

Pathfinder is the workflow kit I use to build software with coding agents. In orchestrator mode each ticket goes through a Developer, an Adversary that tries to break the work, and an independent Tester, then stops at a gate where I decide what gets merged. (More on that here.)

Most of that loop is procedure: which role runs next, whether a report matches the current PR head, whether a claim is done. Code answers those, because they have right answers.

Two places kept asking a different question. Not "what state are we in?" but "what does this evidence mean?" That's where I've started using a model. Not as an orchestrator. As a reader.

model = epistemic authority
code  = procedural authority

The model says what it thinks the evidence means. Code decides what that allows. I decide whether it happens.

A second reader for the evidence

Feature 56 adds an optional Evidence Judge. After the Tester passes a ticket and before it reaches me, Pathfinder can ask Jev, TypeSafe's structured-decision model, one question per verification item: does the recorded evidence support this criterion?

Jev answers with typed choices and confidences, not prose. I can validate a choice from a closed set. I can't validate an essay.

The judge can't PASS or FAIL a ticket, accept work, merge, dispatch an agent or change lifecycle state. Any failure, timeout or malformed answer escalates to me. It's off by default, and an API key in the environment never turns it on.

The hard part was deciding what to send. My first attempt sent the recorded evidence through a secret detector. It went through eight adversarial rounds, and they kept finding syntax it missed. So the boundary became a projection instead of a filter.

Beside their raw evidence, the Tester and Adversary each write a one-line summary in restricted "judge prose": at most 280 characters, only letters, digits, spaces and basic punctuation. Without colons, equals signs, slashes, quotes or newlines, credential pairs, key-value assignments, URLs, headers, JSON and log dumps can't be written at all. "The expired token was rejected with HTTP 401." passes. A curl command doesn't.

Raw commands, logs, file paths, the PR URL and the head SHA stay local. The provider only sees generated ids that map back to local records. The docs are explicit about the limit: this stops accidental leaks of technical payloads and common credentials. It isn't DLP against a short secret written as a normal sentence.

The transport has one rule: validated bytes are transmitted bytes. The provider declares its endpoint and headers up front, and the guarded fetch sends exactly the strings it checked. If you validate one object and serialize another, the check is decoration. Ticket 57.3 found a real case of that, where a serialization hook could swap content after validation, and fixed it by freezing a snapshot first.

A concern that won't resolve

Feature 57 handles a messier moment. Review is complete, but a concern is still open, like "Repeated boundary input lacks an observation." Is that a missing test, an assumption to attack, or a requirement only I can clarify?

Jev can now be asked exactly that, with three possible answers:

  • Tester: gather more evidence
  • Adversary: run a new experiment
  • Human: clarify or decide

There's no Developer route. Only confirmed Tester findings can authorize a repair. There's no continue, dispatch or merge answer either, and a response with any field outside the contract is rejected.

Deterministic policy then maps the validated answer. Tester or Adversary is recommended only when there's no flagged concern, the cited evidence resolves, and confidence is at least 0.8. Everything else goes to Human, including a confident answer that flags security, low confidence, a wrong model, a timeout or a provider failure. Uncertainty can only make the result more conservative.

local workflow evidence      raw logs, commands, SHAs, credentials stay here
        ▼
bounded projection           restricted prose, generated ids
        ▼
guarded Jev transport        pinned model, one request, no retry
        ▼
strict validation            closed contract, refs must resolve
        ▼
deterministic policy         Tester | Adversary | Human
        ▼
human-authorized follow-up   separate command, separate authorization

Every arrow can refuse, and every refusal lands on Human.

Why it's built in slices

  • 57.1: contract and policy, with no network at all. Every combination is table-tested; all 4,094 combined restriction cases resolve to Human.
  • 57.2: a routing-only projection, still offline, independent of the Evidence Judge and its credentials.
  • 57.3: the guarded transport, tested only against a local double.
  • 57.4: a budget of two attempts per ticket that I initialize explicitly. Failures count, nothing is refunded quietly, and nothing retries automatically.
  • 57.5: fixed a bug that predated routing. Asking for more investigation at the same commit left the old Tester PASS eligible, so work could go straight back to done. Follow-up now retires the old reports and requires fresh ones.
  • 57.6: the explicit entrypoint. It needs its own config and per-call consent, and rereads all state after the provider answers. If anything changed, the answer is rejected.

Where it stands

Feature 56 and tickets 57.1 to 57.6 are merged. 57.7 (hostile and interrupted scenarios) is in progress, and 57.8 (docs and acceptance proof) hasn't started. Neither feature is in an npm release yet. No routing test has called Jev live, and the 0.8 thresholds are provisional, not calibrated accuracy.

The tradeoffs are real. Getting one recommendation takes a prepared request, a budget, consent, and then a separate human decision to act on it. That friction is on purpose: a recommendation should be cheap to make and hard to act on by accident. The model sees a thin slice, so some concerns will come back as "insufficient context" and go to me. And a short secret typed as plain prose can still get through.

The point

The goal isn't more AI in the loop. Most of the loop shouldn't have a model in it at all. I use one only where the question is really about meaning, and where a confident wrong answer is cheap because it reaches me as a recommendation, not a state change.

Jev reads. Pathfinder decides what's allowed. I decide what happens.

Pathfinder on GitHub: https://github.com/rikilamadrid/pathfinder

Lamadrid Labs © 2026