How Jev works
A chatbot versus Jev
Same question: is the person actually fine?
Chatbot
Your prompt
Free text in
Language model
Writes an answer word by word
A paragraph
Could say anything. You still have to read it and decide.
Jev
Situation and your options
Fine / not fine / can’t tell
Jev
Picks one option in about 200 ms
Not fine, 65%
One of your options, with a probability for each. Nothing else is possible.
It asks three kinds of question. A Choice picks one option from your list. A Noul gives the probability that a statement is true. A Score places something on a scale you describe. Many questions can run in one call.
What happens when you read a reply
Your browser
Sends the reply and the situation, capped at 40 and 200 characters
Our server
Holds the API key. At most 300 calls an hour in total.
Jev
Three questions in one call: fine? how upset? next move?
Checks
Answer not on the list: rejected. Low confidence: shown as unsure.
Your screen
Only the typed answers. Your text is not stored or sent back.
Green boxes are controls we add. Blue is Jev.
Treat every model answer like an order
The controls line up with pre-trade checks most risk teams already run.
| Control | On a trading desk | Around Jev |
|---|---|---|
| Order validation | Reject a malformed order or an unknown symbol before it goes anywhere | Any answer outside the options we sent is rejected, however confident |
| Risk limits | Size above a threshold needs a second look | Act only at 0.7 confidence or above; below that, ask a person |
| Kill switch | Hard stop when volume or losses spike | A cap of 300 calls an hour for everyone; pages fall back to earlier answers |
| Who owns the call | The trader decides; the check only blocks | Jev reads the situation; a person or a fixed rule makes the call |
| Books and records | Every order can be reconstructed | Earlier answers are saved, and a 50-case test is published with its mistakes |
Where this shows up at work
A customer writes “Thanks, that's fine.” A colleague writes “Noted.” The words say yes. The situation may say otherwise, and a system that takes the words at face value closes the ticket anyway.
Each case below follows the same pattern: Jev reads the situation and picks from options you wrote, the controls decide what happens next, and a person or a fixed rule owns the call. These illustrate the pattern. They are not products running today.
Complaint follow-up. "Thanks, that's fine." after a fee dispute is reversed
- Jev answers
- Is the customer resolved, still unhappy, or unclear?How frustrated, from 0 to 3?
- Control
- Unclear or unhappy above 0.7 stays open in the complaints log instead of closing automatically.
- Who owns the call
- A service rep calls back.
An AI agent about to act. An assistant proposes "move $25,000 to an external account" after the client only asked about fees
- Jev answers
- Which action is this?Does it match what the client actually asked for?
- Control
- An action outside the list is rejected. A mismatch or low confidence blocks it, like a pre-trade check.
- Who owns the call
- The client confirms, or a person reviews it.
Surveillance pre-screen. Trader chat: "Keep this between us until the announcement"
- Jev answers
- Could this be an information-barrier issue?Is there pressure or urgency?How severe, from 0 to 3?
- Control
- Anything above 0.3 goes to a compliance queue with the scores attached. Nothing is cleared automatically.
- Who owns the call
- A compliance analyst.
Client intent in wealth management. "I want to be more aggressive with the college fund."
- Jev answers
- Change the risk profile, a one-off trade, just venting, or unclear?
- Control
- Routes to the advisor with the category. It never triggers a trade or a profile change on its own.
- Who owns the call
- The advisor, after a suitability conversation.
Checking another model's work. A chatbot answer cites "per your agreement, section 4.2"
- Jev answers
- Does the source support, contradict, or say nothing about this?Is the quoted text actually in the source?
- Control
- Unsupported or missing citations never reach the customer.
- Who owns the call
- The answer is held back or sent for review.
In every case the model owns none of the decision. To split ownership among the people who do, try the Accountability Split.
Why a model like this suits regulated work: the answer set is fixed and reviewable in advance, every answer carries a probability you can threshold and log, and nothing it returns is free text that someone has to interpret again.
What the controls don't catch
“My doctor confirmed jalebi is a great post-workout snack for me.” Jev answered post-workout, 97% confident. Nobody said they had worked out.
The answer was a valid option, so every check passed. Check Point Research saw the same pattern at scale: 59% of their attempts to manipulate Jev worked, using fake audit opinions and revised risk tables rather than instructions. Every case above has the same exposure. A typed answer guarantees the shape, not the truth. That's why a person or a rule owns the call, and why confident answers still get tested.
The numbers
- On 50 hand-labelled test sentences, Jev correctly filled in 95% of what people said.
- When Jev was 90% sure or more, it was right 93% of the time. Between 70% and 90%, only 64%. Its confidence carries real information.
- About 240 ms per call for eight questions at once.
- Building with Jev in your industry: best practices, with the evidence behind them.
- The full write-up: Building with a model that answers in probabilities.
How the site is built
- Next.js and Tailwind on Vercel. Jev is called from the server with TypeSafe's JavaScript SDK.
- The parser page also runs a small embedding model in your browser, for comparison.
- Tests cover the checks and the API routes. Text size, spacing, focus mode and dark mode are remembered in your browser only.