Jev best practices
Use it for a first read, not the final call.
Every point names its source. "Vendor" means TypeSafe said it. "Independent" means a third party tested it. "Community project" means open-source work that has not been peer reviewed.
What Jev is good and bad at
Good at
- Cheap. $0.042 per million input tokens, and output is free (early-access pricing).Source: TypeSafe: Models (vendor)
- Fast. Independent tests measured medians of 239 to 474 ms, several times faster than mid-tier LLMs.Source: jev-phishing-bench (independent), Support triage benchmark (DEV Community) (independent)
- It can only answer with the options you wrote, so there is no free text to parse.Source: TypeSafe: API reference (vendor)
Weak at
- Accuracy is mid-tier. It scored about 68% on TypeSafe’s own benchmark, below frontier LLMs.Source: DataCamp (press)
- It can be confident and wrong. Guessing a hidden die roll, it averaged 83% confidence and got 19% right.Source: AI Agents Simplified (die-roll test) (independent)
- It struggles with math, dates, counting, and long inputs full of unrelated text.Source: TypeSafe: Jev 1.13 known issues (vendor)
- It believes fake evidence that looks real. More on that below.Source: Check Point Research (independent)
TypeSafe claims 193.6 times faster and 444.6 times cheaper than frontier models. Those are its own tests, and TypeSafe says the numbers sit at the high end of what you will see.Source: TypeSafe launch post (vendor)
Nine places to use it
ByteByteGo put it well: use the LLM for generations, and use Jev on the decisions around them. Their list of nine places to use Jev is below, with our notes from the research next to each.Source: ByteByteGo, EP227: Top 9 places to use Jev (press)
| Use | What it means | Our note | Source |
|---|---|---|---|
| Model routing | Send each prompt to a cheap or a strong model, based on what it asks | LangChain ships a routing middleware for this. | LangChain: harness with Jev |
| Guardrails | Screen requests for risky content before they reach other systems | Catches crude injection. Fake evidence still gets through, so it cannot be the only check. | Check Point Research |
| Gating tool calls | Sort an agent’s tool calls into allowed, needs approval, or blocked | Run hard limits first. A person confirms anything that cannot be undone. | LangChain: harness with Jev, SEC Rule 15c3-5 |
| Inbox triage | Sort a flood of messages into spam, urgent, archive and so on | An independent support test found it 4 to 7 times faster and 31 to 65 times cheaper than mid-tier LLMs. | Support triage benchmark (DEV Community) |
| Reranking | Order retrieved passages by how well they answer a question | In TypeSafe’s own test, top-1 accuracy rose from 5% to 18%. | TypeSafe: Re-ranking cookbook |
| LLM evals | Grade another model’s answers on a scale you define | Describe each level as a situation. Number-only levels gave confidence 0.33; described levels gave 1.0. | TypeSafe: Score |
| Bulk labeling | Label thousands of table rows quickly and cheaply | Batch your questions: 13 in one call cost 12 times less than 13 calls. | TypeSafe: Parallel questions cookbook |
| Real-time decisions | Make fast choices inside a loop, such as trading | Fixed limits stay in charge. Every robotics demo so far runs in a simulator. | SEC Rule 15c3-5, jev-drone |
| Confidence gates | Act on confident answers and route the rest | Set the bands on your own data. In our test, 0.9 and up was right 93% of the time; 0.7 to 0.9, only 64%. | Our 50-case test, TypeSafe: Confidence |
List of nine uses: ByteByteGo, EP227, September 26, 2026. Descriptions and notes are ours.
Seven rules for building with it
Send only what the question needs.
Put it in named fields. Accuracy drops as unrelated text piles up.
Source: TypeSafe: Jev 1.13 known issues (vendor)Ask many narrow questions in one call.
13 questions in one call cost 12 times less and ran 10 times faster than 13 separate calls. On phishing, one broad question scored 62.6%; five narrow ones combined scored 95.1%.
Source: TypeSafe: Parallel questions cookbook (vendor), jev-phishing-bench (independent)Write the whole question in the instructions.
The model never sees your question IDs, only the text and the options.
Source: TypeSafe: Primitives (vendor)Always offer "none" or "other".
In one bias test, Jev picked a stereotype 79% of the time when forced to choose, and picked "unknown" 95% of the time when it was allowed to.
Source: awesome-typesafe-jev (bias audit) (community project)Let code do math, dates and counts.
Your code finds the candidates. Jev picks among them.
Source: TypeSafe: Jev 1.13 known issues (vendor)Set thresholds by what a mistake costs.
Act, ask someone, or send it to a person. Tune the cut-offs on your own labeled data.
Source: TypeSafe: Confidence (vendor)Pin the model version and log everything.
Keep the question text, the answer, the probabilities, the threshold and the action taken.
Source: TypeSafe: Models (vendor)
What Check Point found
Check Point Research tried to talk Jev into approving a fraudulent company (September 24, 2026).Source: Check Point Research (independent)
- The test: a due-diligence assistant judged a made-up, obviously risky investment firm. The attacker controlled one section of the document.
- Every one of nine setups broke at least once. The strongest attacker won 25 of 27 tries, in about 4 turns, for roughly $0.50 each.
- No "ignore your instructions" tricks. The attacker added a fake audit opinion, a regulatory file number and a revised risk table.
- Structured input and "untrusted" labels did not help. Warnings in the prompt cut the breaks only from 18 to 17 of 27.
What to do about it
Based on Check Point's recommendations, extended for this summary.
- Never let Jev alone approve moving money or changing security when any input came from outside your company.
- Check claimed evidence, like an audit opinion or a filing number, against the real source.
- Screen inputs first, and track where each field came from.
- Watch for someone probing again and again. The attack took about four tries.
- Red-team the whole system, and do it again every time the model version changes.
Check Point put it this way: typed output limits what a model can say, not what it can be convinced of.
Financial services
- No bank, insurer or broker has said publicly that it is piloting Jev (as of September 28, 2026). The public finance examples are community projects.Source: Community survey of Jev projects (community project)
- The Fed’s SR 26-2 guidance (April 2026) covers "non-generative, non-agentic AI models." A Jev classifier likely fits normal model-risk validation. So does the burden of validating a vendor model whose design is not public.Source: Federal Reserve SR 26-2 (regulator)
- Treat question wording like model settings and put it under change control. TypeSafe notes that two wordings of the same scale can behave differently.Source: TypeSafe: Score (vendor)
Where it could fit
These are designs, not live deployments. The source column shows the rule or real-world example each design leans on.
| Use | What Jev does | Who owns the call | Source |
|---|---|---|---|
| Complaints in chats, emails and calls | Flags likely complaints and how serious they are | Complaints team | Full research report |
| Communications surveillance | Ranks messages by policy risk for review | Registered principal | FINRA 2026 oversight report |
| AI agents that take actions | Checks the action matches what the client asked | Risk tech, or a person for exceptions | SEC Rule 15c3-5 |
| AML alerts | Sets priority and likely pattern | AML officer. SARs are always human. | DataRobot on AML alert triage |
| Document checks for KYC, income, claims | Flags fields that look wrong | Operations or underwriter | TypeSafe: Extraction cascade |
| Payment scams before a transfer | Scores scam risk from the stated reason | Fraud team | UK PSR on scam reimbursement |
| Insurance first notice of loss | Sorts claim type and fraud red flags | Adjuster or fraud investigator | NAIC AI bulletin status |
Where it does not fit
- Anything the law says needs specific reasons, like turning down a credit application. The circular was withdrawn in 2025, but the requirement it describes still applies.Source: CFPB Circular 2022-03 (regulator)
- SAR decisions, pricing, and denying claims.Source: Full research report (our analysis)
Ask TypeSafe before you start
- SOC 2 status, data residency, and whether a private deployment is possible. TypeSafe documents zero data retention for enterprise customers.Source: TypeSafe: Models (vendor)
- Whether there is an SLA. None is published.Source: TypeSafe: Models (vendor)
- Whether the price will hold. TypeSafe says it cannot prove the price is not subsidized, so plan for it going up ten times.Source: TypeSafe launch post (vendor)
Other industries
Same pattern everywhere: Jev gives the first read, and each industry has its own line not to cross.
| Industry | Good fit | Line not to cross | Source |
|---|---|---|---|
| Customer support | Routing, priority, refund and churn flags | No auto-closing a ticket without confirming it is actually resolved | Support triage benchmark (DEV Community), Uber COTA (analog) |
| E-commerce | Product categories, search ranking, listing checks | Removing a listing needs a reason you can show the seller | TypeSafe: Hierarchical classification, EU DSA Article 17 |
| HR and recruiting | Routing internal HR tickets | No auto-rejecting candidates. Bias audit and notice before any scoring. | NYC Local Law 144 |
| Healthcare | Message triage, prior-auth completeness, crisis screening | No diagnosis or care denials. No patient data without a signed BAA. | HHS on cloud providers and HIPAA |
| Legal | Citation checks, document review, privilege screens | A lawyer reads the authority. Jev cannot tell if a case exists. | TypeSafe: Citation check |
| Security and trust & safety | Phishing signals, alert triage, first-pass moderation | Never the only security control. People decide appeals. | jev-phishing-bench, Check Point Research |
| AI agents | Model routing, typed tool calls, blocking risky actions | Irreversible actions need a person to confirm | LangChain: harness with Jev |
| Robotics and games | Picking among pre-approved moves | A fixed controller can always overrule it | jev-drone |
Bottom line
Think of Jev as a cheap, fast sensor you can ask twenty narrow questions about every ticket, alert or message. Its readings are trustworthy at the extremes, after you check them on your own data. Anything a customer or counterparty can write into is something they can talk it out of.
All sources on this page (35)
- Full research report (our analysis)
- Our 50-case test (our analysis)
- awesome-typesafe-jev (bias audit) (community project)
- Community survey of Jev projects (community project)
- jev-drone (community project)
- AI Agents Simplified (die-roll test) (independent)
- Check Point Research (independent)
- jev-phishing-bench (independent)
- Support triage benchmark (DEV Community) (independent)
- Uber COTA (analog) (independent)
- ByteByteGo, EP227: Top 9 places to use Jev (press)
- DataCamp (press)
- DataRobot on AML alert triage (press)
- NAIC AI bulletin status (press)
- CFPB Circular 2022-03 (regulator)
- EU DSA Article 17 (regulator)
- Federal Reserve SR 26-2 (regulator)
- FINRA 2026 oversight report (regulator)
- HHS on cloud providers and HIPAA (regulator)
- NYC Local Law 144 (regulator)
- SEC Rule 15c3-5 (regulator)
- UK PSR on scam reimbursement (regulator)
- LangChain: harness with Jev (vendor)
- TypeSafe launch post (vendor)
- TypeSafe: API reference (vendor)
- TypeSafe: Citation check (vendor)
- TypeSafe: Confidence (vendor)
- TypeSafe: Extraction cascade (vendor)
- TypeSafe: Hierarchical classification (vendor)
- TypeSafe: Jev 1.13 known issues (vendor)
- TypeSafe: Models (vendor)
- TypeSafe: Parallel questions cookbook (vendor)
- TypeSafe: Primitives (vendor)
- TypeSafe: Re-ranking cookbook (vendor)
- TypeSafe: Score (vendor)
Want more detail? The full research report cites every number in context. Researched September 28, 2026. Not legal advice.