🔍 Read the full analysis: Using Jev In AI: 24 Paths To Decision Modeling on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A Sept. 29 article by Thorsten Meyer maps 24 possible uses for Jev, a tool that returns typed answers to narrow questions so software can act on them. Meyer says three uses are already live in his publishing operation, 12 meet his fit test, seven need measurement and two are poor fits. The reported costs and accuracy figures are the author’s measurements, not independently verified results.
Meyer describes Jev as a system that receives text or JSON alongside typed questions and returns answers such as a yes probability, a choice among options or a score on an ordered scale. Its output is meant for code to act on, rather than for readers to parse as generated prose. He reports that one call takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens.
The three live examples cover article relevance, language detection and backup topic classification. Meyer says a scan of 78,889 articles cost $2.01 and found 1,576 non-English articles, of which 1,553 were fixed. For classification across 31 topics, he reports 89% overall agreement with a frontier large language model, rising to 97%–99% when Jev’s confidence was at least 0.8. These are figures from Meyer’s own operation and measurement; the source does not provide independent validation.
The guide recommends using Jev only when decisions are high-volume and narrow, errors are inexpensive or uncertain cases can go to a more capable system, and an existing heuristic has been shown to fail. Meyer advises replaying 300–500 past decisions, reviewing disagreements and enabling the system gradually behind a feature flag. His proposed rule is to wire a use case in only if its high-confidence band reaches 95% accuracy.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Small Decisions Add Up
The proposal targets routine decisions that can be too numerous for manual review but still matter to publishing, commerce and operations. If the author’s cost and performance figures hold in other settings, teams could check more items while directing uncertain cases to people or larger models. That could make broad screening affordable without making Jev responsible for consequential decisions on its own.
Meyer’s own criteria also set limits on that case. A cheap model call is not useful just because it can be made: teams need evidence that current rules are failing, and a safe path for uncertain answers. His poor-fit example, event deduplication, found no duplicates in a canary, which he says leaves no demonstrated problem for the tool to solve.
How Meyer Rates Each Use
The 24 proposals span publishing, commerce, software, business operations and home uses, according to the article. Meyer assigns each one a status: live, strong fit, measure first or poor fit. The supplied source details the opening publishing examples, including a thin-source detector marked measure first, a disclosure check marked strong fit, and comment moderation also marked strong fit. It says the complete guide gives each use case a question and a rule for acting on its answer.
For live relevance screening, Meyer says about 10,000 story-and-site pairings were judged over three days, with only 22% clearly on topic. His stated design drops a pairing only when Jev finds the fit clearly low with high confidence; uncertain cases follow the previous publishing path. The article says 88% of the news items he processes begin with a bare headline, motivating a proposed thin-source check, but the source excerpt does not establish that detector’s performance.
“Use Jev only when all four conditions hold: High volume. Narrow question. Cheap errors. A heuristic fails visibly. Measured, not assumed.”
— Thorsten Meyer, author of the Sept. 29 guide
Independent Results Remain Unclear
The reported cost, article scan, classification agreement and confidence results come from Meyer. The source provides no independent replication, detailed evaluation data or comparison with alternative systems, so it is unclear whether the results will generalize to other publishers or decision types. The guide’s own ratings also distinguish proposals needing measurement from those already running.
The source excerpt ends during its commerce and customer operations section. It does not show the full list of 24 applications, the two poor-fit cases beyond the deduplication example, or evidence supporting each proposed use. It is also unclear how performance changes across languages, content types and confidence thresholds outside the reported 31-topic classification measurement.
Measure Before Wider Rollout
Meyer recommends replaying several hundred real past decisions, comparing results overall and by confidence band, and manually reviewing a sample of disagreements. He says teams should activate a use case only after its high-confidence results meet the stated 95% threshold, then test it on 5%–10% of units before a wider rollout. These are recommendations in the guide; the source does not report a schedule for additional deployments or independent evaluation.
For the proposed measure-first applications, the next step is to establish whether the current method makes meaningful errors. Until that evidence exists, their fit remains unproven under Meyer’s own test.
Key Questions
What is Jev, according to the guide?
Meyer describes Jev as a tool that takes text or JSON and typed questions, then returns answers such as probabilities, category choices or scores for software to use.
How many of the 24 proposed uses are already live?
Meyer says three are live in his publishing operation. He rates 12 others strong fits, seven as needing measurement and two as poor fits.
What results does Meyer report from the article scan?
He says a scan of 78,889 articles cost $2.01 and identified 1,576 non-English articles; 1,553 were fixed. These are the author’s reported results, not independently verified figures.
When does Meyer recommend using Jev?
His test calls for high volume, a narrow question, low-cost errors or a route for uncertain cases, and evidence that an existing heuristic fails. He recommends measuring past decisions before rollout.
Have the reported accuracy figures been independently confirmed?
The source gives Meyer’s own measurements but does not cite an independent replication. Whether the figures generalize to other systems and tasks remains unclear.
Source: ThorstenMeyerAI.com
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
