← All posts
Ad Performance
October 8, 2026

Jev for ads: use cases, constraints, and what does it mean for ad automation

by
Elina Minnie
Yellow curly Subscribe to Blog button

Two weeks after TypeSafe AI opened early access to Jev, the marketing feed has filled with one recurring screenshot: thousands of ads judged in seconds, for cents, followed by the word "insane". Some of those posts are real runs by named people who showed their numbers. Some are vendor claims. A measurable share is the same text pasted by different accounts. And since September 29 there are two vendors, because OpenAI announced a Decisions API of its own at DevDay.

This piece sorts the three piles, because they deserve different levels of trust, adds what is and is not known about OpenAI's version, and then asks the question the demos skip: where does a model like this actually fit in how a performance marketing team works?

What Jev actually is

Jev is a decision model. That means:

  • There is no chat window. You do not talk to Jev. Your code, or a tool built on it, sends a question and a list of allowed answers, and Jev sends back one of them. It never writes a sentence.
  • It gives three kinds of answers. Yes or no, a score, or one item from a list you supplied, and each comes with a probability of how sure it is.
  • It is built for one question asked many times. You have a pile of rows and the same question for every one of them: is this search term from a buyer, which format is this ad, does this landing page deliver what the ad promised, which of these hundred photos are worth keeping.

What it does not do is explain itself: you get a label and a confidence, and there is no asking why.

TypeSafe AI came out of stealth on September 15 with a $40M seed round led by DCVC. The founder, Diogo Almeida, co-invented RLHF, the training method behind the instruction-following models everyone uses; Jev is his argument that a lot of what those models do in software did not need the text. Access is a waitlist at typesafe.ai, with API keys issued in batches, and the model exists only as a closed API: no open weights, no self-hosting, which several engineering teams have already named as their blocker.

The price is $0.042 per million input tokens with output free, confirmed across independent gateway listings, over a 32K context. TypeSafe's own homepage illustration prices a single decision call at $0.000081. Latency of 70 to 500 milliseconds and multiples like "444x cheaper than frontier models on tested decision tasks" are the vendor's own figures. On quality, the most honest sentence available comes from an independent model tracker: no unambiguously attributed benchmark results have been recorded for this model yet. There is also no published rate-limit table and no pricing page; the price above is the launch price, and nothing yet says what it looks like at scale.

OpenAI's answer

On September 29, at DevDay, OpenAI announced a Decisions API built on GPT-6 Luna, the smallest and cheapest model in its current lineup. The shape is the same as Jev's: the developer supplies a question, a fixed set of possible answers and some context, and the API returns a selection the code can act on. OpenAI's own description names three jobs, classifying content, routing requests and choosing an agent's next action, and says the context can be text or images. The stated latency is around 150 milliseconds, against 1.6 seconds for the same Luna model answering as a regular chat model, by OpenAI's figures as reported by The New Stack.

We don’t know much about it yet: there is no price, no published limit on how many candidate answers one request can carry, and no word on whether developers can tune it on their own data. Access is a limited preview for selected API customers, with broad release promised within days of the announcement. By the time you read this, all of that might be out. 

The practical reading: decision models are now a category with two vendors, and the buying question shifts from whether the idea works to which one is cheaper and more accurate on your rows. No named run on the Decisions API inside an ad account exists yet. Everything below is Jev.

The runs people showed

What separates the useful posts from the noise: a named person, a number, and enough detail to check. These are the ones that clear that bar, grouped by the job they did. Every figure below is the author's own report, linked, and none has been independently reproduced.

Reading a category's ads at once. Gloria Pu ran 4,367 Meta ads from 38 AI companies through Jev in 39 seconds for $0.27, tagging each by format and checking which were still live months later. Her survival table is the interesting part: product demos outlast everything, talking-head UGC is the most produced and among the fastest to die, and the meme category did not survive at all. Joshua G., who runs acquisition at Function Health, reports a similar pass over 1,000 live ads from 52 brands in under a minute for under six cents.

Sorting search terms. Adriaan Dekker gave it 10,000 Google Ads search terms and got every one classified by intent in 12 seconds for three cents, with the low-intent queries surfacing as negative keyword candidates the same evening. The workflow is three steps: Google Ads to Jev to a spreadsheet. The same job is already productised next door in SEO: Phuc Le runs bulk search-intent classification in his SEO Utils tools on Jev, and says he will add the Decisions API once it is public.

Auditing what happens after the click. The most thorough run we found is Alex G.'s: 29,919 separate judgments over 112 live Meta ads, where the system watched 88 videos end to end and opened every landing page before answering. His headline finding: 80% of the ads it could judge land on a page that drops the ad's promise, and five live ads pointed at a 404.

Choosing assets before production. Deykhan Ten at Higgsfield shows Jev picking one hero product from 17, scoring 100 campaign photos to cut duplicates and weak angles, and then reviewing the twenty generated video ads for product visibility and framing. Jev decides which assets are worth producing; a generative model produces them.

Scoring your own account against your own history. The least flashy run is the most interesting: Caleb Kruse had a coding agent read two years of decisions in his ad account and turn that history into criteria Jev now applies to score each ad set for scaling. Not competitor research, not a demo: his own judgment, made repeatable.

What the numbers cannot say

Every run above shares three limits, and the honest posts admit them.

Longevity is a real signal: an ad still running after sixty days is an ad someone has kept paying for, which usually means it is doing something for them, and that is a fair reason to study its hook. But it’s not a performance number. The Meta Ad Library, which most of these runs read, carries no spend, no CPA and no ROAS for ordinary commercial ads, so a survival table says which hooks were worth keeping alive and nothing about what they returned. Meta's official Ad Library API adds a second constraint: outside political and issue ads, it archives only ads delivered to the EU or UK, and keeps them for a year, so a commercial ad that ran only in the US is not in it at all. And a decision without a rationale cannot be interrogated: when a tag looks wrong there is no why to read, which is why the practical advice circulating is to hand-check a few dozen tags before trusting the rest.

The counter-test worth keeping in view comes from Fedor Kovalev, who ran Jev on 83 contacts for half a cent and found it missed 40% of the ones his reference model marked as worth contacting. His phrasing deserves to travel: cheap AI can help you make bad decisions at extraordinary speed.

Part of the noise around Jev is manufactured: the most viral marketing post about it was published word for word by three different accounts, two of them employees of the same company, and TypeSafe has already collected a public callout for astroturfing. That does not make the named runs above wrong, but it does mean the volume of enthusiasm is not evidence. The same sorting will apply to the wave OpenAI's entry brings.

‍Decision models next to automation rules

If you run ad automation already, the natural question is whether either of these replaces anything. Mostly they do not, because the two act on different things.

A decision model judges rows. Jev and the Decisions API are built for the moment you have thousands of items, each needing the same bounded question answered once: is this search term from a buyer, which format is this creative, does this page deliver what the ad promised. Automation rules act on accounts: they watch live spend and performance on a schedule, apply thresholds you set, and pause, scale or notify with a log you can read afterwards. A rule needs a metric and a threshold; a decision model needs a question and a list of answers. Our guides to automating Meta ads and automating Google Ads cover the rules side, including how to derive thresholds from your own economics.

The two meet in one place, a batch judgment that feeds an action, and the more interesting builds are moving there. Jev flags which of ten thousand search terms look like waste; adding the negatives is an account action, taken by a person or a rule with a person's confirmation. Jev scores every fatiguing creative for replace, refresh or leave; acting on that list is account work. If neither the volume nor the repetition is there, a single high-stakes call with money attached, neither tool is the answer, and the judgment stays with a person. 

Use cases to try

🔵 Tag a category's ads and see what survives: Gloria Pu, Joshua G.

4,367 ads in 39 s for $0.27; 1,000 ads for under $0.06. Survival is a signal, not a performance number: the Ad Library carries no spend or CPA.

🔵 Sort search terms by intent: Adriaan Dekker

10,000 terms in 12 s for $0.03. A missed buyer is invisible; spot-check before trusting.

🔵 Audit ad-to-landing-page match: Alex G.

29,919 judgments across 112 ads for $0.23. Match is not conversion: it finds broken promises, not winners.

🔵 Select assets before production: Deykhan Ten

100 photos scored, 20 clips reviewed. Taste and brand judgment stay yours.

🔵 Score ad sets against your own history: Caleb Kruse

Two years of decisions turned into criteria. Criteria quality decides everything.

FAQs

What is Jev in one sentence?
Plus sign

A model from TypeSafe AI that answers bounded questions, a choice, a score, or a yes and no, with a probability attached, and cannot write text.

Is it accurate enough to run my ads on?
plus sign

Treat it as triage-grade. No independently attributed benchmark results exist yet, and the circulating advice, including from its fans, is to hand-check a few dozen judgments per run, use the confidence score to route uncertain cases to a person, and remember the counter-test above: one tester found it missed 40% of good contacts in his set.

Can it write or generate my ads?
black plus sign

No. It only decides. The builds pairing it with generative models use Jev to pick which assets are worth producing and to review what comes back.

How much does it cost?

The listed price is $0.042 per million input tokens with output free; TypeSafe's own illustration prices a single decision call at $0.000081. There is no published rate-limit table or pricing page yet, so treat this as the launch price rather than the price.

Does it see my campaign performance?

No. It judges whatever rows you feed it. The public runs mostly read the Meta Ad Library, which carries no spend or results data for commercial ads, and the official API archives commercial ads only when they were delivered to the EU or UK.

How do I get access?

A waitlist at typesafe.ai, with API keys issued to approved developers in batches. There is no self-hosted option; the model is API-only.

Elina Minnie

works at the intersection of storytelling, operations, and team strategy in tech. She writes about marketing, remote culture, and product ecosystems, and is a contributor to the Birch Blog.

Get started with Birch product

14
days free trial
Cancel anytime
No credit card required