AudioAdvanced

Jev and Your Transcript: What the Numbers Actually Support

Jev is fast and cheap, but its 67.8% is an agreement rate and its confidence is undefined. What that leaves for transcripts, cue gating and jump cuts.

Applicable Software:

The launch table reads like a scoreboard. Jev: 67.8%, $0.0004 per case, 0.4 seconds. GPT-5.6 Terra: 67.9%, $0.0304, 10.1 seconds. GPT-5.6 Sol: 74.1%, $0.0836, 23.3 seconds. Claude Opus 5: 73.1%, $0.1761, 37.8 seconds.

Take the accuracy column at face value first, because it is not flattering: Jev ties the weakest frontier model on the board, 6.3 points behind Sol and 5.3 behind Opus 5. The pitch is parity with one frontier model at roughly a seventy-sixth of the cost and twenty-five times the speed.

One of those columns means what it looks like. Latency has been reproduced by people with no stake in it. The accuracy column does not measure accuracy. The price is real today, and the vendor declines to promise it will stay there.

The accuracy column is an agreement rate

From TypeSafe's evals site: "Instead of debating the correctness of the harness and labels, we assume that the code is correct, and measure against the current smartest large models. For this eval, the reference labels are generated via an average of the responses of GPT-6 Astra and Claude Fable 5.1, both at high thinking."

The column counts how often Jev matched the averaged opinion of two other models. Nothing in the eval knows what right was — which is also why neither Astra nor Fable appears in the results table.

TypeSafe's own caveat is sharper than the criticism. The reference choice "biases answers towards OpenAI and Anthropic's models," and, on its own account, "We likely underestimate the relative performance of our model and DeepSeek's models." Read that back: the two models topping the accuracy column are an OpenAI model and an Anthropic model, graded against the average of an OpenAI model and an Anthropic model. The column is partly a measure of family resemblance to the graders, and the vendor is the one saying so. The four workflows were also built in-house — "they were made by individuals on our model capabilities team, so some bias could exist."

The structured-output chart is worse than biased: it has no shared denominator. Jev posts 0% schema errors against Claude Haiku 4.5's 45.5%, and the caption explains where each number came from. "The numbers for LLMs are from OpenRouter i.e., there almost certainly is bias here: more complex queries might be routed to better models. Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots." Haiku's figure is third-party production telemetry on unrelated traffic; Jev's is a construction guarantee restated as a measurement. They were never run on the same thing.

That 0% promises the reply will fit your schema, not that it is right. On Hacker News, the account that identifies itself as "CEO here" in the same thread wrote that "because these models are probabilistic, it's also possible to be confidently wrong (and all future models will be smarter still and still have that possibility)." He frames it as a property of every model, which is fair — and which is exactly the failure calibration is supposed to bound.

Speed replicates. Throughput does not.

Four unrelated parties have measured per-call latency and the medians land in or near the claimed 70–500 ms band: ickma2311's pre-registered eval at median 0.44 s, p95 0.71 s on Banking77; Mike Moore's jev-benchmark at p50 421.6 ms, p95 542.0 ms; Codeseys-Labs at 205 ms for one question and 304 ms for fifty against the same state; MinusPod at roughly 0.1 s per podcast episode against Haiku 4.5's 24.2 s.

A fifth test — 6 of 7 planted defects caught at a median 0.35 s a passage, where Claude Fable 5.1 caught all 7 at 8.8 s — comes from Every, written up by Every's head of evals and run by Every's CEO: a hands-on report from an AI publication, not a disinterested audit.

Two runs landed outside the band. Aman Kumar measured a 0.8–0.9 s median across four public datasets with 20 calls in flight, and his worked example took 850 ms from a laptop; a Chinese-language support eval clocked ~890 ms.

And per-call latency is not throughput. ickma2311 hit rate limiting hard enough that 45 of 200 items needed retries, the run took about 3.5 hours of wall clock, and recorded latency "excludes backoff waits" — "throughput under rate limiting is far worse than per-call latency." TypeSafe's published ceiling is 1,200 requests per minute and 250,000 tokens per second; early access in practice was not that. The lever that helps is the one MinusPod used: questions inside one request run in parallel, so batch many cues per call rather than one call per cue.

Accuracy swings by task, and the ambiguous subset cracks first. Moore's 60 hand-labelled tool-call cases: 91.7% overall, 100% clear, 91.7% adversarial, 71.4% ambiguous — and he is careful about his own figure, noting n=60 cannot separate variants and that "nothing here supports or refutes the vendor's speed and cost multipliers." A Chinese eval on 130 hand-labelled support samples hit 97.7%. A 2,000-email phishing bench had Jev at 62.6% against Haiku 4.5's 81.3%, with worse calibration and a five-point swing from wording alone.

Ambiguity is the ordinary condition of a transcript.

The one transcript-shaped test

MinusPod's ad-break detection is the one published test on anything resembling a transcript, and it came out well: across 14 podcast episodes at IoU 0.5, Jev scored F1 0.929 against Haiku 4.5's 0.920, at roughly $0.0016 an episode instead of $0.216.

The request shape is the part worth copying — one question per transcript segment, "is line L0042 advertising?", all in a single request per window, so answers align to segment boundaries that already exist instead of asking a model to invent timestamps.

Three caveats, and the last decides whether you would let it cut. Parameters were tuned on the same 12 ad-bearing episodes they were scored on; cross-validation found the thresholds do generalize, but choosing between prompt variants does not. It is one corpus of clean written transcripts, not ASR output with a real word-error rate. And the precision/recall split runs the wrong way for an editor: precision +0.079 but recall −0.053 against Haiku, which the author describes as erring "toward leaving ads in rather than cutting content." His verdict: "competitive with Haiku at a fraction of the cost and worth a real evaluation — not 'better than Haiku'."

The one pre-registered test against public human labels

One test was pre-registered against public human-labelled data, and it went the other way. A developer publishing as ickma2311 hash-pinned a pre-registration, then ran Jev against Banking77 and CLINC150 — intent sets with human labels, not model consensus.

Banking77 (paired n=208): Jev 0.832, gpt-5.4-nano 0.793, GPT-5.6 Terra 0.875, and a frozen bge-small-en-v1.5 encoder with logistic regression on top, trained on 10,003 examples, 0.933. CLINC150 zero-shot (n=200): Jev 0.870, nano 0.795, Terra 0.915. Both verdicts recorded AMBIGUOUS.

Read the Terra gap with the right error bar: the Banking77 sample was cut from 300 to 208 by Jev's rate limit, and the author refuses to call that harmless — "Terra scores 0.875 on the 208 retained items vs 0.804 on the 92 omitted ones." The retained subset is the easier one, and the comparator is what gained from it.

His conclusion is blunter than anything a critic wrote: "Where labeled data exists, a 9ms supervised encoder wins." Sorting a cue into one of a fixed list of chapter names is intent classification with different vocabulary. If you have hand-labelled thousands of cues from your own channel, the encoder wins and it is not close. Almost nobody has, and that gap — bounded answers, no labelled history — is the honest case for a zero-shot model, narrower than the launch post implies.

Confidence is a shape statistic, and the shape is unpublished

This is the number most people will threshold on, and it is not a probability of being right. TypeSafe describes it as a collapse of the answer's distribution — concentrated on one outcome means confident, spread out means uncertain — and says the confidence property "collapses that shape into a single number from 0 to 1."

How it collapses is not documented. The confidence page ships an interactive widget whose source computes (count * peak - 1) / (count - 1), clamped, and the prose beside it is careful: "This demo uses (3 × largest probability − 1) / 2 to approximate confidence for three options." An approximation, for a demo, at one option count.

Check it against TypeSafe's own published API responses and it does not hold. Two three-option choice examples from the API reference:

probabilities published confidence demo formula
billing 0.08, technical 0.85, sales 0.07 0.82 0.775
billing 0.84, technical 0.159, sales 0.001 0.596 0.76

Nearly the same top probability — 0.85 against 0.84 — and confidence differs by more than two-tenths. A published three-level score peaking at just 0.65 carries confidence 0.78, higher than the 0.84 case. Aman Kumar, who ran about 16,000 calls, hits the same wall from the API side: "TypeSafe does not say how confidence is computed; it is not the top probability (0.93 against 0.97 above)."

So do not invert it. You cannot read confidence 0.70 back into "the winner sat at 0.80," and you cannot assume a cut-off tuned on a five-chapter list transfers to a twelve-chapter one — whatever the real function is, the published examples make it depend on more than the peak. Fit the threshold on your own option list and check it there. And noul, the yes/no primitive, carries no confidence field at all: the single probability is the whole answer.

TypeSafe does market the property. A docs feature card is titled "Calibrated confidence" — "RLCD communicates uncertainty through calibrated probabilities instead of tending toward overconfidence" — and the launch blog's confidence row reads "Calibrated: higher confidence means higher accuracy." What it never publishes is a measurement. Across the 109-entry documentation index, the launch blog and the evals site there is no reliability diagram, no expected calibration error, no Brier score, no per-bin accuracy table. The one scoping sentence it does publish is about groups: "Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct." And "higher confidence means higher accuracy" is monotonicity, strictly weaker — a model reading 0.95 whenever it is right 60% of the time satisfies that sentence and fails calibration.

The strongest external datum here is ickma2311's, and it does not favour the vendor: on CLINC150, Jev's confidence did not rank its own errors better than a chat model's verbalized confidence — AUROC 0.734 against 0.816, paired interval including zero. The number you were going to gate on was no better at spotting its own mistakes than asking an LLM how sure it felt.

The vendor publishes its own failure modes

TypeSafe maintains a per-version jaggedness page, and for jev-1.13 it is more damaging than anything a reviewer produced. The same question about the same ticket, asked two ways, returns noul 0.22 — and, as a yes/no choice, P(yes) = 0.01 at confidence 0.97. A noul and its own negation do not sum to one: "Is the customer asking for a refund?" comes back 0.72, "Is the customer asking for something other than a refund?" comes back 0.47. Total 1.19. On ladders, "jev-1.13's score levels are weak in numerical calibration," so an expectation between two levels can be tested against a threshold but never read as a magnitude. The guidance: "Don't carry a threshold tuned on a Noul over to a Choice, and don't hold the model to arithmetic identities between separate questions."

Every question shape needs its own line, fitted separately, even when two questions ask the same thing in different clothes. And pin the version — jev-latest is an alias, responses report jev-1.13.0.

The same page rules out the obvious transcript workflow. "Accuracy falls as the state grows with content unrelated to the decision. Unrelated detail acts as a distractor." Elsewhere it names the effect: "Jev suffers from context rot, so unrelated material in the state costs you accuracy." There is a hard ceiling too — 64k tokens per request, 32k for state plus the longest question — so an hour of transcript does not fit, and would degrade the answer if it did. Kumar found the same boundary independently: "long input, fuzzy labels, or a rule that needs reasoning: accuracy and confidence fall together," and "never ask it anything that needs the whole document read before the answer exists."

The workable unit is one cue plus a tight window, not the document. Which means the cost arithmetic is per cue: Kumar's 336-token call cost $0.000014, so a 400-cue episode is well under a cent — as 400 questions, cheap in wall-clock only if you batch them.

Where the state actually comes from

Auto Subtitle Generator runs Whisper through transformers.js in your browser — whisper-base is ≈196 MB, fetched once and cached — and it is the only tool here that hands you a machine-readable artefact: SRT, VTT, TXT and a Copy SRT button, every cue editable in place first. That SRT is your state, and cue text is what you can gate on: is this cue an ad read, a chapter boundary, a sponsor mention. Exactly what MinusPod did.

Do not gate on the gaps between cues. The generator requests phrase-level timestamps, not word-level, so cue edges are approximate — and when Whisper returns a chunk with no end timestamp the exporter fills it with the next cue's start, so an inter-cue gap in the SRT is sometimes not a real silence, and sometimes a real silence that has been erased. Treating that boundary as a pause means gating on something the exporter invented.

If you want pauses you need pause timings. Silence Remover detects with silencedetect=noise=-30dB:d=0.5 and Jump Cut Maker runs the same detector at −35 dB / 0.4 s. Both display what they found, and neither offers a copy button, a download or a checkbox. The list is read-only. A machine-readable gap list means running silencedetect yourself, off this site.

Setting the line, and what these tools will do with it

Kumar's write-up is the most detailed outcome-scored account published so far — roughly 16,000 calls, per-item answers and scoring code in the open — and its finding is that the ends are trustworthy and the middle is not. On public sets, 82% of answers came back at confidence 0.9 or higher and 96% of those were correct; below 0.9 it ran 55–72% by band. On 2,559 production gate calls scored against what actually happened, 82% landed under 0.1 and 0.3% of those turned out to matter, the 4% at 0.9 and above mattered 97% of the time, and the 14% in between ran from 6% to 83%. "No single cut-off works there. So trust the ends, and hand the middle to a model that can actually read the thing."

Setting the line is mechanical: take every item you know is a real positive, find the lowest probability Jev gave any of them, and put the line at half of that. Fit on part of your data, check on the rest, re-check as new material arrives. His three production lines came out at 0.12, 0.065 and 0.025 — nowhere near 0.5, or each other — from 187, 124 and 61 known positives. Low hundreds of hand-labels, not thousands.

Then the constraint nobody selling this will mention. The two cutting tools here do not take your list. Silence Remover walks its own silentSegments, Jump Cut Maker its own pauses; each inverts that into a keep-list and concatenates. No per-segment checkbox, no cue-list import, and the only controls are a global dB threshold and a minimum duration. There is no way to keep gap 7 and cut gap 8. After 300 hand-labels and a fitted threshold of 0.065, nothing here will honour a single per-gap verdict.

So give the Jev pass an honest job. Either it decides something about the file — which global dB and duration settings to run, whether this recording is clean enough for an automatic pass at all — or per-cue action means cutting by hand, or in your own ffmpeg script, elsewhere. Audio Cutter is not the escape hatch: one range, -c copy, and audio-only, rejecting an MP4 at the file picker unless you have already pulled the track with Extract Audio.

Review is where the site does help, with one trick. Subtitle Preview turns every parsed cue into a click that seeks the player to that cue's start, highlighting whichever cue is under the playhead. It has no flag, no filter, no jump-to-next-flagged, so a 400-cue episode SRT buries your 40 flagged cues among the rest. Write a flagged-only SRT from your own script and load that instead: the parser accepts any cue set and renumbers from 1, so a 40-cue file really is 40 clickable rows. It draws its own plain black box rather than your libass styling, so it checks words and timing, never the burned-in look.

And nothing comes back to argue with. The whole response is { model, answers, usage } — no rationale, no explanation, no text of any kind; heise named this the transparency problem. A wrong "this cue is an ad read" is a number you cannot question. Nor should you mistake stability for correctness, or assume stability: TypeSafe calls jev-1.13 "extremely consistent," and on easy inputs it is — Kumar found the same page asked twice moved by about 0.01 — but on his hard page type "its probability swung by 0.5 between runs on the same page." On exactly the ambiguous cues where you wanted a second opinion, it is neither stable nor demonstrably calibrated.

The transcript leaves your machine

This site's premise is that nothing uploads: Whisper ran in your tab and the file never moved. The moment you call Jev that changes.

There is no backend here — this site cannot call an API for you. Anything you do with Jev you do from your own script, and the state goes with it. TypeSafe's Services "are hosted in the United States." Its privacy policy retains personal data "for as long as reasonably necessary to provide you with the Services, or otherwise in support of our business or commercial purposes"; zero data retention is an enterprise arrangement, requested via privacy@typesafe.ai under a DPA. On training it is unambiguous: "We will not train or fine tune any artificial intelligence or machine learning models on your prompts or other Input." There is no self-hosted, on-device or open-weight build, and none announced.

For a public podcast, a trade you might shrug at. For an unreleased interview, a client's rough cut, anything under NDA, it is the whole decision — and not one a speed number settles.

What would change this reading

A published reliability diagram with expected calibration error, stating binning scheme and sample size — binned ECE understates true error, so a bare figure is not enough. With genuinely calibrated probabilities you stop fitting thresholds by hand: the cost-optimal line is fixed by the cost ratio alone, C_FP / (C_FP + C_FN). If a wrongly cut line costs nine times a wrongly kept one, act above 0.9 and tune nothing.

A transcript evaluation not tuned in-sample. One transcript test exists, it favoured Jev on F1, and its author tuned on the episodes he scored. What is missing is an independent run on ASR output with a real word-error rate, on more than one 14-episode corpus, reporting recall separately — because recall is the half that decides whether you let a model cut anything.

Any input that is not text. The API takes "Text only. String, JSON object, or array of text values. No image, audio, or video input." It can weigh the words at 04:12; it can never see the frame at 04:12 or hear the waveform under it.

Something to run locally, so the section above stops being a trade at all.

Until then the position is narrow. Per-call speed is established and reproduced. Price today is $0.0004 a case, and TypeSafe's own line is "We can't prove it isn't subsidized; we'll need the long-term to prove the sustainability of our pricing" — they expect it to fall rather than rise, but the arithmetic only holds while it stays put. Accuracy on your material is unmeasured until you measure it.

And that arithmetic is unforgiving at small scale. For one ten-minute video, hand-labelling a couple of hundred cues to fit a threshold costs more than doing the pass by hand. The per-cue call is a fraction of a cent; the labelling, the batching, the review of everything in the middle band, and the fact that no tool here will act on the result are all fixed costs. It only turns over somewhere in a back catalogue.

Sources

Every figure above is someone else's published number, and most of them are the vendor's own. Check them rather than take mine.

Read on 20 September 2026. Jev shipped five days earlier, and TypeSafe versions both the model and its failure-modes page, so treat every number here as dated.

Try it yourself — free in your browser

No upload, no signup, no watermark — these tools run on FFmpeg WebAssembly locally.

Tags:jevtypesafetranscriptionsubtitlessilence removalai editingwhisper