What Jev Can Actually Judge in a Video Edit
Jev is text-only, so it cannot see your footage. What it can judge is your transcript and your gap table - plus the arithmetic, the budget and the caveats.
A decision model shipped on 15 September, and one line of its specification forecloses most of what an editor would want from it. TypeSafe's Models page, on Jev's input: "Text only. String, JSON object, or array of text values. No image, audio, or video input." The State page repeats the restriction — "Images, audio, and video are not supported (yet)" — and the Models page adds that you must "pre-process non-text inputs (images, audio, video, binaries) into text or structured fields before sending them as state." A model that cannot see a frame or hear a waveform cannot judge one.
So the useful question is narrower: which of the texts your edit already produces are worth handing it, which it will quietly mangle, and what the call costs.
The texts an edit here produces
A transcript. Auto Subtitle Generator runs Whisper through transformers.js inside the tab — Balanced is whisper-base, about 196 MB pulled once from the Hugging Face CDN and cached, transcribing in 30-second windows with a 5-second stride. Out comes SRT, VTT or TXT at phrase level, never having left the machine.
A gap table — on screen, not in your clipboard. Silence Remover runs silencedetect and scrapes silence_start, silence_end and silence_duration out of the FFmpeg log with a regex. Its −30 dB and 0.5 s are starting values, not fixtures: both are editable number inputs. Jump Cut Maker runs the same filter behind three presets — Tight at −35 dB / 0.4 s, Natural at −32 dB / 0.7 s, and Custom, which is how you dial it to whatever thresholds your own FFmpeg run used. But both tools render the result as a read-only list: no per-row toggle, no copy button, no export. To get a gap table as text you run silencedetect yourself.
Cue files somebody else wrote, normalised through Subtitle Converter (SRT, VTT, ASS and TXT, all in the page) or shifted through SRT Offset.
Filenames.
That is the complete list. Everything else is perceptual — eyeline, exposure, whether the second angle matches the first on colour, whether the burn covers the speaker's mouth. None of it reaches the model in any form, so none of it can be graded by one: outside the contract, not merely unmeasured.
Two local outputs look API-ready and are the worst possible input. The crop=w:h:x:y line cropdetect settles on in Black Bar Remover is four integers; Video Metadata reports six fields — filename, size, MIME type and last-modified date off the File, duration and resolution off a <video> element — no codec, and most of the rest numbers. TypeSafe's own known-failure-modes page says Jev "is not a calculator", "does not count reliably", and "reads dates as text, not as ordered quantities." Which gap is longest, how many cues exceed 42 characters — that stays in your code.
One job, carried through
A 42-minute recorded talk, one speaker, one camera. Cut the dead air without cutting the beats where he is thinking.
Jump Cut Maker on Natural finds the pauses and removes every one of them; Silence Remover does the same. Neither will take a keep-list back. So the gap table comes from your own FFmpeg run, and what you decide about it has to land somewhere else.
Getting the two files to line up
This is the part worth budgeting for, and the draft version of this workflow that skips it does not run.
Produce both texts from one local pass over the same audio: one ffmpeg -i talk.mp4 -af silencedetect=noise=-32dB:d=0.7 -f null - for the gaps, and the audio you feed Whisper extracted from that same file at the same start. Two decodes of two differently-trimmed files do not share a timebase, and 200 ms of drift is enough to attach the wrong sentence to a pause.
Then the join. For each gap, you want the cue that ends nearest before silence_start and the cue that begins nearest after silence_end. Whisper's chunker splits on pauses too, so most gaps fall cleanly between two cues — but not all. When silence_start lands inside a cue's span, Whisper merged speech across that pause: the "text before" is a fragment you cannot cut out without word timestamps, which need the cross-attention export. Skip those gaps, leave them in the cut, and log how many you skipped. With a slow speaker it can be a quarter of them, and that fraction is the honest ceiling on how much of the job this automates.
The request, once
POST https://api.typesafe.ai/v1/systemone
Authorization: Bearer $TYPESAFE_API_KEY
Content-Type: application/json
{
"model": "jev-1.13.0",
"state": {
"before": "...and that's the entire pipeline running in the browser",
"after": "Right. So the second thing I wanted to show you"
},
"questions": {
"gap_037": {
"type": "noul",
"instructions": "The speaker finished a complete thought before this pause.",
"criteria": {
"true": "The text before ends on a completed sentence or clause.",
"false": "The text before breaks off mid-clause: a restart, a hesitation."
}
}
}
}
Three primitives exist and choosing between them is the design. noul is a gate: criteria optional, one probability back, and no confidence field at all — the docs are explicit that "Noul has no separate confidence", and most summaries of the API get this wrong. choice routes: criteria is a map of at most 255 named options, returning the winner, a probability per option, and a confidence. score is a ladder of ordered levels — ten is the most a Score question takes — returning a probability-weighted value that can land between two of them.
Pin the version. jev-latest is an alias currently resolving to jev-1.13.0, and TypeSafe's own advice is that "if you have tuned confidence thresholds against a specific version, pin that version's ID instead of the alias." And split every question until it is atomic; the docs' prompting guide says to "break down complex or ill-defined questions into separate questions that each evaluate one property." "Did he finish the thought, and is this a good cut" is two questions, and the model cannot answer the second one at all.
What comes back, and what does not
The whole envelope is { model, answers, usage } — for a noul, one number. No rationale, explanation or text field exists anywhere in the response schema; heise raised this as a transparency problem.
This is also where the zero-hallucination headline has to be read literally. What is guaranteed is the envelope: you asked for a noul, you get a float between 0 and 1, keyed by the id you sent, every time, with no parse step that can fail. Nothing is guaranteed about whether 0.81 is the right number for that pause. When a gap comes back marked "finished thought" and plainly was not, the response is still perfectly well-formed — and there is nothing to read back.
So do not put the line at 0.5. TypeSafe's own known-failure-modes page for jev-1.13 publishes a noul of 0.72 for "Is the customer asking for a refund?" and 0.47 for its own negation — a sum of 1.19 — and, on another ticket, one question coming back as noul 0.22 while the same question posed as a yes/no choice returned P(yes) 0.01 at confidence 0.97. Its advice: "don't carry a threshold tuned on a Noul over to a Choice."
The largest independent calibration work published so far — roughly 16,000 calls across four public datasets at 300 items each, plus about 3,000 decisions from the author's own production pipeline — found the ends trustworthy and the middle not. Above 0.9 he measured 97% correct on his own pages and 90 to 99.6% across the public sets; the band in between ran anywhere from 6% to 83% depending on the task. His verdict is the best one-sentence summary of the whole evidence base: there is enough here "to put it in front of a classifier and not enough to put it in place of one."
Everything else stays in your code: the subtractions, the timecodes, the merge of adjacent keeps.
The bill
Say the detect pass returns 180 gaps. Per gap the state is two cue texts — maybe 40 words — plus instructions and criteria. Call it 120 input tokens, noting that nothing published says which tokenizer counts them or whether JSON punctuation counts. 180 × 120 is 21,600 tokens; input is billed at $0.042 per million and output tokens are unmetered, so the whole talk costs $0.0009. Nine hundredths of a cent.
One published neighbour: MinusPod benchmarked podcast ad-break detection over 14 episodes, 12 of them ad-bearing — one yes/no per transcript segment, probabilities aggregated into spans — at roughly $0.0016 and 0.1 s per episode against about $0.216 and 24.2 s for Claude Haiku 4.5, F1 0.929 against 0.920 at IoU 0.5. The F1 parity is not a tie on the same terms: Jev's precision was 0.079 higher and its recall 0.053 lower, so it kept more of the show and missed more of the ads. The author's own verdict is "competitive with Haiku at a fraction of the cost and worth a real evaluation — not 'better than Haiku'", and he notes his parameters were tuned on the same episodes he scored them on.
Time: TypeSafe's own scoreboard puts Jev at 0.4 s a case, MinusPod measured about 0.1 s an episode, and a shell-command gate logged 193 to 642 ms across eleven calls — all inside or near the claimed 70–500 ms band. 180 sequential calls at 0.4 s is 72 seconds; eight in flight is nine. The published ceiling, flagged as liable to change without notice, is 1,200 requests a minute and 250,000 tokens a second, so 180 requests is fifteen percent of one minute.
Then batching, where people expect the money to be. Questions in one request run in parallel and in isolation against the same state: Codeseys-Labs measured 205 ms for one question and 304 ms for fifty, and TypeSafe's parallel-questions cookbook reports "batching: 12.2x cheaper, 10.0x faster" for 13 questions. Read where that came from before you budget on it. The shared state was the Wikipedia article on the GDPR, about 54,000 characters, and in the cookbook's own words "the 13 single-question calls re-send the article 13 times; the batched call sends it once." The multiplier is an artefact of one enormous shared state, not a general discount.
On the talk the arithmetic is much smaller. Send a ten-minute window of transcript as state with forty nouls against it, one per gap: about 1,900 tokens of state plus 1,200 of questions, 3,100 a request, five such windows. Those same forty gaps sent one at a time cost 4,800 tokens instead. Batching saves 1,700 tokens a window — about four ten-thousandths of a dollar across the talk.
The saving is not the point, and pretending it is how people talk themselves into the worse design. Every question in that windowed request reads all ten minutes to judge one pause, and the failure-modes page is blunt about it: "Accuracy falls as the state grows with content unrelated to the decision. Unrelated detail acts as a distractor." Batching earns its keep in the other direction: twenty questions about the same cue block — which chapter, is this a sponsor read, does the line stand alone — for the tokens of nineteen extra questions and almost no extra time.
The ceiling, and which page you believe
Two published numbers for one quantity, and they disagree.
The Models page gives "64k tokens per request; 32k tokens for state plus the longest question", and then, unusually, reconciles its own two figures on the spot: "The 64k budget covers the state plus all questions combined; the 32k budget applies to the state plus the single longest question." The Primitives page, describing that same shared budget, says it is "around 32,000 tokens, roughly 150,000 characters of English text." So the conflict is a single clean one: 32k or 64k for state plus all questions.
For the 42-minute talk it does not matter. Roughly 5,900 words is about 33,000 characters, inside either reading with room to spare. A three-hour stream is where it starts to: at the docs' own ratio of 150,000 characters to 32,000 tokens, three hours of speech is around 141,000 characters, which all but fills the smaller budget before a single question is attached and sits comfortably inside the larger one. That is precisely the case where you need to know which page is right, and you cannot. Window it by chapter and send only the cues the question needs.
What happens when you exceed the budget is not published at all. TypeSafe's documented statuses stop at 401 Unauthorized, 422 Unprocessable Entity, 429 Too Many Requests and 529 Overloaded; there is no over-budget code among them. The Pydantic AI integration names one — the request "fails with a ModelHTTPError (max_tokens_exceeded), which a FallbackModel hands to the model behind Jev like any API error" — but that is the framework's own exception class, not something the endpoint returns. Code against the raw POST above and you will be reading a status line, not catching that name.
Where the honest answer is still to do it by hand
Scrubbing 180 pauses in a timeline is maybe twenty minutes. Writing the script, aligning the two timebases, hand-labelling a few hundred gaps to fit a threshold and checking it on held-out ones costs more than that the first time, and probably the second. It pays back on the tenth episode of a series with the same speaker, mic and room. Not on one video.
The accuracy case is also thinner than the speed case. Speed replicates: several unrelated parties measured it on different tasks and gateways, all inside or near the claimed band. Accuracy does not. The one pre-registered evaluation against real human labels, on Banking77 and CLINC150, put Jev at 0.832 and 0.870 against GPT-5.6 Terra's 0.875 and 0.915, and below a frozen bge-small-en-v1.5 encoder with logistic regression at 0.933 on Banking77; its author logged the verdict AMBIGUOUS and wrote that "where labeled data exists, a 9ms supervised encoder wins."
And the vendor's headline deserves reading twice. On TypeSafe's own scoreboard Jev's 67.8% is a mid-table score: GPT-5.6 Sol is at 74.1%, Claude Opus 5 at 73.1%, and Jev is a statistical tie with GPT-5.6 Terra's 67.9%. Jev leads that board on cost — $0.0004 a case against $0.0836 and $0.1761 — and on speed, 0.4 s against 23.3 s and 37.8 s. It loses on accuracy, on the vendor's own numbers. On top of which it is not accuracy in the ordinary sense: TypeSafe's model capabilities team chose the tasks and wrote the workflows, and the reference labels were the averaged answers of GPT-6 Astra and Claude Fable 5.1, so the figure measures agreement with two other models on a test the vendor set for itself. heise's review flagged the absence of any independent verification, and there is still none.
Every published evaluation, vendor or independent, ran on clean written text. A Whisper transcript is disfluent and carries ASR errors, and its punctuation is itself a guess at where the sentences end — which is the exact judgement you were about to ask Jev to make from that punctuation. Nobody has measured Jev on one. And if the talk is not in English, you are further out still: TypeSafe says English is the primary training language and that other languages, CJK scripts included, are "handled but not equally well."
What leaves your machine
Nothing on this site uploads. Eighty-four tools, FFmpeg compiled to WebAssembly, Whisper through transformers.js, exactly one route handler in the whole application — for /llms.txt — and no server-side fetch anywhere. There is no backend here to call an API for you and there is not going to be one. The JavaScript SDK's dangerouslyAllowBrowser flag defaults to false for the obvious reason: a key in a page is a key given away.
So the moment you write this script yourself, a transcript made on your own machine leaves it. Where it goes: "The Services are hosted in the United States." Retention appears only as "as long as reasonably necessary" in the Privacy Policy and "as long as necessary" in the DPA — no number of days in either, nor in the customer agreement or the Terms. Zero data retention is offered to enterprise customers on request by email to privacy@typesafe.ai and appears in none of those four documents; it is one sentence on a docs index. There is no per-request retention flag in the API, and the SDK docs note that secret headers are redacted from log output while "request and response bodies are not."
Credit where it is written down: the Privacy Policy states that TypeSafe "will not train or fine tune any artificial intelligence or machine learning models on your prompts or other Input," and the Models page repeats that Jev is not trained on customer requests or responses.
What stays local regardless
The cut, the render, the burn, the frame.
Neither Silence Remover nor Jump Cut Maker will accept a keep-list, so the route back into the browser depends entirely on how long yours is — and the arithmetic is unforgiving. Video Splitter in its split-at-these-times mode turns n cut points into n+1 files, and removing a gap takes two cut points. So 180 gaps is about 360 split times and about 361 clips, each one held as a live blob URL, with Download All firing a single save every 400 ms: roughly two and a half minutes of save prompts. Video Merger then wants those clips in order with per-file Move Up and Move Down buttons and no drag-reorder, writes every input into the WebAssembly filesystem before it starts, and always re-encodes with libx264 and AAC, because concat with -c copy produces broken output when the inputs differ. Nobody is hand-ordering 361 clips, and twenty minutes of scrubbing just won the comparison outright.
A 180-entry keep-list's real destination is a file: an EDL, an FCPXML, or a CSV of in and out points, written by the same script that made the calls and imported into Premiere, Resolve or Final Cut, where the trims stay non-destructive and reorderable. Video Splitter plus Video Merger is the right route when the keep-list is short — chapter marks, a dozen cuts, a reel. Subtitle Burner handles the cues either way; it does not care how you got there.
If you never call this API, the transferable part survives. The model is fast because it answers one small typed question whose legal answers you wrote out in advance — and all the schema guarantee buys you is that the answer will be one of them, never that it is the right one. Composing that question — "the speaker finished a thought before this pause," not "is this a good cut" — is most of the work and nearly all of the value. Write a few on paper and you will find half the decisions you meant to automate were never text decisions.
Sources
Every figure above is someone else's published number, and most of them are the vendor's own. Check them rather than take mine.
- TypeSafe — Models (input modality, context length, rate limits, language caveat)
- TypeSafe — Primitives (choice, score, noul and what each returns)
- TypeSafe — State
- TypeSafe — Known failure modes for jev-1.13
- TypeSafe — Parallel questions cookbook
- TypeSafe — Privacy policy
- Independent latency measurements — Aman Kumar
- Independent latency measurements — Codeseys-Labs
- Podcast ad-break comparison — MinusPod
Read on 20 September 2026. Jev shipped five days earlier, and TypeSafe versions both the model and its failure-modes page, so treat every number here as dated.
Try it yourself — free in your browser
No upload, no signup, no watermark — these tools run on FFmpeg WebAssembly locally.