Features
What it actually does
Clip tools mostly fail in the details — a subtitle a beat behind the joke, a crop with the streamer's head out of frame, a clip that ends before the punchline. Most of what is here exists because one of those happened.
Finding the moments
Signal ranking
Sliding windows over the VOD, scored on audio energy, hook phrases, laughter markers, punctuation density and deviation from the VOD's own average speech pace. Weights and thresholds live in a config file, not in the code.
Chat replay as a second opinion
Message-rate spikes and emote density from Twitch chat replay. Independent of the transcript, which matters because it still works on the moments the transcript got wrong.
A ranker that is measured, not assumed
Clips that real viewers made from the same VOD are fetched and clustered into labelled moments, then recall@K is reported for the signal ranking, the language-model ranking, and a blend — all against the same labels. Every change to the ranking is a number that went up or down, which is how six model-based ranking variants came to be rejected in favour of a z-score.
Clip boundaries that respect the bit
A clip cut on a fixed window truncates setup or punchline. The end of a clip is chosen from two independent readings — where the chat's reaction arc closes, and where a local model draws a subject seam — with the chat picking which side of the seam to land on.
Making the clip watchable
Per-speaker subtitle rows
Voice identification against an enrolled voiceprint puts each speaker on their own subtitle row, so a two-person conversation reads as a conversation instead of one run-on caption track.
Word-level subtitle timing
Transcription runs on each clip's own audio, with word timestamps. A single pass over a whole VOD silently drops speech, and those holes are invisible to any check that reads the transcript — so each clip is decoded again on its own.
Scene-aware vertical layouts
Split, full, or avatar-only, chosen per clip from a profile of where the streamer's camera actually sits in that VOD — measured by clustering avatar positions across the whole recording, where camera slots separate from composed scenes by edge proximity alone. A clip that crosses a scene change carries one layout per segment.
Punch-in when the gameplay stops mattering
When the moment is the streamer's reaction rather than what is on screen, the render punches in to the camera with head tracking. Ingesting at 1080p is what makes this hold up: the same camera box from a 720p source needs a 4.0× upscale instead of 2.7×, and it shows.
Burned-in headers, only where they help
Some clips are funny because of what is said, over a picture where nothing happens — invisible to a stranger scrolling. A line baked across the top holds them. On a clip whose payoff you can watch, the same line is clutter that risks spoiling it.
Which clips get offered one comes from a silent scroll test — the cached frames, no transcript, one binary question — combined with a measurement of how much the picture actually changes. When the two disagree the answer is optional, and the app says which one dissented.
Automated review of the finished render
The output file is checked as evidence, not trusted: the whole frame, the avatar band, and the gameplay pane are sampled separately. The check records the criteria it used, so changing them invalidates the old result instead of leaving a stale pass behind.
Running it day to day
One answer to "what now"
Every VOD sits in exactly one of four states — running, your turn, to run, done — derived from files on disk rather than a status column that goes stale. The bar at the top of the app shows them wherever you are.
Interrupted jobs resume
Each stage reports from its own artifacts, so a job that died mid-render picks up at the step it actually reached rather than starting over.
Re-ranking without losing finished work
Raising the number of candidates used to mean deleting every preview, render, trim, caption edit and verdict attached to the old ranking. Re-ranking now diffs candidates by the moment they point at and discards only what moved — and a posted clip is never discarded, whatever the ranking did.
Channel monitoring
clipfarmer can watch a channel and ingest new VODs as they appear, so the pipeline is already at your-turn by the time you sit down.
Giving the disk back
VODs are large. Finished jobs can be archived — the source video released, the decisions and the finished clips kept — and the app reports how much space that reclaimed.
Configuration
Scoring weights, clip length bounds, hook phrase lists, subtitle styling, render settings, ingest resolution, whisper model size, header thresholds and prompt text are all configuration rather than code. The documentation covers the file and the settings that matter most.