How it works
From an eight-hour VOD to a clip worth posting
clipfarmer is a pipeline with a person in the middle of it. Each stage writes its results to disk, so progress survives a crash, a restart, or a change of mind about the ranking.
1. Ingest
You give clipfarmer a Twitch VOD or YouTube URL. It downloads the video, extracts an audio track, and caches keyframes across the whole VOD for later use.
Ingest resolution is a real decision rather than a default, because it sets a ceiling on what the renderer can do later. A punch-in crops the streamer's camera box out of the source frame and scales it to fill a 1080×1920 clip. From a 1080p ingest that box is around 401×713 — a 2.7× upscale, which holds. The same box from a 720p ingest is 268px wide and needs 4.0×, which visibly does not.
2. Chat replay
When the source is a Twitch VOD with chat, the chat log is fetched alongside the video. Message-rate spikes and emote density are a second, independent opinion about where something happened — one that does not depend on the transcript being right.
3. Transcribe
Whisper produces a transcript with word-level timestamps. The timing is what lets subtitles land on the syllable rather than drifting behind the speech, and it is what the candidate scorer reads to find hook phrases and laughter.
Transcription runs against each clip's own audio rather than one pass over the whole VOD. A single long pass silently drops speech in places, and those holes are invisible to any check that reads the transcript — because the transcript is the thing that is missing.
4. Frame cache and scene profile
A clip is rendered as one of three layouts: split (gameplay stacked over an avatar band), full (a single crop of a composed scene), or avatar-only (centred on the streamer, with no invented gameplay pane). Getting this wrong is the most visible defect the renderer can produce — a band cut out of a composed scene shows scenery where a face should be.
Asking a vision model "is this an overlay or a set?" per clip is not reliable, because an idle scene with a game-like backdrop is indistinguishable from real gameplay in any single frame. The separating fact is not in the frame, it is in the VOD: sampling the whole VOD and clustering where the avatar sits collapses it into two or three tight positions. What identifies the webcam slot is edge proximity, and only that — camera clusters sit at edge margin 0.000 while composed scenes sit at 0.18–0.26. Size does not separate them at all.
Every clip's layout is then a lookup against that per-VOD profile, and a clip that crosses a scene change carries one layout per segment.
5. Candidates
Short sliding windows are scored on a weighted blend of signals, all of them cheap and all of them measurable:
| Signal | What it catches |
|---|---|
| Audio energy | RMS loudness spikes — reactions, hype, a room going up |
| Hook phrases | A configurable list: "no way", "let's go", "insane" |
| Laughter markers | Laughter in the transcript, and emote laughs in chat |
| Punctuation density | Exclamation and question marks per word |
| Speech pace | Deviation from the VOD's own average words per second |
6. LLM judge, and measuring it
A local language model re-ranks the candidates. Whether that helps is not taken on faith: the clips real viewers made from the same VOD are clustered into labelled moments and used as ground truth, and recall@K is reported for the signal ranking, the model ranking, and a blend of the two — all against the same labels.
That harness is why the review step exists in the shape it does. Six LLM ranking variants lost to a z-score here. Generation is a model's strength; choosing is taste.
7. Your review
Previews are quick, subtitle-free trims rendered so there is something watchable behind every verdict. You approve or reject; nothing auto-selects.
The review gate sits before rendering, so a clip awaiting a verdict is unrendered by definition. Your verdicts and the pipeline's checks are tracked as separate things, because a clip the pipeline passed is not a clip you approved.
8. Render
Approved clips are rendered as finished vertical video: subtitles with a row per speaker, the layout the scene profile calls for, a punch-in when the gameplay stops mattering, and a burned-in header on the clips that need one.
Which clips need a header is decided by a silent scroll test: the window's cached frames, no transcript, one binary question — with the sound off, do the pictures alone show a stranger something happening, or only a person present in a place? Withholding the words is the whole design. Include the transcript and the answer tracks the transcript instead of the frames.
The pixels get a vote too, as median luma change across the window's cached frames. Neither half decides alone, because each catches the other's mistakes — and when they disagree the verdict is optional, with a note saying which one dissented, rather than a guess.
9. Auto-review
The rendered file is then inspected as evidence rather than trusted: the whole frame, the avatar band, and the gameplay pane are each sampled separately. The review records the prompt and a signature of the criteria used, so changing those settings invalidates the old pass and the app can say which stage stopped a clip and why.
10. Caption, then post
A caption and hashtags are written per clip for you to edit. Approved clips go to the TikTok account you connected — as a draft in your inbox that you finish in the app, or posted directly. The TikTok page covers exactly what that connection does.
What is waiting on whom
A long VOD with a human decision in the middle can sit for days with thirty clips waiting on a person and twenty previews never rendered, and nothing saying so. Every VOD is therefore in exactly one of four states, derived from artifacts:
| State | Meaning | The move |
|---|---|---|
| running | The machine is mid-step | Wait, or watch the detail |
| your turn | Something you can act on this second | Review |
| to run | A machine step is pending; nothing for you yet | Continue |
| done | Every candidate judged, nothing left to run | Start the next VOD |
your turn outranks to run: a VOD with ten watchable previews and twenty more still rendering is your turn, because the render is not blocking you and hiding the work you could be doing is worse than making you wait.
It also only claims your attention for clips you can actually watch. The count of clips waiting on you counts unjudged candidates whose preview file is on disk — which is what separates "30 left for you" from "10 you can actually open".