clipfarmer

How it works

From an eight-hour VOD to a clip worth posting

clipfarmer is a pipeline with a person in the middle of it. Each stage writes its results to disk, so progress survives a crash, a restart, or a change of mind about the ranking.

The pipeline graph: Ingest and Chat replay feed Transcribe and Frame cache, which feed Candidates and Scene profile, then Previews, then Your review, then Render and Auto-review. Each node shows counts of queued, running, done and idle jobs.
The pipeline view, live. Every count is derived from files on disk rather than a status column, which is why an interrupted job can resume at the step it actually reached.

1. Ingest

You give clipfarmer a Twitch VOD or YouTube URL. It downloads the video, extracts an audio track, and caches keyframes across the whole VOD for later use.

Ingest resolution is a real decision rather than a default, because it sets a ceiling on what the renderer can do later. A punch-in crops the streamer's camera box out of the source frame and scales it to fill a 1080×1920 clip. From a 1080p ingest that box is around 401×713 — a 2.7× upscale, which holds. The same box from a 720p ingest is 268px wide and needs 4.0×, which visibly does not.

2. Chat replay

When the source is a Twitch VOD with chat, the chat log is fetched alongside the video. Message-rate spikes and emote density are a second, independent opinion about where something happened — one that does not depend on the transcript being right.

3. Transcribe

Whisper produces a transcript with word-level timestamps. The timing is what lets subtitles land on the syllable rather than drifting behind the speech, and it is what the candidate scorer reads to find hook phrases and laughter.

Transcription runs against each clip's own audio rather than one pass over the whole VOD. A single long pass silently drops speech in places, and those holes are invisible to any check that reads the transcript — because the transcript is the thing that is missing.

4. Frame cache and scene profile

A clip is rendered as one of three layouts: split (gameplay stacked over an avatar band), full (a single crop of a composed scene), or avatar-only (centred on the streamer, with no invented gameplay pane). Getting this wrong is the most visible defect the renderer can produce — a band cut out of a composed scene shows scenery where a face should be.

Asking a vision model "is this an overlay or a set?" per clip is not reliable, because an idle scene with a game-like backdrop is indistinguishable from real gameplay in any single frame. The separating fact is not in the frame, it is in the VOD: sampling the whole VOD and clustering where the avatar sits collapses it into two or three tight positions. What identifies the webcam slot is edge proximity, and only that — camera clusters sit at edge margin 0.000 while composed scenes sit at 0.18–0.26. Size does not separate them at all.

Every clip's layout is then a lookup against that per-VOD profile, and a clip that crosses a scene change carries one layout per segment.

5. Candidates

Short sliding windows are scored on a weighted blend of signals, all of them cheap and all of them measurable:

SignalWhat it catches
Audio energyRMS loudness spikes — reactions, hype, a room going up
Hook phrasesA configurable list: "no way", "let's go", "insane"
Laughter markersLaughter in the transcript, and emote laughs in chat
Punctuation densityExclamation and question marks per word
Speech paceDeviation from the VOD's own average words per second

6. LLM judge, and measuring it

A local language model re-ranks the candidates. Whether that helps is not taken on faith: the clips real viewers made from the same VOD are clustered into labelled moments and used as ground truth, and recall@K is reported for the signal ranking, the model ranking, and a blend of the two — all against the same labels.

That harness is why the review step exists in the shape it does. Six LLM ranking variants lost to a z-score here. Generation is a model's strength; choosing is taste.

7. Your review

Previews are quick, subtitle-free trims rendered so there is something watchable behind every verdict. You approve or reject; nothing auto-selects.

The review gate sits before rendering, so a clip awaiting a verdict is unrendered by definition. Your verdicts and the pipeline's checks are tracked as separate things, because a clip the pipeline passed is not a clip you approved.

Two groups of filter chips: 'you said' — not judged yet, approved, needs re-render, rejected, posted — and 'the pipeline said' — passed checks, cut by pipeline, rendered, not run.
Filters for both sets of verdicts, kept apart.

8. Render

Approved clips are rendered as finished vertical video: subtitles with a row per speaker, the layout the scene profile calls for, a punch-in when the gameplay stops mattering, and a burned-in header on the clips that need one.

Which clips need a header is decided by a silent scroll test: the window's cached frames, no transcript, one binary question — with the sound off, do the pictures alone show a stranger something happening, or only a person present in a place? Withholding the words is the whole design. Include the transcript and the answer tracks the transcript instead of the frames.

The pixels get a vote too, as median luma change across the window's cached frames. Neither half decides alone, because each catches the other's mistakes — and when they disagree the verdict is optional, with a note saying which one dissented, rather than a guess.

9. Auto-review

The rendered file is then inspected as evidence rather than trusted: the whole frame, the avatar band, and the gameplay pane are each sampled separately. The review records the prompt and a signature of the criteria used, so changing those settings invalidates the old pass and the app can say which stage stopped a clip and why.

10. Caption, then post

A caption and hashtags are written per clip for you to edit. Approved clips go to the TikTok account you connected — as a draft in your inbox that you finish in the app, or posted directly. The TikTok page covers exactly what that connection does.

What is waiting on whom

A long VOD with a human decision in the middle can sit for days with thirty clips waiting on a person and twenty previews never rendered, and nothing saying so. Every VOD is therefore in exactly one of four states, derived from artifacts:

StateMeaningThe move
runningThe machine is mid-stepWait, or watch the detail
your turnSomething you can act on this secondReview
to runA machine step is pending; nothing for you yetContinue
doneEvery candidate judged, nothing left to runStart the next VOD

your turn outranks to run: a VOD with ten watchable previews and twenty more still rendering is your turn, because the render is not blocking you and hiding the work you could be doing is worse than making you wait.

It also only claims your attention for clips you can actually watch. The count of clips waiting on you counts unjudged candidates whose preview file is on disk — which is what separates "30 left for you" from "10 you can actually open".

The feature detail Install and run it