Twelve kinds of footage.
Each one edited its own way.
Choosing a type changes how your footage is read. A Q&A is not cut like a travel vlog, and a screen recording is not covered in cards the way a talking head is. The type only nudges the decision — the frames decide.
Eleven project types are listed when you start a project, plus two kinds the Studio recognises on its own: a montage with no speech, and a sound file with no picture. Every status below is the status in the documentation, not a sales position.
Talking head (one person to camera)
One person speaking to the lens. This is the Studio's original purpose, and the path everything else is measured against.
Cuts retakes, abandoned attempts and pauses, and keeps the last good take. Numbers become counting cards, lists become checklists or pills, comparisons become before/after cards. Punch-ins on long face stretches, word-highlight captions, a hook on frame one of a reel and an end card with your link.
One continuous recording is fine — retakes are cut for you. A script makes names exact. Pictures and b-roll named by what they show are placed on the words that mention them.
The episode, or a vertical reel, plus shorts. Audio runs through an adaptive voice chain with a 10 ms crossfade at every cut, at −14 LUFS.
Guard: ten regression clips, “0 of 10 changed” on every release — studio/tests/cut_diff.py, plan_diff.py · docs/studio/genres/talking-head.md
Podcast / interview / two people talking
Two voices, either in one wide shot or on two call panels. The Studio recognises a conversation from the picture and from the words.
Only stutter restarts are cut — in a conversation the repeats stay, because people repeat themselves on purpose. Run-on sentences over 12 seconds are split at their natural breaks. Number cards only for 100 or more, money or percentages, and never the same number twice. A sad moment stays on the face. Captions are placed where no face is.
The call recording, or both files. A script, if you have one, for the names of people and cities — misheard names stay unless the model is sure.
The 29.8-minute call this site quotes became a 29.0-minute episode with 13 cards — about 2 % of the time — every blocking gate passing, and three shorts with their own titles.
Proven on one real 30-minute call (0.5.1). Not proven yet: a studio podcast filmed as one wide shot with two people, three or more speakers, and multi-camera podcasts. There is no speaker detection, so the small face window follows one face, not the person speaking. Source: docs/studio/genres/podcast-interview.md
Reels and shorts
Vertical output, built to be watched without sound and finished in the first second.
Reel format is 9:16 with a hook on frame one and the end card carrying your link. A number card in the opening sentence pops in whole with its label, from the first frame, instead of counting up from zero.
A reel-length recording, or a long one plus “make 5 shorts”. Shorts are asked for at upload, from the finished video, or in the chat — one to ten at a time.
A short is picked for standing on its own, 20–58 seconds, opening on a question or a strong line and ending when the thought ends. Each one is a full vertical reel with captions, a hook and the end card.
The reel format and the shorts rules are measured on the long-video engine; as a project type there is no separate style — the Studio still decides from the frames whether your footage is a talking head or a montage. Source: docs/studio/product/HOW-TO-USE.md:74-78 · shorts
Screen recording / software walkthrough
Software demos, slides and documents, with or without a small face in the corner.
No zoom-ins, because a zoom crops the edges where menus and text live. Fewer cards, because a card covers the thing the viewer is trying to follow. A 9:16 output keeps the whole screen on a blurred fill instead of cropping through the interface; 16:9 stays as recorded.
1080p or more, large fonts and a zoomed browser window — a vertical output shows a landscape screen smaller. A script makes product names exact.
The narration transcribed, retakes and pauses cut, captions on the words and the end card with your link. A screen recording with no narration has nobody speaking, so it is edited as a montage instead.
Built and tested on a narrated test recording (edge case E21). Not proven yet on a real tutorial with your own narration. Detection: at least half the frames look like a screen; in a “tutorial” project, 30 % is enough. Source: docs/studio/genres/screen-recording.md
Travel vlog
Footage with four or more scene changes a minute, or a person who walks in and out of the shot.
No zoom-ins, because the camera already moves and a zoom on moving footage looks shaky. Fewer cards, because the footage is the visual. A 9:16 cut of landscape footage keeps the whole picture on a blurred fill, so scenery is not cropped off and nothing jumps when you leave the frame.
Clips in order, or one assembled recording. Clips without speech are placed by their file names on the words that mention them — “temple.mp4” lands on “temple”. Music goes in as a music file and is ducked under speech.
Where nobody speaks for long stretches the video keeps its own pictures and sound; silent or mixed clips get the Studio's own music bed so the sound does not come and go between cuts.
Built, not proven: no real vlog has gone through the Studio yet. The 0.6.3 loop used Wikimedia travel clips and a silent ten-minute train take. Source: docs/studio/genres/travel-vlog.md
Educational / explainer
One person explaining, slides, or a voice over footage — anything built to be understood rather than only watched.
Enumerations — “first, second, third” — become numbered steps, lists become checklists, numbers count up, comparisons become before/after cards. Each step gets its card as a three-second beat and then the face comes back. A step's title arrives with its number, not seconds later.
One project per topic. The style — captions, colours, pace, zoom — is learned in the account, so two channels taught the same way stay the same even when the topics differ.
Talking-head handling (F3), screen handling (F4) or voice-over handling (F6), whichever the frames match. A finished animation is recognised as animation on its own: no cards, no zooms, captions only.
Uses the proven talking-head rules or the screen and voice-over handling, depending on what is filmed. There is no separate “educational” style yet. Source: docs/studio/genres/educational.md
FAQ / questions and answers
Questions read and answered by one person, or an interviewer asking and a guest answering.
A question of ten words or fewer is shown whole on a card, exactly as it was asked, and the answer stays on the face with captions. With an interviewer present, the conversation rules apply instead.
The whole session, unedited. Then ask for shorts — “make 8 shorts” — and each short opens on its question and ends when the answer ends.
Often this is the best use of the type: one short per answer. In an interview, a question and its answer score as one story, which is why the shorts open on the question.
Uses the proven talking-head rules; a real FAQ recording has not been through the live Studio yet, so this is not proven as its own style. Source: docs/studio/genres/faq-qa.md
News / commentary / tech news
One person to camera, usually with screenshots and figures — the kind of video where a wrong number on screen is worse than no number.
Years are dates, never counting cards. Money and percentages become counting cards, and comparisons like “from 2018 to 2020” become before/after cards. Company and product names are corrected only when the recogniser was unsure and the fix sounds like what was said.
Screenshots of the articles you are talking about, as pictures, named by what they show. A correction is checked against the words that were spoken, so a script fixes spellings, never claims.
Screenshots land on the words that mention them and are never cropped: a picture that is not the shape of the frame sits whole on a card. Every word on screen is a word you said.
Uses the proven talking-head rules (F3); not proven as its own style — a real commentary video has not been through the live Studio yet. Source: docs/studio/genres/news-commentary.md
Product review
A review recording plus your own pictures and short clips of the thing being reviewed.
Each picture and clip is placed on the words that mention it, cuts in on its own word, and then the face comes back. Prices and percentages become counting cards; pros and cons said as a list become a checklist.
Name the product files by what they show — “battery.jpg”, “unboxing.mp4”, “price list.png”. That name is how the Studio knows which sentence the file belongs on.
The review with the product arriving when the story earns it, not at the top. A rating card is drawn only when you say a rating out loud.
Uses the proven talking-head rules (F3) and the proven picture and b-roll placement (0.4, the AR rules). Not proven as its own style. Source: docs/studio/genres/product-review.md
Live session / webinar
Slides with a speaker, or a panel — anything from thirty minutes to three hours.
Measures whether this is slides with a speaker or a panel, and applies the screen or conversation rules. Long recordings go through the fast chunked renderer, which restarts itself with fresh memory so a long render cannot fill the machine.
The whole session. Tick “Also make 5 shorts” at upload, press “Make shorts” when it is done, or write “make 8 shorts” in the chat.
The edited session, and the shorts — which are usually the point. Measured: 30 minutes of footage is about 80 minutes of work and 2.4 GB on disk, of which 1.6 GB is re-creatable work files cleared 14 days after the job finishes. A two-hour webinar is about 5–6 hours.
The long-video engine is proven — a 30-minute recording edited end to end, up to 3 hours allowed, jobs continuing after restarts. The webinar look itself is not proven. Source: docs/studio/genres/live-webinar.md
Montage (no speech)
A shoot with nothing spoken: b-roll, product takes, drone clips, stills and a track. Ad films and “no dialogue” edits live here.
No transcription stage is queued at all. Each clip gives its best moments — six frames a second scored for sharpness and steady motion, with the first 0.6 s and last 0.4 s of a take left out. A clip over 12 seconds gives a second shot, over 24 seconds a third. Anything filmed at 90 fps or more plays at half speed.
Your clips and pictures, and a track you have the rights to. With an example video, the shot length, the hard or soft cuts, the total length, the colour and the loudness all follow that example.
Shots go round-robin through the clips for variety. When the talk is replaced or you ask for a story, one frame of every clip is labelled people, place, detail or product, and the shots follow that order — people first, the product at the end.
The 0.5 montage — a single clip, pictures with music, a clip without speech — is proven (edge cases E04, E07, E08). The 0.6.1 additions are built and tested on test footage and are not proven on a real ad shoot yet; the first one, a nine-clip shoot, is the next real test. Source: docs/studio/genres/montage.md
Voice only
A recording with no video at all: an audio podcast, a voice note, a quote you want as a reel.
The voice is transcribed and cut exactly like a talking head. The picture is a designed, slowly moving canvas, and the words are the picture: big centered captions, four words at a time, plus the planner's number, list and comparison cards.
The audio file — mp3, wav, m4a, aac, ogg, flac or opus — and optionally pictures and clips, which are placed on the words that mention them.
A 9:16 reel unless you choose 16:9. Typical uses: a podcast that exists only as audio, a voice note turned into a reel, an audiogram for one quote.
Proven since 0.5 — edge cases E05 (voice with pictures) and E06 (voice alone) pass on the live server on every release. Source: docs/studio/genres/voice-only.md
Something else
The type you get when nothing else fits — and the one every account starts with, as “My videos”.
Nothing is assumed. The footage is measured — faces, flat screen-like frames, scene changes, shape — and the first rule that matches decides how it is read.
Whatever you have. If the frames are unclear the project type settles it: a “tutorial” project counts a 30 % screen share as a screen recording, and a “travel” project counts unclear footage as a vlog.
The same eight stages, with only the stages that apply. An unknown project type becomes “something else” quietly rather than failing the upload.
Project types, including this one, are listed in studio/accounts.py:24-36; the kind and the reason it was chosen appear in column 3 of the job page and in state.json → footage.
Not sure which one you are?
Pick “Something else” and upload. The Studio measures the frames and tells you what it decided in column 3 — and if the read is wrong, say so in the chat and it is a rule we fix, not a setting you have to find.
Already have an account? Sign in.