Green Room
13 August 2026
Operator mode · pre-OAuth · pre-billing
The product, corrected

Two brains.
One format, owned.

Green Room gives a creator the best possible chance that what they post actually works, and then builds it for them. It reads their Instagram to learn how they really talk and what their content really looks like, researches what is landing in their niche this week, decides what the piece should be and which format it should take, writes the script in their voice, and after they film it off the teleprompter, cuts it to the craft rules for that specific format.

It is not a reel maker with an AI script bolted on. It is a strategy brain and an editing brain wired to each other, that happen to render finished video.

Strategy rules loaded per script4 files, fresh
Craft rules behind one formatthe greenscreen playbook
Reels measured for baselines42 across 10 accounts
Cost per finished reel (measured)~$0.80 simple · ~$2 deep
Target price$34 / mo, 15 reels
Formats it makesone, on purpose
script_gen.py:BRAIN_FILES editing/baselines.json DECISIONS-FOR-BRODY §2
Read this before the rest

The green screen is the output. The brains are the product.

What earlier write-ups said

An AI reel maker. A green-screen editor. Connect Instagram, film a take, get a green-screen reel back.

All true, and all of it describes the output. It parks Green Room in the same aisle as Captions and Submagic, where it loses on funding, headcount and shipping speed, and where the thing that actually makes it good is invisible.

What it actually is

Two intelligence layers, co-equal, wired to each other, and pointed at one format. The first decides what to make, whether green screen actually serves it, and what real thing goes on screen behind every line. The second executes the hardest treatment in short-form: matte the speaker, research a real background per line, place it where their body is not covering it, and lock it to the word that names it.

It does not make talking heads, carousels or music reels. It makes one thing, and it is trying to be the best in the world at it.

social-media-brain/strategy/rules/ social-media-brain/editing/playbooks/ ×6 editing/rules/editing.md (live, overrides craft-rules)
The whole product, in one path

Their feed goes in.
A decision happens twice.

Every competitor has one intelligence layer at most, and it is an editor. The two boxes below are the reason this wins, and the format call between them is the seam nobody else has.

Input
Their Instagram
Top 8–12 posts pulled authenticated, videos downloaded and whisper-transcribed. Their spoken voice, not their captions.
Brain 01
Strategy
What to make, and why it will work in this niche this week.
  • Their voice. Tone, sentence shape, vocabulary, signature moves, learned from how they talk on camera.
  • Their look. Palette, cover style, composition, pulled from the content they already make.
  • Their niche, right now. What formats are landing and which hooks are retaining.
  • The rules and the rubric. Hooks and retention, formats and trends, copy, psychology, weighted against the objective they picked.
4 rule files, loaded fresh per run
The handoff
The fit gate
Is there a receipt? What goes behind each line Reangle if not
Green screen wants something real to show. When an idea has no receipt, the gate reangles the IDEA rather than shipping a weak reel.
Brain 02
Editing
Green screen is the hardest treatment in short-form. It stacks more independent ways to fail than any other, which is exactly why a second intelligence has to own it.
  • The matte. The speaker is cut out of their real room solo, so a bystander walking through does not survive into the reel.
  • Occlusion. Anything behind their body does not exist to the viewer, so the payload lands in the clear band and never under their shoulders.
  • Say it, show it. Every background is verified against the line it sits behind. Say "plastic crates" and get cardboard boxes and the beat is re-sourced.
  • The QA gate. Every render is scored against the craft rules and sent back once. It fails honestly rather than shipping quietly.
greenscreen.md · 127 lines of craft rules
Output
A finished piece
Between the brains: they film it off the teleprompter. Their real face, their real voice, their real words.

Honest note, and it belongs right here: the fit gate runs before a word is written, so an idea that green screen would fight comes back reangled rather than rendered badly. On a live test a personal confession scored 18 out of 100 and came back pointed at the studies and product pages the creator had been copying, keeping the vulnerable frame while the screen carried proof. The gate refuses work. That is the point of it.

fit_check.py · reangles, never switches format everything else lives in the content engine
The argument

Anyone can cut a video. Almost nobody can tell you what to make.

EDITING IS SOLVED

The cut is commodity. The call is not.

Captions, Submagic, OpusClip and CapCut all cut competently and cheaply. None of them will tell a creator that this week, in their niche, this idea should be a 15-second b-roll piece and not the 60-second talking head they were about to film. That decision is where the outcome is actually determined, and it is made before a single frame exists.

PERSONALIZATION, HONESTLY

Everyone else personalizes the pixels.

Their idea of "you" is your face cloned, your voice cloned, your logo and brand colors dropped onto a template. Green Room's idea of "you" is how you actually construct a sentence, which hooks your own audience already rewarded, and what your feed looks like. The script gen is weighted to spoken voice above written captions, because that gap was the biggest accuracy failure in the pilot.

THE SECOND BRAIN

One format, chosen because it is the hardest one.

Anyone can bolt captions onto a talking head. Locking a researched, rights-clean, fact-checked background to the exact word being spoken, while keeping it out from behind the speaker's body, is a different order of problem. That difficulty is the moat: the whole market went generative precisely because research is slow and hard to QA. Doing one thing this well beats doing six things adequately, and the market data says so.

playbooks/talking-head · POV-list · greenscreenspec §2.6 spoken voice above captions
Against the field

They are editors. This one decides.

Every product below can finish a video. The question that separates them is who decides what the video should be, and none of them answer it.

Product Where the idea comes from What "personalized" means Who picks the format Price floor
Captions / Mirage A topic prompt you typeno account ingestion Your face and voice, clonedAI Twin You do, by picking an edit style20+ named styles $24.99/mo20M+ users, $175M+ raised
BIGVU A topic prompt, plus IG connected for analytics onlybrand-kit planning "coming next" Logo, colors, fontsbrand kit You do $12–49/mo12M+ users claimed
Submagic Nowhere. You bring finished footagewrites hook titles, never scripts Caption style presets Not applicable, it finishes what exists $19–69/mo$8M ARR, bootstrapped
OpusClip Your existing long video, slicedvirality score after the fact Clip selection Not applicable $15–29/mo16M+ users
HeyGen A prompt"no cameras, no crew" A synthetic twin of youyou never film Not applicable $29–149/mo$200M ARR, Jun 2026
Green Room The strategy brain, from their own feed plus this week's niche trends How they talk and what their content looks like, learned from their posts The product does, and shows the reasoning $29–39/mo targetabove Captions' floor, deliberately

The market moved the other way on the two things this bets on. Captions rebranded around generative and argues on its own site that AI clips beat real footage. HeyGen sells never filming at all. Green Room keeps the real face, the real voice and, where the format calls for it, real researched backgrounds. That is a position, not a feature, which is why it is harder to copy than a feature.

COMPETITIVE-LANDSCAPE-2026-07-21.md · live pricing pages, fetched that day
What it actually does

Six things have to go right. Every one of them is handled.

Green screen looks simple and is not. A speaker matted out of a real room, over real researched material, timed to their words, is a stack of independent failures. Any one of them left unsolved is what makes a green-screen reel look cheap. This is the stack.

The cutout
Solo matte, feathered edge

No physical green screen needed. They film in their kitchen and get cut out of it, tracked as a person across frames so a partner walking behind them does not survive into the reel.

Real backgrounds
Four provenances, and it says which

Theirs (something they actually did and can prove) beats a receipt (a live capture of the real source) beats found (rights-clean, no faces, no baked-in text) beats made (a graphic nothing else could show). Generation is off by default and capped. The creator sees the mix and can push any beat back to their own footage.

Placement
Measured, not eyeballed

Every matte reports where the body sits and how far the gestures reach. The background's payload goes in the clear band, the cutoff sits flush with the frame, and a pointed-at thing lands on the side they point to.

Say it, show it
Locked to the word

Backgrounds are snapped to the exact word-span that names them, from real speech timings. Then each one is looked at and checked against the line. Say "plastic crates" and get cardboard boxes and the beat gets re-sourced.

Their style, not a template
Learned, then corrected

Voice is read from how they actually talk, then handed back to be corrected, because a profile built from a feed they dislike is a good guess and not a verdict. Craft can be borrowed from accounts they admire. Voice never can. Sounding like someone else is the fastest way to read as a knockoff.

The quality gate
Real, and it does fail things

Every render is scored against the craft rules and sent back once. It genuinely rejects work, including ours. That is the feature, but it also means output quality is still the thing to watch most closely.

Watch this one
playbooks/greenscreen.md · the craft rules references.py · style without voice review_render.py · gate at 8
Why anyone pays for this

Followers are not the product. Customers are.

A professional's income comes from customers. Customers come from people who trust them. Trust comes from showing up, in their own voice, with something worth hearing. That makes content the top of the revenue funnel rather than a line item at the bottom of it, and it is why the people who need it most are exactly the people with no time to make it. Running the business is already the job. Content has always been a second one.

And it is not one wall. It is three, and they fail independently.

Wall one

Knowing what to film

They are an expert at the work, not at what is landing in their lane this week. The blank page is where most of it dies, and no amount of editing software touches it.

Wall two

Actually standing up and filming it

Having the idea and shooting it are two different acts. The second is the one that gets pushed to next week, every week, because there is no script in front of them and no certainty it will be worth the hour.

Why the format helps hereGreen screen is the one treatment where the room does not matter. No tidy background, no good light, no nice desk, no dialled camera. Wake up and shoot it against a wall, because everything behind them is going to be replaced anyway.
Wall three

Everything after the take

The cut, the backgrounds, the captions, the caption copy, the posting. Each one is its own skill, and any one of them being weak makes the other two wasted.

Almost every tool on the market clears exactly one of these. This one is built to clear all three in a single pass, which is the only version of the promise that changes anyone's week.

It also has to serve two different jobs, and confusing them is how creators burn out. Content that travels buys attention and new followers. Content that builds trust turns those followers into customers. They are aimed at different targets and they should never be graded against the same number, because a trust piece measured on views reads as a failure and gets abandoned right when it was doing its actual work. The objective is chosen per reel, and it decides both what gets written and what counts as working.

What the category ships

A synthetic presenter. A cloned voice. A face that never sat down in front of a camera.

It scales, and everyone can tell. The uncanny read arrives before the message does, and for a professional whose entire product is being trusted, that is worse than posting nothing.

What this ships

Them. Their real face, their real voice, in the room they are actually in, saying words written in the register they already use because those words were learned from hours of their own speech.

The only thing that scales is the part nobody sees: the deciding, the writing, the research, the cut. The person on screen is never synthetic, because the person on screen is the entire reason it converts.

Who it is for

Any creator with something real to point at.

Not a niche tool. The one thing it needs is that the creator has a receipt: a number, a headline, a page, a screenshot, a claim someone actually made. Three worked examples of the same engine finding a different receipt.

Dog trainer · 41K · sells a $49 course

"The board-and-train injury numbers just came out"

Reacting to something real and current in her lane. She is on camera being angry about it while the actual report sits behind her.

What goes behind herThe report page, the figure highlighted, the facility listings, a forum thread of owners saying the same thing.
Ceramics studio owner · 12K · sells classes

"Every beginner cracks their first pot the same way"

Teaching over evidence. She has hundreds of clips of the work already, so her own footage carries most of the beats and nothing has to be sourced.

What goes behind herHer own library first: the cracked pot, the wall thickness, the kiln. Her material outranks anything external.
Bookkeeper for freelancers · 8K · sells a service

"The IRS just changed the 1099-K threshold again"

The purest case for the format. The credibility is entirely in visibly showing the actual source, pulled live the day it renders.

What goes behind herThe IRS page itself, scrolled to the paragraph, then the old threshold and the new one side by side.

And when a creator brings an idea with no receipt in it, the gate says so and hands back a reangle rather than making a weak reel. Refusing the wrong work is how the one format stays good.

Where it is today

Real, running, and deliberately unfinished.

Version one runs in operator mode. An installable web app on iOS handles camera, teleprompter, upload and share. Everything heavy runs on one Mac behind a tunnel, gated by a single studio pass.

Built and running

  • The account learnerPulls their posts authenticated, downloads the video past the login wall, whisper-transcribes it, and synthesizes a spoken-voice profile and a visual identity. It reads their best posts by engagement from a recent window, and it says so on screen: that ranking is a proxy, because Instagram does not make saves or shares public.
  • The strategy-brain script writerLoads four rule files fresh per run and writes in their spoken voice against the objective they picked, with a verifiable receipt named for every beat.
  • The fit gate, style direction and their own libraryIdeas without a receipt get reangled. Craft can be learned from accounts they admire without borrowing a word of voice. Their own media persists between reels and carries binding instructions for when it may be used.
  • The full assemble pipelineCaptions on real word timings, take trim, matte, background plan, research fetch, vision QA on every background, one grade family, render, QA gate.
  • A frozen API and a working clientConnect, analyze, script, film, regenerate with feedback, plus a reel shelf. The PWA path held on iOS with no TestFlight needed.

Open, on purpose or otherwise

  • Output quality is the thing to watchThe pipeline runs end to end and the gate is real, but green screen is the hardest treatment in short-form and the gate does still fail renders. Every point of that score is craft work, and it is the work that matters most.
  • Correcting what it learnedThe voice and look profiles are shown but not yet editable in the app, and there is no deep pass that reads the full back catalogue, long-form speech or an uploaded brand doc. Designed, not built.
  • Bringing your own idea or footageThe flow assumes the product supplies the idea. Creators often arrive with one, or with footage, or both. That fork is designed and not built.
  • Other formats are deliberately not hereTalking head, b-roll and text, POV lists, carousels and music reels are out of scope by decision, not by accident. They belong to a separate creator workspace over the same engine. This product is one format.
  • Per-niche trend researchThe tool exists but needs per-creator channel config before it can feed the strategy score. Named as an open seam in the spec.
  • Instagram OAuth, billing, multi-userAll deliberate. OAuth needs a registered Meta app and review. Billing is roughly a day of work behind the existing gate. Cloud GPU triggers with the first paying cohort.
README.md · API.md v1.1 DECISIONS-FOR-BRODY.md §1–4 spec §7 open seams
The honest risks

Four things that could make this not work.

A one-pager that only sells is useless to whoever has to decide. These are the real ones, with what is already being done about each.

Risk 01 · the window

Captions ships Instagram ingestion.

They have 20M+ users, $175M+ raised, a foundation-model team and the entire pipeline live. Reading your feed to learn your voice is an add-on sprint for them, not a rebuild. The half they will not copy is researched real backgrounds, because their money and their rebrand are bet on generative.

What is being doneWhich is why the durable moat has to be the fusion, and why the loops need roughly two quarters of compounding before a feature sprint could blunt the positioning. The clock is running now.
Risk 02 · the cold start

There is no warm audience, by design.

Zero users come from any existing orbit. On the 32-point scorecard this product scored 25, and the single zero was validated demand, because nobody has asked for it. Every other candidate borrows an audience. This one deliberately does not, for reputation shielding. There is no floor: the plan can produce zero, quietly, for months.

What is being doneThe early-warning system is a free voice-analysis tool that has to pull 200 completed profiles from strangers before anything downstream is trusted. It is scheduled first for exactly this reason.
Risk 03 · the gap between doc and code

One format is a bet, and it cuts both ways.

Doing one thing means every creator who wants a talking head or a carousel is not a customer, and it means a single format going out of fashion is an existential problem rather than a bad quarter. The upside is the same fact: focus is what makes the output good enough to be worth paying for, and the market data backs the focused product over the bundle.

What is being doneThe bet is deliberate and the evidence supports it: in this exact category the everything-platform sits in the sub-$500k floor band while the focused product cleared 3,200 paying subscribers at $49 a month in ten months, bootstrapped. Everything not green screen has a home in a separate workspace over the same engine, so the scope decision is not the same as throwing the work away.
Risk 04 · the expensive middle

The two moats are also the two slowest steps.

Corpus learning and real-background research are what nobody else does, and they are the most expensive and least self-serve parts of the pipeline. Everything heavy runs on one Mac that has to be awake, one job at a time. Every competitor is self-serve at scale with credit metering.

What is being doneUnit economics do work at scale: roughly $0.25–0.70 per finished reel on cloud GPU plus LLM, against $29–39 a month. The engineering item is porting the local model calls to API calls.