How Helen Gets Made A making-of, not the live show Watch her live

Workshop · the AI models

What the AI models did on this build, and what they did not

Shown the water fountain with no context, a vision model called it a dark sink. A caption model repeated a wrong story about where Helen came from, in voice, for weeks. One vendor's spend cap kept the channel from posting for a day. A model was in the loop for almost every part of this build (ChatGPT, Cursor, Grok, Claude, Gemini, a few open models, and a 3D generator that turned one photograph into a mesh), and those three days are why none of them decides what gets published.

The cat mesh has its own page; so does which tool did which job.

Helen, a dilute tortoiseshell calico, at the water bowl; a vision model reads a frame like this before every Short is captioned
The frame a vision model reads before every Short. One factual sentence comes out; a second model turns it into the caption. A person wrote the rules both of them work under.
Pick your page — what AI did here
The questionThe page that answers only that
What did the models do, and not do?This page: two calls per Short, the gates a person keeps, the day a vendor took the channel down.
How was the cat made from one photograph?A 3D cat from one photo: two Tripo tasks, 61 MB raw, 2.16 MB shipped.
Which tool did which job?Which AI tool for which job: Cursor, Claude, ChatGPT, Codex, Grok, Gemini, DeepSeek, Tripo.
What is image-to-3D, in plain words?3D and image generation, plainly: what Tripo is fed, what comes back, what it got wrong.
What were the exact Tripo tasks?Tripo multiview: all eight tasks, settings, credits and wall times.
How does a 61 MB mesh become 2 MB?Shippable mesh: the glTF-Transform pass in order, meshopt, the tongue repaint.
Which pictures on the site did a model draw?Generated pictures: what stayed, what went, and who reviewed it.
What did generation cost, and save?Cost: 340 credits, eight tasks, sixteen minutes, three discards.
Where does the generated cat live?The three.js court and the 3D website.

Did AI make it? Parts of it, under supervision, with a lot of rework. Which one? That depended on the job and it changed over the summer, and the decision that mattered most was never which model to use but how to stop any single model from being able to switch the channel off. Which tool did which job and how the cat mesh was made have their own pages.

Two calls per Short. That is the entire creative contribution

The Shorts pipeline makes exactly two model calls, and it makes them when a publishing door opens: 5 am water, 7 am food, 5 pm court, 7 pm face, Eastern. A vision model looks at frames from a Frigate clip and writes one plain sentence about what the cat is doing. A cheaper text model takes that sentence and writes the one-line decree that gets burned onto the video in gold. The cutting, cropping, timing, music bed and upload are ordinary code, described on the pipeline page.

Right now the frame is read by Gemini 2.5 Flash and the line is written by Gemini 2.5 Flash-Lite, both through OpenRouter, in about a second each. When the first vendor is busy, open models on Together AI (Gemma 3n for the frame, Llama 3.3 70B for the line) answer instead, and nobody watching the channel can tell. The daily cost is cents.

What no model is allowed near

  • The decision to publish. The gate is a station allowlist, a person-in-frame purge and a shared-clock check. None of it asks a model, and if the gate cannot decide, nothing goes out. A model's judgement is never on the path between a camera in a house and a public upload.
  • The cameras. Detection is a small OpenVINO model on an Intel iGPU. It says "cat, 0.83" and nothing else.
  • The live streams, the treat, the payments. ffmpeg, Stripe, Home Assistant and a relay. No language model anywhere in that chain.
  • Winter Park history. The narrator prompt carries a hard rule: never invent it. The tribute to Helen Morse and the Woman's Club stays a tribute because a model was told it may not embroider, and because a person checks.
  • Numbers. Captions are generated with digits forbidden. A statistic gets verified or written around. Models invent numbers with total confidence, and a cat channel does not need an invented number to be funny.

The caption, one step at a time

The voice was written down once, from the brand guide, and every caption call starts from it.

The narrator system prompt, as it runs
You are the unseen narrator of Helen the Cat Live, a cat-cam channel that will run for
twenty years. You are Helen's husband, quietly telling the daily story of your cat.
Helen is Her Majesty Helen: a stray young mother rescued at the Woman's Club of Winter
Park and named for Helen Morse, its founding president - who is, in your telling, an
ordinary house cat and also a reigning queen. The eternal joke is the tender gap between
royal grandeur and ordinary cat life.

LEXICON (match helenthecatlive.com, do not invent alternatives): her rooms are the
court, and it is always in session; the four cameras are the Water Cam, the Food Cam,
the Treat Cam and the Court Cam - an overhead of the whole room; the shop is the Royal
Collection. She is a tortoiseshell calico - never call her a tabby.

VOICE: charisma, wisdom, loving, gently emotional. Warm, literate, a little witty, never
maudlin, never purple. FORM: keep it short, one thought per line, at most a single local
touch (a word or place-name). Return only the caption line(s), nothing else - no
preamble, labels, quotes, emoji, or hashtags.

HARD GUARDRAILS: family-friendly always; no politics/religion/real public figures; never
reveal or hint who the narrator is; original lines only, never quote a source; never
invent Winter Park history; no fake people or comments.

The lexicon paragraph is there because a channel that calls its cameras three different things in a week reads as three different writers. The history rule is there because the cat is named after a real person who founded a real club in 1915, and a model that would happily invent a plaque or a date has to be told it may not.

  1. Tell the model which camera it is. Shown the water bowl with no context, the vision model described "a dark sink or basin." Shown the food station, "a garage." The cameras are fixed and point at known things; withholding that did not make the model more objective, it made it guess, and it guessed with confidence. A per-camera table now tells it what it is looking at and what the cat probably came for.
    The vision prompt as rendered for the Water Cam
    This is a fixed camera on a pet drinking fountain with running water, on a table. It is
    called the Water Cam, and the cat is most likely drinking, or about to drink. Use that -
    do not guess a different setting, and do not describe it as a sink, a bathroom or a
    garage.
    
    Describe, in one plain factual sentence, what the cat in this frame is doing right now
    -- just the observable action and setting.
  2. Look at three frames, not one. Frigate's snapshot is the single best-scoring frame of an event. A cat walking past the food bowl scores highest at the bowl, so a walk-past used to caption as dinner. The pipeline now samples three frames spread across the clip and asks for the whole arc. If that fails for any reason it falls back to the single frame, so it can never cost a caption.
  3. Describe first, write second. The observer never writes and the writer never sees the image. The factual sentence can be checked on its own, and the writer cannot invent from pixels because it was never shown any.
    The caption prompt
    WHAT HELEN JUST DID:
        the cat laps from her water bowl
    
    Write ONE very short thought or observation about what Helen just did -- a brief line,
    ideally 5-9 words and NEVER more than 10, warm and a little wry. You MAY fold in this
    single local touch or omit it: 'her water bowl'. No history. Return only the line, no
    period needed.
    Five to nine words, because the line is hand-lettered in gold at the bottom of a vertical video and has to draw on and finish before the clip ends. "No history," because the local touch is a place name, not a lecture.
  4. Filter the line. Do not trust it. Every generated line passes a plain regular expression (war, death, crime, elections, named public figures, disease, disasters, tragedy, abuse) and a separate digits check, because "no invented numbers" is easier to enforce as "no numbers." A line that fails is logged and thrown away, and the next model in the chain is tried. Nothing is published "anyway." A wrapper strips the code fences, "Caption:" labels and quotation marks that small models like to add, and never changes the words.
    helen_models.py, the check every line passes
    def check_content(text, allow_digits=True):
        t = text or ""
        if not t.strip():
            return "empty"
        if BANNED.search(t):              # war, death, crime, elections, disease, disasters ...
            return "content filter (political/tragic/contested)"
        if not allow_digits and re.search(r"\d", t):
            return "contains a number (no invented statistics)"
        return None
  5. If every model is down, use a human line. The last link in the chain is a bank of sixty-odd hand-written lines in the voice, a dozen per station plus a few general ones, rotated so an outage does not repeat itself. "Her Majesty takes the waters, as is her custom." The upload still happens. A viewer cannot tell which lines were generated and which came from the bank, and that is the standard the generated ones are held to.

The day one vendor took the channel down

Until 2026-08-23 the caption function called one vendor directly, with a model name written into the code. That morning the account hit its spend cap and started returning HTTP 400. The frame description had a fallback; the caption call had none, so every visit died with an opaque error and the live uploader published nothing all day. The backfill kept working only because its lines come from a hand-written hooks file.

The fix is a few hundred lines that any pipeline could copy. Nothing in the code names a model; it asks for a role. A small resolver turns the role into an ordered chain that crosses more than one vendor. A lane that refuses is remembered for the rest of the run and logged once with the vendor's own message, not retried twenty-five times. If every lane is down the caption degrades to a bank line and the Short still uploads.

The chains, trimmed
# helen_models.py -- nothing in the pipeline names a model. It asks for a role.
DEFAULT_ROLES = {
    "vision": [                       # must be able to see the frame
        "openrouter:google/gemini-2.5-flash",
        "openrouter:google/gemini-2.5-flash-lite",
        "together:google/gemma-3n-E4B-it",
        "gemini:gemini-2.5-flash",
        "anthropic:claude-haiku-4-5-20251001",   # the original; now last
    ],
    "caption": [                      # short creative text, cheap tier first
        "openrouter:google/gemini-2.5-flash-lite",
        "openrouter:google/gemini-2.5-flash",
        "together:meta-llama/Llama-3.3-70B-Instruct-Turbo",
        "gemini:gemini-2.5-flash-lite",
        "anthropic:claude-haiku-4-5-20251001",
    ],
}

# helen_pipeline.py
VISION_MODEL = "role:vision"      # a bare model name here raises StackLockError

Two defaults sit side by side in that design and they point in opposite directions. Captions fail soft: a missing line is not a reason to skip a Short. Publishing fails closed: a gate that cannot decide does not publish. Getting those the wrong way round is how a cat channel ends up either silent or unsafe.

The day the prompt itself was wrong

Until 2026-08-17 the narrator prompt described Helen as a shelter cat. She was never in a shelter; she was a stray young mother rescued at the Woman's Club of Winter Park with two kittens. The model had been reproducing the error, in voice, for weeks, and doing it well. The brand guide now carries the correct sentence in bold with a note about the mistake, and the prompt matches the site. A model will write your errors beautifully. The thing that catches them is a person reading the output against the facts.

Where the models saved weeks

Code. The detector server, the ffmpeg restream scripts, the Stripe webhook verifier, the cam proxy, the three.js court, the Shorts pipeline: every one was drafted with an assistant and then argued with. Cursor did most of the bulk implementation, and the model it runs today is Grok. Claude did the planning, the measuring and the careful patches to files a wrong line would take down. ChatGPT and Codex did front-end and visual work. Grok, Gemini 2.5 Pro and DeepSeek were used as reviewers on designs the operator had already committed to, because a reader with no stake in the code finds what the author cannot. That is the operator's account of who did what; the tool page says which parts a file can prove.

The cat. One front photograph became a walk-around mesh with fur in two API tasks and seventy credits. By hand that is a week for someone who knows Blender, which the operator does not. The mesh then needed compression, an orientation fix, a scale fix and one small hand repair before it was fit for the court; that page has the sizes and every fix.

Captions at volume. Four Shorts a day, every day, each with a line that sounds like the same person, for the price of a stamp. Nobody could sustain that by hand and stay funny.

The gate a person keeps

Anything longer than a caption is a draft until someone reads it. Drafts are unpublished by default, so forgetting the review step produces nothing rather than something wrong; a person reads, flips the status, commits, deploys. A name on a piece of writing means a person read it; if nobody did, it does not get a name. Certain days have the decree voice switched off entirely (Rainbow Bridge day, Pet Memorial Day, Veterans Day) by an empty entry in a hooks file that someone wrote in advance. The model does not get a vote on which days those are.

The pages you are reading were made the same way: a person and an assistant working from the actual config files, the actual scripts and live checks, with a review before anything shipped. Where a claim could not be verified, the page says so or leaves it out.

The models made this possible for one person with a day job. They also invented a sink, called a walk-past a meal, and took the channel dark for a day over a billing limit. The generated cat, which tool did which job, or the pipeline the two calls sit inside.

The 3D and image side now has its own section: what image-to-3D does and gets wrong, every Tripo task and credit, packing 61 MB into 2 MB, which pictures a model drew and what it cost.

The website that renders and sells those files has its own section too, with the models' part named at each step: the 3D website, and inside it the page shell and the checks a model wrote and a reviewer broke.