A year of course updates with the same presenter, and what kept breaking
I teach a technical course that I sell directly, and about fourteen months ago I stopped filming myself and switched to a generated presenter.
The reason was not that filming is difficult. It is that my course changes constantly, and every change meant setting up again, matching the lighting to material shot in a different season, and re recording a passage where my voice sounded noticeably unlike the passage before it.
The first video was easy. The tenth was where the actual problem lives, and nobody writes about the tenth.
Course material changes more than anyone budgets for
I looked back at my revision log before writing this. Over fourteen months I made substantive changes to the course on twenty three occasions.
Most of those were small: a tool changed its interface, a recommendation became outdated, a section needed reordering because students kept asking the same question at the same point. Four were large enough to require a new module.
With filmed video, a small change is not small. Re recording ninety seconds requires the room, the light, the same shirt, the same microphone position, and a voice that is having the same kind of day. The realistic response is to batch small changes and do them annually, which is what I did in the first year and which meant that for most of that year parts of the course were wrong.
The whole reason for switching was to make a ninety second change cost ninety seconds.
The three things that break consistency
Once you are past the tenth video, consistency is the entire problem, and it fails in three places.
The presenter drifts. Not the face, which stays the same, but everything around it. I built videos four and five with a different background than videos one to three because I was experimenting, and the result was a course that visibly changed appearance halfway through. Students notice this and read it as a quality signal about the whole product.
The pacing drifts. I write scripts at different times of day and in different moods, and the generated narration follows the script. Segments written late in the evening came out longer winded than segments written in the morning, and across a course that reads as inconsistency in the teaching rather than in the writing.
The vocabulary drifts. Six months in I was using a term I had not used at the start, without ever having introduced it. This is a documentation problem that predates any video tooling and video makes it harder to spot, because you cannot search a video for the first use of a word.
What I standardised
I made these decisions once, wrote them down, and have not revisited them since.
One presenter, one background, applied globally. The AI talking head generator I use lets the avatar background be set to none, a solid colour or an image, and lets an avatar be replaced across every scene at once. I picked a plain colour on day one for exactly one reason: it will still look the same in two years, which no image or set will.
A brand theme applied to every project. Logo, two colours, one typeface, saved once and applied rather than rebuilt. This eliminated an entire category of drift for about ten minutes of setup.
A script template. Every segment opens the same way, states what the viewer will be able to do, and closes the same way. This is what fixed the pacing drift, because the template constrains the writing regardless of when I write it.
A fixed narration register. There are eight styles available and I use Explanatory for everything. Switching between them across a course is the audio equivalent of changing the background.
A glossary file. Every term I use, with the segment where it is first defined. This is the least interesting item on the list and the one that has saved me the most.
A fixed segment length target. Which brings me to the constraint that shaped everything else.
The sixty second cap, and why it turned out to be useful
The expressive engine, which produces the natural head movement and gesture rather than basic lip sync, is capped at sixty seconds per video. When I found this out I treated it as a limitation to work around.
It is a limitation. It also forced me into a structure I should have chosen anyway.
A course built as sixty second units cannot contain a nine minute exposition, so it contains a sequence of single ideas. Each one has a title describing what it covers. Students can find the one thing they need to rewatch, and I can revise one unit without touching anything around it.
That last property is the whole argument. When a tool changes its interface, I regenerate one unit. Under my previous structure, where a module was one continuous eleven minute recording, the same change meant re recording eleven minutes or leaving an obvious edit seam.
There is a second engine available that handles basic lip sync without the sixty second constraint, and I use it for the few segments where the presenter is on screen only briefly and the movement matters less.
The design principles behind chunked course structure are not something I invented, and the standards published by Quality Matters on course design cover the same ground from a pedagogical direction rather than a production one.
What the question panel told me
The share page carries a chat panel where a viewer can ask questions about the video. I ignored it for months.
When I finally read the accumulated questions, two things were obvious. The first is that a large number of questions were about one specific segment, which told me that segment was not working. I rewrote it and the questions stopped.
The second is more uncomfortable. Several questions were things I had answered, in a different segment, that students had not connected. That is a course structure problem, not a video problem, and I would not have found it from completion data alone, because those students finished everything.
Completion rate did tell me one useful thing: the drop off is not gradual, it is concentrated at specific segments, and those segments are always the ones where I assumed knowledge.
Languages, which changed the economics
About a third of my students are not native English speakers, and until last year my answer to that was captions.
A finished video can be translated into another language, generating a separate version rather than replacing the original, with the on screen text translated along with the narration. The presenter's mouth movement is regenerated to match the new audio, which is what stops it looking like a dub.
I now maintain the course in three languages. This would have been unthinkable when each version meant a separate recording, and it is the single largest change to my business in the last two years.
I pay someone to check each translation. This is not optional. Technical terminology and anything idiomatic come through badly, and a course that teaches a term wrongly in translation is worse than a course that does not exist in that language.
Captions stay on in every version. Students watch on phones, often in shared spaces, and captions for prerecorded video are a WCAG success criterion rather than a nice extra.
What I still record myself
I do not expect any of this to change.
The course introduction. People buy from a person. The introduction is where a prospective student decides whether they want to spend twenty hours with me, and a generated presenter answers that question badly.
Anything where I am demonstrating something on my own screen. Screen capture is a separate job done with a separate tool, and the generated presenter has nothing to do with it. I record the screen, then place that clip inside the project alongside the generated segments.
Anything responding to a specific student or a current event. The value of that message is that I chose to make it, and automating it removes the reason it works.
Questions I get from other instructors
Does it look obviously synthetic? Somewhat, and I say so up front, in the course description and in the introduction, which I record myself. Nobody has objected. Several people have said they preferred it once they understood why.
How long does a segment take now? Twenty minutes of writing, five minutes of generation, ten minutes of checking. Against filming, that is the difference between a revision I will do and one I will postpone.
What about your voice? Voice cloning is available and needs a short clean sample. I use it, because a course narrated by a stock voice and introduced by me is jarring. Cloning anyone else's voice is a different matter entirely and requires their agreement.
Was the switch worth it? The measure I use is how out of date the course is at any given moment. Previously it was months. Now it is days. That is the whole return, and it does not appear in any production cost comparison.
If I were starting again
I would decide the six standards first, before making a single video, because retrofitting consistency across a course is far more work than imposing it from the start.
And I would build in sixty second units from the beginning rather than discovering the constraint at video eleven, because the structure it forces is better than the one I would have chosen on my own.