An A/B test that teaches you something
The useful version of a test is one deliberate change at a time, mostly to the packaging rather than the film.
Version B did better. You do not know why, and you cannot repeat it.
That is the ordinary outcome, and the cause is almost always the same. Version B had a new title, a new thumbnail and a different opening, so the result tells you that some combination of three things worked, which is not a finding you can carry into the next video.
Change one thing. Decide before you start what result would make you change your mind. Run it long enough that a quiet week cannot decide it for you. Everything below is the mechanism behind those three sentences.
Why did your last test teach you nothing?
Because the variables were entangled, and entangled variables cannot be separated afterwards by thinking about them harder.
If the title and the thumbnail both changed, a win could mean the title was better, or the thumbnail was better, or one was worse and the other rescued it. All three explanations fit the same result. You now have a preference, not evidence, and the next video inherits both changes on the strength of it.
There is a second trap underneath. Video A and video B are usually different videos, on different days, about different subjects. That is not a test of anything. It is two publications with a comparison written over the top.
What is worth testing when your numbers are small?
The packaging, not the film.
The title and the thumbnail are cheap to change, they can be changed after the video is published, and on a channel with enough traffic the platform will now rotate a few versions of each and report which won. The opening line belongs with them, it decides whether anyone stays past the first seconds, but unlike the other two it is set before you upload and cannot be swapped afterwards, so you test it across videos rather than on one. Between them they decide whether the video is watched at all, rather than how good it was once it started. Testing an edit style is testing the most expensive variable you own, with the slowest feedback and the least chance of a clean result.
Test the things that sit between the viewer and the video. The video itself is improved by editing it better, which is a different activity and does not need a control group.
What would change your mind?
Write it down before you look at anything.
Say out loud what result would make you keep the new thumbnail, and what result would make you go back. It can be plain: if the new one does clearly better across a full publishing cycle, it becomes the house style. If the two are close, the change did nothing you can measure, and the old one stays.
This step feels like paperwork and it is the whole discipline. Without it, you will read whatever number arrives as confirmation of whatever you already wanted, and every test you run will agree with you.
How long should a test run?
Longer than one video, and across like for like.
Compare the same slot, the same kind of subject and the same day of the week, because a video that went out on a bank holiday is not evidence about anything except bank holidays. A run of several publications with one variable held steady tells you more than a single head to head, and it costs you nothing except patience.
The honest version at a small scale is not really a test at all. It is a slow series of single deliberate changes, one per video, each written down with the date and the reason, so that after a few months you have a list of things you have tried rather than a hunch you have repeated.
When is a difference just noise?
More often than anybody would like, and the smaller your channel, the truer that is.
At low view counts, the spread between two videos is dominated by things you did not control. A single share by somebody with an audience, a slow news week, the platform testing you against a different set of viewers. Two numbers that differ a little are usually the same number twice.
Which leads to the expensive advice. If the two results are close, do not pick a winner. Record that the change did nothing measurable and keep the version you would rather look at, because at that scale taste is a legitimate tie breaker and pretending otherwise builds a house style on noise.
What does this look like in a studio?
Three thumbnail concepts per video rather than one, made as genuinely different ideas rather than three colourways of the same idea, so that what comes back is information about which approach the audience responds to.
When the question is about the words rather than the pictures, two alternative voiceovers are recorded over the same cut. The pictures, the pacing, the music and the length are identical, so the only thing that can explain a difference is the script. That is a controlled test, and it is only possible because everything else was deliberately held still. Where the platform offers its own built-in experiment, that is the one case where the allocation is randomised for you rather than judged after the fact; anywhere else, a sequence of single changes is an observation, and worth calling one.
Neither of those requires a studio. They require deciding what the question is before making the thing, which is the part that is genuinely hard.
What can you run this week?
Take a video that has been up long enough to settle, change only its thumbnail, and note the date. Leave the title alone even though you want to change it too. Look again after a full publishing cycle and write one sentence about what happened, including the sentence nothing happened if that is what you find.
The next decision is about appetite. A test that could be conclusive requires you to hold everything else still, and holding everything still is boring. If that is not something you want to do, run the slow series instead and stop calling either of them an experiment.
Keep reading
This is a service, and the method is written down.
Everything above came out of doing the work rather than writing about it. If you want the method instead of the story, it runs in order on one page.
Start with a free audit
Tell us where your content is now. We will come back with what we would change and what result to expect.
A person reads the channel and writes the audit by hand: a considered read typically takes three working days. That is the usual shape, not a promised turnaround. We use these details only to reply to you: no lists, no lurking.
What you will get
A fit snapshot: where your channel stands, and whether we are a match.
Two to three opportunities: specific, prioritised, yours to keep.
A recommended next step, even if that step is not us.
The audit is free and commits you to nothing: nobody follows up with a call you did not ask for.