ChatGPT Voice gets 90 minutes, three blind reviewers and a rescue threshold. The result is a real hire-or-fire test, not a feature demo.
$32,700 in savings. 10 months of runway. non-technical founder building a real SaaS to $10K/month using only AI tools - or going broke trying.
this playlist is my full journey from $1,940 MRR to $10K MRR, filmed in real time with real numbers. the wins. the fuck-ups and the dollars. no code experience. no safety net. just AI tools and a deadline.
new videos every week. subscribe to see how it ends.
# Final transcript: AI ON PAYROLL
- Source: `source/video.mp4`
- Source SHA-256: `ef1728b88ea4f47de92d0be2b2182b0bcea7dd30fc2a1a818b9e00a2fa6a0ed9`
- Retranscription model: `medium.en`
- Language: `en`
## 00:00:00.000
ChatGPT, start three blind reviewers. The clock is running. Do not tell me which AI made either
## 00:00:06.720
edit. Understood. I'm checking the launch details now. Within 90 minutes, three blind reviews land
## 00:00:13.580
in their files with Aster and Cobalt anonymous. Any identity leaked? Armed.
## 00:00:26.860
Three reviews, one decision, more than three rescues and you're fired. And now what I'm
## 00:00:33.340
going to do is get the new voice mode inside of the ChatGPT app to essentially hire some reviewers
## 00:00:41.220
to look at both edits and to basically determine which one makes the best YouTube video. I will be
## 00:00:47.420
the one making the final decision at the end, but that's what we're going to be doing. Strap in,
## 00:00:52.080
let's fucking go. So why are we doing this today? OpenAI's voice mode came out in the new ChatGPT
## 00:00:59.300
app about eight hours ago as of the time of this recording. This is what it should look like when
## 00:01:06.080
it eventually comes into your ChatGPT app. You update the app and you should see something that
## 00:01:11.060
looks like this, where it just has try ChatGPT voice, start voice. And when you click on that,
## 00:01:18.720
you'll get a pop-up that looks something like this, meet ChatGPT voice. What you might have
## 00:01:23.160
noticed in the intro is that it could look at my screen, which means that you don't have to put a
## 00:01:29.560
screenshot in the prompt. You can just ask Chat voice, read the thing on my screen and tell me
## 00:01:36.420
what to do next. Now this is AI on payroll. It's a type of video format I do. And what
## 00:01:43.280
that means is we're going to hire ChatGPT voice. If all three reviews work, the names stay hidden
## 00:01:51.200
and I rescue voice three times or less, which means I just have to interfere with what it
## 00:01:57.920
was doing. And look, I'd say there was one time where I had to interfere here. So call that once.
## 00:02:06.080
Four or five rescues means probation. It's in timeout. And then six means that AI,
## 00:02:13.020
ChatGPT voice is fired. See in those two videos that there were two different approaches
## 00:02:19.440
in terms of taste and style to the video. The second one was more dramatic.
## 00:02:30.440
More dramatic. The first one had no music, but had its own distinct visual style as well.
## 00:02:45.440
And what it's doing here is it's spawning some sub-agents to go and review those three videos.
## 00:02:54.640
And each one has a different task. So one reviewer judges story and retention.
## 00:03:01.740
That's this one over here. The other reviewer checks proof and technical defects. And then
## 00:03:08.580
the third reviewer checks or attacks the obvious taste. And it looks like we have our second rescue.
## 00:03:16.020
As I said earlier, after rescue at six times, it means it's fired. So far, we've had this
## 00:03:23.620
one instance here and then another instance where it finished working and it's not doing anything
## 00:03:30.060
yet, but it's not given me an output. So here, I'm guessing I'll have to prompt it to continue.
## 00:03:36.840
So that's two rescues. You just stopped working randomly. Continue your task.
## 00:03:46.020
The three reviews and anonymous decision package are complete.
## 00:03:49.360
The only hold is your anonymous lock. Say my anonymous choice is Aster or Cobalt.
## 00:03:59.720
Cobalt. Aster and Cobalt are the code names for the two videos. So Aster is this one.
## 00:04:11.290
Look, so I've watched the full one of that and it's it's quite an emotional video about
## 00:04:18.190
second brains, which was not my intention. But whichever model decided to create that
## 00:04:22.150
decided to go for to go for the emotional vibes, whatever. So that is Aster. Cobalt is this one.
## 00:04:42.070
I don't know about you, but I actually like Aster, the first one better than Cobalt,
## 00:04:47.230
the one that was just on the screen, mainly because from what I can see, Cobalt includes
## 00:04:51.570
a few word gaps and retakes that ideally I didn't want to end up in the final video. So
## 00:04:58.490
I'm going to let it know. My anonymous choice is Aster. Lock it.
## 00:05:06.050
Okay, locking that in now. Aster is now locked as your anonymous choice.
## 00:05:10.890
I haven't opened any model map.
## 00:05:17.290
Lock Aster and reveal the sealed experiment folder and not seeing a model map in it.
## 00:05:22.750
Which exact source should I open for the reveal?
## 00:05:27.010
I don't know. You tell me you're the one who's running this.
## 00:05:31.850
The documented sealed map is missing, so I'm recovering the documented deterministic mapping
## 00:05:36.990
from the two delivery manifests without creating or changing any identity file.
## 00:05:47.150
There is a mute button. All right. It was right there. I just thought that was going
## 00:05:51.430
to activate the normal speech to text feature. But the mute button is there,
## 00:05:55.490
which is great because I was wondering why on earth that would not have existed
## 00:05:59.890
earlier, because it just made a lot of sense. OK, so I'm going to deduct the rescue from here.
## 00:06:09.750
Reveal verified. Candidate Aster is Fable 5. Candidate Cobalt is GPT-5.6 Sol.
## 00:06:17.250
The original sealed map was absent, but the deterministic assignment
## 00:06:21.150
and the delivered media hashes map.
## 00:06:29.070
So Fable 5 was the one that edited the video that we chose.
## 00:06:34.170
All right. Cool. What else can you do for me? Best next demo:
## 00:06:39.970
Give me a messy live production problem. One constraint, conflicting inputs, a deadline.
## 00:06:44.850
I don't know what any of that means. Just speak simply and give it to me in simple words.
## 00:06:50.990
Sure, I'll put it plainly. In simple terms, I can be your live producer. You talk. I keep you
## 00:06:57.470
on track, give you the next shot, hand jobs to other AIs, bring their answers back and
## 00:07:03.070
help you choose what to do next. So really, this can be something that you just have always on.
## 00:07:10.190
And because you can actually mute it, I think there's a lot of power in having this as an
## 00:07:15.510
always on AI agent, almost like a chief of staff, because I can also get it to connect.
## 00:07:23.550
So coming back to our video, I just thought of another question I had.
## 00:07:27.250
What was the rationale or the thought process behind picking ChatGPT-5.6? And
## 00:07:32.210
why do you think I preferred Fable's when the three reviewers all preferred ChatGPT-5.6?
## 00:07:39.250
Don't you think there's a bit of favoritism going on there because the three reviewers
## 00:07:43.550
are ChatGPT-5.6 itself? Let me check.
## 00:07:48.050
Smart correction, they didn't all pick 5.6 overall. For the long video, two picked Cobalt.
## 00:07:54.650
That's GPT-5.6's edit, basically because it was shorter, clearer and more polished.
## 00:08:00.970
The technical reviewer picked Aster, which is Fable's edit, and two preferred Aster for
## 00:08:05.890
Shorts because they felt more direct and it had a full five-video set.
## 00:08:10.390
Blind labels help, but if all reviewers are GPT based, that's not fully independent.
## 00:08:16.850
A stronger test would include human scoring and reviewers from different models.
## 00:08:27.250
All right, fair enough. So where does that leave us in terms of the AI video editing between
## 00:08:33.790
ChatGPT-5.6 Sol and Fable 5? We've already seen that the long form video,
## 00:08:39.630
my own personal judgment here is that Fable 5 did a better job. But what about the Shorts?
## 00:08:44.110
This is, well, these are the Shorts that Fable 5 edited. I'm just going to pick
## 00:08:48.150
a random one and we'll take a look. I found Obsidian's breaking point. 44,331 notes opened.
## 00:08:57.070
Note 44,000. Not bad. It all makes sense. And it's removed the retakes and it seems pretty coherent.
## 00:09:04.710
Then for GPT-5.6 Sol, crashed nine times while I tried to squeeze the full vault
## 00:09:10.370
into its graph view. I'm not going to do crash number 10 for some dramatic.
## 00:09:15.430
That one's harder to pick. I'm going to have a look at another one from GPT Sol.
## 00:09:19.930
Most second brains look pretty, but do fuck all. And if yours can't help with a real decision.
## 00:09:24.870
One more from Fable. A log from March. One of my second brain logs from March 6th.
## 00:09:33.690
Oh, it's really hard. Fable 5 included a retake at the start, which it shouldn't have. I didn't
## 00:09:39.430
see any of those mistakes from GPT Sol. I like the motion graphics better from
## 00:09:43.890
Fable though. I don't know. I'll let you decide.
## 00:09:47.930
I think where I'm leaning is honestly, it's very close.
## 00:09:51.730
Fable wins for the long forms. And I would say it's a tie for the short forms. And that would
## 00:09:59.430
bring Fable as the winner overall. That is the final verdict for the AI videos.
## 00:10:05.930
But what is the final verdict for ChatGPT voice mode? If it, as we've already said,
## 00:10:11.050
if I had to rescue it four or five times, it's going to be on probation rescue six times.
## 00:10:16.070
It's going to be fired. I only had to rescue at once. So I consider this to be something that is
## 00:10:20.810
worth hiring. ChatGPT voice mode is now on payroll.
## 00:10:29.210
And as you also saw those three reviewers, the ChatGPT voice spawned as sub agents,
## 00:10:36.490
they were looking at the evidence of the videos, but then there's this whole other dimension of
## 00:10:41.530
taste. And this is where I see the interaction between ChatGPT voice and humans.
## 00:10:49.590
Honestly, a massive moment for AI in general. I think some other people likened this to
## 00:10:56.090
the open call moment, but for voice mode and I can see that just right now, I can see this
## 00:11:01.730
is already a permanent fixture on my screen. And if that isn't a defining moment in AI,
## 00:11:06.690
then I don't know what is, but the point I'm trying to make here is that the interaction
## 00:11:11.830
needs to be as follows. You can get the voice mode to manage evidence and you know, files,
## 00:11:18.570
plugins, get it to do stuff that you would normally get AI to do.
## 00:11:24.150
But, and this is a big but, make sure that your own tastes and judgment is still being used.
## 00:11:31.990
Because if I just went off of what ChatGPT voices three sub agents suggested, then
## 00:11:40.170
I would be defaulting and deferring my judgment to the AI. And what this test has shown is chat
## 00:11:45.850
GPT voice is interesting only if it manages real work without making me its full-time babysitter.
## 00:11:53.530
The human boundary is clear. Voice can manage the evidence and it doesn't get my taste vote.
## 00:11:59.010
That stays with me. So if you want to do a test like this with ChatGPT voice mode,
## 00:12:03.830
here are four rules. Give it a real job, something that you are actually going to do.
## 00:12:07.890
I was actually going to compare both videos that the different models created a real job. Set a
## 00:12:13.670
clock. It's not going to be as helpful if it can only do it in 24 hours versus in 10 minutes or
## 00:12:19.170
whatever we did. Then count every rescue. The rescues are the attempts where you have to go
## 00:12:23.230
and babysit it because it's kind of defeating the purpose of what you want this thing to do
## 00:12:27.930
when you have to constantly babysit it. And then keep one human line, meaning name the decision it
## 00:12:33.670
can't make. For the video editing, it was my final call on which video I preferred best.
## 00:12:39.110
If you see the power of ChatGPT voice, but you don't know how you can leverage it in your own
## 00:12:43.610
business, invite Lennox in. I'm Lennox, go check the link in the description below.
## 00:12:48.870
Can't wait to see you soon.