WEBVTT
Kind: captions
Language: en

00:00:00.400 --> 00:00:02.070 align:start position:0%
 
If<00:00:00.560><c> we</c><00:00:00.719><c> look</c><00:00:00.800><c> at</c><00:00:00.960><c> how</c><00:00:01.120><c> the</c><00:00:01.280><c> intelligence</c><00:00:01.760><c> of</c><00:00:01.920><c> an</c>

00:00:02.070 --> 00:00:02.080 align:start position:0%
If we look at how the intelligence of an
 

00:00:02.080 --> 00:00:04.550 align:start position:0%
If we look at how the intelligence of an
AI<00:00:02.480><c> model</c><00:00:02.879><c> grows</c><00:00:03.600><c> as</c><00:00:03.840><c> we</c><00:00:04.080><c> increase</c><00:00:04.319><c> its</c>

00:00:04.550 --> 00:00:04.560 align:start position:0%
AI model grows as we increase its
 

00:00:04.560 --> 00:00:07.110 align:start position:0%
AI model grows as we increase its
compute,<00:00:05.440><c> we</c><00:00:05.600><c> can</c><00:00:05.759><c> see</c><00:00:05.920><c> that</c><00:00:06.080><c> we</c><00:00:06.319><c> are</c><00:00:06.480><c> limited</c>

00:00:07.110 --> 00:00:07.120 align:start position:0%
compute, we can see that we are limited
 

00:00:07.120 --> 00:00:09.509 align:start position:0%
compute, we can see that we are limited
by<00:00:07.359><c> our</c><00:00:07.600><c> data.</c><00:00:08.480><c> It</c><00:00:08.639><c> doesn't</c><00:00:08.880><c> matter</c><00:00:09.040><c> how</c><00:00:09.280><c> big</c>

00:00:09.509 --> 00:00:09.519 align:start position:0%
by our data. It doesn't matter how big
 

00:00:09.519 --> 00:00:11.589 align:start position:0%
by our data. It doesn't matter how big
we<00:00:09.679><c> will</c><00:00:09.920><c> make</c><00:00:10.080><c> the</c><00:00:10.320><c> model.</c><00:00:10.800><c> If</c><00:00:11.040><c> we're</c><00:00:11.280><c> using</c>

00:00:11.589 --> 00:00:11.599 align:start position:0%
we will make the model. If we're using
 

00:00:11.599 --> 00:00:13.910 align:start position:0%
we will make the model. If we're using
supervised<00:00:12.320><c> fine</c><00:00:12.719><c> tuning,</c><00:00:13.200><c> which</c><00:00:13.440><c> is</c><00:00:13.679><c> the</c>

00:00:13.910 --> 00:00:13.920 align:start position:0%
supervised fine tuning, which is the
 

00:00:13.920 --> 00:00:16.230 align:start position:0%
supervised fine tuning, which is the
naive<00:00:14.400><c> training</c><00:00:14.719><c> method,</c><00:00:15.440><c> the</c><00:00:15.679><c> performance</c>

00:00:16.230 --> 00:00:16.240 align:start position:0%
naive training method, the performance
 

00:00:16.240 --> 00:00:18.550 align:start position:0%
naive training method, the performance
is<00:00:16.400><c> only</c><00:00:16.640><c> going</c><00:00:16.800><c> to</c><00:00:16.960><c> be</c><00:00:17.039><c> limited</c><00:00:17.600><c> to</c><00:00:17.840><c> imitating</c>

00:00:18.550 --> 00:00:18.560 align:start position:0%
is only going to be limited to imitating
 

00:00:18.560 --> 00:00:21.269 align:start position:0%
is only going to be limited to imitating
perfectly<00:00:19.279><c> what</c><00:00:19.520><c> the</c><00:00:19.760><c> model</c><00:00:19.920><c> is</c><00:00:20.160><c> trained</c><00:00:20.400><c> on.</c>

00:00:21.269 --> 00:00:21.279 align:start position:0%
perfectly what the model is trained on.
 

00:00:21.279 --> 00:00:23.750 align:start position:0%
perfectly what the model is trained on.
In<00:00:21.520><c> this</c><00:00:21.760><c> video,</c><00:00:22.080><c> I</c><00:00:22.240><c> want</c><00:00:22.400><c> to</c><00:00:22.560><c> show</c><00:00:22.720><c> you</c><00:00:23.119><c> and</c><00:00:23.519><c> to</c>

00:00:23.750 --> 00:00:23.760 align:start position:0%
In this video, I want to show you and to
 

00:00:23.760 --> 00:00:26.390 align:start position:0%
In this video, I want to show you and to
visualize<00:00:24.400><c> intuitively</c><00:00:25.359><c> an</c><00:00:25.680><c> algorithm</c><00:00:26.160><c> that</c>

00:00:26.390 --> 00:00:26.400 align:start position:0%
visualize intuitively an algorithm that
 

00:00:26.400 --> 00:00:28.630 align:start position:0%
visualize intuitively an algorithm that
allows<00:00:26.640><c> us</c><00:00:26.880><c> to</c><00:00:27.119><c> surpass</c><00:00:27.599><c> the</c><00:00:27.840><c> intelligence</c>

00:00:28.630 --> 00:00:28.640 align:start position:0%
allows us to surpass the intelligence
 

00:00:28.640 --> 00:00:31.429 align:start position:0%
allows us to surpass the intelligence
that<00:00:29.039><c> exists</c><00:00:29.439><c> in</c><00:00:29.679><c> our</c><00:00:29.920><c> data.</c><00:00:30.880><c> What</c><00:00:31.039><c> we</c><00:00:31.199><c> call</c>

00:00:31.429 --> 00:00:31.439 align:start position:0%
that exists in our data. What we call
 

00:00:31.439 --> 00:00:34.229 align:start position:0%
that exists in our data. What we call
reinforcement<00:00:32.079><c> learning</c><00:00:32.880><c> and</c><00:00:33.200><c> specifically</c>

00:00:34.229 --> 00:00:34.239 align:start position:0%
reinforcement learning and specifically
 

00:00:34.239 --> 00:00:37.750 align:start position:0%
reinforcement learning and specifically
GRPO.<00:00:35.600><c> If</c><00:00:35.760><c> we</c><00:00:35.920><c> look</c><00:00:36.079><c> at</c><00:00:36.160><c> how</c><00:00:36.399><c> naive</c><00:00:37.200><c> supervised</c>

00:00:37.750 --> 00:00:37.760 align:start position:0%
GRPO. If we look at how naive supervised
 

00:00:37.760 --> 00:00:40.310 align:start position:0%
GRPO. If we look at how naive supervised
fine-tuning<00:00:38.480><c> works,</c><00:00:39.200><c> we</c><00:00:39.440><c> can</c><00:00:39.600><c> see</c><00:00:39.840><c> that</c><00:00:40.160><c> we</c>

00:00:40.310 --> 00:00:40.320 align:start position:0%
fine-tuning works, we can see that we
 

00:00:40.320 --> 00:00:43.430 align:start position:0%
fine-tuning works, we can see that we
just<00:00:40.559><c> take</c><00:00:40.879><c> human</c><00:00:41.280><c> data</c><00:00:41.680><c> and</c><00:00:41.920><c> our</c><00:00:42.160><c> model</c><00:00:43.120><c> and</c>

00:00:43.430 --> 00:00:43.440 align:start position:0%
just take human data and our model and
 

00:00:43.440 --> 00:00:46.630 align:start position:0%
just take human data and our model and
every<00:00:43.840><c> word</c><00:00:44.559><c> split</c><00:00:44.960><c> a</c><00:00:45.280><c> piece</c><00:00:45.520><c> of</c><00:00:45.680><c> text</c><00:00:46.239><c> between</c>

00:00:46.630 --> 00:00:46.640 align:start position:0%
every word split a piece of text between
 

00:00:46.640 --> 00:00:50.069 align:start position:0%
every word split a piece of text between
all<00:00:46.879><c> of</c><00:00:47.039><c> the</c><00:00:47.200><c> words</c><00:00:47.440><c> before</c><00:00:47.840><c> it</c><00:00:48.399><c> and</c><00:00:49.039><c> that</c><00:00:49.440><c> word</c>

00:00:50.069 --> 00:00:50.079 align:start position:0%
all of the words before it and that word
 

00:00:50.079 --> 00:00:52.630 align:start position:0%
all of the words before it and that word
and<00:00:50.320><c> we</c><00:00:50.480><c> train</c><00:00:50.719><c> the</c><00:00:50.960><c> model</c><00:00:51.440><c> that</c><00:00:52.000><c> the</c><00:00:52.239><c> context</c>

00:00:52.630 --> 00:00:52.640 align:start position:0%
and we train the model that the context
 

00:00:52.640 --> 00:00:54.630 align:start position:0%
and we train the model that the context
of<00:00:52.800><c> all</c><00:00:52.960><c> the</c><00:00:53.120><c> words</c><00:00:53.440><c> before</c><00:00:53.920><c> would</c><00:00:54.160><c> make</c><00:00:54.320><c> it</c>

00:00:54.630 --> 00:00:54.640 align:start position:0%
of all the words before would make it
 

00:00:54.640 --> 00:00:57.110 align:start position:0%
of all the words before would make it
statistically<00:00:55.280><c> more</c><00:00:55.600><c> likely</c><00:00:56.320><c> for</c><00:00:56.719><c> that</c>

00:00:57.110 --> 00:00:57.120 align:start position:0%
statistically more likely for that
 

00:00:57.120 --> 00:00:59.910 align:start position:0%
statistically more likely for that
target<00:00:57.600><c> word</c><00:00:57.920><c> to</c><00:00:58.160><c> be</c><00:00:58.320><c> selected</c><00:00:58.800><c> next.</c><00:00:59.520><c> And</c><00:00:59.680><c> if</c>

00:00:59.910 --> 00:00:59.920 align:start position:0%
target word to be selected next. And if
 

00:00:59.920 --> 00:01:02.229 align:start position:0%
target word to be selected next. And if
we<00:01:00.079><c> do</c><00:01:00.160><c> it</c><00:01:00.320><c> for</c><00:01:00.640><c> trillions</c><00:01:01.120><c> of</c><00:01:01.359><c> times,</c><00:01:02.000><c> the</c>

00:01:02.229 --> 00:01:02.239 align:start position:0%
we do it for trillions of times, the
 

00:01:02.239 --> 00:01:04.869 align:start position:0%
we do it for trillions of times, the
model<00:01:02.559><c> will</c><00:01:02.879><c> end</c><00:01:03.120><c> up</c><00:01:03.520><c> understanding</c><00:01:04.400><c> language</c>

00:01:04.869 --> 00:01:04.879 align:start position:0%
model will end up understanding language
 

00:01:04.879 --> 00:01:12.789 align:start position:0%
model will end up understanding language
in<00:01:05.119><c> a</c><00:01:05.280><c> way.</c>

00:01:12.789 --> 00:01:12.799 align:start position:0%
 
 

00:01:12.799 --> 00:01:15.190 align:start position:0%
 
But<00:01:13.040><c> this</c><00:01:13.360><c> training</c><00:01:14.080><c> doesn't</c><00:01:14.479><c> make</c><00:01:14.640><c> the</c><00:01:14.880><c> model</c>

00:01:15.190 --> 00:01:15.200 align:start position:0%
But this training doesn't make the model
 

00:01:15.200 --> 00:01:17.429 align:start position:0%
But this training doesn't make the model
correct<00:01:15.600><c> about</c><00:01:15.840><c> anything.</c><00:01:16.479><c> It</c><00:01:16.720><c> make</c><00:01:16.880><c> it</c><00:01:17.119><c> good</c>

00:01:17.429 --> 00:01:17.439 align:start position:0%
correct about anything. It make it good
 

00:01:17.439 --> 00:01:19.990 align:start position:0%
correct about anything. It make it good
at<00:01:17.759><c> imitating</c><00:01:18.640><c> what</c><00:01:18.960><c> was</c><00:01:19.119><c> in</c><00:01:19.360><c> our</c><00:01:19.600><c> training</c>

00:01:19.990 --> 00:01:20.000 align:start position:0%
at imitating what was in our training
 

00:01:20.000 --> 00:01:22.789 align:start position:0%
at imitating what was in our training
data.<00:01:21.040><c> It</c><00:01:21.280><c> can</c><00:01:21.439><c> look</c><00:01:21.600><c> like</c><00:01:21.759><c> it</c><00:01:22.000><c> reasons</c><00:01:22.400><c> and</c><00:01:22.560><c> it</c>

00:01:22.789 --> 00:01:22.799 align:start position:0%
data. It can look like it reasons and it
 

00:01:22.799 --> 00:01:25.190 align:start position:0%
data. It can look like it reasons and it
can<00:01:22.960><c> look</c><00:01:23.360><c> like</c><00:01:23.600><c> it</c><00:01:24.240><c> thinks</c><00:01:24.560><c> about</c><00:01:24.799><c> how</c><00:01:24.960><c> to</c><00:01:25.119><c> get</c>

00:01:25.190 --> 00:01:25.200 align:start position:0%
can look like it thinks about how to get
 

00:01:25.200 --> 00:01:26.870 align:start position:0%
can look like it thinks about how to get
the<00:01:25.439><c> right</c><00:01:25.600><c> answer,</c><00:01:26.159><c> but</c><00:01:26.400><c> it's</c><00:01:26.640><c> just</c>

00:01:26.870 --> 00:01:26.880 align:start position:0%
the right answer, but it's just
 

00:01:26.880 --> 00:01:28.870 align:start position:0%
the right answer, but it's just
imitating<00:01:27.520><c> how</c><00:01:27.840><c> humans</c><00:01:28.240><c> look</c><00:01:28.479><c> when</c><00:01:28.720><c> they</c>

00:01:28.870 --> 00:01:28.880 align:start position:0%
imitating how humans look when they
 

00:01:28.880 --> 00:01:31.030 align:start position:0%
imitating how humans look when they
reason.

00:01:31.030 --> 00:01:31.040 align:start position:0%
reason.
 

00:01:31.040 --> 00:01:33.830 align:start position:0%
reason.
And<00:01:31.200><c> this</c><00:01:31.439><c> is</c><00:01:31.520><c> where</c><00:01:31.920><c> GRPO</c><00:01:32.720><c> for</c><00:01:32.960><c> reasoning</c>

00:01:33.830 --> 00:01:33.840 align:start position:0%
And this is where GRPO for reasoning
 

00:01:33.840 --> 00:01:36.149 align:start position:0%
And this is where GRPO for reasoning
comes<00:01:34.079><c> into</c><00:01:34.400><c> the</c><00:01:34.560><c> image.</c><00:01:35.600><c> Let's</c><00:01:35.759><c> say</c><00:01:35.920><c> we</c><00:01:36.079><c> have</c>

00:01:36.149 --> 00:01:36.159 align:start position:0%
comes into the image. Let's say we have
 

00:01:36.159 --> 00:01:38.950 align:start position:0%
comes into the image. Let's say we have
the<00:01:36.320><c> historical</c><00:01:37.040><c> weather</c><00:01:37.360><c> data</c><00:01:37.920><c> of</c><00:01:38.240><c> places</c>

00:01:38.950 --> 00:01:38.960 align:start position:0%
the historical weather data of places
 

00:01:38.960 --> 00:01:41.030 align:start position:0%
the historical weather data of places
and<00:01:39.200><c> we</c><00:01:39.439><c> want</c><00:01:39.600><c> to</c><00:01:39.680><c> train</c><00:01:39.920><c> a</c><00:01:40.159><c> model</c><00:01:40.479><c> to</c><00:01:40.720><c> be</c><00:01:40.799><c> able</c>

00:01:41.030 --> 00:01:41.040 align:start position:0%
and we want to train a model to be able
 

00:01:41.040 --> 00:01:43.429 align:start position:0%
and we want to train a model to be able
to<00:01:41.200><c> reason</c><00:01:41.680><c> about</c><00:01:42.000><c> weather</c><00:01:42.479><c> and</c><00:01:42.799><c> predict</c><00:01:43.119><c> the</c>

00:01:43.429 --> 00:01:43.439 align:start position:0%
to reason about weather and predict the
 

00:01:43.439 --> 00:01:45.670 align:start position:0%
to reason about weather and predict the
temperature<00:01:43.840><c> of</c><00:01:44.079><c> a</c><00:01:44.240><c> future</c><00:01:44.560><c> day.</c><00:01:45.200><c> We</c><00:01:45.439><c> can</c><00:01:45.520><c> take</c>

00:01:45.670 --> 00:01:45.680 align:start position:0%
temperature of a future day. We can take
 

00:01:45.680 --> 00:01:48.149 align:start position:0%
temperature of a future day. We can take
a<00:01:45.759><c> naively</c><00:01:46.320><c> trained</c><00:01:46.720><c> model</c><00:01:47.119><c> and</c><00:01:47.439><c> ask</c><00:01:47.680><c> it</c><00:01:47.920><c> to</c>

00:01:48.149 --> 00:01:48.159 align:start position:0%
a naively trained model and ask it to
 

00:01:48.159 --> 00:01:50.389 align:start position:0%
a naively trained model and ask it to
predict<00:01:48.640><c> the</c><00:01:48.960><c> temperature</c><00:01:49.360><c> tomorrow</c><00:01:50.079><c> based</c>

00:01:50.389 --> 00:01:50.399 align:start position:0%
predict the temperature tomorrow based
 

00:01:50.399 --> 00:01:52.870 align:start position:0%
predict the temperature tomorrow based
on<00:01:50.560><c> the</c><00:01:50.799><c> temperature</c><00:01:51.200><c> today.</c><00:01:52.159><c> This</c><00:01:52.399><c> model</c>

00:01:52.870 --> 00:01:52.880 align:start position:0%
on the temperature today. This model
 

00:01:52.880 --> 00:01:56.149 align:start position:0%
on the temperature today. This model
might<00:01:53.520><c> seem</c><00:01:53.840><c> like</c><00:01:54.000><c> it</c><00:01:54.240><c> is</c><00:01:54.399><c> reasoning</c><00:01:55.680><c> but</c><00:01:55.920><c> at</c>

00:01:56.149 --> 00:01:56.159 align:start position:0%
might seem like it is reasoning but at
 

00:01:56.159 --> 00:01:58.230 align:start position:0%
might seem like it is reasoning but at
this<00:01:56.320><c> point</c><00:01:56.560><c> it</c><00:01:56.799><c> just</c><00:01:57.119><c> imitate</c><00:01:57.246><c> [music]</c><00:01:57.759><c> how</c>

00:01:58.230 --> 00:01:58.240 align:start position:0%
this point it just imitate [music] how
 

00:01:58.240 --> 00:02:01.510 align:start position:0%
this point it just imitate [music] how
human<00:01:58.640><c> reason</c><00:01:59.680><c> in</c><00:02:00.079><c> the</c><00:02:00.399><c> training</c><00:02:00.799><c> data.</c><00:02:01.280><c> The</c>

00:02:01.510 --> 00:02:01.520 align:start position:0%
human reason in the training data. The
 

00:02:01.520 --> 00:02:03.590 align:start position:0%
human reason in the training data. The
important<00:02:01.920><c> thing</c><00:02:02.240><c> is</c><00:02:02.479><c> that</c><00:02:02.719><c> at</c><00:02:02.960><c> the</c><00:02:03.119><c> end</c><00:02:03.360><c> the</c>

00:02:03.590 --> 00:02:03.600 align:start position:0%
important thing is that at the end the
 

00:02:03.600 --> 00:02:08.309 align:start position:0%
important thing is that at the end the
model<00:02:03.920><c> produces</c><00:02:04.880><c> one</c><00:02:05.360><c> verifiable</c><00:02:06.399><c> answer.</c>

00:02:08.309 --> 00:02:08.319 align:start position:0%
model produces one verifiable answer.
 

00:02:08.319 --> 00:02:10.630 align:start position:0%
model produces one verifiable answer.
But<00:02:08.560><c> now</c><00:02:08.800><c> let's</c><00:02:09.119><c> not</c><00:02:09.360><c> stop</c><00:02:09.599><c> here</c><00:02:10.000><c> and</c><00:02:10.239><c> let</c><00:02:10.479><c> the</c>

00:02:10.630 --> 00:02:10.640 align:start position:0%
But now let's not stop here and let the
 

00:02:10.640 --> 00:02:13.110 align:start position:0%
But now let's not stop here and let the
model<00:02:10.959><c> generate</c><00:02:11.520><c> another</c><00:02:12.319><c> reasoning</c>

00:02:13.110 --> 00:02:13.120 align:start position:0%
model generate another reasoning
 

00:02:13.120 --> 00:02:16.390 align:start position:0%
model generate another reasoning
imitation<00:02:13.760><c> example</c><00:02:15.200><c> that</c><00:02:15.680><c> converges</c><00:02:16.160><c> to</c>

00:02:16.390 --> 00:02:16.400 align:start position:0%
imitation example that converges to
 

00:02:16.400 --> 00:02:21.110 align:start position:0%
imitation example that converges to
another<00:02:16.800><c> answer</c><00:02:17.920><c> and</c><00:02:18.319><c> another</c><00:02:18.720><c> one.</c>

00:02:21.110 --> 00:02:21.120 align:start position:0%
another answer and another one.
 

00:02:21.120 --> 00:02:23.350 align:start position:0%
another answer and another one.
And<00:02:21.360><c> now</c><00:02:21.599><c> we</c><00:02:21.840><c> have</c><00:02:21.920><c> a</c><00:02:22.160><c> group</c><00:02:22.319><c> of</c><00:02:22.560><c> answers</c><00:02:23.120><c> which</c>

00:02:23.350 --> 00:02:23.360 align:start position:0%
And now we have a group of answers which
 

00:02:23.360 --> 00:02:26.630 align:start position:0%
And now we have a group of answers which
is<00:02:23.520><c> where</c><00:02:23.760><c> the</c><00:02:24.000><c> G</c><00:02:24.319><c> in</c><00:02:24.560><c> grpo</c><00:02:25.360><c> comes</c><00:02:25.680><c> from.</c><00:02:26.239><c> For</c>

00:02:26.630 --> 00:02:26.640 align:start position:0%
is where the G in grpo comes from. For
 

00:02:26.640 --> 00:02:29.190 align:start position:0%
is where the G in grpo comes from. For
training<00:02:27.040><c> the</c><00:02:27.280><c> model</c><00:02:27.520><c> with</c><00:02:27.840><c> GRPO</c><00:02:28.800><c> let's</c><00:02:29.040><c> say</c>

00:02:29.190 --> 00:02:29.200 align:start position:0%
training the model with GRPO let's say
 

00:02:29.200 --> 00:02:32.550 align:start position:0%
training the model with GRPO let's say
we<00:02:29.440><c> have</c><00:02:30.080><c> eight</c><00:02:30.879><c> different</c><00:02:31.520><c> reasoning</c><00:02:32.160><c> and</c>

00:02:32.550 --> 00:02:32.560 align:start position:0%
we have eight different reasoning and
 

00:02:32.560 --> 00:02:35.589 align:start position:0%
we have eight different reasoning and
eight<00:02:33.040><c> different</c><00:02:33.519><c> answers.</c><00:02:34.959><c> We</c><00:02:35.200><c> can</c><00:02:35.360><c> take</c>

00:02:35.589 --> 00:02:35.599 align:start position:0%
eight different answers. We can take
 

00:02:35.599 --> 00:02:37.750 align:start position:0%
eight different answers. We can take
each<00:02:35.760><c> of</c><00:02:36.000><c> those</c><00:02:36.239><c> answers</c><00:02:36.640><c> and</c><00:02:36.959><c> compare</c><00:02:37.280><c> it</c><00:02:37.519><c> to</c>

00:02:37.750 --> 00:02:37.760 align:start position:0%
each of those answers and compare it to
 

00:02:37.760 --> 00:02:40.309 align:start position:0%
each of those answers and compare it to
the<00:02:38.000><c> real</c><00:02:38.400><c> true</c><00:02:38.800><c> data</c>

00:02:40.309 --> 00:02:40.319 align:start position:0%
the real true data
 

00:02:40.319 --> 00:02:43.350 align:start position:0%
the real true data
and<00:02:40.560><c> this</c><00:02:40.720><c> will</c><00:02:40.959><c> give</c><00:02:41.120><c> us</c><00:02:41.360><c> the</c><00:02:41.840><c> error</c><00:02:42.239><c> for</c><00:02:42.640><c> each</c>

00:02:43.350 --> 00:02:43.360 align:start position:0%
and this will give us the error for each
 

00:02:43.360 --> 00:02:46.150 align:start position:0%
and this will give us the error for each
one<00:02:43.920><c> of</c><00:02:44.080><c> the</c><00:02:44.239><c> reasoning</c><00:02:44.720><c> paths.</c><00:02:45.599><c> Now</c><00:02:45.840><c> let's</c>

00:02:46.150 --> 00:02:46.160 align:start position:0%
one of the reasoning paths. Now let's
 

00:02:46.160 --> 00:02:48.150 align:start position:0%
one of the reasoning paths. Now let's
sort<00:02:46.480><c> all</c><00:02:46.720><c> of</c><00:02:46.800><c> the</c><00:02:47.040><c> answers</c><00:02:47.519><c> from</c><00:02:47.760><c> the</c>

00:02:48.150 --> 00:02:48.160 align:start position:0%
sort all of the answers from the
 

00:02:48.160 --> 00:02:51.110 align:start position:0%
sort all of the answers from the
smallest<00:02:48.720><c> error</c><00:02:49.040><c> to</c><00:02:49.280><c> the</c><00:02:49.519><c> highest</c><00:02:50.000><c> error.</c><00:02:50.800><c> Now</c>

00:02:51.110 --> 00:02:51.120 align:start position:0%
smallest error to the highest error. Now
 

00:02:51.120 --> 00:02:53.270 align:start position:0%
smallest error to the highest error. Now
we<00:02:51.280><c> can</c><00:02:51.440><c> do</c><00:02:51.599><c> something</c><00:02:51.840><c> very</c><00:02:52.160><c> simple.</c><00:02:53.040><c> Take</c>

00:02:53.270 --> 00:02:53.280 align:start position:0%
we can do something very simple. Take
 

00:02:53.280 --> 00:02:55.350 align:start position:0%
we can do something very simple. Take
each<00:02:53.519><c> of</c><00:02:53.760><c> those</c><00:02:54.000><c> answers</c><00:02:54.480><c> and</c><00:02:54.879><c> compare</c><00:02:55.200><c> it</c>

00:02:55.350 --> 00:02:55.360 align:start position:0%
each of those answers and compare it
 

00:02:55.360 --> 00:02:58.070 align:start position:0%
each of those answers and compare it
with<00:02:55.680><c> the</c><00:02:56.000><c> average</c><00:02:56.400><c> of</c><00:02:56.640><c> all</c><00:02:56.959><c> answers.</c><00:02:57.840><c> This</c>

00:02:58.070 --> 00:02:58.080 align:start position:0%
with the average of all answers. This
 

00:02:58.080 --> 00:03:01.430 align:start position:0%
with the average of all answers. This
will<00:02:58.319><c> tell</c><00:02:58.480><c> us</c><00:02:58.800><c> how</c><00:02:59.440><c> relatively</c><00:03:00.239><c> good</c><00:03:01.040><c> each</c><00:03:01.280><c> of</c>

00:03:01.430 --> 00:03:01.440 align:start position:0%
will tell us how relatively good each of
 

00:03:01.440 --> 00:03:04.070 align:start position:0%
will tell us how relatively good each of
the<00:03:01.599><c> responses</c><00:03:02.239><c> are.</c><00:03:02.959><c> And</c><00:03:03.120><c> we</c><00:03:03.360><c> can</c><00:03:03.519><c> assume</c>

00:03:04.070 --> 00:03:04.080 align:start position:0%
the responses are. And we can assume
 

00:03:04.080 --> 00:03:06.790 align:start position:0%
the responses are. And we can assume
that<00:03:04.560><c> if</c><00:03:04.879><c> the</c><00:03:05.120><c> answer</c><00:03:05.440><c> is</c><00:03:05.680><c> more</c><00:03:05.920><c> correct,</c><00:03:06.560><c> the</c>

00:03:06.790 --> 00:03:06.800 align:start position:0%
that if the answer is more correct, the
 

00:03:06.800 --> 00:03:09.030 align:start position:0%
that if the answer is more correct, the
reasoning<00:03:07.519><c> that</c><00:03:07.760><c> the</c><00:03:08.000><c> model</c><00:03:08.239><c> did</c><00:03:08.480><c> to</c><00:03:08.720><c> get</c><00:03:08.879><c> to</c>

00:03:09.030 --> 00:03:09.040 align:start position:0%
reasoning that the model did to get to
 

00:03:09.040 --> 00:03:11.030 align:start position:0%
reasoning that the model did to get to
this<00:03:09.200><c> answer</c><00:03:09.519><c> is</c><00:03:09.840><c> more</c><00:03:10.080><c> correct.</c><00:03:10.607><c> [music]</c><00:03:10.800><c> And</c>

00:03:11.030 --> 00:03:11.040 align:start position:0%
this answer is more correct. [music] And
 

00:03:11.040 --> 00:03:12.470 align:start position:0%
this answer is more correct. [music] And
now<00:03:11.200><c> we</c><00:03:11.360><c> can</c><00:03:11.519><c> train</c><00:03:11.680><c> the</c><00:03:11.920><c> model</c><00:03:12.080><c> on</c><00:03:12.319><c> the</c>

00:03:12.470 --> 00:03:12.480 align:start position:0%
now we can train the model on the
 

00:03:12.480 --> 00:03:15.670 align:start position:0%
now we can train the model on the
responses<00:03:12.959><c> that</c><00:03:13.280><c> it</c><00:03:13.599><c> itself</c><00:03:14.239><c> produces.</c><00:03:15.440><c> For</c>

00:03:15.670 --> 00:03:15.680 align:start position:0%
responses that it itself produces. For
 

00:03:15.680 --> 00:03:17.750 align:start position:0%
responses that it itself produces. For
the<00:03:15.920><c> good</c><00:03:16.239><c> responses,</c><00:03:16.959><c> it's</c><00:03:17.280><c> going</c><00:03:17.360><c> to</c><00:03:17.519><c> be</c>

00:03:17.750 --> 00:03:17.760 align:start position:0%
the good responses, it's going to be
 

00:03:17.760 --> 00:03:23.830 align:start position:0%
the good responses, it's going to be
just<00:03:18.080><c> like</c><00:03:18.400><c> supervised</c><00:03:18.959><c> fine-tuning.</c>

00:03:23.830 --> 00:03:23.840 align:start position:0%
 
 

00:03:23.840 --> 00:03:25.509 align:start position:0%
 
And<00:03:24.000><c> for</c><00:03:24.239><c> the</c><00:03:24.400><c> negative</c><00:03:24.720><c> ones,</c><00:03:25.040><c> it's</c><00:03:25.280><c> going</c><00:03:25.360><c> to</c>

00:03:25.509 --> 00:03:25.519 align:start position:0%
And for the negative ones, it's going to
 

00:03:25.519 --> 00:03:27.990 align:start position:0%
And for the negative ones, it's going to
be<00:03:25.599><c> the</c><00:03:25.840><c> same,</c><00:03:26.560><c> just</c><00:03:26.720><c> the</c><00:03:26.959><c> loss</c><00:03:27.280><c> is</c><00:03:27.519><c> multiplied</c>

00:03:27.990 --> 00:03:28.000 align:start position:0%
be the same, just the loss is multiplied
 

00:03:28.000 --> 00:03:30.710 align:start position:0%
be the same, just the loss is multiplied
by<00:03:28.159><c> a</c><00:03:28.319><c> negative</c><00:03:28.720><c> number.</c><00:03:29.599><c> So</c><00:03:30.000><c> it's</c><00:03:30.400><c> less</c>

00:03:30.710 --> 00:03:30.720 align:start position:0%
by a negative number. So it's less
 

00:03:30.720 --> 00:03:32.949 align:start position:0%
by a negative number. So it's less
likely<00:03:31.040><c> to</c><00:03:31.280><c> produce</c><00:03:31.840><c> similar</c><00:03:32.239><c> responses</c><00:03:32.799><c> in</c>

00:03:32.949 --> 00:03:32.959 align:start position:0%
likely to produce similar responses in
 

00:03:32.959 --> 00:03:35.487 align:start position:0%
likely to produce similar responses in
the<00:03:33.120><c> future.</c>

00:03:35.487 --> 00:03:35.497 align:start position:0%
the future.
 

00:03:35.497 --> 00:03:38.229 align:start position:0%
the future.
[music]

00:03:38.229 --> 00:03:38.239 align:start position:0%
 
 

00:03:38.239 --> 00:03:41.509 align:start position:0%
 
Now<00:03:38.480><c> if</c><00:03:38.720><c> we</c><00:03:38.879><c> repeat</c><00:03:39.280><c> this</c><00:03:39.680><c> many</c><00:03:40.000><c> many</c><00:03:40.319><c> times</c><00:03:41.280><c> we</c>

00:03:41.509 --> 00:03:41.519 align:start position:0%
Now if we repeat this many many times we
 

00:03:41.519 --> 00:03:44.070 align:start position:0%
Now if we repeat this many many times we
can<00:03:41.680><c> see</c><00:03:41.840><c> that</c><00:03:42.080><c> the</c><00:03:42.319><c> accuracy</c><00:03:42.879><c> will</c><00:03:43.280><c> increase</c>

00:03:44.070 --> 00:03:44.080 align:start position:0%
can see that the accuracy will increase
 

00:03:44.080 --> 00:03:45.830 align:start position:0%
can see that the accuracy will increase
as<00:03:44.319><c> the</c><00:03:44.560><c> model</c><00:03:44.667><c> [music]</c><00:03:44.799><c> learns</c><00:03:45.200><c> to</c><00:03:45.440><c> actually</c>

00:03:45.830 --> 00:03:45.840 align:start position:0%
as the model [music] learns to actually
 

00:03:45.840 --> 00:03:48.149 align:start position:0%
as the model [music] learns to actually
reason<00:03:46.480><c> and</c><00:03:46.799><c> get</c><00:03:47.040><c> more</c><00:03:47.200><c> and</c><00:03:47.440><c> more</c><00:03:47.599><c> accurate</c>

00:03:48.149 --> 00:03:48.159 align:start position:0%
reason and get more and more accurate
 

00:03:48.159 --> 00:03:51.509 align:start position:0%
reason and get more and more accurate
answers.

00:03:51.509 --> 00:03:51.519 align:start position:0%
 
 

00:03:51.519 --> 00:03:53.589 align:start position:0%
 
But<00:03:51.760><c> there</c><00:03:52.000><c> are</c><00:03:52.080><c> a</c><00:03:52.239><c> few</c><00:03:52.400><c> failure</c><00:03:52.799><c> modes</c><00:03:53.200><c> worth</c>

00:03:53.589 --> 00:03:53.599 align:start position:0%
But there are a few failure modes worth
 

00:03:53.599 --> 00:03:56.390 align:start position:0%
But there are a few failure modes worth
discussing<00:03:54.319><c> because</c><00:03:54.879><c> even</c><00:03:55.280><c> this</c><00:03:55.760><c> approach</c>

00:03:56.390 --> 00:03:56.400 align:start position:0%
discussing because even this approach
 

00:03:56.400 --> 00:03:58.630 align:start position:0%
discussing because even this approach
would<00:03:56.720><c> fail</c><00:03:57.040><c> most</c><00:03:57.280><c> of</c><00:03:57.360><c> the</c><00:03:57.599><c> time.</c><00:03:58.159><c> The</c><00:03:58.400><c> first</c>

00:03:58.630 --> 00:03:58.640 align:start position:0%
would fail most of the time. The first
 

00:03:58.640 --> 00:04:00.710 align:start position:0%
would fail most of the time. The first
failure<00:03:59.040><c> mode</c><00:03:59.439><c> is</c><00:03:59.680><c> that</c><00:03:59.920><c> because</c><00:04:00.239><c> the</c><00:04:00.480><c> model</c>

00:04:00.710 --> 00:04:00.720 align:start position:0%
failure mode is that because the model
 

00:04:00.720 --> 00:04:04.070 align:start position:0%
failure mode is that because the model
is<00:04:01.040><c> unconstrained</c><00:04:01.920><c> to</c><00:04:02.239><c> human</c><00:04:02.560><c> language,</c><00:04:03.760><c> the</c>

00:04:04.070 --> 00:04:04.080 align:start position:0%
is unconstrained to human language, the
 

00:04:04.080 --> 00:04:06.470 align:start position:0%
is unconstrained to human language, the
more<00:04:04.400><c> we</c><00:04:04.640><c> train</c><00:04:04.959><c> it,</c><00:04:05.439><c> the</c><00:04:05.599><c> more</c><00:04:05.760><c> it</c><00:04:05.920><c> can</c><00:04:06.080><c> drift</c>

00:04:06.470 --> 00:04:06.480 align:start position:0%
more we train it, the more it can drift
 

00:04:06.480 --> 00:04:08.710 align:start position:0%
more we train it, the more it can drift
away<00:04:07.120><c> from</c><00:04:07.439><c> being</c><00:04:07.680><c> similar</c><00:04:08.000><c> to</c><00:04:08.159><c> its</c><00:04:08.400><c> training</c>

00:04:08.710 --> 00:04:08.720 align:start position:0%
away from being similar to its training
 

00:04:08.720 --> 00:04:11.509 align:start position:0%
away from being similar to its training
data<00:04:09.519><c> and</c><00:04:10.239><c> starting</c><00:04:10.640><c> to</c><00:04:10.799><c> look</c><00:04:11.040><c> like</c><00:04:11.280><c> a</c>

00:04:11.509 --> 00:04:11.519 align:start position:0%
data and starting to look like a
 

00:04:11.519 --> 00:04:12.630 align:start position:0%
data and starting to look like a
language<00:04:11.760><c> that</c><00:04:12.000><c> is</c><00:04:12.177><c> [music]</c><00:04:12.239><c> completely</c>

00:04:12.630 --> 00:04:12.640 align:start position:0%
language that is [music] completely
 

00:04:12.640 --> 00:04:15.509 align:start position:0%
language that is [music] completely
different<00:04:13.120><c> from</c><00:04:13.680><c> English.</c>

00:04:15.509 --> 00:04:15.519 align:start position:0%
different from English.
 

00:04:15.519 --> 00:04:17.509 align:start position:0%
different from English.
The<00:04:15.840><c> second</c><00:04:16.079><c> failure</c><00:04:16.479><c> mode</c><00:04:16.799><c> that</c><00:04:17.120><c> doesn't</c>

00:04:17.509 --> 00:04:17.519 align:start position:0%
The second failure mode that doesn't
 

00:04:17.519 --> 00:04:20.949 align:start position:0%
The second failure mode that doesn't
always<00:04:17.919><c> happen</c><00:04:18.160><c> with</c><00:04:18.479><c> GRPO</c><00:04:19.600><c> but</c><00:04:19.840><c> can</c><00:04:20.160><c> happen</c>

00:04:20.949 --> 00:04:20.959 align:start position:0%
always happen with GRPO but can happen
 

00:04:20.959 --> 00:04:24.469 align:start position:0%
always happen with GRPO but can happen
is<00:04:21.280><c> that</c><00:04:21.519><c> the</c><00:04:21.840><c> diversity</c><00:04:22.960><c> becomes</c><00:04:23.520><c> so</c><00:04:23.840><c> low</c>

00:04:24.469 --> 00:04:24.479 align:start position:0%
is that the diversity becomes so low
 

00:04:24.479 --> 00:04:27.430 align:start position:0%
is that the diversity becomes so low
that<00:04:24.800><c> the</c><00:04:25.040><c> model</c><00:04:25.360><c> has</c><00:04:25.680><c> no</c><00:04:26.080><c> way</c><00:04:26.320><c> to</c><00:04:26.560><c> learn</c>

00:04:27.430 --> 00:04:27.440 align:start position:0%
that the model has no way to learn
 

00:04:27.440 --> 00:04:29.270 align:start position:0%
that the model has no way to learn
because<00:04:27.919><c> the</c><00:04:28.240><c> good</c><00:04:28.400><c> examples</c><00:04:28.880><c> are</c><00:04:29.040><c> too</c>

00:04:29.270 --> 00:04:29.280 align:start position:0%
because the good examples are too
 

00:04:29.280 --> 00:04:31.749 align:start position:0%
because the good examples are too
similar<00:04:29.600><c> to</c><00:04:29.759><c> the</c><00:04:29.919><c> bad</c><00:04:30.160><c> examples.</c><00:04:31.440><c> Something</c>

00:04:31.749 --> 00:04:31.759 align:start position:0%
similar to the bad examples. Something
 

00:04:31.759 --> 00:04:33.590 align:start position:0%
similar to the bad examples. Something
that<00:04:32.000><c> can</c><00:04:32.160><c> help</c><00:04:32.479><c> both</c><00:04:32.720><c> of</c><00:04:32.800><c> those</c><00:04:33.120><c> failure</c>

00:04:33.590 --> 00:04:33.600 align:start position:0%
that can help both of those failure
 

00:04:33.600 --> 00:04:37.749 align:start position:0%
that can help both of those failure
modes<00:04:34.080><c> is</c><00:04:34.720><c> KL</c><00:04:35.280><c> divergence</c><00:04:35.919><c> loss,</c><00:04:37.040><c> which</c><00:04:37.440><c> acts</c>

00:04:37.749 --> 00:04:37.759 align:start position:0%
modes is KL divergence loss, which acts
 

00:04:37.759 --> 00:04:40.790 align:start position:0%
modes is KL divergence loss, which acts
as<00:04:38.000><c> a</c><00:04:38.240><c> spring</c><00:04:38.880><c> that</c><00:04:39.360><c> pulls</c><00:04:39.680><c> the</c><00:04:39.919><c> model</c><00:04:40.320><c> closer</c>

00:04:40.790 --> 00:04:40.800 align:start position:0%
as a spring that pulls the model closer
 

00:04:40.800 --> 00:04:43.590 align:start position:0%
as a spring that pulls the model closer
to<00:04:41.040><c> the</c><00:04:41.280><c> original</c><00:04:41.680><c> one,</c><00:04:42.560><c> but</c><00:04:42.960><c> it</c><00:04:43.199><c> doesn't</c><00:04:43.360><c> pull</c>

00:04:43.590 --> 00:04:43.600 align:start position:0%
to the original one, but it doesn't pull
 

00:04:43.600 --> 00:04:46.150 align:start position:0%
to the original one, but it doesn't pull
it<00:04:43.840><c> closer</c><00:04:44.479><c> parameter-wise.</c><00:04:45.583><c> [music]</c><00:04:45.919><c> It</c>

00:04:46.150 --> 00:04:46.160 align:start position:0%
it closer parameter-wise. [music] It
 

00:04:46.160 --> 00:04:49.110 align:start position:0%
it closer parameter-wise. [music] It
makes<00:04:46.400><c> sure</c><00:04:46.639><c> that</c><00:04:47.040><c> the</c><00:04:47.360><c> output</c><00:04:47.840><c> distribution</c>

00:04:49.110 --> 00:04:49.120 align:start position:0%
makes sure that the output distribution
 

00:04:49.120 --> 00:04:54.310 align:start position:0%
makes sure that the output distribution
is<00:04:49.919><c> at</c><00:04:50.080><c> least</c><00:04:50.960><c> x%</c><00:04:51.919><c> let's</c><00:04:52.160><c> say</c><00:04:52.400><c> 5%</c><00:04:53.040><c> similar.</c><00:04:54.080><c> And</c>

00:04:54.310 --> 00:04:54.320 align:start position:0%
is at least x% let's say 5% similar. And
 

00:04:54.320 --> 00:04:56.469 align:start position:0%
is at least x% let's say 5% similar. And
this<00:04:54.479><c> is</c><00:04:54.560><c> how</c><00:04:54.720><c> we</c><00:04:54.960><c> can</c><00:04:55.199><c> get</c><00:04:55.680><c> essentially</c>

00:04:56.469 --> 00:04:56.479 align:start position:0%
this is how we can get essentially
 

00:04:56.479 --> 00:04:59.350 align:start position:0%
this is how we can get essentially
unlimited<00:04:57.120><c> intelligence</c><00:04:58.479><c> even</c><00:04:58.800><c> if</c><00:04:59.040><c> we're</c>

00:04:59.350 --> 00:04:59.360 align:start position:0%
unlimited intelligence even if we're
 

00:04:59.360 --> 00:05:01.830 align:start position:0%
unlimited intelligence even if we're
limited<00:04:59.680><c> by</c><00:04:59.840><c> our</c><00:05:00.000><c> data.</c><00:05:00.720><c> as</c><00:05:01.040><c> long</c><00:05:01.199><c> as</c><00:05:01.440><c> we</c><00:05:01.680><c> can</c>

00:05:01.830 --> 00:05:01.840 align:start position:0%
limited by our data. as long as we can
 

00:05:01.840 --> 00:05:05.030 align:start position:0%
limited by our data. as long as we can
verify<00:05:02.400><c> how</c><00:05:02.720><c> good</c><00:05:03.040><c> the</c><00:05:03.280><c> response</c><00:05:03.759><c> is.</c><00:05:04.560><c> Now,</c><00:05:04.880><c> if</c>

00:05:05.030 --> 00:05:05.040 align:start position:0%
verify how good the response is. Now, if
 

00:05:05.040 --> 00:05:06.870 align:start position:0%
verify how good the response is. Now, if
some<00:05:05.280><c> of</c><00:05:05.360><c> you</c><00:05:05.520><c> want</c><00:05:05.759><c> to</c><00:05:06.000><c> understand</c><00:05:06.320><c> it</c><00:05:06.560><c> even</c>

00:05:06.870 --> 00:05:06.880 align:start position:0%
some of you want to understand it even
 

00:05:06.880 --> 00:05:09.029 align:start position:0%
some of you want to understand it even
deeper,<00:05:07.600><c> I</c><00:05:07.840><c> would</c><00:05:08.000><c> really</c><00:05:08.240><c> recommend</c><00:05:08.639><c> you</c><00:05:08.800><c> to</c>

00:05:09.029 --> 00:05:09.039 align:start position:0%
deeper, I would really recommend you to
 

00:05:09.039 --> 00:05:11.830 align:start position:0%
deeper, I would really recommend you to
play<00:05:09.280><c> with</c><00:05:09.440><c> it</c><00:05:10.320><c> and</c><00:05:10.720><c> to</c><00:05:11.039><c> start</c><00:05:11.280><c> interacting</c>

00:05:11.830 --> 00:05:11.840 align:start position:0%
play with it and to start interacting
 

00:05:11.840 --> 00:05:14.230 align:start position:0%
play with it and to start interacting
with<00:05:12.160><c> problems</c><00:05:12.639><c> related</c><00:05:13.039><c> to</c><00:05:13.199><c> this</c><00:05:13.440><c> field.</c><00:05:14.080><c> And</c>

00:05:14.230 --> 00:05:14.240 align:start position:0%
with problems related to this field. And
 

00:05:14.240 --> 00:05:15.990 align:start position:0%
with problems related to this field. And
for<00:05:14.479><c> that,</c><00:05:14.720><c> I</c><00:05:14.880><c> can</c><00:05:15.039><c> highly</c><00:05:15.360><c> recommend</c><00:05:15.759><c> the</c>

00:05:15.990 --> 00:05:16.000 align:start position:0%
for that, I can highly recommend the
 

00:05:16.000 --> 00:05:18.150 align:start position:0%
for that, I can highly recommend the
sponsor<00:05:16.400><c> of</c><00:05:16.560><c> this</c><00:05:16.800><c> video,</c><00:05:17.360><c> Brilliant.</c>

00:05:18.150 --> 00:05:18.160 align:start position:0%
sponsor of this video, Brilliant.
 

00:05:18.160 --> 00:05:20.310 align:start position:0%
sponsor of this video, Brilliant.
Watching<00:05:18.560><c> videos</c><00:05:18.960><c> like</c><00:05:19.280><c> this</c><00:05:19.520><c> one</c><00:05:19.680><c> is</c><00:05:19.919><c> rarely</c>

00:05:20.310 --> 00:05:20.320 align:start position:0%
Watching videos like this one is rarely
 

00:05:20.320 --> 00:05:22.310 align:start position:0%
Watching videos like this one is rarely
enough.<00:05:20.960><c> And</c><00:05:21.120><c> in</c><00:05:21.360><c> Brilliant,</c><00:05:21.840><c> you</c><00:05:22.080><c> will</c><00:05:22.160><c> be</c>

00:05:22.310 --> 00:05:22.320 align:start position:0%
enough. And in Brilliant, you will be
 

00:05:22.320 --> 00:05:25.528 align:start position:0%
enough. And in Brilliant, you will be
able<00:05:22.479><c> to</c><00:05:22.639><c> learn</c><00:05:22.960><c> with</c><00:05:23.440><c> interactive</c><00:05:24.160><c> lessons.</c>

00:05:25.528 --> 00:05:25.538 align:start position:0%
able to learn with interactive lessons.
 

00:05:25.538 --> 00:05:27.350 align:start position:0%
able to learn with interactive lessons.
[music]<00:05:25.919><c> I</c><00:05:26.160><c> would</c><00:05:26.320><c> specifically</c><00:05:26.960><c> recommend</c>

00:05:27.350 --> 00:05:27.360 align:start position:0%
[music] I would specifically recommend
 

00:05:27.360 --> 00:05:29.189 align:start position:0%
[music] I would specifically recommend
you<00:05:27.520><c> to</c><00:05:27.680><c> try</c><00:05:27.840><c> the</c><00:05:28.080><c> course</c><00:05:28.400><c> about</c><00:05:28.639><c> clustering</c>

00:05:29.189 --> 00:05:29.199 align:start position:0%
you to try the course about clustering
 

00:05:29.199 --> 00:05:31.590 align:start position:0%
you to try the course about clustering
and<00:05:29.520><c> classification</c><00:05:30.639><c> as</c><00:05:30.880><c> it</c><00:05:31.039><c> covers</c><00:05:31.360><c> a</c>

00:05:31.590 --> 00:05:31.600 align:start position:0%
and classification as it covers a
 

00:05:31.600 --> 00:05:34.230 align:start position:0%
and classification as it covers a
diverse<00:05:32.160><c> set</c><00:05:32.400><c> of</c><00:05:32.560><c> topics</c><00:05:33.199><c> directly</c><00:05:33.840><c> related</c>

00:05:34.230 --> 00:05:34.240 align:start position:0%
diverse set of topics directly related
 

00:05:34.240 --> 00:05:36.950 align:start position:0%
diverse set of topics directly related
to<00:05:34.400><c> both</c><00:05:34.639><c> diffusion</c><00:05:35.280><c> and</c><00:05:35.520><c> language</c><00:05:36.000><c> models.</c>

00:05:36.950 --> 00:05:36.960 align:start position:0%
to both diffusion and language models.
 

00:05:36.960 --> 00:05:38.390 align:start position:0%
to both diffusion and language models.
So<00:05:37.120><c> if</c><00:05:37.360><c> you</c><00:05:37.440><c> want</c><00:05:37.600><c> to</c><00:05:37.680><c> become</c><00:05:37.919><c> a</c><00:05:38.160><c> better</c>

00:05:38.390 --> 00:05:38.400 align:start position:0%
So if you want to become a better
 

00:05:38.400 --> 00:05:41.110 align:start position:0%
So if you want to become a better
problem<00:05:38.720><c> solver</c><00:05:39.440><c> by</c><00:05:39.840><c> building</c><00:05:40.320><c> skills</c><00:05:40.720><c> in</c>

00:05:41.110 --> 00:05:41.120 align:start position:0%
problem solver by building skills in
 

00:05:41.120 --> 00:05:44.390 align:start position:0%
problem solver by building skills in
many<00:05:41.600><c> many</c><00:05:42.000><c> scientific</c><00:05:42.639><c> subjects</c><00:05:43.840><c> crafted</c>

00:05:44.390 --> 00:05:44.400 align:start position:0%
many many scientific subjects crafted
 

00:05:44.400 --> 00:05:47.670 align:start position:0%
many many scientific subjects crafted
from<00:05:44.800><c> worldclass</c><00:05:45.680><c> teachers</c><00:05:46.400><c> from</c><00:05:46.880><c> Harvard,</c>

00:05:47.670 --> 00:05:47.680 align:start position:0%
from worldclass teachers from Harvard,
 

00:05:47.680 --> 00:05:50.310 align:start position:0%
from worldclass teachers from Harvard,
Stanford<00:05:48.320><c> and</c><00:05:48.639><c> MIT,</c><00:05:49.600><c> you</c><00:05:49.840><c> should</c><00:05:50.000><c> go</c><00:05:50.160><c> to</c>

00:05:50.310 --> 00:05:50.320 align:start position:0%
Stanford and MIT, you should go to
 

00:05:50.320 --> 00:05:52.870 align:start position:0%
Stanford and MIT, you should go to
brilliant.org/galahhat

00:05:52.870 --> 00:05:52.880 align:start position:0%
brilliant.org/galahhat
 

00:05:52.880 --> 00:05:55.430 align:start position:0%
brilliant.org/galahhat
and<00:05:53.199><c> try</c><00:05:53.360><c> it</c><00:05:53.520><c> for</c><00:05:53.759><c> free</c><00:05:54.000><c> for</c><00:05:54.400><c> 30</c><00:05:54.639><c> days.</c><00:05:55.199><c> And</c>

00:05:55.430 --> 00:05:55.440 align:start position:0%
and try it for free for 30 days. And
 

00:05:55.440 --> 00:05:57.350 align:start position:0%
and try it for free for 30 days. And
because<00:05:55.680><c> you</c><00:05:55.919><c> used</c><00:05:56.160><c> my</c><00:05:56.400><c> code,</c><00:05:56.800><c> if</c><00:05:57.039><c> you</c><00:05:57.199><c> do</c>

00:05:57.350 --> 00:05:57.360 align:start position:0%
because you used my code, if you do
 

00:05:57.360 --> 00:05:59.189 align:start position:0%
because you used my code, if you do
choose<00:05:57.840><c> to</c><00:05:58.080><c> get</c><00:05:58.240><c> a</c><00:05:58.289><c> [music]</c><00:05:58.479><c> subscription</c><00:05:58.960><c> one</c>

00:05:59.189 --> 00:05:59.199 align:start position:0%
choose to get a [music] subscription one
 

00:05:59.199 --> 00:06:04.390 align:start position:0%
choose to get a [music] subscription one
day,<00:05:59.520><c> you</c><00:05:59.680><c> will</c><00:05:59.840><c> get</c><00:06:00.160><c> 20%</c><00:06:00.880><c> off.</c>

00:06:04.390 --> 00:06:04.400 align:start position:0%
 
 

00:06:04.400 --> 00:06:06.870 align:start position:0%
 
Something<00:06:04.639><c> about</c><00:06:04.880><c> how</c><00:06:05.120><c> GRPO</c><00:06:05.919><c> works</c><00:06:06.560><c> reminds</c>

00:06:06.870 --> 00:06:06.880 align:start position:0%
Something about how GRPO works reminds
 

00:06:06.880 --> 00:06:09.590 align:start position:0%
Something about how GRPO works reminds
me<00:06:07.039><c> of</c><00:06:07.199><c> how</c><00:06:07.600><c> evolution</c><00:06:08.880><c> doesn't</c><00:06:09.199><c> really</c><00:06:09.440><c> need</c>

00:06:09.590 --> 00:06:09.600 align:start position:0%
me of how evolution doesn't really need
 

00:06:09.600 --> 00:06:14.150 align:start position:0%
me of how evolution doesn't really need
a<00:06:09.840><c> clear</c><00:06:10.160><c> target</c><00:06:10.800><c> to</c><00:06:11.840><c> improve</c><00:06:12.639><c> organisms.</c><00:06:14.000><c> And</c>

00:06:14.150 --> 00:06:14.160 align:start position:0%
a clear target to improve organisms. And
 

00:06:14.160 --> 00:06:15.590 align:start position:0%
a clear target to improve organisms. And
it's<00:06:14.400><c> interesting</c><00:06:14.800><c> to</c><00:06:15.039><c> think</c><00:06:15.199><c> about</c><00:06:15.360><c> the</c>

00:06:15.590 --> 00:06:15.600 align:start position:0%
it's interesting to think about the
 

00:06:15.600 --> 00:06:18.390 align:start position:0%
it's interesting to think about the
parallels<00:06:16.160><c> of</c><00:06:16.400><c> the</c><00:06:16.639><c> failure</c><00:06:17.039><c> modes</c><00:06:17.759><c> like</c><00:06:18.160><c> lack</c>

00:06:18.390 --> 00:06:18.400 align:start position:0%
parallels of the failure modes like lack
 

00:06:18.400 --> 00:06:21.510 align:start position:0%
parallels of the failure modes like lack
of<00:06:18.639><c> genetic</c><00:06:19.039><c> diversity.</c><00:06:20.479><c> Anyway,</c><00:06:21.120><c> I</c><00:06:21.360><c> would</c>

00:06:21.510 --> 00:06:21.520 align:start position:0%
of genetic diversity. Anyway, I would
 

00:06:21.520 --> 00:06:23.189 align:start position:0%
of genetic diversity. Anyway, I would
love<00:06:21.680><c> to</c><00:06:21.840><c> hear</c><00:06:22.000><c> your</c><00:06:22.319><c> thoughts</c><00:06:22.720><c> in</c><00:06:22.960><c> the</c>

00:06:23.189 --> 00:06:23.199 align:start position:0%
love to hear your thoughts in the
 

00:06:23.199 --> 00:06:25.110 align:start position:0%
love to hear your thoughts in the
comments.<00:06:23.834><c> [music]</c><00:06:23.840><c> And</c><00:06:24.319><c> if</c><00:06:24.560><c> you</c><00:06:24.720><c> like</c><00:06:24.880><c> this</c>

00:06:25.110 --> 00:06:25.120 align:start position:0%
comments. [music] And if you like this
 

00:06:25.120 --> 00:06:26.950 align:start position:0%
comments. [music] And if you like this
video,<00:06:25.440><c> maybe</c><00:06:25.600><c> you</c><00:06:25.759><c> like</c><00:06:26.000><c> other</c><00:06:26.400><c> videos</c><00:06:26.639><c> I</c>

00:06:26.950 --> 00:06:26.960 align:start position:0%
video, maybe you like other videos I
 

00:06:26.960 --> 00:06:29.350 align:start position:0%
video, maybe you like other videos I
make.<00:06:27.280><c> So</c><00:06:27.600><c> check</c><00:06:27.759><c> out</c><00:06:27.919><c> my</c><00:06:28.240><c> channel.</c><00:06:29.039><c> It</c><00:06:29.199><c> is</c>

00:06:29.350 --> 00:06:29.360 align:start position:0%
make. So check out my channel. It is
 

00:06:29.360 --> 00:06:31.830 align:start position:0%
make. So check out my channel. It is
still<00:06:29.680><c> small,</c><00:06:30.080><c> so</c><00:06:30.800><c> YouTube</c><00:06:31.120><c> is</c><00:06:31.360><c> probably</c><00:06:31.600><c> not</c>

00:06:31.830 --> 00:06:31.840 align:start position:0%
still small, so YouTube is probably not
 

00:06:31.840 --> 00:06:34.070 align:start position:0%
still small, so YouTube is probably not
going<00:06:32.000><c> to</c><00:06:32.800><c> show</c><00:06:32.960><c> you</c><00:06:33.120><c> another</c><00:06:33.440><c> video</c><00:06:33.759><c> if</c><00:06:34.000><c> you</c>

00:06:34.070 --> 00:06:34.080 align:start position:0%
going to show you another video if you
 

00:06:34.080 --> 00:06:36.230 align:start position:0%
going to show you another video if you
don't<00:06:34.240><c> click</c><00:06:34.479><c> on</c><00:06:34.560><c> it</c><00:06:34.800><c> now.</c><00:06:35.280><c> And</c><00:06:35.600><c> I</c><00:06:35.840><c> hope</c><00:06:35.919><c> to</c><00:06:36.080><c> see</c>

00:06:36.230 --> 00:06:36.240 align:start position:0%
don't click on it now. And I hope to see
 

00:06:36.240 --> 00:06:41.160 align:start position:0%
don't click on it now. And I hope to see
you<00:06:36.319><c> again.</c><00:06:37.280><c> See</c><00:06:37.520><c> you</c><00:06:37.600><c> next</c><00:06:37.840><c> time.</c><00:06:38.160><c> Bye-bye.</c>

