Shadowing Practice: OpenAI fights back - Learn English Speaking with Video

Creando lección...
1
they're back okay let me catch you guys up quick last week anthropic dropped a new model opus 5 .5
2
and it was unbelievably good it was so unbelievably good
3
that openai rushed out two model drops gpt6 soul
4
and gpt6 luna you might have noticed i didn't do a video on those models there's a reason they weren't
5
that good i was not particularly impressed with either of them
6
and didn't really have much to say
7
but there was one other model i happened to get early access to
8
that is now available for all that model is called gpt 6 .1 sol
9
and this model made it very very hard to film
10
that sonnet 5 .5 video because i knew something was coming something surprisingly cheap surprisingly capable
11
and most surprisingly better than astra
12
so yeah i got a lot to say about this one i'm filming this at two in the morning, right after filming my Sonnet 5 .5 videos.
13
So pardon me for stumbling over a few words here and there.
14
I'm doing my best to get this out as reasonably quickly as possible because I want to have some coverage.
15
And I'll be real, it is also quite fun to cover these things before they are out.
16
So you're getting my true honest take and not the distilled version of what everyone else is saying.
17
I'm sure this model is going to cause some pretty crazy waves.
18
So it'll be nice to have my take out initially separately first.
19
As always, I feel obligated to remind you i do have early access
20
but i'm not being paid in any way shape
21
or form openai has no influence over what
22
and how i say things just
23
when they politely asked me to wait until the model is out to talk about it
24
which makes a lot of sense
25
but i have to wait for one other thing first today's
26
sponsor in order to build good software with agents they need to get feedback
27
and let's be real they're getting a lot of
28
that feedback from our ci that's why we've all been seeing our ci bills skyrocket
29
and also why we've been getting more
30
and more frustrated with github actions today's sponsor is depot
31
and they're here to solve all of this and more Not only can they make your CI up to 10 times faster, as well as your Docker builds up to 40 times faster, especially when they're downloading cache,
32
they're also cheaper, and they give better feedback for your agents.
33
All this is possible due to Depot Metal.
34
They're running their own bare metal with AMD EPYC processors
35
that are way faster than what you get from traditional CI providers like, of course, GitHub Actions.
36
If you want it to be a drop -in replacement, it absolutely can be, but Depot's APIs are so much better
37
that you should probably use those instead they enable parallelization
38
and most importantly resilience
39
when github inevitably goes down randomly for no good reason we've had our releases get blocked
40
because we weren't using depot and i'm so thankful
41
that i've been moving more and more stuff over for example
42
when ben moved picthing over to bun we immediately had some
43
ci failures normally this would be obscure piles of text
44
that our agents parse through for us but
45
when we use depot it becomes way easier to see they'll even analyze the failures
46
and give suggestions to make it much simpler to get this feedback back to our agents this is especially useful
47
when you tell your agents that you can use depot
48
because they'll no longer have to push changes and wait for
49
that to trigger a build they can just run the cli to trigger the exact same ci
50
that you'd be triggering through github instead no longer do you
51
have to file prs with broken code just to get feedback
52
to your agents they could just run a tool instead your
53
agents will also get way better breakdowns of what is taking
54
so long in your actual ci runs so
55
that you can figure out how to improve them and make them faster and more reliable.
56
You and your agents deserve faster Docker, faster build times, faster CI, better results, and ideally a cheaper price.
57
Get all of that and more at soydiv .link.
58
Let's talk about this model a bit
59
because it is not quite what I expected and it's probably not what you guys expected either.
60
Especially when you consider the GPT -6 sole just came out like a week ago.
61
It'll be around a one week gap from 6 .0 sole to 6 .1 sole.
62
I also want to disclose the numbers i'm currently showing on my screen are unlikely to be exactly accurate
63
because i am running terminal bench for myself in the first
64
two times i ran it i screwed things up the third
65
one seems to be doing much better i didn't run it on medium initially
66
so there's a miss it there but low high x high
67
and max although the max run is incomplete
68
so i'm currently backfilling scores from x high for the ones
69
that max either got wrong or in a previous run
70
or didn't do yet because it takes like eight plus hours on some of these tasks.
71
This bench is nuts.
72
I'm more skeptical of benchmarks than ever now
73
that I've been running a lot more of them myself in order to get the coverage I want to give here.
74
For what it is worth, Terminal Bench 4 is state -of -the -art score here, as is Deep SWE, although this one's weirder
75
because it goes down on X high and max and stays even on low and high, but those even low and high scores are scoring around what Astra did on high.
76
The difference being it's doing it for comically cheaper switch over
77
to the log scale you'll see what i mean this model
78
on low costs 21 cents versus aster on low costing a dollar 46
79
and opus 5 .5 on max getting the same score for 14 .65
80
while i will gladly admit
81
that deep sweet is far from a perfect measure of how
82
good a model is at day -to -day code work the fact
83
that 6 .1 soul is scoring the same as opus and is also 73x cheaper.
84
is at least worth noticing.
85
Here's where I'm going to say some of the things that I probably shouldn't.
86
Considering that GPT -6 Sol came out last week on Tuesday, and this model's coming out this week on Tuesday,
87
I think it's reasonable to infer that 6 .1 Sol was not meant to be 6 .1 Sol.
88
There are things I'm not supposed to say, and I'm definitely walking a thin line here by sharing it, so I hope this proves I'm not paid off by OpenAI,
89
because I'm about to give you guys info that I.. yeah, just let me get through this.
90
First and foremost, 6 .1 is significantly smarter of the GPT -6 Sol, thereby indicating this isn't just a .1 bump, there is something fundamentally different here.
91
Next point is that it has meaningfully slower tokens per second.
92
That tends to indicate the model is bigger, hard to know for sure.
93
Seems like this model might be different.
94
Most importantly, we have Tebow's tweet.
95
What am I referring to there?
96
Well, right before I started filming, Tebow dropped quite a wall of text.
97
The thing I want to emphasize here is first off that the pro $200 subscription is back.
98
But more importantly, more importantly is this sentence, they are changing how they calculate the usage in the sub.
99
In effect, if you do the math, it will net out at half the dollar in API spend compared to the old pro $200 plan.
100
Why in the world would they do this, especially right now where there's allegedly an internal code red
101
because opus 5 .5 is so unbelievably good
102
and has made the 200 cloud code sub such an unbelievable
103
value the only reason in the world tebow would post this
104
right now is uh i don't know maybe a new model is coming where their margins aren't as good
105
so the ability to subsidize has gone down
106
because remember you can get eight to nine thousand dollars of usage in a month on the 200
107
cloud code plan
108
and you can get over 12 grand on the 200 codex
109
plan i did actually run a lot of numbers before this
110
and the amount you could get on Astra did go down slightly closer to like eight grand
111
or so hard to know for sure because they differ for everyone everywhere
112
and it's hard to log all of this stuff
113
but for my math roughly nine grand a month the usage
114
and that's where the price for this model comes in this model is two dollars per million input tokens
115
and ten dollars per million out this makes it way cheaper
116
than 5 .6 soul was at launch half the price of 5 .6 Sol after discounts
117
and the same price as GPT -6 Sol, one -fifth the price of Astra.
118
However, this is not the whole story because cash reads matter.
119
And the cash read price for this model is going to be 10 cents per mil in.
120
That's a big deal.
121
OpenAI has not changed cash read price.
122
As far as I know ever before, it's always been exactly 10 % of the normal read price.
123
That makes it a 90 % discount.
124
And now it's a 95 % discount.
125
That means they cut the cash read cost in half massively reducing the cost for real world agentic use
126
which to be clear is what we're using these for most of the time
127
so this makes the model absurdly cheap for doing real world code work
128
that also means
129
that they are almost certainly cutting into their margins historically these
130
margins are rumored to be as high as 95 percent like
131
for every 10 you spend they only have to spend 50 cents
132
and as crazy as that sounds it makes a lot of sense especially
133
when you consider how expensive it is to make and train these models.
134
But that also gives them wiggle room to change things around a bit, which appears to be what's happening here.
135
That also means that if they were to keep subsidizing the same level that they were on the subscriptions, that your electricity cost for your sub would be more than you're paying.
136
So I get why they have to change this.
137
They've kind of just left the details out there for us to reverse engineer.
138
So take this as you will 6 .1 coming
139
so fast seems to indicate it is not just a new snapshot of gpt6
140
so let's talk more about this model as i was showing
141
earlier seems really good at terminal bench every time we refresh the numbers change
142
because new runs come in and it looks like max failed some things
143
that x high passed which is why i just dropped a bit
144
but again pretty much all of these even high
145
and x high are scoring higher than anything else ever has
146
and this is for me running this benchmark on random VMs on my network.
147
So not the best suite to test against.
148
I also had to drop three particular tasks from it because they expected an H100 to work against, which I make decent money.
149
I don't make H100 money, okay?
150
But none of this is real world code work.
151
So let's talk a bit about that.
152
Obviously, we'll have all the fun things like fish slap near the end.
153
So stay tuned for that.
154
But I just want to fixate a bit on the costs here.
155
Because the most expensive run with 6 .1 sol for me was about $1 .38 per task.
156
And the cheapest run with Opus 5 .5 was $5 .12.
157
That's a 4 to 5x gap from the cheapest Opus to the most expensive Sol.
158
So at this point, I would imagine you are hoping
159
and praying this model is good and that it can actually replace Opus 5 .5 for day -to -day work.
160
And I promise we'll get some good answers to that in a bit.
161
But first, we need to talk a bit about model behaviors here, because this model is a part of the GPT -6 family,
162
which means it has behaviors that are worth talking about.
163
I know I cite this diagram a lot, but there's a reason for it.
164
The thing that made me so frustrated with GPT -6 Astra wasn't
165
that it was less intelligent than the best models from Anthropic, because it was more intelligent than the best models from Anthropic, and I would argue in many ways still is.
166
But there is a problem.
167
It is also dumb.
168
It is smart and dumb at the same time.
169
GWT -6 Astra would just randomly spike into the dumbest bullshit I've seen a model do this year, even worse than some of the small open -weight models I play with.
170
It's still so deeply frustrating that Astra does this, because on the other end, when it does well, it's unbelievable.
171
But these spikes got to the point where I effectively churned.
172
I was only using my codec subs for computer use, and I ended up just leaning on to Fable 5 .1
173
and obviously now Opus 5 .5 for almost all of my day -to -day work.
174
So have they addressed the spikiness?
175
Has GPT 6 .1 Sol fixed the problems that I was so frustrated about with Astra?
176
I would say mostly.
177
Not entirely, but for the most part, yeah, this is a much better model.
178
Its peaks are not as high.
179
This is not the incredible revolutionary 3D capabilities that we saw with Astra.
180
In fact, I would put it slightly below 5 .5 Opus in most of those types of things.
181
It is not as good at computer use as Astra, although it is close enough to the point where I have
182
been happy using it for all of my day -to -day work.
183
I actually had 6 .1 Sol go through all of my emails
184
and find invoices that I had forgotten to pay or was behind on, mostly like investing stuff, and set up new tabs in Chrome for every investment I needed to wire,
185
fill out all the details for me, and just leave me to hit send.
186
It didn't get a single thing wrong and called out additional stuff
187
that I absolutely would have missed if I was doing this work myself
188
so I'm literally trusting this model to wire money for me it's trustworthy enough for
189
that and honestly I don't know if I would have trusted Astra with
190
that due to the spikiness six point one soul much much
191
less spiky from what I've heard from the other testers they
192
seem to agree with this analysis I know for a fact that Julius
193
and Ben who have also been testing have had a much
194
better experience with this than Astra in terms of the spikiness
195
Julius called the model incredible Ben called it incredibly boring I
196
think that's the best place you can be for a model drop like this
197
but as i had mentioned before its peaks are not as impressive
198
while it does quality work the majority of the time there are some tasks
199
that are just at the edge of its capability
200
that it will start to do weirder stuff on for the most part it's fine
201
but i i'm still reaching for opus a decent bit we'll
202
talk more about the comparison later i do default to this
203
model for a bunch of stuff though first off as i
204
mentioned before computer use i can't wait for ultra fast to be like an actual thing you can use with OpenAI models, because when it is, this model is going to be crazy on it,
205
because it can already figure out how to navigate computer use totally fine.
206
If it can suddenly do it six times faster, it's going to be unbelievably fun.
207
Still not quite as good as Astra, but more than good enough that for the price difference, I wouldn't even think twice about it.
208
But as I mentioned before, there are certain things I would still occasionally use Astra for that I am more than happy to use Sol for.
209
One of those things is deep code reviews.
210
I have still found OpenAI models
211
and the like Rottweiler nature where they'll dig into a problem
212
and shake it and tear it to pieces until they find every single thing wrong with it.
213
I find 6 .1 Sol to be incredibly capable in this particular way.
214
So as you can probably guess, I had 6 .1 Sol do some deep audits on Orchestrator V2 and other parts of my real world code bases.
215
In my Orchestrator V2 audit, it performed nearly identically to Astra.
216
I do believe it was slightly higher a score.
217
Okay, not in this analysis, but in my other analysis, it did actually score very, very slightly higher, but it did it at about half the price, $297 versus $584.
218
Sonnet was still cheaper, and Opus was slightly cheaper as well.
219
The difference being, neither of these models were anywhere near as thorough with their analysis.
220
6 .1 Sol dug deep to find things, which is why it was able to get a score comparable to Astra, although it did admittedly burn way more tokens.
221
Another task I've had a lot of fun testing with is asking the model to find opportunities to improve a code base.
222
In this case, to improve T3 code.
223
This is the one where Grok 4 .7 scored strangely well.
224
Of course, Astro scored way better at an 83 .8 versus the 80 .7 from Grok 4 .7.
225
But GBD61 Sol hit it out of the park with an 87 .4.
226
I didn't save all the prices for these runs.
227
It's been a bit, okay?
228
But for Opus 5 .5, it cost five bucks
229
and for sonnet 5 .5 it cost almost nine dollars with gbd61 solid was two dollars
230
and 15 cents that's the difference this model's token price is cheaper than sonnet
231
but its token utilization is still maintaining openai's usual efficiency which results in just crazy price to performance.
232
This whole thread was particularly fun
233
because I had Opus 5 .5 review this model with a different name obviously so I didn't know what it was.
234
I went and edited the history after.
235
And it concluded very quickly this was a frontier tier model.
236
Its reviews and bug repros matched the fixes that later merged.
237
Its first draft code had real bugs which review bots caught.
238
Four reviewers are still running.
239
The local T3 code reviews back frontier tier again in a
240
blind 10 model bench on the same prompt 6 .1 souls
241
placed first of the a7 .4 it found the fish slop runs
242
and compared those two it did say 6 .1 souls quality
243
output was slightly below astra's as well as the two open ai models with opus 55
244
and sona 55 which we will absolutely show you in a bit
245
but i do want to call out the price here
246
because it only cost seven dollars to run versus 15 for sona 55
247
and 50 for opus opus's honest tier call was
248
that this model is incredible for scoped work top of the frontier refine what's wrong
249
and tell me the truth i would choose it over astra
250
and about level with opus for a long unattended building this was below frontier follows its process rules even
251
when they stop all progress
252
and it does not ask for help this i absolutely noticed i had mentioned before well a few times now
253
that my ts rust port
254
that i'm making with opus 5 .5 is going way better than
255
when i was working on that same port using Astra and Sol in the past.
256
I had that port running for a while with this model and it made no progress.
257
It burned a shitload of tokens, but it didn't actually improve the compiler at all.
258
Opus was able to from scratch restart it
259
and get it working in a day after I had spent months
260
and hundreds of thousands of dollars in tokens with this model as well as with Astra and 56 Sol.
261
Opus did it in like a grand in like a night with just two subscriptions with the cloud plan
262
so for unattended long like heavy rewrite type stuff anthropic is just comically far ahead right now
263
and it also didn't have great judgment
264
when i was using it for managing my fleet
265
and for those wondering my fleet is all the computers i use for running all my agents
266
and code because one computer is far from enough i don't run any of them on this macbook now
267
so when i use this model to manage the fleet it made a couple dumb mistakes here
268
and there to be fair so is opus astra is the only one
269
that hasn't really made too many of those dumb mistakes
270
but like I'm gonna be
271
so real I am entirely done using Astra for this model
272
after I had Opus do all of this review I asked
273
it how much does it think this model should cost it
274
guessed $5 per mil in 50 cents cashed in 30 per
275
mil out putting it at Opus's prices roughly it said that
276
because it's performing like Opus its speed should add a premium
277
because it is quite fast and it's not a pro model
278
which is where it expects those higher like hundred dollar out tiered pricing things to come.
279
ProModel isn't really a thing anymore.
280
We just use Fable and Aster, but you get the idea.
281
This is the funniest part of the whole thread, though.
282
If OpenAI wants people to adopt it, I would expect $3 per mil in, $0 .30 for cashed, and $20 instead.
283
That would still be a fair price for what it does.
284
To which I responded, if I told you it was $2 in, $10 out, and $0 .10 per mil cash read, what would you think?
285
I'd call that very aggressive pricing.
286
For how you use it, it costs is about a quarter of what I guessed.
287
The cash price does most of the work.
288
Agent workloads are 96 % cash reads.
289
So 10 cents for cash reads matters more than the $2 and $10 headline prices.
290
When I looked at all of my sessions, its price guess would have been $5 ,700.
291
But after looking at these new prices, it redid the math and it would have been $1 ,550.
292
That is a massive decrease.
293
And for all my PR review type tasks, it was expecting those to be up to $10 and it's actually only up to $3.
294
And that's where like heavy PRs with 10s of 1000s of lines of code, according to Opus.
295
So don't blame me blame Opus for saying this.
296
First off Opus says it becomes the default model for scoped work.
297
Second off, it says that bloated system prompts barely matter anymore because of the cash pricing.
298
It's just noise.
299
Third, it says long loops are still a bad idea, but not because of money.
300
It's because according to it, the TSRL support wasted four days and made no progress at all.
301
And Opus even said they'd be suspicious of it lasting.
302
They expect this price to go up in the future.
303
I cannot fathom OpenAI ever increasing the price for a model, but Opus thinking they will is hilarious and shows just how good a value the model is.
304
I love this call out here.
305
GBD 6 .1 Sol did three rounds of work for about half the cost of Sonnet's single round.
306
A lot of this comes down to how context was managed, both because 6 .1 Sol is much more efficient so it's not doing as many calls that bloat the context.
307
It's not outputting as many tokens that are like building up over time.
308
So the average number of tokens being read per request was only around 110 ,000 tokens versus 360 ,000 for Sonnet 5.
309
The result is that Solout used under half as many input tokens as Sonnet making this model significantly more efficient.
310
Speaking of efficiency I want to talk about these deep SWE scores a tiny bit more
311
because this is a weird bench for me to have forked
312
and include in these things i actually did it for a different reason not to compare against 6 .1 sol
313
but to compare against a new release from open router
314
jev router open router added jev router to try
315
and optimize costs with your requests
316
and i thought it was an incredibly stupid idea once i
317
started running it against benchmarks i confirmed it's an incredibly stupid idea it turns out a model
318
that cannot reason
319
that is given a prompt into no context cannot make a good decision around how hard the problem is
320
and jev router ended up being deep seek v 4 .1
321
flash router for the vast majority of its runs around 60
322
of all the requests went straight to deep seek 4 .1 flash
323
so didn't like it that much it also routes to other smarter models
324
which should give it more of an advantage
325
but it ended up being more expensive than gpt6 astro was on low
326
while also taking four to five times longer
327
because six astro low took 4 .6 minutes
328
and jev router took 20 jev router's average task took 104
329
steps whereas gpt6 astros took 19 you get the idea it wasn't very good
330
but the whole point of jev router is
331
that it would be as cheap as possible to get a certain score
332
that was the promise on the tin whether
333
or not you believe them is up to you not me
334
i think it's bullshit regardless jev router was routing to deep
335
seek 4 .1 flash for the majority of its requests despite
336
jev router routing to the cheapest possible small open weight models
337
from whatever provider will give it away for free 6 .1
338
soul on low got the same score for an eighth the price.
339
OpenAI is here to destroy any wins anyone else has in terms of efficiency.
340
Completing this bench in 4 .8 minutes for 21 cents with
341
the second highest score I've ever seen on it is a massive achievement, tying Opus 5 .5, which took 50 minutes per task on max,
342
a tenth the time and a 70th the price for the same score.
343
If your work fits within the things 6 .1 Sonnet does well, you should probably use it for everything.
344
But if your work doesn't fit in it particularly well, you should probably keep using Opus, and maybe give Opus the ability to call 6 .1 Sol when it should for various tasks.
345
I'm almost certainly going to be setting things up so
346
that Opus 5 .5 can call 6 .1 Sol to do investigation work to try and root cause bugs,
347
to do analysis of codebases to figure out what things need to be touched and why, to help me triage real world work, to help me review the work that Opus does and more.
348
I'm kind of spoiling the ending here, aren't I?
349
I'm going to keep using Opus 5 .5 for now.
350
Before I explain why, let me do the thing that I'm most excited for.
351
Fish Slop!
352
The first thing you might have noticed is the inclusion of Slop in Fish Slop.
353
This model did the horrible thing I hate, where it surrounded the game in a bunch of absolutely garbage UI
354
and coming to this right after the 5 .5 sonnet demo hurts me deeply
355
because sonnet 5 .5 did not make graphics anywhere near this good looking
356
but at least it made a ui that was nowhere near this awful
357
and man do i wish the bad ui is where the
358
issue stopped i will turn on the sound oh god it's blaring
359
it is stunning looking the fish are some of the best
360
the model for the sub is way better the propellers work way better i'm gonna mute the sound
361
because that is looking pretty bad i haven't even heard it honestly
362
but damn like looks beautiful but
363
if you actually are playing it one of the first things you'll notice is
364
that the movement feels significantly worse than it does in either the Opus
365
or the Sonnet versions that I have demoed in the past.
366
Yeah, it moves jank.
367
It also has a significantly worse frame rate than the versions from the other models.
368
It does have higher graphic fidelity, so that makes sense.
369
The models here with the plants are significantly better than they were with the Sonnet version.
370
The derives of the fidelity of the extras in the tank is absolutely hilarious.
371
Like, yeah.
372
But goddamn, I'm so tired of the unnecessary text everywhere.
373
This model doesn't worse than almost any I've ever seen before.
374
We got take a breather, paused.
375
Your little world can wait.
376
Back to the reef.
377
Start start a new tank, slop01, feeder submarine, a little underwater chaos,
378
the big little goal, your little ecosystem, little fish become big earners, four meals, and they're all grown up.
379
Make the family a little bigger, a little golden overachiever.
380
There's so many of these.
381
There's like 20 plus of them.
382
And I promise you guys, as soon as I saw this, I took a screenshot, i sent it to open ai
383
and i crashed out in the slack
384
because i cannot fathom how they haven't fixed this problem this
385
model is unacceptably garbage at ui it has regressed again and
386
if you're looking for a model that can make front ends that don't suck
387
go spend your money somewhere else
388
because it should not be spent here this model sucks at
389
front end it sucks at design it has no taste
390
and you're gonna have to bring your taste yourself still
391
but it is admittedly really good at blender
392
if you give it like a screenshot of a thing you want it to model in 3d
393
and say hey you have blender over the cli go make
394
this it will it'll do a pretty damn good job
395
but i would never have it make the actual mechanics for my games
396
because it feels awful to play it also has like nowhere
397
near as much gameplay loop in fact the first time i
398
tried demoing this before filming it just randomly game overed as
399
i was like getting started in the first 30 seconds
400
and never like said why actually i think i technically beat it
401
i also want to like take this version and hand it to opus
402
or sonnet and say hey can you make this play better
403
because the graphics are good but the game sucks but
404
when you combine how cheap it was to make this
405
because like this was five dollars i think to generate that's pretty insane
406
and if you combine that with like ultra fast if that ever happens
407
suddenly you're going to be able to make a game in
408
a few minutes on demand we're actually now getting to
409
that threshold where game development is about to flip upside down
410
because of how models are finally understanding three -dimensional space
411
and the tooling necessary to do these types of things it's
412
happening as per usual i was not allowed to put the
413
code i wrote with this model inside of t3 code
414
or other open source projects during the testing window
415
so i had to use it exclusively on my internal projects like lakebed as well as for auditing other work
416
which means i mostly uses for auditing other work and I was very impressed.
417
This is a real PR I was working on to fix
418
a bug where my new little work tree like setup window
419
that would appear in a new thread in T3 code would disappear if you left and came back.
420
I had cloud code work on this but this problem went pretty deep
421
so I wanted to make sure that whatever solution I came up with was very very very well vetted.
422
While I personally still do not trust this model to write
423
the code I'm trying to land I absolutely trust it to review things.
424
Ignore the GPT -6 soul there, just a placeholder.
425
So when I had 6 .1 soul look through this, it found real problems that were entirely missed by Fable and by Opus.
426
First, it called out that follow -up messages can stay blocked after the agent starts, which is very annoying if you want to queue a message.
427
And also that recovered setup progress was disappearing too early.
428
It figured all of these things out with a combination of reading the code and analyzing it as well as computer use.
429
And it's able to prevent me from merging a real regression in t3 code.
430
So I literally just copy pasted those things to Claude and then told it to take another look.
431
Said better, but I'll still fix two small gaps.
432
Remember, I can't use this model of the code for this project at the time.
433
So I copy pasted that again over to Claude and it eventually got it good enough.
434
And then I finally merged.
435
But that is what I like this model for.
436
And I cannot wait to push its limits for actually coding.
437
Although I will say from the code that I did have the misfortune of reading, it is harder to justify merging this code than it is
438
for code from opus normally i would make you guys wait for the opus versus soul video
439
or the sonnet versus soul video
440
but i'll just spoil the details now i like using them in tandem
441
because i find soul to be way better at reviewing
442
and digging into the details but i find opus a more pleasant collaborator
443
and significantly better at actually implementing code without getting blocked constantly throughout its work
444
and even now with the rust rewrite of typescript i find
445
myself in a similar pattern where i have opus 5 .5 just going
446
and going and going making the code base work and work well.
447
And then I had Sol come in and do an audit.
448
And this is the funniest part.
449
I remember before I said that I had Sol and Astra working on that TSROS port for effectively months, I told Sol to come in,
450
and it got it from 83 .7 % to 100 % in under a day.
451
I was blown away by that, that it had somehow unblocked the work that Astra and Sol were doing as well as 6 .1 Sol.
452
And I was absolutely blown away by that, that it had taken the work that 5 -6 -Soul, 6 -1 -Soul, and Astra had done over months
453
and got it unblocked where it had been stuck for weeks and finished it.
454
I was much more blown away when I had 6 -1 -Soul take a look at that work and critique it.
455
And what it brought up was that of the 1 .8 million lines of code, 1 .3 million were not being used.
456
The reason was because Opus concluded all of the code from all the other agents was useless slop
457
that had no chance of being recovered.
458
And it chose to rewrite it from scratch itself in another crate.
459
So on one hand, the only reason the code worked was Opus.
460
But on the other hand, the only reason the slop was still around was also Opus.
461
So I had to have this model come in
462
and clean up the mess that other OpenAI models had made because Opus didn't even notice the mess was still there.
463
What I'm trying to say is this model absolutely has a place in your workflows.
464
It could probably even be your default coding model and you wouldn't have too many issues with it.
465
But I still find Opus to be a better collaborator overall.
466
That said, I have almost no reason to use Sonnet anymore because this will effectively take its place.
467
And you bet your butt the moment this model comes out, I'll be going and making adjustments inside of my clod config
468
because I already have it set up so
469
that I can use Sol inside of clod code
470
because I want to make sure Opus knows this is the model to have review its work
471
and investigate the things going on in the code base.
472
This is a damn good model, and I'm really happy to have it.
473
I wish we had something bigger, smarter, and more capable overall.
474
I was really hoping for something to truly dethrone Opus 5 .5 as my daily driver.
475
This isn't it, and I'm not planning on canceling any of my Claude subs as a result of this release, but I am planning on taking a lot more advantage of my Codex subs in my day -to -day work,
476
admittedly in Claude code.
477
This is an awesome release and an unbelievable price for what you're getting, but this does potentially mark the start of the end of the subsidization era, so make sure you're subscribed so that you can be here
478
when I cover all of that
479
and more god I hope this doesn't get me canceled online
480
I have no idea how others feel about this beyond like
481
a handful of early access testers I've talked to I legitimately don't know
482
if you are going to love it or hate it
483
or land somewhere between I will know in a few hours I guess
484
because it's a yeah it's three in the morning I am
485
going to go to bed now this is a tiring one
486
hopefully I did a good job let me know in the comments
487
and until next time peace nerds god I'm so dead All right

Vocabulario y notas de pronunciación para esta lección

Esta lección de conversación de nivel C1 se basa en el vídeo “OpenAI fights back”. Las palabras que más se repiten: model, opus, code, sol, price. Este vídeo tiene 487 frases y 6233 palabras para practicar shadowing. La parte hablada dura 30:54. El hablante habla rápido, unas 202 palabras por minuto, así que espera sonidos enlazados y reducidos. Solo el 81 % de las palabras está entre las 3.000 más comunes del inglés, por lo que el vocabulario es exigente.

Vocabulario clave de este vídeo

Las 15 palabras más avanzadas del vídeo, con su pronunciación y significado:

PalabraPronunciaciónSignificado
sonnet sustantivo/ˈsɒnɪt/soneto
token sustantivo/ˈtoʊkən/seña
router sustantivo/ˈɹaʊ.tɚ/router, enrutador
slop sustantivo/slɑp/puaj
depot sustantivo/ˈdiːpoʊ/depósito
anthropic adjetivo/ænˈθɹɒp.ɪk/antrópico
unbelievably adverbioincreíblemente
merge verbo/mɝd͡ʒ/mergear
frontier sustantivo/ˈfɹʌntɪə/frontera
audit sustantivo/ˈɔːdɪt/auditoría, auditaje
importantly adverbio/ɪmˈpɔɹ.tənt.li/importantemente
astrology sustantivo/əˈstɹɒlədʒi/astrología
frustrate verbo/ˈfɹʌsˌtɹeɪt/frustrar
rewrite verbo/ɹiˈɹaɪt/reescribir
admittedly adverbiola verdad es que, lo cierto es que

Phrasal verbs que vas a escuchar

PalabraPronunciaciónSignificado
call out verboespecificar
come out verbo/ˌkʌm ˈaʊt/revelarse, salir a la luz
end up verboterminar, ir a parar
figure out verboaveriguar, captar
bring up verbolevantar, alzar
clean up verbolimpiar
come up with verboingeniar, ingeniarse
fill out verbo/fɪl ˈaʊt/rellenar, cumplimentar

Frases que vale la pena repetir

Frases cortas y completas del vídeo que puedes reutilizar en la conversación diaria:

  • What am I referring to there?
  • I didn't save all the prices for these runs.
  • I'd call that very aggressive pricing.
  • I'm kind of spoiling the ending here, aren't I?
  • Said better, but I'll still fix two small gaps.

Gramática en este vídeo

Las estructuras que más usa el hablante, con las palabras exactas del vídeo:

EstructuraEn el vídeo
Present perfect continuous have/has been + -ing — una acción que empezó antes y continúawe've been getting · i've been moving · I've been running
Oraciones condicionales if + oración, will/would + verbo — una condición y su resultadoif you do the math, it will net
Present perfect have/has + participio pasado — una acción pasada que sigue importando ahorahas made · has gone · has not changed
Voz pasiva be + participio pasado — importa lo que ocurre, no quién lo haceis called · being paid · was not meant

Pronunciación a tener en cuenta

El hablante usa 115 contracciones y formas reducidas, como I'm, didn't, I've. Dilas en su forma corta, tal como las oyes.

  • Los sonidos de “th”: anthropic /ænˈθɹɒp.ɪk/, fathom /ˈfæðəm/
  • Los sonidos de “sh” y “zh”: benchmark /ˈbɛn(t)ʃmɑːk/, screenshot /ˈskriːn.ʃɑt/, subscription /səbˈskɹɪpʃən/
  • Palabras largas — cuida el acento: importantly /ɪmˈpɔɹ.tənt.li/, astrology /əˈstɹɒlədʒi/, unbelievable /ˌʌnbɪˈliːvəbl̩/, surprisingly /səˈpɹaɪzɪŋli/, collaborator /kəˈlæbəɹeɪtɚ/

Sonidos difíciles para hispanohablantes:

  • s + consonante al inicio — sin añadir una “e” delante: slop /slɑp/, screenshot /ˈskriːn.ʃɑt/, sponsor /ˈspɒn.səː/
  • /dʒ/ — no es la “y” ni la “ch”: merge /mɝd͡ʒ/, astrology /əˈstɹɒlədʒi/

Cómo practicar con este vídeo

  1. Escucha el vídeo entero una vez sin hablar y anota las palabras que no conoces.
  2. Empieza a velocidad 0,75×, haz shadowing frase por frase y vuelve a la velocidad normal cuando te resulte fácil.
  3. Grábate y compara con el original, prestando atención a palabras como sonnet, token, router.

¿Qué es la Técnica de Shadowing?

Shadowing es una técnica de aprendizaje de idiomas respaldada por la ciencia, desarrollada originalmente para la formación de intérpretes profesionales y popularizada por el políglota Dr. Alexander Arguelles. El método es simple pero poderoso: escuchas audio en inglés nativo y lo repites en voz alta de inmediato, como una sombra que sigue al hablante con solo 1-2 segundos de retraso. A diferencia de la escucha pasiva o los ejercicios de gramática, el shadowing obliga a tu cerebro y músculos de la boca a procesar y reproducir simultáneamente patrones de habla reales. Las investigaciones muestran que mejora significativamente la precisión de la pronunciación, la entonación, el ritmo, el habla conectada, la comprensión auditiva y la fluidez al hablar, convirtiéndola en una de las metodologías más efectivas para la preparación del IELTS Speaking y la comunicación en inglés en el mundo real.

Técnica de shadowing: lee la guía completa paso a paso →