Shadowing Practice: Why Coding Agents Keep Making Your Codebase Worse - Learn English Speaking with Video

Ders oluşturuluyor...
1
Oh my god, I ran out six Codex subscriptions in a week.
2
Isn't that amazing?
3
What did you ship, dude?
4
A slop code bench?
5
Oh, that's incredible.
6
Listen, have you seen this?
7
Do not mortgage your code base in the hopes that in a year, the model will come out that's smart enough to fix all your problems.
8
I'm never letting anyone merge anything with this pattern in it again.
9
Am I 99th percentile?
10
Am I being left behind?
11
Do less things in parallel.
12
The same thing
13
that was always true about engineering is still true in all
14
work today i'm joined by dex horthy ceo of human layer
15
the man who coined the term context engineering we talk about
16
most of what you see online for agents is completely fake
17
why token maxing doesn't work in the real world how benchmarks work
18
and which one you should look at how much faster you should actually aim to be with agents
19
and it's not 100x
20
and at the end he shares his agentic engineering workflow what is automated what is a software factory
21
and where does he still spend this time i saw a bunch of people commenting as well
22
that like astra got much better at computer use it got much better at like blenders
23
and like game slop and this kind of stuff um
24
but it didn't actually get that much better at like writing maintainable code
25
or like controlling complexity or like you know removing things
26
or like changing abstractions or whatever it is uh
27
and uh i think i think i think
28
that was my feeling a lot and like a lot of people had the vibe and
29
so one of the things we're talking about tomorrow is sort
30
of this uh benchmark i discovered uh called slop code bench
31
it was made by uh the lab of um gabe orlansky
32
he's like a researcher at the university of wisconsin have you
33
seen this no a slop code bench oh it's incredible uh it's an entire benchmark designed around proving
34
that like models make code bases worth worse over time really yeah okay
35
so up until
36
that point there have been like i would say like three
37
flavors of benchmarks there's like the core like sweet bench benchmark
38
which is like here's a tiny pr from a real open
39
source repo uh we'll rewind the repo to before the pr
40
was made we'll like pull the tests out of the pr basically say hey model here's the issue
41
that was filed yeah 10 years ago on this random ruby project
42
or whatever it is uh go like fix the issue
43
and then we would take basically any changes it made to the test we would throw them out
44
and we would drop new test like the test in from the benchmark case
45
and then we would run those tests against the models code
46
and then if it passed it would get a reward and
47
if it failed then it also has to like do it without breaking any other tests uh
48
but it's very very like small targeted like can you get
49
the test to pass yeah um obviously these are benchmarks they're different from verifiers
50
but it's like all the benchmark oh this model score is 95 on sweet bench coding is solved
51
and it's like that's not what real coding looks like
52
so you have this other chapter of things where it's like
53
okay cool instead of solving this one little pr it's like
54
rebuild all of it microsoft excel from scratch with every single feature
55
or rebuild all of redis from the specs with every single feature um
56
so have these like marathon benchmarks um those two are from
57
one called uh sweet marathon actually okay it was like four to 400 hour tasks uh
58
and then cognition introduced this idea of like yes okay we're gonna use much bigger tasks
59
and also we are going to have a judge model run afterwards
60
and give like contribute to the score on on the exercise
61
of like did this thing break any of our code base rules
62
or conventions or patterns or whatever it is
63
so it's sort of like telling it like hey look in
64
c++ we can't do logging like this it has to be like this
65
uh but none of those really map to how real software is built
66
which is they all give the model the entire problem up front yeah here's the thing solve this thing
67
and what slop code bench does is really interesting is it it gives the model a thing to build greenfield
68
and then it says cool now take that code base
69
and add this feature now take that code base
70
and add this feature now take that and
71
so they have these like challenges each has like three to ten checkpoints in it yeah and
72
so the problem is like what it's trying to demonstrate is
73
like as models continue to work unattended in the same code
74
base without knowing the whole problem up front it's like
75
if you know the whole problem up front you can architect for it
76
but without knowing the whole problem up front you have this
77
you have this problem where like the code gets worse
78
and worse over time and it kind of shows
79
that like the number of failing tests just like trends up for all the models
80
and like i think gpt 5 .5 codex
81
when they ran it in may got like a 15 on this oh yeah
82
and i just ran it i've been doing some like independent research
83
because they're they're a lot they're not super funded
84
so we've been kind of sponsoring some research there
85
and we did it for astra
86
and astra got like 16 .5 okay one point higher the ability to maintain like complex code bases over time
87
so anyway we're gonna talk about that because it used to be a vibe
88
but now there's like okay this i mean it's you know
89
it's one benchmark out of one lab it hasn't really been
90
adopted by a lot of model cards I think partially because it's really slow and expensive to run
91
but I think it's a fascinating sort of
92
you know proof
93
that models are still not great at maintaining like a code base over time do you expect
94
that to change that models will get better at code bases over time
95
or will we solve that with regards to harness engineering well
96
so here's the crazy thing is like we used to do
97
this thing in the like claude 3 .5 3 .7 maybe
98
like opus 4 days where it's like cool we're just gonna yes i know the code is sloppy
99
but by the time that matters that'll be gpt7's problem we're on gpt6
100
and we haven't made a lot of progress there
101
so it's like you can't guess when
102
that will happen it will probably eventually happen
103
but it seems like the labs are much more interested in
104
like hill climbing on other sorts of things like computer use
105
and other like frontiers
106
that like the last 10 percent of software maintainability may be like as hard
107
or harder than the whole first 90 percent
108
that we spent the last three
109
or four years on yeah interesting what are some of the
110
problems where you think the industry is headed like i heard
111
a lot of conversations yesterday about remote agent execution a lot
112
about multiplayer i still don't know exactly how that's going to
113
play out in practice like the theory makes a lot of sense yeah i get
114
that get it running off my machine get it running somewhere locally
115
or centrally where other colleagues can also look at the context
116
and contribute in essence but i haven't seen a company kind of succeed with
117
that i feel that's quite frontier stuff yeah yeah
118
and i mean a lot of the cloud agent vendors like you look at cursor cloud agents
119
and cognition and devon and things like this they've made good progress there um
120
but i think uh the the challenge we we work with
121
a lot of like larger enterprises our whole thing is like okay
122
if you have a hundred
123
or more engineers that's where like we have plenty of five person teams
124
that love human layer but
125
if you have a hundred plus engineers that's where human layer has like the biggest impact
126
because it is all driven around this collaboration
127
and those teams tend to hit a lot of friction
128
when building with cloud agents because they have 10 years of history
129
and hundreds of repos and lots of like opinions and asterisks
130
and special things about how dev environments are run
131
and how can the agent test its work and things like this
132
that become quite complex
133
so I do think the future is kind of remote background agents both for the collaboration angle
134
and also for the like hey I want to be able to close my laptop
135
and do this or it's like for small stuff
136
if i can kick off a copy change from slack
137
and i don't have to bother a developer with it like i
138
but i think i think setting up a like ticket to pr
139
or like slack to pr workflow is pretty straightforward for the small stuff right like ramp
140
and brex and all these companies they've all like published papers i'm like hey we built this background agent
141
and it solves 50 of our prs
142
or whatever um i actually think that's like not that exciting
143
because i've talk to people who've built that in you know a weekend
144
or a hack week or something the the bigger question
145
that i think um i'm especially excited about trying to figure
146
out a human layer is like how do we enable people to collaborate meaningfully on the bigger pieces of work
147
um where it's like yeah i could tell the agent to do this thing
148
and then i would get a pr out
149
and then it would probably be somewhat right
150
but it would have a bunch of issues then i'm talking back
151
and forth with the agent on pr comments or back in slack thread
152
or wherever wherever it is but it's like okay you went
153
and wrote the code and it's not good
154
and it's a lot harder to like basically like fix what
155
was already built agents kind of already committed to a thing there's a lot of context to think through um
156
and so our whole thing is like what teams have been doing for 30
157
or 40 years back when building a feature took hours or days not five minutes
158
is they would like get together
159
and they would align on what's being built either with the agent
160
or with each other
161
or both i mean the metaphorical whoever's actually going to build
162
the pr is going to talk to the senior engineer staff engineer
163
and just like have a conversation
164
and like okay here's the approaches we're thinking about this usually happens in like google docs
165
or confluence or whatever it is
166
but this whole shift left idea of like can we
167
and like you're never going to one shot it right there's
168
no there's no amount of effort where it's like okay we're
169
going to spend like five hours in a meeting building a perfect plan
170
and then we're going to give that to the agent
171
and the thing
172
that comes out is going to be perfect it's like no
173
there's still going to be issues with it like you can't capture everything uh
174
and the source of truth should be the code not some specification or plan
175
but can you in this planning process eliminate as many failure
176
modes as possible can you just you know from a 25
177
000 foot view of the problem like a one pager can i eliminate like 50 of the bad things
178
that the agent might have done not necessarily going to do
179
but could have done and then can we zoom into you know the 10 000 foot view
180
and eliminate another like 80 of bad solutions to this problem um
181
so that by the end
182
when the code comes out it's more likely to be you
183
know zero to ten percent rework required rather than you know
184
25 to 50 to 75 percent rework which is common with like
185
if you try to one -shot most larger agent prs like
186
people don't just say hey claude go do this thing they're sitting there working back
187
and forth if it's large which ingredients make it a successful plan in the end
188
because you mentioned okay what we've already been doing
189
when features are kind of more expensive or more time intensive to build we're doing that now more and more
190
I feel like people are leaning towards or experimenting with spec -driven frameworks, which I think is good as well.
191
But when we're talking about, okay, is this problem actually worth solving, is that then part of the plan or how we solve it from an interface perspective?
192
Is that part of the plan or more part of the spec?
193
Where are the lines?
194
Yeah, I mean, I think the words plan and spec are like kind of mean different things to everybody.
195
So I won't, I don't want to get up
196
and be like the spec is this and here's what it looks like and here's what it's for.
197
And the plan is this and here's how it should be formatted
198
and here's how it should be used like i i i
199
don't think i think every team's gonna kind of come to
200
their own kind of taste on this i think we have
201
some opinions on like what a good starting point would be
202
but we always
203
when we work with customers we always end up like okay
204
cool like let's add a little bit of customization to your
205
we call it a design doc rather than a plan
206
because it's like we separate out like okay cool where are
207
we going what's the end state versus like how do we
208
actually get there what are the changes we're gonna make um
209
but our design doc often is like we'll put some customization in there for for users
210
when we're kind of forward deploying with them
211
and just say like okay cool your design doc also needs
212
these four questions answered at the top before we get into like everything else
213
that it kind of comes by default yeah um i don't know
214
if that really answers your question i think personally it does as long as it's fine right
215
if if it's fine that teams will eventually define what is the plan
216
and what is the spec
217
and what works for our environment rather than following some standard
218
yeah in essence they have to be effective as a team
219
and i think it's very context dependent also organizationally
220
so in the end i think that's fine you mentioned humiliator is good
221
when the engineers are like in the hundreds
222
or even beyond at a bigger scale and for me i am wondering
223
when agentic engineering kind of breaks or at what scale so
224
if we're talking about 100 engineers
225
or my context i have a thousand engineers we're trying to
226
figure out how to kind of enable people in terms of
227
agentic engineering yeah one of the pain points i see is that we make models available.
228
We make onboarding as seamless as possible, give people access to internal services through their agentic tools.
229
And what I still have a challenge with, and I don't know how to solve yet, is how do I prevent everyone from using the biggest, most expensive model for everything?
230
Not just for where it's actually necessary, but for everything.
231
Yeah, we work with plenty of teams where Fable has just been banned.
232
Figure out how to make it work with Opus, and you can submit a special request for like temporary access to fable for like a specific thing
233
but you kind of have to sell it it's straight to
234
jail fable yeah um yeah i mean i think the reason
235
why we care about this space is one like i've been
236
building software factories my whole career basically uh for my first
237
job i was like i was just a back -end platform engineer
238
and i immediately identified like oh
239
that guy over there like one of the first engineers of the company
240
that maintains the deployment system and the rollout system
241
and the staging environment
242
and how like engineers create preview environments of like hey here's
243
my feature work in progress like hey pm here's a url
244
on the internal network like you can go look at what i'm working on kind of thing um
245
that seemed to be the most like impactful person in the company
246
and so that was kind of the work i was always drawn towards um
247
and what we're finding is like a lot of tools work great in this space uh
248
that if you're just building like an xjs app
249
or you have a mono repo with everything in it like there's a lot less friction
250
but i think like there's a lot more like flexibility
251
and thought called for in the case where it's a larger team with a lot of baggage
252
and opinions and dev environments are complicated and hundreds of repos
253
and like i don't know we we had a we had a i won't call it a complaint
254
but like a pain point
255
that some people were having is like well we really want
256
to get our our pms using ai as well product managers to build better prds
257
and of course like someone who doesn't have a lot of skill with ai they say go write this thing
258
and they get a 10 page document out
259
and they're like amazing here let me go give it to
260
the team the whole team is like oh my god what is this like ai psychosis slop
261
that you're making me try to read through yeah um
262
and part of that is just like giving the model better steering
263
and skills of like hey here's how to keep it concise
264
make sure you're writing for humans not for whatever uh and
265
so make it clear and make it visual
266
but it's also uh it would find
267
that a lot of the things proposed in these prds were
268
so detached from reality in terms of like how does the code base work
269
and what are the systems and so like a weird like tactical thing we found
270
that works really really well is like give a product person
271
who's technical enough to like check out the repos on their box
272
and things like this or like set up github mcp to go get repo contents um
273
and have them you know drop in a two sentence prompt
274
and then let let like a kind of pipeline take over
275
like cool here's the thing we want to build go go
276
research everything go find across all 200 repos find the seven that matter to this product feature, pull all that context together in one thing.
277
And then like interrogate the user, the product person in this case of like, cool, what problem are we solving?
278
Because all the things that go into a good product requirements, what problem are we solving?
279
How are we going to measure success?
280
Uh, like, how are we going to know that it's working or if we should throw it out and abandon this thing?
281
Cause everything in these bigger companies, everything kind of is an experiment.
282
Right.
283
Uh, and then going through like, cool, how should it work?
284
Should it work like this?
285
Or should it work like this?
286
Sort of the like grill me style thing.
287
Um, and so like, this is just one example of like.
288
that's not really built in and like some teams will build
289
that internally and they'll come up with these things slowly over time
290
and what we're trying to help people do is like speed run
291
that process it's just like
292
if you don't have any opinions about how to build good product specs with ai here's a thing
293
that will work for you know 80 of teams no matter
294
how many repos you have no matter how old your code base is um and
295
so it's it's basically like it's as much as anything it's
296
about raising the floor that's why like staff principal engineers they come to human layer and and say like, oh, like, I like this.
297
Like, I can kind of see this maps onto kind of how I was working, but I had no idea how I would ever be able to take how I was working
298
and give it to everybody else on the team.
299
Yeah, yeah.
300
I like that you highlight that.
301
The decision -making in the end, that's still what a human does, right?
302
You bring all the context and it's baked in the harness or people are self -solutioning kind of something similar.
303
And in the end, it tries to provide whoever is the user with all the context necessary to try and make decisions.
304
And likely, if they don't know the decision, they can still kind of spar or ask questions or kind of, in a grill me way, get to an answer in the end.
305
For me, when I heard that dark factories is kind of in my bubble, a lot of organizations are striving for dark factories.
306
How do we get there?
307
How do we do that?
308
And I still don't know if that's like a realistic outcome.
309
Because especially when someone gives me a magical answer and it's not what I want, I'm just going to say, well, it's not what I wanted.
310
But if I created this magical answer based on my decision, like I'm building my own conviction that this is the answer yeah
311
so i have way the buy -in is very different i
312
was sitting in a meeting yesterday it's all hype it's everyone
313
wants to impress each other everyone wants to claim they've figured
314
out the dark factory thing i i think most of it
315
is is bs okay i was sitting a meeting yesterday um
316
with like the head of a department for a very old company you know thousands
317
and thousands of engineers um
318
and talking with kind of principal engineers on the ground who
319
are like drowning in slop prs from people who you know
320
either just don't care don't have a strong software engineering background
321
and uh he's like so you're telling me that the code
322
that is written by these people using ai is bad
323
and the two principal engineers are like oh yeah
324
and he pulls up a pr it's like look at this 30 000 lines of slop
325
or whatever it is like all right well i gotta go talk to my friends at like really big company
326
and other really big company who have claimed
327
that like 90 of this org's code is written by ai now
328
and i got to figure out what what they're talking about
329
because you're telling me that that's like completely infeasible
330
and like i have to figure out what they're doing that's special
331
or i have to like figure out why what is the
332
agenda anyways the point is like there's this whole like i
333
don't know uh it's like a posturing game among all the top executives at all the top companies
334
and everyone in in big enterprises everyone's like am i 90th percentile am i 99th percentile am i being left behind
335
and this is like a thing that this conversation at every executive meeting
336
that happens is people are just like everyone writes down their
337
number of like where do you think we are from zero
338
to 100 among all other companies in our kind of zone
339
or size yeah um i talked to andre brestloff today actually yeah
340
and he coined something which i'm gonna steal he said this is kind of the instagram
341
phase yeah right now everything is fake god i love it everything is bells
342
and whistles and if you say well actually for me it doesn't work people say skill issue
343
because how dare you kind of try
344
and disprove my accuracy of what i'm portraying online as i disagree that's incredible i love
345
that i mean a lot of it happens on twitter
346
but yeah it's the same it's the same thing it's like
347
everyone is running around talking about how cool their factory is
348
and how many things they have
349
and how much adversarial review they're doing posting videos of the cool thing they built
350
and I don't know for a
351
while there was a thing of like how many tokens there
352
was people have gone off on this a little bit
353
but it was like the token maxing thing of like oh
354
my god I ran out six codex subscriptions in a week yeah isn't that amazing
355
and it's like what did you ship dude it's like oh I made like a bunch of games for myself
356
that no one will ever play
357
and I'm like yeah that's what I thought I've also done
358
that I mean it's fine it's
359
and it's trying to be like hey I made this cool
360
thing check it out I don't think it's okay to be like
361
if you're not maxing out six codex subs then you're a
362
bad engineer i think that's like hurtful to the industry
363
and it's like it's already there's already a lot of noise
364
and hype and confusion
365
and my my sort of crusade has always been like okay what actually kind of works what is practical
366
and like what can you what understanding can you build up
367
from like first principles rather than like working backwards from what you see on the internet yeah
368
if we give productivity a multiplier player.
369
Is 10x then realistic, or when should people be happy?
370
Is it like 2x, 5x, 3x?
371
Do you have a number?
372
I mean, if you think about it, like in terms of like economic impact in the world, I think every engineer going 2 to 4x faster would be transformational to human society.
373
The amount of value that we would get, the extent to which our cities would improve, the value of things would improve,
374
like just life would get better would be insane.
375
But everyone's going to miss out on that because they're too busy trying to get to 100x.
376
And I think 10x is possible.
377
Sorry, what were you going to say?
378
I was saying people are trying to go infinite.
379
We don't need engineers anymore.
380
Yeah, people are trying to go infinite or 100x or 10x or whatever it is.
381
And it's like, figure out how to go two to three times faster.
382
And then you can go do two to three times on top of that and two to three times up on that.
383
But if you try to throw out everything we've done, I mean, there's something to be said for like, don't get stuck and shackled to your traditions and like some things
384
that you used to do may not be necessary anymore but
385
if you go and throw out everything that we've been doing
386
and try to like rebuild all of engineering from scratch it's not going to work it's just like
387
and when an engineer sets to do a project
388
and they're like ah yes i'll architect it to be able to do this this this this
389
and this when all you needed was this one small thing but
390
and what is it the jagney you ain't gonna need it
391
it's the same thing is it's just the same story over
392
again is like people want to have this perfect match
393
and there is a time
394
when you're like cool we've kind of like taped this together enough
395
but now we understand the problem now we can go build the next thing
396
and too many people are trying to skip
397
that step trying to skip the like okay let's just get faster
398
and faster and faster incrementally over time because that compounds exponentially
399
but um I was talking to Dario this morning
400
and he said from his perspective what makes his team at
401
AMP really productive has nothing to do with AI is they
402
don't do pull requests anymore yeah I heard about this yeah everyone just pushes domain exactly
403
so no pull request fatigue no no one wants to do
404
code reviews anymore still do code reviews i'm like huh maybe
405
it's time for me to run an experiment to see
406
if it's actually now's the time to do something like
407
that i think it's worth trying i mean the thing about
408
pull requests is like every team's pull request culture is different
409
like on my first job everyone had merged on main
410
and there was no pull request requirement
411
that was enforced it was a cultural thing it was like
412
you don't merge without getting a review from somebody else on the team
413
and then culture beats process every single time and so
414
if you have a team where like okay we have a
415
cultural thing of like you don't push slop to main
416
and it's like you don't you don't shirk responsibly well the reviewer reviewed it too
417
so we're both guilty
418
if it's bad it's like no you own it you are the only person who owns it
419
and if someone finds it and like realizes that it's bad
420
or you realize
421
that it's bad later like it's still on you it's just
422
uh we you know we're changing we're changing the process of like how do we guarantee quality
423
and what are the outcomes and how do we manage risks
424
and things like this i feel like to get to
425
that culture that could be a fun experiment right you remove the pull request
426
and see if because you're doing
427
that you can get to a culture of we own what we push to production
428
or to main yeah we don't push any slop especially with
429
agents how easy it is nowadays you take ownership see if
430
that kind of is a cycle
431
that then starts to roll do you did he mention like is is it
432
that pull requests aren't required
433
or like pull requests are not just not a thing it's is killed
434
because i could i could see a world where like okay look i don't require review
435
but i want to review on this yeah we have this
436
for doc review right like i don't recommend anybody say like
437
oh design doc reviews are required before you write the code
438
if you didn't get your design doc review the pr is not relevant
439
and it's like no i'm incentivized to send my co -founder
440
who's the cto is like the kind of slop the slop gatekeeper
441
and make sure that like the code base stays maintainable over time
442
and things like this i send him my design docs
443
because i know it's a lot faster for him to be like no no no on it's easier for him
444
and easier for me for us to have
445
that discussion on a one page or a two pager
446
and then when we get to pr time like after i've polished it
447
and i've got the test working and i've maybe like done some stuff
448
that wasn't in the plan of like okay let's just change this ui
449
and like riff a little bit over here um
450
and then he says oh this entire architecture is wrong
451
or like i don't like that this was a separate table
452
and it should be this instead like
453
so I think there are incentives to get feedback from people as early in the process as possible
454
and like it's easier to fix things before you merge it
455
than after yeah once it's in there then it's like okay
456
cool now we're like fixing bugs in production yeah 100 percent yeah how's your thinking evolved
457
when people said we're no longer looking at the code
458
or doing code reviews specifically we have so many deterministic tools
459
that look at code conventions
460
and now we even have built -in verification layers where the agent just sends me screenshots
461
or sends me videos about the functionality that it's being built
462
so i don't necessarily look at the code anymore well there's
463
two reasons to look at the code one is to make sure it's correct
464
and one is to make sure it's maintainable i mean there's
465
a bunch of other reasons there's like alignment with your team
466
and things like this but like in that world like the the video
467
and the screenshots tell you whether it's correct in terms of
468
like the features there doesn't tell me anything about whether the code is maintainable
469
and i have seen many many deterministic litters we use a
470
bunch of I'm going to talk about some of them actually tomorrow during my talk.
471
that help like block large categories of slop um
472
and you can do things like cyclomatic complexity
473
that is like tries to like keep you from creating functions
474
that are too chaotic and hard to grok for a human
475
but i have never seen a prompt or a deterministic tool that can like drive code -based
476
maintainability and fight the model's tendency to write unmaintainable code in any meaningful way without a human in the loop.
477
Do you think we can close that gap with deterministic tooling?
478
That would be, you would make a lot of money if you could figure that out.
479
And too, this is an interesting company in that space, right?
480
Because they're doing deterministic tooling for correctness, but not for maintainability.
481
No. But you want maintainability.
482
You want maintainability uh i use uh matt pokok has a skill called improve code base architecture yeah
483
which i use quite a bit and it does it does help um
484
and there is there are times where i'll just like run this skill
485
and then like cool i'm not even going to read the
486
the discussion just go make the first three fixes there
487
because i'm sure it's probably better than what was there in the first place
488
but i think in general like that the thing that makes
489
that skill useful is is having a human in the loop
490
the ability to be like no i don't like i don't want to refactor it
491
that way i want to refactor it this way and Like I think humans are good at this,
492
uh, skilled experience software engineers are good at this because they have like been bitten by bad architecture.
493
So many, they've been up at three in the morning trying to fix an issue and been like, Oh my God, I hate this pattern.
494
I'm never letting anyone merge anything with this pattern in it again.
495
And you do that hundreds and hundreds of times over the course of, you know, five, 10 year career, you build up this taste.
496
Uh, and we haven't figured out how to put that taste into models. And
497
so for now humans have to do it we might try
498
to figure it out uh literally the last model came out the astro came out
499
and my buddy ian who ran a built a coding agent startup
500
that got acquired by expo uh he posted he's like i
501
am begging the rl researchers at the lab please start including tasks
502
that have like deletion of code as the goal
503
or like hey let's take these three things
504
and refactor out a good abstraction like it's just so obvious
505
that like either we don't know how to do it
506
or we don't care to do it in the in the rl like layer uh
507
so i it could get cracked i mean anything can happen uh
508
but again also is like cool there's all this hype
509
and what is the future and no one knows
510
when it's going to happen
511
and my advice is always like do not mortgage your code base in the hopes
512
that in a year or six months
513
or 18 months a model will come out that's smart enough
514
to fix all your problems yeah someone listening might think i fully agree with the taste aspect
515
but i don't have 10 years under my belt I don't have 15 years under my belt pre -agentic tools
516
or pre -agentic era how do people still build taste from your perspective nowadays and how can they do that
517
I mean I think I think it comes from from burning yourself
518
and stubbing your toe I think it can also come through mentorship
519
and this is a thing I think about a lot
520
and I don't have a good answer for
521
which is like how do we bring on the next generation
522
of software engineers it used to be through pull request review right it's like
523
and like you think about like the goal of a the
524
worst pull request comment is like this is wrong i rewrote
525
it for you here you go like the person learns nothing uh
526
and they learn that
527
if they write bad code someone will just come fix it for them yeah
528
so a good pull request for a good senior engineer
529
or manager will be like hey i don't think this is
530
quite right can you like look at how we do this over here yeah
531
and like do that pattern or why did you do this yeah like try
532
and probe exactly what's the logic behind this why why why
533
not you know whatever whether you give them the answer
534
or not it's like okay why do we do this can
535
you is there any other places in the code base where and
536
that forces you to learn right you have to go read
537
other code oh that's okay cool that's cool i like that
538
and it like you reinforce uh and now if you give comments like
539
that that used to be really helpful they're not
540
because i i mean i talk to people all the time who are like yeah i'm commenting on this pr
541
and i know that the person is not they don't understand the code they wrote
542
and they're not really going to understand the comment
543
and they're just going to paste the comments to their ai
544
and it's going to make the fixes and at
545
that point like i'm doing every part of your job
546
that matters you didn't think about or understand the code before you committed it
547
and shipped it you just love the agent do it
548
and then like you're not even thinking
549
and under you you have the option some people will
550
but like you have the option to sometimes it just feels like i'm just talking to your agent yeah at
551
which point like i might as well just have my agent do it zero zero value on top yeah um
552
so i don't know i've seen mitchell hashimoto do this actually
553
a couple times on on ghosty prs where he doesn't even say hey this needs to be like
554
that over there he says ask your clod why this is wrong that's it he gives no hints
555
or anything yeah he's just like I will tell you
556
that this is wrong I won't tell you why I'm forcing you to whether by understanding the software
557
and building taste in software engineering or by understanding agentic coding
558
and being good at using agents to understand the code base
559
whatever it is like you need to go you need to go do some more homework yeah So I think, I think that's directionally correct.
560
Um, I also think there's going to be a lot more the same way we collaborate on code
561
and pull requests right now as the, as, as design docs and plans become more standardized and formalized, or at least their purpose,
562
parts of them become more universal.
563
There will be a lot of training and coaching that happens during design because design doc is basically a prompt.
564
And so two people, five people commenting on a design doc, you're teaching the author how to prompt better it's like okay
565
cool like I don't like this section like you know you need to go find other patterns
566
or whatever it is and that
567
that there's still a flavor of this where it's like okay I'm just gonna give it to my agent
568
and we're gonna like I'm not gonna learn anything
569
but like I think like learning how to build a really good design doc
570
which again learning how to build a really good prompt for the whatever session is gonna be building
571
or in that session um is uh is going to emerge as well as like just as important
572
if not more important than code review i feel like i
573
learned a lot through kind of defining solutions in a team early in my career
574
because i came from not a traditional computer science education i got rolled into a team
575
because i said i want to do engineering not just operations
576
because i was responsible for operations and then i learned the process of refinement
577
and in that team specifically we had the convention of we also define the solution direction not always
578
but most of the time and sometimes as part of executing on the work
579
that needs to be done the person executing would define the solution direction
580
and align on the on the way but for the ones where we did
581
that in advance i love just listening and learning
582
and asking questions on y x y and z
583
and thinking back now likely whatever we captured back then in
584
jira would have been great food for an agent to then execute on
585
and even the person picking up that task of the work
586
that needed to be done didn't have to do too much thinking anymore of exactly kind of how to execute
587
or what to execute maybe how still because you always have new findings
588
but mainly why and what was answered which i loved
589
and i learned a lot from now it's kind of up to the individual i feel like
590
and my fear is that because it's
591
so easy to be on an island with your agent
592
and go into execution mode create slop
593
or create something valuable it's really up to the user that you lose that that kind of collaborative effort and also interaction,
594
I feel like we'll miss that if we don't put extra effort in that.
595
Yeah, I mean we're talking about building a feature in the
596
human layer where you can call someone on your team
597
and do a voice call and then the transcript becomes an artifact attached to the task because that's the last piece, right?
598
I think the AI native forge is much more than GitHub that manages code.
599
It's going to be the code, it's going to be the artifacts
600
and it's going to be the sessions themselves all kind of linked together at like the top level
601
and the last part is like the human conversation and
602
so like maybe that's part number four is like can you reify those things for the agent
603
but yeah I totally agree it's like a lot of the learning happens
604
in those human conversations with the people who have that learning
605
and like if we can't rely on the models to absorb that taste that
606
things then we have to find other ways to give them
607
to people yeah yeah as a last thought i wanted to
608
zoom in on kind of what your agentic engineering workflow looks like nowadays yeah
609
because i feel like people listening they might see something online all instagramified i'm going to keep using
610
that and keep echoing that because i think it's quite accurate and
611
if they ask a question someone might say skill issue i
612
have no clue what's real what's not real what works for you nowadays um yeah
613
so we do
614
and i i did a pretty fairly deep breakdown on this
615
on um i'll give you a link to put in the show notes
616
or whatever but i basically have a page
617
that is all the podcasts we've ever done all the all the written content
618
and video content we've done
619
and i've done a couple deep dives on our process with screen share
620
and stuff awesome i would say probably at least 50 to
621
60 percent of our work is like basically one -shotted uh not one -shotted like to perfection
622
but like we don't do any planning up front
623
and it's literally just like
624
if it's going to be under a couple hundred lines then um then we just do
625
that about 10 percent of our prs are merged by you know call them like janitor bots
626
or whatever that run overnight my co -founder gave a talk on like control loops
627
and how we build all of this but we basically have like things
628
that we want to be true about the code base like desired end state
629
and then we have like a process to progress the code base
630
that direction it's kind of like a thermostat right okay i want it to be 71
631
or whatever the european version of 71 is uh
632
and it's 69 cool we're going to turn on the heat um and
633
so it's like we wake up every morning to a handful
634
of prs along different dimensions uh of cleaning up something in
635
the code refactoring fixing a react doctor rule migrating an endpoint
636
from the old framework to the new one everything we do automatically in the factory is like very small
637
because we want it to be digestible
638
if the pr is too big it never gets merged um
639
that's probably like 10 of our prs again about 50 are
640
just like hey i want to make this small fix go do it
641
and then i'll test it and maybe like send a couple more prompts on it um Um...
642
Does your day then also start?
643
Wake up, look at the PRs that ran overnight were created then?
644
Yep.
645
We don't merge them every day.
646
It has like a Kanban thing where like if it doesn't get merged, the thing doesn't run the next day.
647
There's always only for each type of PR of which there are like three or four or five, maybe six now.
648
There's only a max one open at any time.
649
And then we use labels to basically say like, okay, if there's already open PR with that label, don't try to fix that thing
650
because you might just try to fix the same thing again yeah um
651
and then probably about 30 of our 30 of our like code shipped it's probably fewer prs
652
but about 30 of what we ship is done doing through a like specific process
653
that we use called rpi
654
or like the new version of it is like crispy although
655
like now it's like qr do i cordoy i don't know anyways we've lost uh i stopped using acronyms
656
but it's all i would count it all in like the rpi family
657
which is like go do a bunch of code based research
658
um take what the user asked for go build a query plan then do another session
659
that doesn't know what the ticket was have that do the research
660
and then like one or two or three sets of like planning sessions uh
661
if it's a very ambiguous like large thing we'll do like
662
a product requirements first otherwise we'll get straight into the technical design
663
and then if it's fairly small
664
if it's like it's it's too big to one shot
665
because it just needs a lot of context from a lot of places
666
but it's like fairly small we might skip the actual like
667
here's the steps it's like cool just go implement this design uh
668
or if it's quite large
669
and i want to check it along the way i'll work with a model in a third type of document
670
that is like breaking down what order are we going to do this in
671
and like how am i going to check it along the way
672
because like it's a lot easier to ship 10 000 lines of code
673
if you check it every thousand lines uh
674
and like before it gets out because
675
if you get to 10 000 lines of code
676
and it's broken uh then you have a massive surface area to debug across
677
but if you can get it you can get a 90 in the first thousand
678
and like fix that
679
and then do the next thousand okay that's okay the next
680
thousand fix ten percent it's like much less overall work
681
and i think much less mental effort as you go you're
682
like checkpointing along the way to get to the final solution
683
sometimes we even do them as individual prs like well what's the plan for stacked prs on this
684
and like what's the first mergeable bit that we can do yeah uh
685
and then we go implement
686
and like use we use sub agents we sometimes we'll use like you know fable input orchestrating soul
687
or we use soul orchestrating fable
688
or we use just opus as the as the main thing
689
and opus as the sub agent like kind of depends on
690
the task i use soul for almost everything now for user
691
choosers um yeah the user chooses what the parent model is
692
and then there's a couple there's a couple like less obvious
693
hacks you can do to do like cross provider uh like
694
if you want fabled orchestrate sole sub agents you have to
695
tell it to use the codex cli right now um we do have a custom harness
696
that will support that soon uh because we just think it's a cool pattern
697
that works well um but uh yeah
698
and then from implementation you know it's review the code as you go
699
or review the code at the end uh make sure it's
700
good we have a kind of an internal cultural rule is like do not send anyone code for review
701
that you yourself have not read like do a pass on
702
it don't just make me read your agent slop sometimes i read the code
703
and i have to change 50 percent of it sometimes i have to read the code
704
and i actually throw it all out and go back to the design
705
because i'm like oh kyle's never going to approve this okay what's a simpler way to do this yeah
706
and sometimes i read the code and it's
707
and it's perfect yeah i think people underestimate
708
that you can just throw out the output
709
and try again well that's why also i love this like incremental like zooming in planning process like product tech
710
and then like the plan of the steps that we're going to do
711
because at any point you're you're getting more
712
and more specific at any point you can be like oh
713
i don't like this okay throw this out we're gonna go back to design we're gonna change some stuff
714
because if we have to do that then then i don't want to do it
715
do you ever see a world where
716
because this is also kind of in my bubble what people
717
that are in leadership teams are envisioning yeah
718
that a lot of the day is working on more of the planning
719
or specification how that is defined within the org or the team and then overnight the work executes
720
yeah i mean i just i i don't want to wake
721
up to 10 000 lines of coke it's like not like
722
if i'm if i'm
723
if i'm planning on reading the code i think the work
724
is a lot more human in the loop than than
725
that um they're very specific things that lend themselves to hill climbing overnight um
726
but i think a lot of the work you do you're
727
just gonna you're going to enjoy it more you're going to
728
get more done you're actually going to have more through because
729
that that also requires you to do a ton of paralyzation of course it's like plan 10 things
730
and then kick them all off overnight and then
731
when i'm gonna spend the first four hours of my day
732
reviewing code like no i'm gonna plan two things i'm going
733
to start them now i'm going to like watch them as they go
734
and i'm going to get changes landed in production today it's
735
much better to do two things two things two things than like okay i have 10 things in flight
736
and they're all kind of going and oh i forgot about two of them
737
and by four days later i did ship two of them
738
but the other three are still in flight it's just the same thing
739
that was always true about engineering is still true and all work is like
740
do less things in parallel it feels more efficient to like uh
741
I don't know how to say like to like try to like multitask really hard uh
742
but your end -to -end throughput of like value delivered to the customer will be faster
743
if you do less things like a little bit slower
744
and just like take a little bit more care with them interesting I wonder
745
if the way we get access to models like
746
if I look at myself my own license with cloud code
747
yeah I spent 200 bucks a month yep at least uh
748
I have weekly limits for my models yep able has a little bar
749
that I want full
750
or close otherwise i don't feel like i squeezed enough juice out of my license yep
751
so when it's friday and i know it's saturday
752
and it's reset time um
753
and i still have a lot of work to do that's really
754
when i kick out these night processes where sometimes i don't
755
really care about the code i want to parallelize i want to maximize my usage yeah um i wonder
756
if that way of thinking or that way of execution has also primed then
757
and kind of catapulted us into maybe this is what the
758
future of engineering is going to look like the only way we got here is
759
because of these limitations on how we can access models in the first place well
760
and this is like also like the biggest disconnect of like
761
everyone who's posting their like maxed out all six subs on
762
twitter versus like the real world is like no enterprise has subscription plans they're all paying per token
763
and so there's this whole cultural thing of like
764
if you can get all three thousand dollars of spend out
765
of your two hundred dollar thing then you're winning i don't know what you built
766
and like what what value it is and who's going to use it
767
but like that's how you win
768
and it's like no the skill issue is not like couldn't use enough tokens like uh
769
because that doesn't apply to most people who are doing serious work we have companies like customers with 15 engineers
770
and they are all they're all on non -subscription plans
771
because they need security requirements they need data retention requirements they need all this stuff
772
that you can only get on enterprise plan which means you
773
and get paying per token uh
774
so yeah i i i think the we're going to look
775
back on the era of subscription maxing i don't know
776
if it's like you know the intelligence is going to get too cheap to meter it's going to get
777
so cheap that they won't need subscriptions you'll just pay for it
778
and it'll be good value
779
or the subscriptions are going to go away at some point uh
780
because all it is is a loss leader
781
and like a way for them to get data on you
782
and like train the next model of course um
783
so i i think we'll probably look back
784
and be like oh yeah all that token maxing stuff was just
785
because of how the model i think you're right it's like
786
because of how we access the models
787
because how of how they're distributed to like especially like independent hackers
788
and builders that it became this thing that like seemed much more important
789
and got a lot of attention
790
and then one day it was just like nobody cared yeah
791
everyone like it feels like there's a lot of fomo
792
if you don't max out whatever plan you have yeah
793
and the best way to do so
794
if i work during the day is going to be
795
that i do my thinking and i let it run overnight yep
796
but i don't think that's how actually i want my day
797
to also be i do think there's there's value in like
798
blocks of thinking and blocks of coding right like my thought is like you know
799
when you run human layer it does this research step it's
800
slow it can in a large code base it can take 10 15 20 minutes
801
and so i might kick off two or three of those things
802
and let it run and go get a cup of coffee
803
or something but it's like rather than me going back
804
and forth with the model
805
and telling okay now go look over here okay now go look
806
and then i noticed oh you didn't understand
807
that thing yeah we also have to go look over there
808
and see how that works
809
and like you're constantly like catching the model's lack of understanding
810
versus like i'm willing to like async for 15 20 minutes
811
and just guarantee that it's like gonna have 90
812
or 99 of the context that it needs to do this task almost certainly
813
and anything that it missed is like very easy to go find
814
that is worth a little bit of async um and then
815
while those are running i'm like reading another plan and going back and forth on that.
816
Um, but I try not to paralyze too hard.
817
And when I do paralyze, I have one most important thing and it's not like, okay, launch three tasks.
818
Okay.
819
Get this one to design.
820
Now get this one to design.
821
Now get this one to design.
822
Okay.
823
Now move this one forward.
824
Now move this.
825
It's like when the top, when the most important thing is unblocked, we drop whatever we're doing.
826
Even if it's like, I'm in the middle of writing a prompt, I will go back to the most important one because the goal should be like land one thing.
827
You should always have one most important thing.
828
And your top goal is to get that thing landed.
829
And the other stuff you're working on is just like stuff to fill the time
830
while you're waiting for
831
that thing to be unblocked yeah you should always know what
832
the most important thing is as has always been true in engineering
833
and everything really the feeling of parallelization is very cool though
834
and i haven't done any i haven't done any measurements
835
but overall if you look at the most important thing
836
that needs to go out and you parallelize along the way i wonder
837
if that thing would go faster if you just wait
838
and just focus on
839
that one thing i want like traditionally i want to see
840
a stud it would be very zen yeah i want to
841
see a study of like the you know they hook up like the brain scans
842
and like watch someone's brain as it's working i want to
843
see someone running six clods in parallel i want to see one person running one coding agent at a time
844
and then i want to see someone sitting in a casino cranking the j the thing
845
and i want to slot machine yeah exactly
846
uh and then maybe maybe the person doing slot machine
847
and then maybe someone else at like the poker table
848
or it's like a poker you play one hand at a time
849
and you focus yeah good stuff thanks
850
so much for coming on dex this is great yeah man
851
this is super fun i'm really excited for the show i think you've had some really dope people on here
852
so hopefully hopefully i don't uh i don't tank your tank
853
your rep man should be good thanks thanks again patrick that's it

Bu Ders Hakkında

Gölgeleme Tekniği Nedir?

Gölgeleme, başlangıçta profesyonel tercüman eğitimi için geliştirilen ve çok dilli Dr. Alexander Arguelles tarafından popüler hale getirilen, bilim destekli bir dil öğrenme tekniğidir. Yöntem basit ama güçlüdür: ana dili İngilizce olan bir sesi dinler ve hemen yüksek sesle tekrar edersiniz — konuşmacıyı 1-2 saniye gecikmeyle takip eden bir gölge gibi. Pasif dinleme veya dilbilgisi alıştırmalarının aksine, gölgeleme beyninizi ve ağız kaslarınızı gerçek konuşma kalıplarını eşzamanlı olarak işlemeye ve yeniden üretmeye zorlar. Araştırmalar, telaffuz doğruluğu, tonlama, ritim, bağlı konuşma, dinleme anlama ve konuşma akıcılığını önemli ölçüde geliştirdiğini göstermektedir — bu da onu IELTS Konuşma hazırlığı ve gerçek dünya İngilizce iletişimi için en etkili yöntemlerden biri yapar.

Shadowing tekniği: adım adım eksiksiz rehberi okuyun →