Shadowing Practice: Why Software Factories Fail - Learn English Speaking with Video

Creando lección...
1
What's up everybody, how are we doing?
2
Guys, give it up for all the great speakers today so far.
3
All right, this is harness engineering is not enough and why software factories fail.
4
And we're gonna click, maybe.
5
Oh, that's way too many slides, hold on guys.
6
Okay, so we're all racing to put AI coding into production.
7
and there's been lots been said about loop engineering and we should probably write more loops and yeah, I don't know, I guess we're doing loops now.
8
StrongDM built a Lightsoft software factory where nobody even reads the code
9
and the prevailing narrative is we should just spend more tokens, you are the bottleneck, the models are good enough,
10
code is free, just ship more stuff.
11
But at the same time, we are starting to see the cracks.
12
Our friend Mario at AI Engineer Europe begged us to slow down
13
because companies that should not be having outages because of coding agents are having outages due to coding agent mishaps.
14
Code bases are falling apart faster than they ever have before.
15
And our friends at Pharos AI actually even did a report
16
since we all picked up all these AI coding tools in January, maybe February.
17
Pull request, code review quality is way down we're having more comments, longer comments, and tons of PRs being merged without any review at all.
18
Incidents are way up, bugs per developer are way up, and many people will tell you that you're holding it wrong.
19
That's the only reason.
20
You're not.
21
Well, maybe you are, but that's not the point.
22
I've spoken a lot about how to hold it better when it comes to working with AI.
23
Probably a million views on YouTube at this point across a bunch of different talks.
24
And the basic thing is, as engineers, we've been told that if token maxing isn't working, then it's a skill issue.
25
You just need to spend more tokens.
26
Let go of reading the code.
27
That with enough harness engineering, if we maybe sprinkle some magic words, adversarial review, on enough of our PR bots, that we can get the best of both worlds.
28
10 to 100x faster, high quality, and nobody has to do that thing we all hate called code review you.
29
I'm here to convince you today that this is in fact not a skill issue,
30
that no amount of harness engineering or loops maxing can solve what is fundamentally a model training issue.
31
That's why we say the harness is not enough.
32
And to understand this, we kind of have to grapple and dig into how coding models are trained.
33
I'm going to talk about what I think the shortcomings are with some of the current benchmarks
34
and what better ones might look like, and we'll talk about how to move faster safely in the meantime.
35
It's going to sound like a rant, but there is hope here.
36
I'm going to talk about our journey and a bunch of the landmines we've hit building in this world, a bunch of exciting new techniques that we've been working with a lot of our users and customers to develop,
37
and I think how we all as a community get to
38
the next chapter of agentic engineering after whatever this thing that we're in.
39
So we use a lot of words here.
40
I'm going to zoom out a little bit.
41
I want to give you kind of like a brief history of the software factory factory, and it's actually, I just learned this last week,
42
the term software factory was defined at a NATO conference in 1968.
43
We're going to start around 2022, right before AI started coming around.
44
And basically, in a typical 2022 software factory, you will have some people building stuff.
45
You'll have engineers, you'll have PMs, maybe you have some sort of leadership team that is driving the vision here, and they all decide that stuff needs to get done.
46
And so you put it in a tracker, a linear, a JIRA, a beads, some sort of state machine that tracks what needs to be done.
47
And then someone goes and grabs something off of there and they build the thing.
48
And there may be some automated testing in that process, maybe some manual testing in that process.
49
At a certain point, we make this pull request thing.
50
It says, okay, cool, we've got to run a bunch of checks, automated stuff, a human's going to review the change and review the code, and perhaps we might even have a human pull it down and test it somehow.
51
And if anything goes wrong here, we loop back to someone builds the thing, and eventually we're ready for prod.
52
And so we ship it to production, and once it's in prod, it makes contact with our users.
53
And users do a thing that we all love.
54
Users love to complain.
55
I love our users, but yeah, they're gonna ask for things, they're gonna find bugs, they're gonna file feature requests, and that goes back to your team.
56
You might also add monitoring.
57
And so, you know, what do we want more than anything else?
58
We want to wake up engineers at three in the morning when something breaks, so they can get dragged out of bud to try to go fix it.
59
And we go on and on in this loop, and we ship a bunch of code.
60
And one thing that we noticed here is that Teams figured this out decades ago,
61
is that this someone -builds -the -thing step is usually going to take hours or days in most cases.
62
And the review part will also take hours or days for large things.
63
And so teams started doing these upfront planning, architecture proposals, sprint planning, and they would collaborate on these things as a team with the hopes
64
that we might decrease the percent chance that something would need to be reworked, that we would be able to reduce the time spent in reviewing every line of code
65
because we aligned on everything ahead of time.
66
This brings us to the agentic software factory.
67
Every company and their mother is talking about how they built a coding agent factory
68
that shifts 75 % of their code now.
69
Literally everybody.
70
And so if we look at the software factory from 2022, we just replace someone builds the thing with an agent builds the thing.
71
And we have an orchestration and a harness and a sandbox and a model and computer use, and I'm not gonna get into the details of that.
72
You can watch 100 talks about that this week, I'm sure.
73
But now the building part takes minutes or hours, but this human part still takes hours or days if you're gonna review the code and you're gonna test the changes.
74
And so we bring an agentic code review.
75
And we bring in agentic regression testing and it makes this part faster, but it's probably still the bottleneck.
76
But we can do more loops here.
77
Why not?
78
Let's do some more loops.
79
So we can route all incidents straight into the factory.
80
Why does someone need to get woken up and try to fix it
81
when they could just wake up to a pull request?
82
And maybe that fixes the issue for you.
83
You can take all the user feedback
84
and just stick it straight into the factory so
85
that people ask for stuff and it gets built
86
and now your only job is how much things can you stuff into the queue of stuff to do
87
and how fast can you review and test the changes
88
which brings us of course to I'm sure you know the
89
lights off software factory where basically Dan Shapiro coined this is
90
we no longer read the code we say you know what this is going great
91
that code review thing no thanks we're just not going to do
92
that anymore and we invest into all these other parts of the system you're testing you're monitoring your rollout, everything else.
93
We just write more code and build those systems better.
94
And now our job really is just how much stuff can we ask the agent to build.
95
I am going to posit that this does not work.
96
And this is why software factories fail.
97
As an aside, what I'm going to say has nothing to do with vibe coding.
98
So Addy had this great post, I'm just going to literally take his quote verbatim, a developer vibe coding a side project a dozen people will ever run
99
and a team keeping a 10 year old enterprise system alive for another quarter share almost no constraints worth naming.
100
And most of what you hear on the internet is one
101
of these groups of people telling the other group of people how to live their lives.
102
So if you love vibe coding, please go on.
103
At HumanLayer what we care about is how do we help people solve hard problems in complex code bases.
104
We use the word brownfield a lot, which historically has meant like some 10 -year -old Java thing.
105
I actually think agents really start to struggle after maybe three to six months, especially with the pace at which we can ship now.
106
You can ask me how I know this, and I will tell you that it is because in July 2025, we tried this.
107
We went full lights off, and if you have tried this seriously for a number of months, you probably found at least one issue that the agent couldn't solve.
108
Even with your most advanced prompting, you do research, you do reproductions, you just, you have to go and dig into that code base
109
that you stopped reading three months ago to try to figure out what's broken.
110
And in the meantime, your site was down, your users were pissed, and if you were like me, you were probably miserable reading all this slop code that you let slip into your system.
111
And what I want to get to is basically models have a shortcoming.
112
They can't maintain and improve code base quality over time, not without a decent amount of human steering.
113
And when When I say maintainability, I'm basically talking about issues like it becomes really, really hard to make a change in one part of the code base without breaking other parts of the code base.
114
This is Martin Fowler's Shotgun Surgery textbook code smell.
115
I'm not going to say much more about maintainability, there's a bunch of books that you can go read about it.
116
In fact, John Osterhut is actually here speaking this week, so you can go ask him in person about the philosophy of software design if you want to.
117
But it brings us to this question of like, why can't models do software maintainability? quality.
118
And you may also be saying, but Dex, you know, surely the models have gotten much better since then.
119
They've gotten better in some ways, but they're still about the same in others.
120
If you want to solve one -off problems or Vibe code a new marketing site, yes, they got way better since 2025 and 2024.
121
But as far as improving code -based quality, I think they have not gotten much better.
122
Now, I cannot prove this because there are no good benchmarks for a model's ability to maintain code -based quality, and I'll get into, like, where we're going with that,
123
but if you've worked with coding agents for a while, a lot of people are posting about this, it's just like.
124
You probably have this vibe that they generally make things worse over time and make the code base harder to work.
125
And to figure out why this happens, I want to zoom out to the first great coding agent.
126
Why did Cloud Code go from nothing to $4 billion, and I think now they're at $9 billion in revenue in under a year?
127
Because there were great CLI agents before Cloud Code.
128
You had Ader, you had CodeBuff.
129
There was a bunch of tools in this category.
130
They had all the same tools, read, write, edit, grab, bash.
131
So what was the difference?
132
The difference was that this was the first time
133
that a model lab trained a model against the harness that they were going to distribute it to users in.
134
And it got really, really good, and this is just some of the tools, but it got really, really good at calling these sorts of tools in an agentic loop.
135
In fact, the OpenAI team did a talk in November about basically
136
if you are a harness builder and you don't own the model weights and you can't RL the model in your harness,
137
you will always be at a disadvantage compared to somebody who owns both the model and the harm.
138
And I'm gonna slide a couple slides from my buddy Calvin
139
French -Owen who was a MCS on Codex during the initial launch, but LMs are just next token predictors.
140
This is a slide from over a year ago where basically as you're doing your agentic loop, context window goes in, next step comes out.
141
And we're gonna try to do this, I haven't actually timed this, but we're gonna see if we can do coding agent reinforcement learning in 60 seconds.
142
So So what we're going to do is we want to train a model to get better at tool calling, better at solving software problems, we're going to give it a problem and we're going to generate a bunch of traits, try to solve the problem a bunch of different times,
143
we're going to score them all on correctness and the test tasks and all this stuff, and then we're going to reinforce, we're going to make the bad behavior less likely, and we're going to update the weights to make the good behavior more likely.
144
One of the classic ones here is Sweebench Multilingual, they're about 15 minute tasks, they're from open source repos like Redis, JQ, and Django and all this stuff, and they have binary one
145
or zero rewards on did you fix the problem you were trying to fix
146
and did you do it without breaking anything else
147
and we look at actually a real problem for one of these benchmarks this is fast lane
148
which is a Ruby project basically there was some issue where we weren't checking for nil
149
and we have a stack trace blow up
150
because you have a null point
151
and in this in this benchmark you have a base commit
152
that we're gonna check out before the issue was solved by
153
a human in the past we're gonna give it a test patch
154
that says here's what the behavior should be afterwards we have
155
a golden patch both these are hidden from the model and
156
so we have the agent go try to solve the problem
157
we store its patch we undo all the changes it made to any test files
158
so I'm sure you've seen models comment out tests to get things working
159
and then we're going to apply our golden test patch
160
and then we're gonna run the test old tests
161
and did the new test pass and
162
if they both pass then then we get the reward otherwise we don't
163
and so models are trying to get the test to pass there's There's no way in this system
164
that we can penalize it for poor program design or for eroding the maintainability of our system.
165
That's why we get things like this.
166
Try catches around things that probably don't need a try catch.
167
Or things like this.
168
I think Vaibov gave us this example earlier of casting things to other things just
169
so the model just wants to get the test to pass.
170
And so if you can't verify the maintainability of the code, it gets way harder to train on this.
171
If you remember this picture, verifying code quality and maintainability is orders of magnitude harder than the code runs because the cost function of bad architects
172
is measured in months and years if you have a coding episode
173
and then you only find out months later
174
that like oh somebody vibed this a little bit too hard it's really hard to propagate
175
that reward signal back across the gap
176
and now the frontier is getting better slowly
177
and since I know someone's gonna be in the YouTube comments about this yes I know benchmarks
178
and verifiers are different and they actually have to be separate data sets
179
but they're shaped the same and the the structure of these benchmarks is directionally correct
180
when we look at these as like what is the future of evaluating code maintainability
181
there's a really cool one called three marathon from abundant ai where they do like 400 hour tasks
182
that like clone all of microsoft excel every single feature
183
and they have some sophisticated reward channel stuff deep sweep from
184
data curve is also like large tasks on oss repos
185
that are not actually in the training set because they were never actually built in the real world
186
And then you have Frontier Code from Cognition, which is multi -PR tasks.
187
They do interesting things like, hey, if the model writes tests that don't fail on the pre -patch code, then it gets penalized.
188
And then we have a judge model that says, okay, did this follow all of our code quality rules?
189
So we're getting better.
190
But I think models judging quality can only go so far, because if the new model, if the model knew what good code looks like, it would probably write it in the first place.
191
And review agents and throwing more tokens at the problem, it can raise the floor, but we're still constrained by what we can teach during RL.
192
And so I will posit that for now we're stuck reading the code, but we can still move pretty fast.
193
And of course, there's a world where this should solve in the future, and if you want to just keep YOLOing prompts until you get to GPP7, you don't have to think about this, by all means, please.
194
But bitter lesson be damned, we've got some problems to solve, so let's engineer our way out of it.
195
So, turning the lights back on, we're gonna put the code review back.
196
We're gonna embrace this approach of like, how do we plan up front to reduce the chance that we have a long or difficult review process.
197
We're gonna find leverage, we're gonna use AI to help with this.
198
The first thing we're gonna do is we're gonna do some sort of product review, understanding what problem we're solving, what's the desired behavior, maybe looking at mockups.
199
Here's a product review I was working on yesterday with a mockup of a new feature.
200
Once we have our product review, we're gonna, by the way, small stuff still just goes straight straight to the agent, but once we have the product review, we're gonna also do architecture, system architecture.
201
A lot of people have been doing this for a while, component contracts, data models, constraints.
202
This is an example of a doc
203
that we build to understand how these systems are gonna fit together and what's like the high level picture of it.
204
From there we do something that I think is really under emphasized in agent encoding these days, which is program design.
205
I think people assume that once you get the architecture right, the model can just cook.
206
But we often look into the types and the method signatures, the program layout and the call stacks.
207
So here's some examples, I don't think you'll be able to read this one, but this is like the level of abstraction we're at, is how are we actually gonna lay this stuff out and how are these systems gonna interact?
208
Dylan Mulroy from Cloudflare talks a lot about how he's using these call graphs as part of his planning process, I think this is exactly right.
209
And then once we've done the program design, we can do this thing called vertical slices, which is the order of implementation, multi -repo coordination, how are we gonna build this across our entire system
210
and how are we gonna check it along the way.
211
I've talked a little bit about how models have horizontal plans.
212
I won't go too deep into it.
213
If you want to learn more about this, you can go watch our talk from AI Engineer Miami.
214
Couple shots of a doc like this, going through the tests and the steps in between each phase.
215
The main idea here is 30 minutes over here in pre -planning and alignment can save you hours in review.
216
And so it's actually feasible to still read every line of code.
217
Let me skip this part.
218
Basically, the summary here is like you don't have too many PRs.
219
If you're drowning in PRs, you actually have too many bad PRs.
220
Because a good PR is a joy to review.
221
You're just reading through it like, yep, this is great.
222
This is what we discussed.
223
This is what we talked about.
224
But even if a PR needs 20 % rework, which is generous for a lot of AI vibe -coded slop,
225
It's an emotional and intellectual burden on both the reviewer and the submitter.
226
And so if you use model -assisted planning and alignment, your alignment is shorter because you used AI to get all the information at once, your code review is faster because you aligned up front,
227
and your coding is faster because AI did it, and so now you're actually really moving faster, but you're still reading everything and you're still owning the code.
228
So closing advice, it's easy to hear all this and be a little bummed out.
229
I really like the world where we just YOLO everything and we can just not have to ever read code ever again, but we're engineers and these are just constraints
230
and models are good at certain things and they're not good at other things
231
and so go figure out how to solve problems given a set of constraints.
232
Use loops, they're great.
233
Go solve hard problems, seek leverage.
234
If you wanna help with this, we're building a human layer.
235
Human layer is an AI IDE and collaboration platform.
236
It's building blocks for your software factory and soon to be better verifiers for software quality.
237
We've got sort of a Figma for cloud code and codex style collaborative workspace.
238
It walks you through the workflows for doing this sort of work.
239
And we are talking to design partners.
240
We are hiring founding engineers here in San Francisco.
241
And these slides are live.
242
You can go get them right now.
243
You can try human layer at humanlayer .com.
244
It's free for small teams.
245
Go solve hard problems in complex code bases.
246
Thank you all for your energy.

Sobre esta lección

¿Qué es la Técnica de Shadowing?

Shadowing es una técnica de aprendizaje de idiomas respaldada por la ciencia, desarrollada originalmente para la formación de intérpretes profesionales y popularizada por el políglota Dr. Alexander Arguelles. El método es simple pero poderoso: escuchas audio en inglés nativo y lo repites en voz alta de inmediato, como una sombra que sigue al hablante con solo 1-2 segundos de retraso. A diferencia de la escucha pasiva o los ejercicios de gramática, el shadowing obliga a tu cerebro y músculos de la boca a procesar y reproducir simultáneamente patrones de habla reales. Las investigaciones muestran que mejora significativamente la precisión de la pronunciación, la entonación, el ritmo, el habla conectada, la comprensión auditiva y la fluidez al hablar, convirtiéndola en una de las metodologías más efectivas para la preparación del IELTS Speaking y la comunicación en inglés en el mundo real.

Técnica de shadowing: lee la guía completa paso a paso →