쉐도잉 연습: WWDC26: Run local agentic AI on the Mac using MLX | Apple - 영상으로 영어 말하기 배우기

레슨 만드는 중...
1
Hi, I'm Angelos, an engineer on the MLX team.
2
Today I'm going to show you how to build and run agentic AI workflows entirely on your Mac using MLX.
3
No cloud, no API keys, just your hardware doing the work.
4
Over the past year, AI agents have gone from research prototypes to everyday productivity tools.
5
But before we talk about agents, let's look at what we had before.
6
Here's the chat experience you're familiar with.
7
You send a prompt to the language model.
8
The model sends a response back.
9
If you need to act on that response, run a command, check a file, or fix an error, that's on you.
10
But now you're talking to an agent.
11
The agent talks to the model to decide what to do.
12
Then it calls tools to actually do it: running commands, reading files, hitting APIs — It observes the results and goes back to the model to figure out the next step.
13
User to agent.
14
Agent to model.
15
Agent to tools.
16
This is the agentic loop.
17
And it keeps cycling until your task is done.
18
What makes this particularly exciting on Apple silicon is that the entire loop can run locally.
19
Your data stays on your machine; AI is available anywhere at any time and there are no usage costs.
20
Let me now show you what this looks like in practice.
21
Here I have an agent running locally on my Mac.
22
On my screen you can see the setup: on the left, MLX running the model, and on the right the OpenCode agent I am interacting with.
23
I asked it to fetch the recent pull requests from our MLX repository, summarize the changes, and identify anything that needs my attention.
24
The model reasons about the request, calls the GitHub CLI to fetch PR data, reads through the diffs, and produces a concise summary.
25
All of this is happening locally, the model runs on my hardware and only the git commands reach the network.
26
Well it seems like I have a lot of work to do after finishing this video.
27
Now that you've seen what's possible, let me walk you through how we'll get there today.
28
We'll start by introducing the local agentic AI stack, the four layers that make all of this work, from MLX at the foundation all the way up to the agent.
29
Then I'll show you step-by-step how to set up your own local agent.
30
After that, we'll look at how MLX gets the most out of your hardware to make agents fast.
31
And finally, we'll go through more live demos, including building a SwiftUI app from scratch and fixing a bug in Xcode.
32
Let's start with the stack.
33
The stack that powers local agentic AI on the Mac has four layers.
34
Let me walk you through each one, starting from the bottom.
35
At the bottom is MLX, our open-source array framework purpose-built for Apple silicon.
36
It handles all the low-level computation, Metal acceleration, and memory management.
37
This is the foundation everything else is built on.
38
One level up, we have the language model layer.
39
MLX-LM provides everything you need to load, run, quantize, and fine-tune large language models.
40
It supports thousands of models from HuggingFace and gives you both CLI tools and a Python API.
41
If you saw our sessions last year, this is what we covered in depth.
42
But to serve an agent, we need something more: a persistent server with a standard API.
43
That's where MLX-LM Server comes in.
44
This is an OpenAI-compatible HTTP server that exposes your local model through a standard API.
45
It supports structured tool calling so the model can invoke functions reliably, and reasoning models that can analyze complex problems step-by-step before responding.
46
It's a drop-in replacement for any cloud LLM API.
47
And at the top of the stack, we have the agent itself.
48
This can be any framework or tool that speaks the OpenAI chat completions protocol: Xcode, OpenCode, Pi agent, a custom script, or anything else.
49
Because MLX-LM Server provides a standard interface, any agent framework works out of the box.
50
And it's not just us building on this stack.
51
Several popular apps and tools build on MLX and MLX-LM.
52
Ollama, LM Studio, and vLLM are just a few of the most popular ones.
53
The ecosystem is broad and growing, and if you're using one of these tools, chances are you're already running on MLX.
54
So that's the stack.
55
Let me now show you how to set everything up yourself.
56
It only takes three steps to go from zero to a fully local agentic workflow.
57
Step one: install MLX-LM.
58
A single pip install gets you everything you need.
59
Step two: start the server.
60
Run mlx_lm.server with a model that supports tool calling.
61
Starting with a small model to test your set-up is always a good idea.
62
The server starts up, loads the model, and is ready to accept requests on local host.
63
Step three: point your agent at the local server.
64
In most agent frameworks, you just set the base URL to your local server's address and you're done.
65
The agent doesn't know or care that the model is running on your Mac rather than in the cloud.
66
Let me show you a concrete example.
67
Here's the configuration for OpenCode.
68
We define a local provider.
69
In particular, we set the URL to local host and set the model name the server expects.
70
We also tell OpenCode to use this local model for everything.
71
That's it. Now every interaction runs through your local model.
72
Now that we have an agent talking to MLX, let's look at how MLX gets the most out of your hardware and addresses the key challenges of running agents locally.
73
The first challenge is prompt processing.
74
In an agentic workflow, every time the model receives tool output, it has to process all that new context before it can reason about the next step.
75
This happens over and over throughout the agentic loop, and it adds up fast.
76
Agentic sessions usually comprise hundreds of thousands of tokens and most of those are not generated.
77
The M5 chip introduces dedicated Neural Accelerators, and MLX can target them for exactly this kind of work.
78
Specifically, Neural Accelerators make matrix multiplication four times faster on M5 compared to M4.
79
And with the specialized multiplication and attention kernels in MLX this translates almost exactly to prompt processing speedup.
80
Reducing prompt processing time means your agents can read your codebase or process tool results almost four times faster.
81
And the best part?
82
Taking advantage of Neural Accelerators requires no special arguments or code changes on your part, MLX selects the best kernel for the available hardware and it just works.
83
Let's now talk about the second challenge, concurrency.
84
In practice, agents rarely work alone.
85
A common pattern is for an agent to spawn several subagents, each tackling a different part of the problem in parallel.
86
One might be reading documentation, another searching code, and a third writing tests; all at the same time.
87
That means multiple requests hitting your local model simultaneously.
88
MLX-LM Server handles this with continuous batching.
89
Instead of processing requests one at a time, it dynamically groups incoming requests into batches and processes them together on the GPU.
90
New requests can join a batch in progress without waiting for the current one to finish.
91
The result is that your subagents don't stall waiting in a queue.
92
They all get served concurrently, which keeps the entire agentic workflow moving.
93
Finally, the third challenge is model size.
94
Sometimes a single machine, even one with 512GB of RAM, just isn't enough because the model is too large to fit in memory.
95
The most recent DeepSeek model for instance has a whopping 1.6 trillion parameters and requires more than 800GB of memory just for the weights.
96
MLX's distributed support lets you spread a model across multiple Macs connected over Thunderbolt or Ethernet.
97
For agents, this is powerful in two ways.
98
First, it lets you run much larger, more capable models that wouldn't fit on a single machine.
99
Second, it parallelizes prompt processing across devices, which directly speeds up the agentic loop since the model can process tool results faster.
100
Setting up distributed inference with MLX-LM Server is fairly straightforward.
101
You launch the server using mlx.launch and a hostfile that contains information about the nodes and the type of connection.
102
The model is automatically sharded across all available devices and everything else just works.
103
Starting with macOS 26.2, we have support for Thunderbolt RDMA, which provides low-latency, high-bandwidth communication over Thunderbolt.
104
As a result, distributed inference with MLX has seen significant speed-ups: up to three times with four nodes.
105
To learn how to set up your Macs for distributed inference with MLX, check out our session "Explore distributed inference and training with MLX".
106
Remember our PR summary demo from earlier?
107
That was a simple read-and-report task.
108
Let's now push things further and see what happens when we ask an agent to write an entire project from scratch and then fix a bug in an existing one.
109
In this demo, I'm going to ask the agent to build a small SwiftUI application from scratch.
110
I have started with a blank Xcode project and I am asking the agent to build a drawing app for the iPad.
111
And off it goes.
112
The agent first looks at the current directory to find out the existing project structure, makes a plan to guide its implementation, and gets on to writing the code.
113
Using an agent means we don't need to copy anything or even build the project.
114
The agent writes the file then builds the app, fixing any errors it encounters along the way.
115
And here we are: the model is done, it only took a couple of minutes to create the first version of the app.
116
At the same time, I have the project open in Xcode and I am launching the app in the simulator.
117
Let's have a look at what the agent created.
118
It seems that we have a fully functional drawing app.
119
That's really nice for something that was built in 2 minutes.
120
With agentic coding, however, we can keep iterating until we are happy with the result.
121
For instance, I prefer rounded end caps.
122
I think they look much better.
123
Let's ask the agent to add them.
124
The agent will edit the code and recompile the app until it compiles without errors.
125
Let's test the new version.
126
We now have rounded end caps.
127
This is cool indeed.
128
It is even more cool that all of this happened locally, the model ran through MLX-LM server on this Mac and the agent used standard development tools like xcodebuild to verify and build its work.
129
For our final demo, let's look at something that integrates directly with your development environment.
130
Here I have the same drawing app project open in Xcode.
131
Let's connect Xcode to our already running MLX server.
132
We open the settings and navigate to the Intelligence tab.
133
We click on Add Chat Provider... and select a Locally Hosted provider.
134
We set the Port to 8080 or whichever port we selected when launching our MLX server and we're done.
135
Now Xcode can talk to our local model.
136
I have introduced a bug to our previously working app and now we can ask the model to fix it.
137
Within seconds, it identifies the bug and inspects the code around it.
138
Finally, it writes a fix and we can now build and run our app.
139
This shows how a locally running agent can integrate with your existing development workflow in Xcode, reading project files, understanding build errors, and making targeted fixes.
140
Local AI means your code never leaves your Mac.
141
Today, we showed you the full stack for running agentic AI locally on your Mac, from MLX all the way up to the agent, and how Neural Accelerators, continuous batching, and distributed inference make it fast.
142
To get started, install MLX-LM, launch the server, and point your favorite agent at it.
143
Everything we showed today is open-source and available right now.
144
Thank you for watching and I'm excited to see what you build with local agentic AI on the Mac.

이 비디오로 말하기 연습을 해야 하는 이유는 무엇인가요?

이 비디오는 인공지능 에이전트를 이용한 새로운 작업 방식에 대한 소개를 담고 있습니다. 이러한 주제를 다루는 비디오에서 연습하면 실제 업무 환경에서도 필요한 영어 회화 연습에 큰 도움이 됩니다. 특히, 기술과 관련된 어휘 및 구문을 통해 실용적인 영어 지식을 쌓을 수 있습니다. 유튜브 영어 공부를 통해 이 비디오를 보면, 실시간 대화처럼 정보를 주고받는 과정을 이해하고, 자연스러운 의사소통 능력을 기를 수 있습니다.

문맥 속의 문법과 표현

  • “agent talks to the model”: 이 표현은 주어와 동사, 그리고 목적어의 명확한 구조를 가지고 있어, 영어 문장의 기본적인 구성을 이해하는 데 도움이 됩니다.
  • “calls tools to actually do it”: 'calls'와 'to'의 구조를 통해 행위의 목적을 명확히 하는 문장을 연습할 수 있습니다. 이 구조는 비즈니스 환경에서도 자주 사용됩니다.
  • “runs locally”: 이 단어는 기술적인 맥락에서 자주 등장하며, shadow speak와 같은 문맥에서 응용할 수 있는 표현입니다.
  • “decide what to do”: 선택적 의사결정 과정을 설명할 때 유용한 구문으로, 상황에 따른 다양한 선택지를 표현하는 연습을 할 수 있습니다.

일반적인 발음 함정

이 비디오에서 자주 등장하는 기본 단어들은 영어 학습자가 발음하기 어려운 단어들로 이루어져 있습니다. 특히, “model”, “agentic”과 같은 단어는 초급 학습자에게 낯설 수 있으며, 올바른 발음 연습이 필요합니다. 또한, “commands”의 경우 발음 시 약간의 억양을 주의해야 합니다. 이러한 발음상의 트랩을 피하기 위해 shadowspeaks를 활용한 연습은 매우 효과적입니다. 실제 비디오를 시청하면서 이 단어들을 반복적으로 연습하는 것은 말하기 능력을 키우는 데 중요한 방법입니다.

이 영상의 문법

화자가 가장 많이 쓰는 문형을 영상 속 실제 표현과 함께 정리했습니다.

문형영상 속 표현
현재완료 have/has + 과거분사 — 과거의 일이 지금도 관련이 있을 때have gone · you've seen · has seen
수동태 be + 과거분사 — 누가 하는지보다 무슨 일이 일어나는지에 초점is done · is built · are not generated

쉐도잉이란? 영어 실력을 빠르게 키우는 과학적 방법

쉐도잉(Shadowing)은 원래 전문 통역사 훈련을 위해 개발된 언어 학습 기법으로, 다언어 학자인 Dr. Alexander Arguelles에 의해 대중화된 방법입니다. 핵심 원리는 간단하지만 매우 강력합니다: 원어민의 영어를 들으면서 1~2초의 짧은 지연으로 즉시 소리 내어 따라 말하는 것——마치 '그림자(shadow)'처럼 화자를 따라가는 것입니다. 문법 공부나 수동적인 청취와 달리, 쉐도잉은 뇌와 입 근육이 동시에 실시간으로 영어를 처리하고 재현하도록 훈련합니다. 연구에 따르면 이 방법은 발음 정확도, 억양, 리듬, 연음, 청취력, 말하기 유창성을 크게 향상시킵니다. IELTS 스피킹 준비와 자연스러운 영어 소통을 원하는 분들에게 특히 효과적입니다.

섀도잉 방법: 단계별 전체 가이드 읽기 →