Shadowing Practice: WWDC26: Run local agentic AI on the Mac using MLX | Apple - Learn English Speaking with Video

Creating lesson...
1
Hi, I'm Angelos, an engineer on the MLX team.
2
Today I'm going to show you how to build and run agentic AI workflows entirely on your Mac using MLX.
3
No cloud, no API keys, just your hardware doing the work.
4
Over the past year, AI agents have gone from research prototypes to everyday productivity tools.
5
But before we talk about agents, let's look at what we had before.
6
Here's the chat experience you're familiar with.
7
You send a prompt to the language model.
8
The model sends a response back.
9
If you need to act on that response, run a command, check a file, or fix an error, that's on you.
10
But now you're talking to an agent.
11
The agent talks to the model to decide what to do.
12
Then it calls tools to actually do it: running commands, reading files, hitting APIs — It observes the results and goes back to the model to figure out the next step.
13
User to agent.
14
Agent to model.
15
Agent to tools.
16
This is the agentic loop.
17
And it keeps cycling until your task is done.
18
What makes this particularly exciting on Apple silicon is that the entire loop can run locally.
19
Your data stays on your machine; AI is available anywhere at any time and there are no usage costs.
20
Let me now show you what this looks like in practice.
21
Here I have an agent running locally on my Mac.
22
On my screen you can see the setup: on the left, MLX running the model, and on the right the OpenCode agent I am interacting with.
23
I asked it to fetch the recent pull requests from our MLX repository, summarize the changes, and identify anything that needs my attention.
24
The model reasons about the request, calls the GitHub CLI to fetch PR data, reads through the diffs, and produces a concise summary.
25
All of this is happening locally, the model runs on my hardware and only the git commands reach the network.
26
Well it seems like I have a lot of work to do after finishing this video.
27
Now that you've seen what's possible, let me walk you through how we'll get there today.
28
We'll start by introducing the local agentic AI stack, the four layers that make all of this work, from MLX at the foundation all the way up to the agent.
29
Then I'll show you step-by-step how to set up your own local agent.
30
After that, we'll look at how MLX gets the most out of your hardware to make agents fast.
31
And finally, we'll go through more live demos, including building a SwiftUI app from scratch and fixing a bug in Xcode.
32
Let's start with the stack.
33
The stack that powers local agentic AI on the Mac has four layers.
34
Let me walk you through each one, starting from the bottom.
35
At the bottom is MLX, our open-source array framework purpose-built for Apple silicon.
36
It handles all the low-level computation, Metal acceleration, and memory management.
37
This is the foundation everything else is built on.
38
One level up, we have the language model layer.
39
MLX-LM provides everything you need to load, run, quantize, and fine-tune large language models.
40
It supports thousands of models from HuggingFace and gives you both CLI tools and a Python API.
41
If you saw our sessions last year, this is what we covered in depth.
42
But to serve an agent, we need something more: a persistent server with a standard API.
43
That's where MLX-LM Server comes in.
44
This is an OpenAI-compatible HTTP server that exposes your local model through a standard API.
45
It supports structured tool calling so the model can invoke functions reliably, and reasoning models that can analyze complex problems step-by-step before responding.
46
It's a drop-in replacement for any cloud LLM API.
47
And at the top of the stack, we have the agent itself.
48
This can be any framework or tool that speaks the OpenAI chat completions protocol: Xcode, OpenCode, Pi agent, a custom script, or anything else.
49
Because MLX-LM Server provides a standard interface, any agent framework works out of the box.
50
And it's not just us building on this stack.
51
Several popular apps and tools build on MLX and MLX-LM.
52
Ollama, LM Studio, and vLLM are just a few of the most popular ones.
53
The ecosystem is broad and growing, and if you're using one of these tools, chances are you're already running on MLX.
54
So that's the stack.
55
Let me now show you how to set everything up yourself.
56
It only takes three steps to go from zero to a fully local agentic workflow.
57
Step one: install MLX-LM.
58
A single pip install gets you everything you need.
59
Step two: start the server.
60
Run mlx_lm.server with a model that supports tool calling.
61
Starting with a small model to test your set-up is always a good idea.
62
The server starts up, loads the model, and is ready to accept requests on local host.
63
Step three: point your agent at the local server.
64
In most agent frameworks, you just set the base URL to your local server's address and you're done.
65
The agent doesn't know or care that the model is running on your Mac rather than in the cloud.
66
Let me show you a concrete example.
67
Here's the configuration for OpenCode.
68
We define a local provider.
69
In particular, we set the URL to local host and set the model name the server expects.
70
We also tell OpenCode to use this local model for everything.
71
That's it. Now every interaction runs through your local model.
72
Now that we have an agent talking to MLX, let's look at how MLX gets the most out of your hardware and addresses the key challenges of running agents locally.
73
The first challenge is prompt processing.
74
In an agentic workflow, every time the model receives tool output, it has to process all that new context before it can reason about the next step.
75
This happens over and over throughout the agentic loop, and it adds up fast.
76
Agentic sessions usually comprise hundreds of thousands of tokens and most of those are not generated.
77
The M5 chip introduces dedicated Neural Accelerators, and MLX can target them for exactly this kind of work.
78
Specifically, Neural Accelerators make matrix multiplication four times faster on M5 compared to M4.
79
And with the specialized multiplication and attention kernels in MLX this translates almost exactly to prompt processing speedup.
80
Reducing prompt processing time means your agents can read your codebase or process tool results almost four times faster.
81
And the best part?
82
Taking advantage of Neural Accelerators requires no special arguments or code changes on your part, MLX selects the best kernel for the available hardware and it just works.
83
Let's now talk about the second challenge, concurrency.
84
In practice, agents rarely work alone.
85
A common pattern is for an agent to spawn several subagents, each tackling a different part of the problem in parallel.
86
One might be reading documentation, another searching code, and a third writing tests; all at the same time.
87
That means multiple requests hitting your local model simultaneously.
88
MLX-LM Server handles this with continuous batching.
89
Instead of processing requests one at a time, it dynamically groups incoming requests into batches and processes them together on the GPU.
90
New requests can join a batch in progress without waiting for the current one to finish.
91
The result is that your subagents don't stall waiting in a queue.
92
They all get served concurrently, which keeps the entire agentic workflow moving.
93
Finally, the third challenge is model size.
94
Sometimes a single machine, even one with 512GB of RAM, just isn't enough because the model is too large to fit in memory.
95
The most recent DeepSeek model for instance has a whopping 1.6 trillion parameters and requires more than 800GB of memory just for the weights.
96
MLX's distributed support lets you spread a model across multiple Macs connected over Thunderbolt or Ethernet.
97
For agents, this is powerful in two ways.
98
First, it lets you run much larger, more capable models that wouldn't fit on a single machine.
99
Second, it parallelizes prompt processing across devices, which directly speeds up the agentic loop since the model can process tool results faster.
100
Setting up distributed inference with MLX-LM Server is fairly straightforward.
101
You launch the server using mlx.launch and a hostfile that contains information about the nodes and the type of connection.
102
The model is automatically sharded across all available devices and everything else just works.
103
Starting with macOS 26.2, we have support for Thunderbolt RDMA, which provides low-latency, high-bandwidth communication over Thunderbolt.
104
As a result, distributed inference with MLX has seen significant speed-ups: up to three times with four nodes.
105
To learn how to set up your Macs for distributed inference with MLX, check out our session "Explore distributed inference and training with MLX".
106
Remember our PR summary demo from earlier?
107
That was a simple read-and-report task.
108
Let's now push things further and see what happens when we ask an agent to write an entire project from scratch and then fix a bug in an existing one.
109
In this demo, I'm going to ask the agent to build a small SwiftUI application from scratch.
110
I have started with a blank Xcode project and I am asking the agent to build a drawing app for the iPad.
111
And off it goes.
112
The agent first looks at the current directory to find out the existing project structure, makes a plan to guide its implementation, and gets on to writing the code.
113
Using an agent means we don't need to copy anything or even build the project.
114
The agent writes the file then builds the app, fixing any errors it encounters along the way.
115
And here we are: the model is done, it only took a couple of minutes to create the first version of the app.
116
At the same time, I have the project open in Xcode and I am launching the app in the simulator.
117
Let's have a look at what the agent created.
118
It seems that we have a fully functional drawing app.
119
That's really nice for something that was built in 2 minutes.
120
With agentic coding, however, we can keep iterating until we are happy with the result.
121
For instance, I prefer rounded end caps.
122
I think they look much better.
123
Let's ask the agent to add them.
124
The agent will edit the code and recompile the app until it compiles without errors.
125
Let's test the new version.
126
We now have rounded end caps.
127
This is cool indeed.
128
It is even more cool that all of this happened locally, the model ran through MLX-LM server on this Mac and the agent used standard development tools like xcodebuild to verify and build its work.
129
For our final demo, let's look at something that integrates directly with your development environment.
130
Here I have the same drawing app project open in Xcode.
131
Let's connect Xcode to our already running MLX server.
132
We open the settings and navigate to the Intelligence tab.
133
We click on Add Chat Provider... and select a Locally Hosted provider.
134
We set the Port to 8080 or whichever port we selected when launching our MLX server and we're done.
135
Now Xcode can talk to our local model.
136
I have introduced a bug to our previously working app and now we can ask the model to fix it.
137
Within seconds, it identifies the bug and inspects the code around it.
138
Finally, it writes a fix and we can now build and run our app.
139
This shows how a locally running agent can integrate with your existing development workflow in Xcode, reading project files, understanding build errors, and making targeted fixes.
140
Local AI means your code never leaves your Mac.
141
Today, we showed you the full stack for running agentic AI locally on your Mac, from MLX all the way up to the agent, and how Neural Accelerators, continuous batching, and distributed inference make it fast.
142
To get started, install MLX-LM, launch the server, and point your favorite agent at it.
143
Everything we showed today is open-source and available right now.
144
Thank you for watching and I'm excited to see what you build with local agentic AI on the Mac.

About This Lesson

In this lesson, you will practice your English listening and speaking skills by engaging with a transcript from a video about running local agentic AI on a Mac. You will learn to articulate technical jargon while enhancing your pronunciation and fluency through the effective shadowing technique. By following along with the speaker’s pace and intonation, you’ll improve your ability to communicate about complex topics clearly and confidently.

Key Vocabulary & Phrases

  • Agentic AI - Refers to artificial intelligence systems that can act autonomously based on user prompts.
  • MLX - A framework for handling machine learning tasks specifically built for Apple silicon.
  • Local model - A language model that runs entirely on a device without needing cloud services.
  • API - An application programming interface that allows different software components to communicate.
  • Agent loop - The cycle of interaction between the user, the agent, and the tools required to complete tasks.
  • Persistent server - Provides ongoing service to manage requests continuously.
  • Quantization - The process of reducing the number of bits that represent a number, improving efficiency in computation.

Practice Tips

To make the most out of this lesson, use the shadow speech method as you work through the transcript. Begin by playing short segments of the video and repeating them immediately after, mimicking the speaker’s tone, rhythm, and speed. Since the speed of delivery in the video is moderate, you can focus on grasping complex phrases. If you find any part challenging, use the shadowspeaks technique, slowing down the playback to better absorb and reproduce the sounds.

Additionally, try to emphasize technical terms—like “local model” and “agent loop”—as they represent key concepts. Practice these terms in context, imagining how you might explain them to someone else. By consistently utilizing the shadowing site for focused practice, you can enhance your ability to discuss advanced topics in English, gaining both confidence and comprehension along the way.

Grammar in this video

The structures the speaker uses most, with the exact words from the video:

StructureIn the video
Present perfect have/has + past participle — a past action that still matters nowhave gone · you've seen · has seen
Passive voice be + past participle — the focus is on what happens, not who does itis done · is built · are not generated

What is the Shadowing Technique?

Shadowing is a science-backed language learning technique originally developed for professional interpreter training and popularized by polyglot Dr. Alexander Arguelles. The method is simple but powerful: you listen to native English audio and immediately repeat it out loud — like a shadow following the speaker with just a 1–2 second delay. Unlike passive listening or grammar drills, shadowing forces your brain and mouth muscles to simultaneously process and reproduce real speech patterns. Research shows it significantly improves pronunciation accuracy, intonation, rhythm, connected speech, listening comprehension, and speaking fluency — making it one of the most effective methods for IELTS Speaking preparation and real-world English communication.

Shadowing technique: read the full step-by-step guide →