Shadowing Practice: The Engineering that Runs the Digital World šŸ› ļøāš™ļøšŸ’» How do CPUs Work? - Learn English Speaking with Video

Creating lesson...
1
Inside every desktop computer, smartphone,Ā  gaming console, laptop, or practically any other device you use on a daily basis is a CPUĀ  or Central Processing Unit, and in this video, we’re going to see how they work. A typicalĀ  processor for a powerful laptop like this one is built from billions of nanoscopic transistorsĀ  connected together using dozens of layers of wires and is essentially the brain of the device. But before we explore the microprocessor and all its complexity, let’s travel to the early daysĀ  of personal computers and video game consoles and compare the Apple 2e from 1983 to the modern-dayĀ  MacBook Pro. Inside the Apple 2e we find a chip called a 6502, which is considered one of theĀ  great-grandparents of all modern processors.
2
This chip is built from 4528 transistors and canĀ  perform around 430 thousand calculations a second.
3
While it could only run primitive applicationsĀ  and video games with simple graphics, this chip was the backbone of a generation of earlyĀ  computers and video game consoles such as the NES, the Commodore 64, and the Atari. Compare that toĀ  the MacBook Pro’s M1 processor which is built from 16 billion transistors capable of performingĀ  around 3 trillion calculations a second, thus enabling it to generate expansiveĀ  3D worlds with immersive graphics.
4
Despite these two chips being released around 45Ā  years apart, the underlying principles of how they work are rather similar. In a way, you can thinkĀ  of these devices as sharing a common section of technological DNA. In fact, if we were to openĀ  up a desktop computer and grab the CPU or the GPU inside the graphics card, or teardown a NintendoĀ  Switch or Smartphone and find the system on a chip or SoC, or even if you could make your way intoĀ  an AI Data Center and grab a state-of-the-art AI chip, you’d find that all of these processorsĀ  operate using the same underlying principles.
5
In other words, both the Oregon Trail ofĀ  50 years ago and advanced AI algorithms run on processors with similarĀ  technological DNA, but of course, one of these chips is 10 billion times moreĀ  computationally powerful than the other.
6
So, in this video, we’re going to take apartĀ  this microprocessor and find out exactly what the shared technological DNA is and how itĀ  enables CPUs to work. And just to be clear, the technological DNA is not transistors and it’sĀ  not logic gates, but rather it’s an architectural design and basic operational principle that’sĀ  fundamental to microprocessors and differentiates these chips from other integrated circuits.Ā  So, stick around, and let’s dive right in.
7
This video is sponsored by Brilliant. Let’s begin with a quick 3D animated teardown of this Macbook Pro. When we open it up,Ā  we find a range of different components such as the touchpad, battery cells, speakers, a coolingĀ  fan, and the motherboard in the center. Mounted to the motherboard are the solid state drive or SSDĀ  Storage chips where all your files are saved and a range of other chips. Underneath the heat pipe,Ā  we find the DRAM which is the short term working memory and central processing unit. Let’s desolderĀ  the DRAM and CPU and open it up. Inside we find three parts: on the top is a protective cover thatĀ  conducts and dissipates heat, on the bottom is an interposer with thousands of connection pointsĀ  on either side and wires running inside of it, and soldered onto the interposer is the integratedĀ  circuit or IC which is also called a die and is the functional part of the CPU. On the die we canĀ  see the complex design of billions of transistors and wires organized into different sectionsĀ  such as the 4 high-performance computational cores, 4 energy-efficient cores, graphicsĀ  processing cores, cache memory and many other sections. Let’s zoom in on one of the performanceĀ  cores where we find that it’s separated into different functional blocks which we’ll add labelsĀ  to and then reorganize into an architectural diagram. This diagram illustrates how data andĀ  instructions move around a single processing core in the CPU and, although it’s rather complicated,Ā  you’ll understand how it works by the end of this video. But for now, let’s zoom in even furtherĀ  to get a nanoscopic view of a massive multilayer labyrinth of wires with the transistors atĀ  the very bottom. Here we see a group of 6 transistors that are wired together to build anĀ  AND logic gate, and in this view we can see around 650 transistors out of the total 16 billionĀ  that make up the overall chip. Understanding how billions of transistors work together toĀ  build a CPU capable of playing video games, watching movies or browsing the internet will takeĀ  a bit of work, so let’s start with an analogy.
8
You may have heard that CPUs are like superĀ  powerful calculators. This analogy is only around 20% complete as it’s missing some criticalĀ  parts, so let’s add them in to make a more accurate analogy. First, we’ll add a table for theĀ  calculator to sit on, along with a pencil and a sheet of paper. Next, we’ll add rows and rows ofĀ  bookshelves containing thousands of books along with a cart that can carry a small stack of booksĀ  between the shelves and the table. And finally, we’ll add an automated robot which we’ll callĀ  a control unit or controller. The controller can grab books from the bookshelves, move them toĀ  the cart and onto the table, and it can put them back. The controller can also read the contentsĀ  of each book, write on the paper and in the books, and use the calculator. You can think of theĀ  controller as a super-fast human, but we’re bad at animating humans, so it’s a robot instead. AndĀ  with that, we have all the parts for our analogy.
9
Now let’s see how each part of our CPU analogyĀ  works. To start, the bookshelves are the storage devices in your computer, such as the SSD chips ,Ā  the cart represents the DRAM , and the table and its contents represent the CPU. On the CPU Table,Ā  there’s a small space for a single open book which is similar to the very limited capacityĀ  of the cache memory inside the CPU itself.
10
Next, the single sheet of paper representsĀ  the Registers which are used for storing values or numbers that are actively being used.Ā  Specifically, on it are four general-purpose registers and a few more special locations whichĀ  we’ll discuss in a little bit. Additionally, the pencil is there to write and eraseĀ  things on the paper and in the books.
11
Finally, the calculator represents the ArithmeticĀ  Logic Unit or ALU. This ALU calculator works using binary so there are only the digits zero andĀ  one and it can do simple functions like add, subtract or multiply two numbers. TheĀ  ALU calculator has many more functions that you may be unfamiliar with but are stillĀ  rather simple. For example, it can increase or decrease a number by one or the ALU can performĀ  bit shifts which is essentially taking a number and adding a zero to the end of it. In decimal,Ā  bit shifting is like multiplying a number by 10 but in binary it’s equivalent to multiplying aĀ  number by two. The ALU can also perform logic functions on two numbers such as the logicalĀ  AND, OR, or Exclusive OR operations. For example, here is the logical AND operation for 2 binaryĀ  inputs and you can see that the output is the logical AND for each place value of the 2 inputs. However, more importantly, the ALU calculator can perform comparisons. For example, you canĀ  input two numbers and hit the comparison button to test whether the numbers are equal toĀ  one another and, if they are, then an equals flag goes up while the other comparison flags likeĀ  less than or greater than stay down. Finally, the ALU calculator’s display that outputs theĀ  result has a special name called the accumulator.
12
So now that we’ve explained the various parts ofĀ  this analogy, how does it all work together? Well, the first step is to load a program that we wantĀ  to run, which is like moving a set of books from the bookshelves to the cart, and then moving aĀ  single book to the table and opening it up. It’s important to note that the DRAM cart and cacheĀ  memory on the table are both temporary and limited capacity locations, whereas the SSD bookshelvesĀ  can hold a lot more and are semi-permanent long term storage. Additionally, when the computer isĀ  turned off, there are no books in the DRAM or the CPU, but when the computer turns on, the cart andĀ  table are very actively shuffling books around.
13
Let’s take a look at the contents of one of theĀ  books. Essentially, there are two types of pages: instructions and data. You can think of theĀ  instructions as the directions in a cookbook, with each step numbered sequentially across the pages.Ā  And, the pages of data contain a list of addresses with values stored at each address and are likeĀ  the ingredients that go into the recipe itself.
14
Similar to cooking, you need both theĀ  recipe and ingredients to make it work, and just a few ingredients can be combinedĀ  in dozens of different ways using different recipes. But to not use an analogy insideĀ  another analogy, let’s drop the cooking one and focus on the books, table and calculator. Let’s start at the beginning of this program and flip to page one instruction one, which isĀ  called a ā€˜Load’ and is the most common type of instruction. This ā€˜load’ has us open the pagesĀ  of data and find a specific address. We then copy and write down the value stored at that addressĀ  into one of the general purpose registers on the sheet of paper. With instruction one completed, weĀ  move to instruction two, which is to increment the value in register zero by 1. So, we plug the valueĀ  into the ALU calculator and hit plus 1. The third instruction, called a ā€˜store’ instruction, isĀ  used to save or store the output of the calculator found in the accumulator display into the pagesĀ  of data in the same address it was in before.
15
These simple yet very common instructions areĀ  equivalent to this line of code. Next we move onto instruction 4 and complete it and thenĀ  instruction 5 and so on, moving through the list of instructions which goes on and on and on. In order to keep track of which instruction is the next one to be completed, we use one of theĀ  special locations on the sheet of paper that we mentioned earlier called the Program Counter orĀ  PC, also called an Instruction Address Register or Instruction Pointer. Since the PC currentlyĀ  has a value of 5, we find instruction 5, complete it, and increase the program counterĀ  by one. Therefore, the next instruction to be completed will be instruction 6. However, what ifĀ  after completing instruction 6, we want to jump directly to instruction 42? Well, to do thisĀ  we use a jump instruction at 7 which directly sets the value in the program counter to a newĀ  number and in this case it’s 42. As a result, the sequence of instructions will be 5, 6, 7 which isĀ  the jump instruction, then 42, 43, 44 and so on.
16
A similar set of instructions is called aĀ  conditional branch which is used for implementing IF statements, loops, and other conditional code.Ā  Let’s use an example of a FOR loop with a few simple lines of code inside of it. Quite simply,Ā  this loop is used to repeat the code inside of it 4 times. Here are the corresponding instructionsĀ  of the FOR loop along with the instructions for the code inside of it and we color coded each ofĀ  the elements to keep track of which specific lines of code result in the corresponding instructions.Ā  We’ll discuss how compilers turn code into instruction later in this video, but for nowĀ  let’s focus on the FOR loop and its instructions.
17
Specifically, here’s where ā€˜i’ gets set to 0Ā  and stored in an address in the pages of data, here’s where ā€˜i’ is loaded from that addressĀ  and incremented by one, and here’s the contents of the loop. At the top, you can see the threeĀ  instructions, Load, Compare, and Branch greater than or Equal to. The Load first grabs the valueĀ  for ā€˜i’ and places it into register 0. Compare feeds ā€˜i’ stored in register 0 and a value of 4Ā  into the ALU and compares them, resulting in the applicable comparison flags being triggered. Next,Ā  branch greater than or equal to checks whether either the greater than or the equals flag is on,Ā  and if it is, it sets the program counter to 23, which corresponds to completing and leavingĀ  the loop. However when ā€˜i’ is less than 4 th ose flags aren’t triggered and the loop continues,Ā  until it hits the jump instruction at address 22, where the jump sets the program counter to 6Ā  which corresponds to the top of the for loop. As a result the loop will repeat a total of 4 times. Note that this code on the left is in C++, whereas the actual instructions completed by yourĀ  CPU are in a binary language called machine code, and the semi-readable version of the instructionsĀ  is called assembly, but we modified this assembly a little bit to make it more readable. One interesting note is that you may think that with everything a computer can doĀ  there must be tens of thousands of different instructions. Well actually the 6502 processorĀ  in the Apple 2e from 1983 could only complete 56 different instructions whereas the modern M1 chipĀ  in the MacBook Pro can complete 354 instructions.
18
Here’s the list of all the instructions eachĀ  chip can execute and if you take a good look, most of these instructions are rather simple.Ā  Let’s just think about that for a second. Every single thing you do on your computer can beĀ  constructed using only various sequences of 354 different instructions. However many programs haveĀ  millions upon millions of lines of instructions, and hopefully, there aren’t any bugs in them. So now that we’ve discussed the range of possible instructions, let’s further explore how CPUsĀ  work. In order to complete an instruction there are always three key steps: Fetch, Decode,Ā  and Execute. The first step is Fetch and is where the controller uses the value in theĀ  program counter to search through the pages of instructions in the book for the correspondingĀ  instruction address. The controller then copies the instruction found at that addressĀ  into a special location called the current instruction register or CIR. At the same time theĀ  controller increases the program counter by 1.
19
The second step is decode, and in this stepĀ  the current instruction is fed into a circuit called the instruction decoder. In our analogyĀ  from earlier, this decoder is a key part of the controller, and in essence it’s the circuitry thatĀ  reads in an instruction and both interprets what the machine code of an instruction actually does,Ā  and simultaneously produces the control signals to properly execute that instruction. Specifically,Ā  this instruction decoder circuit uses the binary values of the instruction and an incrediblyĀ  complex arrangement of logic gates to produce the corresponding control signals which areĀ  then sent to the different elements in the CPU.
20
Instruction decoders are one of the moreĀ  complicated parts of the CPU but here’s an example along with a simplified explanation.Ā  Let’s say we have this ADD instruction in the current instruction register or CIR and it’s fedĀ  into the instruction decoder. The first part of the binary instruction specifies we want to useĀ  the ALU. With the ALU selected, the next 3 bits specify that we want to use the ADD function,Ā  and then the last 4 bits of the instruction indicates we want the values in register 0 andĀ  register 1 to be routed and sent to the ALU.
21
Instruction decoding is considerably moreĀ  complicated than that but let’s move onto the third step which is Execute. DuringĀ  execute, using our example instruction, the control signals from the instruction decoderĀ  and an intricate set of electrical timing signals are used to first send the value in registerĀ  0 and then the value in register 1 to the ALU.
22
The timing signals are used to accommodate theĀ  time it takes electricity to travel from the registers to the ALU and for transistors and logicĀ  gates to change their state, thereby ensuring the correct result at the output. After the values areĀ  input, the ALU adds the two numbers together, and a subsequent timing signal saves the result intoĀ  the accumulator, thus completing the Execute step.
23
These three steps, Fetch, Decode, and Execute areĀ  used to complete a single line of instructions and once it’s completed, these steps repeat butĀ  using the new value in the program counter. In essence Fetch, Decode, and Execute form a cycleĀ  so let’s run through it again. During Fetch, the controller uses the program counter’s value toĀ  fetch the corresponding instruction and places it in the CIR and the program counter increasesĀ  by 1. Next during Decode the instruction’s binary is fed into the instruction decoderĀ  where a complex set of logic gates generate the correct electrical control signals forĀ  that instruction. Finally, during Execute, the instruction is completed using the controlĀ  signals and timing signals, and in this case, the value in the accumulator is stored backĀ  into a memory address which, using the analogy, is like writing the value from the calculatorĀ  display into a data location in the book. Then the Fetch, Decode, Execute cycle repeats againĀ  using the next program counter’s value and so on.
24
The Fetch Decode Execute Cycle is used in everyĀ  processor no matter whether it’s the 6502 in the Apple2e or the M1 in the MacBook Pro. But ofĀ  course there are many differences such as the size of the cache, the registers, or functions on theĀ  ALU calculator and much more. We’ll explore the exact differences in a few minutes, but forĀ  now, one important detail is that the Fetch Decode Execute cycle uses your computer’s clockĀ  to regulate its pace. The 6502 chip had a One Megahertz clock which ticked away at a millionĀ  times a second and thus each step in the Fetch, Decode, Execute cycle took a microsecond.Ā  Additionally the 6502 was an 8-bit processor meaning the size of the registers and the ALU’sĀ  input and output were 8-bits wide. On the other hand, the M1 chip is a 64-bit processor, so itĀ  can handle much larger numbers and it uses a 3.2 Gigahertz clock and therefore each step takes aĀ  third of a nanosecond. Additionally, the M1 chip, along with all modern chips, uses a techniqueĀ  called pipelining where multiple instructions are queued up resulting in fetch, decode, and executeĀ  for different program counter values and different instructions being completed at the same time. There are many other optimizations in modern processors that we’ll soon discussĀ  but it’s important to understand that from the second you turn on your laptop,Ā  smartphone, gaming console, GPU or AI Server, to the second you shut it off, the processorĀ  is continuously cycling through Fetch, Decode, Execute over and over using programs filled withĀ  instructions and data along with the CPUs clock to regulate its pace. In essence this Fetch,Ā  Decode, Execute cycle is the common section of technological DNA that has powered everyĀ  single processor built over the past 50 years.
25
This cycle of steps is incredibly powerful,Ā  capable of performing trillions to quadrillions of mathematical operations every second inĀ  a single chip. But you may be wondering, are there alternatives to the Fetch Decode ExecuteĀ  cycle? Well, there’s a world of different kinds of microchips, but specifically, alternativesĀ  include Application Specific Integrated Circuits or ASICs such as these microchips found in thisĀ  bitcoin mining computer, or Field Programmable Gate Arrays or FPGAs which are the main chipsĀ  in a number automotive computers and cameras.
26
Both ASICs and FPGAs don’t use the fetch andĀ  decode steps, but rather they perform repetitive operations by flowing data through a set patternĀ  of logic gates and execution units, making them highly optimized, but very inflexible. AndĀ  then even further from these chips, are Quantum Computers which are based on Qubits and QuantumĀ  circuits which we’ll discuss in future videos.
27
But, now that we’ve covered the Fetch, Decode,Ā  Execute cycle, it’s important to discuss two more steps which are Memory and Writeback. Memory isĀ  analogous to moving books from the SSD bookshelves onto the DRAM cart and then onto the table orĀ  Cache Memory, and Writeback is like writing data into the books, and when space for a new book isĀ  needed on the table, the old book is placed back on the DRAM cart and eventually returned to theĀ  bookshelves. These two steps use another special location called the memory address register andĀ  are critical to a functioning computer, but they typically take a lot longer to complete than theĀ  Fetch Decode Execute steps, and therefore in some architectures and textbooks they’re includedĀ  in the cycle and sometimes they aren’t. We’re working on a separate video on how data movesĀ  around these memory locations, so stay tuned.
28
Now that we’ve uncovered the technologicalĀ  DNA inside all processors, it’s important to note that, similar to the DNA found in theĀ  nucleus of the cell and there being multiple layers of biological organization and structureĀ  for all living things, there are many layers of complexity or abstraction between the FetchĀ  Decode Execute cycle and a computer running a video game or browsing the internet. IfĀ  you want to dive into some of the other layers and understand more about how computersĀ  work, we recommend you check out Brilliant which is the sponsor of this video. BrilliantĀ  has a massive library of interactive courses that include subjects like calculus, scientificĀ  thinking, circuits, programming in python, logic, data analysis, and many more topics that wouldĀ  take far too long to list. However, Brilliant is much more than a list of courses, rather it’sĀ  as if your favorite teacher who makes classes engaging is combined with your favorite video gameĀ  and then mixed with the knowledge from countless textbooks. The result would be Brilliant. Their mission is to create a world of better problem solvers, and every one of theirĀ  courses focuses on critical thinking through interactive games and lessons. Furthermore,Ā  with technology progressing faster than ever, Brilliant continuously updates their lessonsĀ  to anticipate what you need to know for your education and career. For example, they have aĀ  new course on AI and Large Language Models that explains how Generative AI works far betterĀ  than any other textbook or video out there.
29
Develop your knowledge by learning a little everyĀ  day. You can start today by signing up for free using the link: Brilliant.org/BranchEducation,Ā  or by scanning the QR code on screen, and you’ll then have access to the wide range ofĀ  courses throughout their catalog. If you enjoy their content and decide to stay, the link in theĀ  description below will also save you 20% off an annual premium subscription, which will give youĀ  unlimited daily access to everything on Brilliant.
30
Ok, so let’s quickly run through slightlyĀ  more advanced topics to finish up this video.
31
Earlier we mentioned that the MacBook Pro’s M1Ā  chip can complete 354 different instructions.
32
This set of instructions is called ArmV8.4 andĀ  it’s categorized as a RISC architecture or Reduced Instruction Set Computer. For example, here’s aĀ  simple game of Snake using 145 lines of C++ code.
33
It’s the job of a compiler, which is a separateĀ  piece of software, to take this code along with ArmV8.4’s 354 RISC instructions, and generateĀ  a list of 676 assembly instructions equivalent to the machine code instructions that wouldĀ  be found in a book or program named snake.app The other common architecture found in Intel andĀ  AMD chips is called CISC, or Complex Instruction Set Computer and is composed of thousands ofĀ  different possible instructions. For example, here’s the equivalent snake program thatĀ  is compiled to run on an Intel or AMD Chip using CISC and x86-64bit instructions, andĀ  you can see it’s only 560 instructions now.
34
A few key differences between RISC and CISCĀ  are that each RISC instruction is relatively simple and is executed at a consistently fastĀ  execution rate. Additionally, RISC architectures are more energy efficient and thus used inĀ  all smartphones, whereas CISC architectures have thousands of different instructions andĀ  pack a lot more into a single instruction.
35
Additionally, the CISC instructionĀ  decoder is much more complicated, and individual instructions have a variableĀ  execution rate sometimes taking multiple clock cycles to execute. There are many additionalĀ  pros and cons to RISC vs CISC which we’ll save for yet another video, but we thought it worthĀ  mentioning these simplified differences here.
36
Computer architecture is incredibly complicatedĀ  with many different facets and layers of complexity and we have plans to make more videosĀ  that dive into each of these topics, but it’s important to note that each video we make takesĀ  close to a combined 1100 hours of researching, script writing, modeling, animating and editing.Ā  For example, we spent over 250 hours tearing down these non-working computers we bought from Ebay,Ā  and meticulously rebuilding each of the 3D models in Blender. So, if you could take a few seconds toĀ  like this video, subscribe if you haven’t already, share this video with someone who might beĀ  curious as to how CPUs work, and most importantly, write a quick comment below it would help us outĀ  immensely. Just a few seconds of your time helps us far more than you think. So, thank you. In the final section of this video we’ll discuss this diagram we showed earlier and theĀ  architecture of modern processors such as the M1. In contrast, the analogy we’ve laid outĀ  is rather simple, and you’re probably thinking that there must be more components in an actualĀ  CPU. In fact, this analogy is actually pretty close to what’s happening inside an Apple2eĀ  Computer. Specifically, the floppy drives, are the bookshelves, and then when we open upĀ  this computer we see the DRAM chips, which are the cart, and then going inside the 6502 processorĀ  we find an integrated circuit or die which has the corresponding sections that we’ll organize intoĀ  an internal architectural diagram. In this diagram you can see the instruction decoder, ALU, theĀ  Program Counter, Current Instruction Register, the other registers and a few other sections.Ā  Specifically, here’s where the program counter is used to fetch an instruction, and here’s whereĀ  the instructions and data from the DRAM chips are bussed in and out. Finally, here’s whereĀ  the instructions are decoded and the control signals are generated. One note is that there’sĀ  no cache in the 6502 because the DRAM chips in the 80s were just as fast as the instructions,Ā  so the table in the analogy is even smaller.
37
As we said at the beginning of this video,Ā  the 6502 chip is made from 4528 transistors, so let’s see what an M1 chip with 16 billionĀ  transistors would look like. To start, we have to significantly increase the size ofĀ  this table. Next, we have to section off areas for each of the performance and energy efficientĀ  cores, the GPU, and other areas. When we focus on one of the performance cores we see the complexĀ  diagram from earlier, so let’s discuss how this diagram compares to our analogy. Specifically,Ā  there’s a separate set of 64 kilobyte data and instruction caches. As mentioned earlier there’sĀ  a pipeline to queue 8 instructions per clock cycle and additional sections like a branch predictorĀ  to reduce issues with conditional branching and help the pipeline run smoothly. Here you canĀ  see the pipelined instruction decoder, and 32 general purpose registers. One key difference isĀ  that the calculator is broken up into 8 separate smaller calculators each handling a few functions.Ā  Additionally, there’s a special section for load and store instructions. This is the layout of justĀ  a single core out of the 8 and there are entirely different architectures in the Graphics ProcessingĀ  Unit as well as inside the Neural Processing Unit.
38
One important note is that the inclusionĀ  of these 3 types of processors along with hardware accelerators such as the mediaĀ  engine makes this M1 chip closer to a system on a chip or SoC than a traditional CPU.Ā  Similarly, all the processors in these devices, including the CPU in your desktop computer canĀ  be considered SoCs and therefore the difference is more a marketing term than a technical one. On a separate note it’s important to mention that the M1 along with all modern processors areĀ  proprietary designs and knowledge, and therefore the diagrams we’ve shown are close approximationsĀ  that we built using input from industry experts.
39
Let’s finally discuss our analogy in the termsĀ  of GPU chips found in graphics cards. We have a separate video covering how graphics cardsĀ  work, but with respect to this analogy, a GPU CUDA core is actually very similar in complexityĀ  to the architecture of the 6502. Therefore with 10,000 to 20,000 CUDA cores in a single GPU chip,Ā  it’s like having a massive array of 6502 cores.
40
The difference is that GPUs typically useĀ  32-bit ALU calculators and perform single instruction multiple thread or SIMT calculationsĀ  where a single instruction is fetched, decoded, and then distributed to a batch of cores, andĀ  then those cores execute that instruction using different addresses and data. However, there areĀ  many more nuances to SIMT and GPU architecture, so let’s wrap up this video on how CPUs work. We’re thankful to all our Patreon and YouTube Membership Sponsors for supporting our videos.Ā  If you want to financially support our work, you can find the links in the description below. This is Branch Education, and we create 3D animations that dive deeply into the technologyĀ  that drives our modern world. Watch another Branch video by clicking one of these cards or clickĀ  here to subscribe. Thanks for watching to the end!

Vocabulary and speaking notes for this lesson

This video has 40 sentences and 5002 words to shadow. The speech runs for 36:22. The speaker talks at a steady 138 words per minute, a comfortable pace for shadowing. Only 76% of the words are among the 3,000 most common in English, so the vocabulary is demanding.

Key vocabulary in this video

15 words from the video worth learning, with pronunciation and meaning:

WordPronunciationMeaning
instruction noun/ÉŖnˈstɹʌkŹƒÉ™n/The act of instructing, teaching, or providing with information or knowledge.
chip noun/t͔ʃɪp/A small piece broken from a larger piece of solid material.
execute verb/ĖˆÉ›ksɪˌkjuːt/To kill, especially as punishment for a capital crime.
fetch verb/fɛt͔ʃ/To retrieve; to bear towards; to go and get.
processor noun/ˈpÉ¹É‘ĖŒsɛsɚ/A person or institution who processes things (foods, photos, applications, etc.).
analogy noun/É™ĖˆnƦləd͔ʒi/A relationship of resemblance or equivalence between two situations, people, or objects, especially when used as a basis for explanation or extrapolation.
calculator noun/kƦl.kjə.leÉŖ.tɚ/A mechanical or electronic device that performs mathematical calculations; (now usually) an electronic one specifically.
loop noun/luːp/A length of thread, line or rope that is doubled over to make an opening.
controller noun/kənˈtɹoʊlɚ/One who controls something.
cart noun/kɑɹt/A small, open, wheeled vehicle, drawn or pushed by a person or animal, often with two wheels on one axle, more often used for transporting goods than…
circuit noun/ˈsɝ.kÉŖt/The act of moving or revolving around, or as in a circle or orbit; a revolution
correspond verb/ˌkoÉ¹É™Ėˆspɑnd/To be equivalent or similar in character, quantity, quality, origin, structure, function etc.
gate noun/ˈɔeɪ̯t/A doorlike structure outside a house.
graphics noun/ˈɔɹæfɪks/The making of architectural or design drawings.
binary adjective/ˈbaÉŖ.nə.ɹi/Being in one of two mutually exclusive states.

Phrasal verbs you will hear

WordPronunciationMeaning
move around verbTo relocate to new homes repeatedly; to not live in any one place for long.
open up verbTo open.
run through verbTo summarise briefly.
check out verbTo record the departure or withdrawal of someone or something (such as guests, employees, books, etc.).
find out verbTo discover, as by asking or investigating.
finish up verbTo complete the last details of a task.
go up verbTo move upwards.
lay out verb/ˌleɪ ˈaʊt/To expend or contribute money to an expense or purchase.

Pronunciation to watch

The speaker uses 34 contractions and reduced forms, such as we'll, we've, we're. Say them the short way, as you hear them.

  • The ā€œthā€ sounds: thread /ĖˆĪøÉ¹É›d/, mathematical /ˌmæθ(.ə)ˈmƦt.ÉŖ.kəl/, thereby /Ć°É›É¹ĖˆbaÉŖ/
  • The ā€œshā€ and ā€œzhā€ sounds: instruction /ÉŖnˈstɹʌkŹƒÉ™n/, cache /kæʃ/, sheet /ʃit/, efficient /ɪˈfÉŖŹƒÉ™nt/, essentially /ɪˈsɛnŹƒÉ™li/
  • Long words — get the stress right: analogy /É™ĖˆnƦləd͔ʒi/, calculator /kƦl.kjə.leÉŖ.tɚ/, architecture /ĖˆÉ‘Ė.kɪˌtɛk.tĶ”ŹƒÉ™/, technological /ˌtɛk.nÉ™Ėˆlɑ.dŹ’ÉŖ.kəl/, complicated /ˈkɑm.plɪˌkeÉŖ.tÉŖd/

How to practise with this video

  1. Listen to the whole video once without speaking and note the words you do not know.
  2. Shadow it sentence by sentence at normal speed, repeating each one until your rhythm matches the speaker.
  3. Record yourself and compare with the original, paying attention to words like instruction, chip, execute.

Grammar in this video

The structures the speaker uses most, with the exact words from the video:

StructureIn the video
ā€œUsed toā€ used to + verb — a past habit or state that is no longer trueused to save Ā· used to complete
Passive voice be + past participle — the focus is on what happens, not who does itis built Ā· being released Ā· is sponsored
Relative clauses who / which + clause — extra information about a person or thingprocessor which is Ā· DRAM which is Ā· book which is
Present perfect have/has + past participle — a past action that still matters nowwe've explained Ā· we've discussed Ā· has powered

What is the Shadowing Technique?

Shadowing is a science-backed language learning technique originally developed for professional interpreter training and popularized by polyglot Dr. Alexander Arguelles. The method is simple but powerful: you listen to native English audio and immediately repeat it out loud — like a shadow following the speaker with just a 1–2 second delay. Unlike passive listening or grammar drills, shadowing forces your brain and mouth muscles to simultaneously process and reproduce real speech patterns. Research shows it significantly improves pronunciation accuracy, intonation, rhythm, connected speech, listening comprehension, and speaking fluency — making it one of the most effective methods for IELTS Speaking preparation and real-world English communication.

Shadowing technique: read the full step-by-step guide →