쉐도잉 연습: Natural Language Processing - Tokenization (NLP Zero to Hero - Part 1) - 영상으로 영어 말하기 배우기

레슨 만드는 중...
1
Hi, and welcome to this series on Zero2Hero for Natural Language Processing using TensorFlow.
2
If you're not an expert on AI or ML, don't worry.
3
We're taking the concepts of NLP and teaching them from first principles.
4
In this first lesson, we'll talk about how to represent words in a way that a computer can process them, with a view to later training a neural network that can understand their meaning.
5
This process is called tokenization.
6
So let's take a look.
7
Consider the word listen, as you can see here.
8
It's made up of a sequence of letters.
9
These letters can be represented by numbers using an encoding scheme.
10
A popular one called ASCII has these letters represented by these numbers.
11
This bunch of numbers can then represent the word listen.
12
But the word silent has the same letters, and thus the same numbers, just in a different order.
13
So it makes it hard for us to understand sentiment of a word just by the letters in it.
14
So it might be easier, instead of encoding letters, to encode words.
15
Consider the sentence, I love my dog.
16
So what would happen if we start encoding the words in this sentence instead of the letters in each words?
17
So for example, the word I could be one.
18
And then the sentence, I love my dog, could be one, two, three, four.
19
Now, if I take another sentence, for example, I love my cat, how would we encode it?
20
Now we see I love my has already been given one, two, three.
21
So all I need to do is encode a cat.
22
I'll give that the number five.
23
And now if we look at the two sentences, they are one, two, three, four and one, two, three,
24
five, which already show some form of similarity between them.
25
And it's a similarity you'd expect because they're both about loving a pet.
26
Given this method of encoding sentences into numbers, now let's take a look at some code to achieve this for us.
27
This process, as I mentioned before, is called tokenization, and there's an API for that.
28
We'll look at how to use it with Python.
29
So here's your first look at some code to tokenize these sentences.
30
Let's go through it line by line.
31
First of all, we'll need the tokenizer APIs, and we can get these from TensorFlow Keras like this.
32
We can represent our sentences as a Python array of strings like this.
33
It's simply the I love my dog and I love my cat that we saw earlier.
34
Now the fun begins.
35
I can create an instance of a tokenizer object.
36
The numWords parameter is the maximum number of words to keep.
37
So instead of, for example, just these two sentences, imagine if we had hundreds of books to tokenize.
38
But we just want the most frequent 100 words in all of that.
39
This would automatically do that for us when we do the next step.
40
And that's to tell the tokenizer to go through all the text and then fit itself to them like this.
41
The full list of words is available as the tokenizer's word index property.
42
So we can take a look at it like this and then simply print it out.
43
The result will be this dictionary showing the key being the word and the value being the token for that word.
44
So, for example, my has a value of three.
45
The tokenizer is also smart enough to catch some exceptions.
46
So, for example, if we updated our sentences to this by adding a third sentence, noting that dog here is followed by an exclamation mark.
47
The nice thing is that the tokenizer is smart enough to spot this and not create a new token.
48
It's just dog.
49
And you can see the results here.
50
There's no token for dog exclamation, but there is one for dog.
51
And there's also a new token for the word you.
52
If you want to try this out for yourself, I've put the code in a colab here.
53
Take it for a spin and experiment.
54
You've now seen how words can be tokenized and the tools in TensorFlow that handle that tokenization for you.
55
Now that your words are represented by numbers like this, you'll next need to represent your sentences by sequences of numbers in the correct order.
56
You'll then have data ready for processing by a neural network to understand or maybe even generate new text.
57
You'll see the tools that you can use to manage this sequencing in the next episode.
58
So don't forget to hit that subscribe button.

NLP의 기초: 토큰화, 영어 학습과의 연결

이 비디오는 자연어 처리(NLP)의 기초 개념인 토큰화에 대해 설명하는데, 영어 학습자에게도 유용한 통찰을 제공합니다. 컴퓨터가 언어를 이해하려면 단어를 숫자로 변환하는 과정이 필요한데, 이는 우리가 영어 단어의 의미와 사용법을 익히는 과정과도 닮아 있습니다. 비디오에서는 "I love my dog"와 "I love my cat" 같은 간단한 문장을 예로, 단어별로 토큰을 부여하는 방법을 보여줍니다. 이를 통해 컴퓨터가 문장의 유사성을 파악할 수 있다는 점은, 영어 회화에서 문맥을 이해하는 것과 같은 원리입니다.

일상 대화에 유용한 5가지 표현

  • "I love my dog" - 가장 기본적인 소유와 감정 표현, 동물을 사랑한다는 말로 자주 사용됩니다.
  • "I love my cat" - "dog" 대신 "cat"을 사용한 변형, 같은 구조로 다른 동물 이름을 넣어 쉽게 응용할 수 있습니다.
  • "tokenization" - NLP 용어이지만, 영어로 "단어를 토큰으로 변환하다"는 의미로도 사용될 수 있습니다.
  • "word index" - 단어와 토큰의 대응 관계를 나타내는 용어, 학습에서 단어 목록을 정리할 때 참고할 수 있습니다.
  • "fit to text" - 토크나이저가 텍스트에 맞게 설정된다는 뜻, 영어 학습에서도 "텍스트에 맞춰 연습하다"는 의미로 사용할 수 있습니다.

Shadowing 연습 가이드: 비디오를 활용한 단계별 방법

이 비디오는 영어 발음과 문장 구조를 연습하기 좋은 자료입니다. shadowspeak 또는 shadow speech 기법을 적용해 보세요. 다음 단계를 따라해 보세요:

  1. 비디오를 1~2회 듣고 전체 내용을 이해합니다. 주요 용어인 "tokenization"과 "word index"의 의미를 파악합니다.
  2. 문장을 짧게 나누어 반복해서 듣고 따라합니다. "I love my dog"와 "I love my cat"을 여러 번 연습하여 발음과 리듬을 익힙니다.
  3. 토크나이저가 처리하는 과정을 설명하는 부분을 shadowing합니다. "The tokenizer is smart enough to catch some exceptions"와 같은 문장은 복잡하지만, 천천히 따라하며 문법 구조를 분석합니다.
  4. 자신의 목소리를 녹음하여 원본과 비교합니다. 발음의 차이점을 찾고, 부족한 부분을 보완합니다.
  5. 비디오의 내용을 바탕으로 자신만의 문장을 만들어 봅니다. "I love my bird"나 "I love my rabbit"과 같이 변형해 사용해 보세요.

이렇게 연습하면 영어 회화 실력뿐만 아니라, NLP의 기초 개념도 함께 학습할 수 있습니다. shadowing site에서 더 다양한 자료를 찾아 연습하면 효과가 더 좋습니다.

쉐도잉이란? 영어 실력을 빠르게 키우는 과학적 방법

쉐도잉(Shadowing)은 원래 전문 통역사 훈련을 위해 개발된 언어 학습 기법으로, 다언어 학자인 Dr. Alexander Arguelles에 의해 대중화된 방법입니다. 핵심 원리는 간단하지만 매우 강력합니다: 원어민의 영어를 들으면서 1~2초의 짧은 지연으로 즉시 소리 내어 따라 말하는 것——마치 '그림자(shadow)'처럼 화자를 따라가는 것입니다. 문법 공부나 수동적인 청취와 달리, 쉐도잉은 뇌와 입 근육이 동시에 실시간으로 영어를 처리하고 재현하도록 훈련합니다. 연구에 따르면 이 방법은 발음 정확도, 억양, 리듬, 연음, 청취력, 말하기 유창성을 크게 향상시킵니다. IELTS 스피킹 준비와 자연스러운 영어 소통을 원하는 분들에게 특히 효과적입니다.

섀도잉 방법: 단계별 전체 가이드 읽기 →