Shadowing Practice: What is Statistics? A Beginner's Guide to Statistics (Data Analytics)! - Learn English Speaking with Video

Creating lesson...
1
If you want to finally understand statistics, this is the place to be.
2
After this video, you will know what statistics is, what descriptive statistics is, and what inferential statistics is.
3
So, let's start with the first question: What is statistics?
4
Statistics deals with the collection, analysis and presentation of data.
5
An example: We would like to investigate whether gender has an influence on the preferred newspaper.
6
Then gender and newspaper are our so -called variables that we want to analyze.
7
In order to analyze whether gender has an influence on the preferred newspaper, we first need to collect data.
8
To do this, we create a questionnaire that asks about gender and preferred newspaper.
9
We will then send out the survey and wait two weeks.
10
Afterwards, we can display the received answers in a table.
11
In this table, we have one column for each variable one for gender and one for newspaper.
12
On the other hand, each row is the response of one surveyed person.
13
The first respondent is male and stated "New York Post".
14
The second is female and stated "USA Today" and so on and so forth.
15
Of course, the data does not have to be from a survey.
16
The data can also come from an experiment in which you, for example, want to study the effect of two drugs on blood pressure.
17
Now the first step is done.
18
We have collected data and we can start analyzing the data.
19
But what do we actually want to analyze?
20
We did not survey the entire population, but we took a sample.
21
Now the big question is: do we just want to describe the sample data
22
or do we want to make a statement about the whole population?
23
If our aim is limited to the sample itself, i .e we only want to describe the collected data, We will use descriptive statistics.
24
Descriptive statistics will provide a detailed summary of the sample.
25
However, if we want to draw conclusions about the population as a whole, Inferential statistics are used.
26
This approach allows us to make educated guesses about the population based on the sample data.
27
Let us take a closer look at both methods, starting with descriptive statistics.
28
Why is descriptive statistics so important?
29
Let's say a company wants to know how its employees travel to work.
30
So, the company creates a survey to answer this question.
31
Once enough data has been collected, this data can be analyzed using descriptive statistics.
32
But what is descriptive statistics?
33
Descriptive statistics aims to describe and summarize a dataset in a meaningful way.
34
But it is important to note that descriptive statistics only describe the collected data without drawing conclusions about a larger population.
35
Put simply, just because we know how some people from one company get to work,
36
we cannot say how all working people of the company get to work.
37
This is the task of inferential statistics, which we will discuss later.
38
To describe data descriptively, we now look at the four key components: measures of central tendency,
39
measures of dispersion, frequency tables and charts.
40
Let's start with the first one: Measures of Central Tendency.
41
Measures of Central Tendency are for example the mean, the median and the mode.
42
Let's first have a look at the mean.
43
The arithmetic mean is the sum of all observations divided by the number of observations.
44
An example.
45
Imagine we have the test scores of 5 students.
46
To find the mean score, we sum up all the scores and divide by the number of scores.
47
The mean test score of these 5 students is therefore 86 .6.
48
What about the median?
49
When the values in a dataset are arranged in ascending order, the median is the middle value.
50
If there is an odd number of data points, The median is simply the middle value.
51
If there is an even number of data points, the median is the average of the two middle values.
52
It is important to note that the median is resistant to extreme values or outliers.
53
Let's look at this example: No matter how tall the last person is, The person in the middle remains the person in the middle,
54
so the median does not change.
55
But if we look at the mean, it does have an effect on how tall the last person is.
56
The mean is therefore not robust to outliers.
57
Let's continue with the mode.
58
The mode refers to the value or values that appear most frequently in a set of data.
59
For example, if 14 people travel to work by car, 6 by bike, 5 walk and 5 take public transport,
60
Then "car" occurs most often and is therefore the mode.
61
Great, let's continue with the measures of dispersion.
62
Measures of dispersion describe how spread out the values in a dataset are.
63
Measures of dispersion are for example the variance and standard deviation, the range and the interquartile range.
64
Let's start with the standard deviation.
65
The standard deviation indicates the average distance between each data point and the mean.
66
But what does that mean?
67
Each person has some deviation from the mean.
68
Now we want to know how much the persons deviate from the mean value on average.
69
In this example, the average deviation from the mean value is 11 .5 cm.
70
To calculate the standard deviation, we can use this equation: sigma is the standard deviation, n is the number of persons,
71
xi is the size of each person and x bar is the mean value of all persons.
72
But attention, there are two slightly different equations for the standard deviation.
73
The difference is that we have once 1 /n and once 1 / .
74
To keep it simple, if our survey doesn't cover the whole population, we always use this equation to estimate the standard deviation.
75
Likewise, if we have conducted a clinical study, then we also use this equation to estimate the standard deviation.
76
But what is the difference between the standard deviation and the variance?
77
As we now know, the standard deviation is the quadratic mean of the distance from the mean.
78
The variance now is the squared standard deviation.
79
If you want to know more details about the standard deviation and the variance, please watch our video.
80
Let's move on to range and interquartile range.
81
It is easy to understand.
82
The range is simply the difference between the maximum and minimum value.
83
Interquartile range represents the middle 50 % of the data.
84
It is the difference between the first quartile, Q1, and the third quartile, Q3.
85
Therefore, 25 % of the values are smaller than the interquartile range and 25 % of the values are larger.
86
The interquartile range contains exactly the middle 50 % of the values.
87
Before we get to the last two points, let's briefly compare measures of central tendency and measures of dispersion.
88
Let's say we measured the blood pressure of patients.
89
Measures of central tendency provide a single value that represents the entire dataset,
90
helping to identify a central value around which data points tend to cluster.
91
Measures of dispersion, like the standard deviation, the range and the interquartile range, indicate how spread out the data points are,
92
whether they are closely packed around the center or spread far from it.
93
In summary, while measures of central tendency provide a central point of the dataset,
94
Measures of dispersion describe how the data is spread around the center.
95
Let's move on to tables.
96
Here we will have a look at the most important ones: frequency tables and contingency tables.
97
A frequency table displays how often each distinct value appears in a dataset.
98
Let's have a closer look at the example from the beginning.
99
A company surveyed its employees to find out how they get to work.
100
The options given were car, bicycle, walk and public transport.
101
Here are the results from 30 employees: the first answered car, the next walk, and so on and so forth.
102
Now we can create a frequency table to summarize this data.
103
To do this, we simply enter the four possible options: car,
104
bicycle, walk and public transport in the first column and then count how often they occurred.
105
From the table, it is evident that the most common mode of transport among the are with 14 employees preferring it.
106
The frequency table thus provides a clear and concise summary of the data.
107
But what if we have not only one but two categorical variables?
108
This is where the contingency table, also called crosstab, comes in.
109
Imagine the company doesn't have one factory but two: one in Detroit and one in Cleveland.
110
So, we also ask the employees at which location they work.
111
If we want to display both variables, we can use a contingency table.
112
A contingency table provides a way to analyze and compare the relationship between two categorical variables.
113
The rows of a contingency table represent the categories of one variable, while the columns represent the categories of another variable.
114
Each cell in the table shows the number of observations that fall into the corresponding category combination.
115
For example, the first cell shows that Carr and Detroit were answered six times.
116
And what about the charts?
117
Let's take a look at the most important ones.
118
To do this, let's simply use datadap .net.
119
If you like, you can load this sample data set with the link in the video description.
120
Or you just copy your own data into this table.
121
Here below you can see the variables distance to work, mode of transport and site.
122
Datadap gives you a hint about the level of measurement, but you can also change it here.
123
Now, if we only click on Mode of Transport, we get a frequency table and we can also display the percentage values.
124
If we scroll down, we get a bar chart and a pie chart.
125
Here on the left we can adjust further settings.
126
For example, we can specify whether we want to display the frequencies or the percentage values,
127
or whether the bars should be vertical or horizontal.
128
If you also select "Side", we get a cross table here, and a grouped bar chart for the diagrams.
129
Here we can specify whether we want the chart to be grouped or stacked.
130
If we click on Distance to Work and Mode of Transport,
131
we get a bar chart where the height of the bar shows the mean value of the individual groups.
132
Here we can also display the dispersion.
133
We also get a histogram, a box plot, a violin plot and a rainbow plot.
134
If you would like to know more about what a box plot, a violin plot and a rainbow plot are, take a look at my videos.
135
Let's continue with inferential statistics.
136
At the beginning, we briefly go through what inferential statistics is and then I'll explain the six key components to you.
137
So, what is inferential statistics?
138
Inferential statistics allows us to make a conclusion or inference about a population based on data from a sample.
139
What is the population and what is the sample?
140
The population is the whole group we're interested in.
141
If you want to study the average height of all adults in the United States, then the population would be all adults in the United States.
142
The sample is the smaller group we actually study, chosen from the population.
143
For example, 150 adults were selected from the United States.
144
And now we want to use the sample to make a statement about the population.
145
And here are the six steps how to do that: 1.
146
Hypothesis.
147
First, we need a statement, a hypothesis, that we want to test. For example,
148
you want to know whether a drug will have a positive effect on blood pressure in people with high blood pressure.
149
But what's next?
150
In our hypothesis, we stated that we would like to study people with high blood pressure.
151
So, our population is all people with high blood pressure in, for example, the US.
152
Obviously, we cannot collect data from the whole population.
153
So, we take a sample from the population.
154
Now, we use this sample to make a statement about the population.
155
But how do we do that?
156
For this, we need a hypothesis test.
157
Hypothesis testing is a method for testing a claim about a parameter in a population, using data measured in a sample.
158
Great, that's exactly what we need.
159
There are many different hypothesis tests.
160
And at the end of this video I will give you a guide on how to find the right test.
161
And of course you can find videos about many more hypothesis tests on our channel.
162
But how does a hypothesis test work?
163
When we conduct a hypothesis test, we start with a research hypothesis, also called alternative hypothesis.
164
This is the hypothesis we are trying to find evidence for.
165
In our case, the research hypothesis is: the drug has an effect on blood pressure.
166
But we cannot test this hypothesis directly with a classical hypothesis test,
167
So, we test the opposite hypothesis: that the drug has no effect on blood pressure.
168
But what does that mean?
169
First, we assume that the drug has no effect in a population.
170
We therefore assume that in general, people who take the drug and people who don't take the drug have the same blood pressure on average.
171
If we now take a random sample and it turns out that the drug has a large effect in a sample,
172
then we can ask: how likely it is to draw such a sample
173
or one that deviates even more if the drug actually has no effect?
174
So, in reality, on average, there is no difference in the population.
175
If this probability is very low, we can ask ourselves: maybe the drug has an effect in the population?
176
have enough evidence to reject the null hypothesis that the drug has no effect.
177
And it is this probability that is called the p -value.
178
Let's summarize this in three simple steps: 1.
179
The null hypothesis states that there is no difference in the population.
180
2. The hypothesis test calculates how much the sample deviates from the null hypothesis.
181
3. The p -value indicates the probability of getting a sample
182
that deviates as much as our sample or one that even deviates more than our sample, is true.
183
But at what point is the p -value small enough for us to reject the nile hypothesis?
184
This brings us to the next point: statistical significance.
185
If the p -value is less than a predetermined threshold, the result is considered statistically significant.
186
This means that the result is unlikely to have occurred by chance alone
187
and that we have enough evidence to reject the null hypothesis.
188
This threshold is often 0 .05.
189
Therefore, a small p -value suggests that the consistent with the null hypothesis.
190
This leads us to reject the null hypothesis in favor of the alternative hypothesis.
191
A large p -value suggests that the observed data is consistent with the null hypothesis and we will not reject it.
192
But note: there is always a risk of making an error.
193
A small p -value does not prove that the alternative hypothesis is true.
194
It is only saying that it is unlikely to get such a result
195
or a more extreme when the null hypothesis is true.
196
And again, if the null hypothesis is true, there is no difference in a population.
197
And the other way around,
198
a large p -value does not prove that the null hypothesis to get such a result or a more extreme,
199
when the null hypothesis is true.
200
So, there are two types of errors which are called type 1 and type 2 error.
201
Let's start with the type 1 error.
202
In hypothesis testing, a type 1 error occurs when a true null hypothesis is rejected.
203
So, in reality, the null hypothesis is true, but we make the decision to reject the null hypothesis.
204
In our example, it means that the drug actually had no difference in blood pressure whether the drug is taken or not,
205
the blood pressure remains the same in both cases.
206
But our sample happened to be so far off the true value that we mistakenly thought the drug was working.
207
And a type 2 error occurs when a false null hypothesis is not rejected.
208
So, in reality, the null hypothesis is false, But we make the decision not to reject the null hypothesis.
209
In our example, this means: the drug actually did work, there is a difference between those who have taken the drug and those who have not,
210
but it was just a coincidence that the sample taken did not show much difference
211
and we mistakenly thought the drug was not working.
212
And now I'll show you how Datadap helps you to find a suitable hypothesis test and, of course, calculates it and interprets the results for you.
213
data .net and copy your own data in here.
214
We will just use this example dataset.
215
After copying your data into the table, the variables appear down here.
216
DataTab automatically tries to determine the correct level of measurement, but you can also change it up here.
217
Now we just click on "Hypothesis testing" and select the variables we want to use for the calculation of a hypothesis test.
218
Datadab will then suggest a suitable test.
219
For example, in this case, a chi -square test or, in that case, an analysis of variance.
220
Then you will see the hypotheses and the results.
221
If you're not sure how to interpret the results, click on "Summary in words".
222
Further, you can check the assumptions and decide whether you want to calculate a parametric or a non -parametric test.
223
You can find out the difference between parametric and non -parametric tests in my next video.
224
Thank you for watching and I hope you enjoyed the video!

About This Lesson

In this lesson, you will delve into the fascinating world of statistics, learning about its definition, the importance of descriptive and inferential statistics, and how to summarize data effectively. This session will enhance your understanding of how statistics can be applied in real-world scenarios, such as analyzing survey results or helping companies make informed decisions based on data collection. By the end of this practice, you will not only be familiar with statistical concepts but also improve your English speaking skills through engaging with the content.

Key Vocabulary & Phrases

  • Statistics: The study of the collection, analysis, and presentation of data.
  • Descriptive Statistics: A branch of statistics that summarizes and describes features of a dataset.
  • Inferential Statistics: A method to make predictions or generalizations about a population based on a sample.
  • Measures of Central Tendency: Statistical measures that describe the center of a dataset (mean, median, mode).
  • Data Collection: The process of gathering information for analysis, often through surveys or experiments.
  • Sample: A smaller, representative portion of a population used for statistical analysis.
  • Frequency Tables: A way to display how often each value occurs in a dataset.
  • Charts: Visual representations of data that facilitate understanding and interpretation.

Practice Tips

To maximize your learning and speaking skills, consider utilizing the shadowing technique as you watch the video. This method involves listening to the speaker and repeating their words simultaneously, which helps improve pronunciation, intonation, and fluency. Pay attention to the speed and tone of the speaker in the video, as it is vital for mimicking their style accurately.

For effective shadow speaking, try to:

  • Start by listening to a short section of the video, then pause and repeat what you've heard right after. This will help you internalize the pronunciation of difficult words.
  • Record your voice while shadowing, then compare it to the original speaker to identify areas for improvement.
  • Focus on the phrases related to statistics and data analysis, as practicing these will not only expand your vocabulary but also prepare you to discuss this topic more confidently.
  • Use a slower playback speed if you're struggling to keep up, gradually increasing the speed as you become more comfortable with the material.

By engaging actively with the content through shadowing, you will not only learn statistics but also enhance your English speaking skills effectively. Embrace the challenge and enjoy learning English with YouTube!

What is the Shadowing Technique?

Shadowing is a science-backed language learning technique originally developed for professional interpreter training and popularized by polyglot Dr. Alexander Arguelles. The method is simple but powerful: you listen to native English audio and immediately repeat it out loud — like a shadow following the speaker with just a 1–2 second delay. Unlike passive listening or grammar drills, shadowing forces your brain and mouth muscles to simultaneously process and reproduce real speech patterns. Research shows it significantly improves pronunciation accuracy, intonation, rhythm, connected speech, listening comprehension, and speaking fluency — making it one of the most effective methods for IELTS Speaking preparation and real-world English communication.

Shadowing technique: read the full step-by-step guide →