Free YouTube Transcribe

Video transcript

02 Descriptive Statistics and Frequencies in SPSS – SPSS for Beginners

Research By Design · 1,869 words · 9 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:06Welcome to the second video in SPSS for Beginners from RStats Institute

0:12at Missouri State University. In our first video, we learn how to create

0:17variables in SPSS. The next step is to add some data, and we're going to begin

0:22in Data View. Here in Data View these are the same four variables that we created

0:30in the first video. So now we can add some numbers. Pause the video and enter

0:40these same numbers into your SPSS spreadsheet.

0:54Now that we have numbers, it's important to understand just what these data

0:58represent. The first column is a random identification number. It stands in for

1:05the names of the participants and keeps our data anonymous. This second variable

1:12is gender. And these last two columns represent the height and the weight for

1:19each participant. Even after you've named a variable, it's possible to change the

1:24variable names. Double-click on a variable name to change it. When you do,

1:30you will be taken to Variable View, which is where you actually will make the

1:34changes. We set the measure for each variable previously. The ID variable is

1:41nominal because it stands for a number. It stands in for a participant's name. The

1:48variable "gender" is also nominal, and we are going to code gender as 1 and 2,

1:54for Male and Female. When a categorical variable has only two categories, we call

2:01it "dichotomous" So the 1 and the 2 are categories. You can be in one category or

2:06the other. You can't be in both; you can't be in neither. These last two columns

2:14represent height and weight. Height and weight are both quantitative variables,

2:19not categorical. They are measuring something. They both have fixed intervals

2:25between the scores, and they both have a meaningful zero. Both height and weight

2:32are set to Scale because they are both ratio level. Before we begin analyzing

2:38these numbers, there is one other thing that we should do. For a variable like

2:43gender - where we did not code the 1 and the 2 - we don't want to get confused with

2:48who was male who was female; what number stood for what. And so we are going to

2:54assign value labels for each level of this categorical variable.

3:00Click on Values. I'm going to tell SPSS to represent all of the 1's as Male and

3:09all of the 2's as Female. Of course, I could make 2 = Female or

3:150 = Female; really any number that I wanted to, depending on the coding.

3:20The 1 does not mean that males are "first place." The 2 does not mean that females

3:25are twice as good the number is only a placeholder; it does not indicate an order

3:30or a quantity. So, now click OK. In fact, when I return to Data View, you can

3:37see the numbers, but watch this: you see this button? Click it and you can toggle

3:44between numbers and value labels. Let's leave this set with the value labels on.

3:49It's just easier that way. Now we can look at our data. The height of our

3:55participant was measured in inches, and we have values between 60 and 70 inches,

4:00which is between five and six feet tall (1.5 to 1.8 meters). Weight was measured in

4:08pounds. Of course it might be easier to see the range if we sorted these data.

4:13Ctrl-click on Mac or right-click on PC and choose "Sort Ascending." All of these

4:22participants were between 116 and 153 pounds. Just for illustration, I'm going to

4:29pretend that there were two participants for whom we did not get their height or

4:33their weight; both of them female. Notice that when the numbers are toggled on, all

4:40that I need to do is type "2"; however, when the value labels are toggled on,

4:47I need to double-click and select "Female." So now I think we're ready to analyze

4:55these data. One of the simplest things that we can do is to count up how often

5:00things occur. For example, we want to know how many males and females were in

5:05our sample. We want their frequencies. This is easy enough to do in SPSS. We're

5:11going to use the Analyze menu. Whenever you

5:16run an analysis in SPSS, you use the Analyze menu. We can see that there are lots of

5:22options, each with their own sub-menus and sub-sub-menus. The one that we want

5:30is Analyze -> Descriptive Statistics -> Frequencies. This window pops up and you

5:40will see lots of windows of this type in SPSS. All of the variables that we have

5:45in our dataset are on the left, and the variables that we want to analyze go on

5:50the right. You can select a variable for analysis by clicking on its name and

5:55then clicking on this arrow between the boxes. Alternatively, you can also

6:01drag-and-drop, and in some cases you can double click. Let me show you just how

6:07easy it is to use SPSS: click OK. What we are seeing now is the output window, and

6:16here is something very important to know about SPSS, especially compared with

6:19other types of statistical software: SPSS will give you copious amounts of output,

6:25often more than you really need, and you need to know how to interpret that

6:31output. In SPSS, it is easy to run an analysis, but it takes some education to

6:38learn how to interpret the output. First, we see a summary of the variables in the

6:48box labeled "Statistics." We have 12 valid scores for gender with no missing data,

6:54but for height, we only have scores for 10 people, with 2 missing values.

7:01The valid sample size is the number of participants for whom we actually have

7:06scores. This first frequency table is for gender. The total tells us that we have

7:1312 valid scores. We see that there are 5 males and 7 females.

7:20Notice the columns for "Percent" and "Valid Percent." They are exactly the same. They are the

7:25same because we have no missing values for gender. This second frequency table

7:31is for height. Remember that we have missing values for height for two of our

7:37participants, so we see the valid total is 10. Two values are missing in the data

7:43set - called system missing - and the total is 12. We see that the Percent column is

7:50different than the Valid Percent column. The Percent column is calculated based

7:56on the total sample size of 12; the Valid Percent is calculated on the valid

8:02n of 10 people for whom we actually have data. I recommend reporting the valid

8:09percent column unless you have a specific reason why you need to report

8:13Percent. Well, this is a good start, but we can do better. Let's make some pictures

8:21of our data. I am going to run another analysis and I want you to see that you

8:26do not need to go back to the data set. You can run a new analysis from the

8:31output window, as well. Just click on Analyze -> Descriptive Statistics ->

8:39-> Frequencies. You can see that our previous analysis is still in the window.

8:45We could clear it by clicking on this Reset button, but let's just continue

8:50with these data. So this time, click on Charts, and then under Chart Type, click

8:58on Bar Charts. Let's change Chart Values to Percentages. Click Continue, but

9:05before you click OK, let's turn off the frequency tables

9:09because we already have those. Now, click OK.

9:14In the output window, we see that the chart for gender looks really good.

9:19We have two distinct bars, one for male one for female, and we can estimate the

9:25percentages of each. But when we look at the bar chart for height, the options

9:31just don't look as good. We can definitely do better. Let's run

9:36another analysis. Click on Analyze -> Descriptive Statistics -> Frequencies.

9:45This time, click on Charts, and then under Chart Type, click on Histogram. Let's also

9:53choose "Show normal curve on histogram." Notice that the chart values are now

9:58gray, because we don't need them. Click continue, but before clicking OK, let's do

10:04one more thing. Click on Statistics. Here we can choose other options like the

10:09mean, the standard deviatio,n the minimum, and maximum. We could also get the

10:14variance, the standard error of the mean, and the sum is good, too. As you see, we

10:21can pick as many of these options as we would like. If we change our mind, we can

10:25unselect them, too. Click continue and then OK. In the output

10:34window, we see all of the statistics that we asked for. For example, the average

10:39height was 65.8 inches. The tallest person? 70 inches tall. The shortest? 62

10:46inches tall. If we added up all of their heights, they would total 658 inches.

10:55But notice this first histogram for gender. It just doesn't look good, not like it

11:00did with the bar chart. The bars are connected, but gender is supposed to be

11:04discrete categories. We no longer see the labels for males and females. And the

11:10normal curve makes absolutely no sense. On the other hand, the histogram for

11:17height is much improved. The bars touch, indicating that the data are connected,

11:24and the superimposed normal curve makes sense with these data. We can see that

11:30the shape of the data match reasonably well with a normal distribution. The

11:37important thing to learn here is that you should choose the statistics and the

11:42graphs that are appropriate to your data. A nominal variable like gender should be

11:49reported with frequencies and a bar chart. Scale variables like height should

11:56be reported with a mean, standard deviation, and a histogram. We know that

12:03the average height for all participants is 65.8 inches, but

12:08what if we want to split that by males and females? Let me show you how. Click on

12:14Analyze, but instead of Descriptive Statistics, choose Compare Means and this

12:20first option, simply labeled, "Means." Here we have the options for dependent

12:27variables. "Layers" refers to the independent variable, or categorical

12:32variable. We did not really assign people to the condition called gender, so gender

12:38would really be what is called a "quasi independent variable." Still, we will use

12:44gender as our independent variable. We want to examine differences in height, so

12:51height will be the dependent variable. Now click OK. We can see the means and

12:59the standard deviations from males and females separately and together. There

13:06are 5 each for males and females, 10 total. We see that males were a few

13:12inches taller on the average than females.

13:15The total mean and standard deviation here are the same as the values that we

13:20got earlier using the Frequencies command.

13:25Overall frequency counts, charts, and descriptive statistics are a great way

13:30to take a peek at your data and see just what you have. It's a good idea to do

13:36this before running any other kind of analysis. In our next video, we will

13:41look a little bit more at these descriptive statistics and how to

13:46convert raw scores into z-scores. I'll see you then.

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.