Full transcript
0:06Welcome to the second video in SPSS for Beginners from RStats Institute
0:12at Missouri State University. In our first video, we learn how to create
0:17variables in SPSS. The next step is to add some data, and we're going to begin
0:22in Data View. Here in Data View these are the same four variables that we created
0:30in the first video. So now we can add some numbers. Pause the video and enter
0:40these same numbers into your SPSS spreadsheet.
0:54Now that we have numbers, it's important to understand just what these data
0:58represent. The first column is a random identification number. It stands in for
1:05the names of the participants and keeps our data anonymous. This second variable
1:12is gender. And these last two columns represent the height and the weight for
1:19each participant. Even after you've named a variable, it's possible to change the
1:24variable names. Double-click on a variable name to change it. When you do,
1:30you will be taken to Variable View, which is where you actually will make the
1:34changes. We set the measure for each variable previously. The ID variable is
1:41nominal because it stands for a number. It stands in for a participant's name. The
1:48variable "gender" is also nominal, and we are going to code gender as 1 and 2,
1:54for Male and Female. When a categorical variable has only two categories, we call
2:01it "dichotomous" So the 1 and the 2 are categories. You can be in one category or
2:06the other. You can't be in both; you can't be in neither. These last two columns
2:14represent height and weight. Height and weight are both quantitative variables,
2:19not categorical. They are measuring something. They both have fixed intervals
2:25between the scores, and they both have a meaningful zero. Both height and weight
2:32are set to Scale because they are both ratio level. Before we begin analyzing
2:38these numbers, there is one other thing that we should do. For a variable like
2:43gender - where we did not code the 1 and the 2 - we don't want to get confused with
2:48who was male who was female; what number stood for what. And so we are going to
2:54assign value labels for each level of this categorical variable.
3:00Click on Values. I'm going to tell SPSS to represent all of the 1's as Male and
3:09all of the 2's as Female. Of course, I could make 2 = Female or
3:150 = Female; really any number that I wanted to, depending on the coding.
3:20The 1 does not mean that males are "first place." The 2 does not mean that females
3:25are twice as good the number is only a placeholder; it does not indicate an order
3:30or a quantity. So, now click OK. In fact, when I return to Data View, you can
3:37see the numbers, but watch this: you see this button? Click it and you can toggle
3:44between numbers and value labels. Let's leave this set with the value labels on.
3:49It's just easier that way. Now we can look at our data. The height of our
3:55participant was measured in inches, and we have values between 60 and 70 inches,
4:00which is between five and six feet tall (1.5 to 1.8 meters). Weight was measured in
4:08pounds. Of course it might be easier to see the range if we sorted these data.
4:13Ctrl-click on Mac or right-click on PC and choose "Sort Ascending." All of these
4:22participants were between 116 and 153 pounds. Just for illustration, I'm going to
4:29pretend that there were two participants for whom we did not get their height or
4:33their weight; both of them female. Notice that when the numbers are toggled on, all
4:40that I need to do is type "2"; however, when the value labels are toggled on,
4:47I need to double-click and select "Female." So now I think we're ready to analyze
4:55these data. One of the simplest things that we can do is to count up how often
5:00things occur. For example, we want to know how many males and females were in
5:05our sample. We want their frequencies. This is easy enough to do in SPSS. We're
5:11going to use the Analyze menu. Whenever you
5:16run an analysis in SPSS, you use the Analyze menu. We can see that there are lots of
5:22options, each with their own sub-menus and sub-sub-menus. The one that we want
5:30is Analyze -> Descriptive Statistics -> Frequencies. This window pops up and you
5:40will see lots of windows of this type in SPSS. All of the variables that we have
5:45in our dataset are on the left, and the variables that we want to analyze go on
5:50the right. You can select a variable for analysis by clicking on its name and
5:55then clicking on this arrow between the boxes. Alternatively, you can also
6:01drag-and-drop, and in some cases you can double click. Let me show you just how
6:07easy it is to use SPSS: click OK. What we are seeing now is the output window, and
6:16here is something very important to know about SPSS, especially compared with
6:19other types of statistical software: SPSS will give you copious amounts of output,
6:25often more than you really need, and you need to know how to interpret that
6:31output. In SPSS, it is easy to run an analysis, but it takes some education to
6:38learn how to interpret the output. First, we see a summary of the variables in the
6:48box labeled "Statistics." We have 12 valid scores for gender with no missing data,
6:54but for height, we only have scores for 10 people, with 2 missing values.
7:01The valid sample size is the number of participants for whom we actually have
7:06scores. This first frequency table is for gender. The total tells us that we have
7:1312 valid scores. We see that there are 5 males and 7 females.
7:20Notice the columns for "Percent" and "Valid Percent." They are exactly the same. They are the
7:25same because we have no missing values for gender. This second frequency table
7:31is for height. Remember that we have missing values for height for two of our
7:37participants, so we see the valid total is 10. Two values are missing in the data
7:43set - called system missing - and the total is 12. We see that the Percent column is
7:50different than the Valid Percent column. The Percent column is calculated based
7:56on the total sample size of 12; the Valid Percent is calculated on the valid
8:02n of 10 people for whom we actually have data. I recommend reporting the valid
8:09percent column unless you have a specific reason why you need to report
8:13Percent. Well, this is a good start, but we can do better. Let's make some pictures
8:21of our data. I am going to run another analysis and I want you to see that you
8:26do not need to go back to the data set. You can run a new analysis from the
8:31output window, as well. Just click on Analyze -> Descriptive Statistics ->
8:39-> Frequencies. You can see that our previous analysis is still in the window.
8:45We could clear it by clicking on this Reset button, but let's just continue
8:50with these data. So this time, click on Charts, and then under Chart Type, click
8:58on Bar Charts. Let's change Chart Values to Percentages. Click Continue, but
9:05before you click OK, let's turn off the frequency tables
9:09because we already have those. Now, click OK.
9:14In the output window, we see that the chart for gender looks really good.
9:19We have two distinct bars, one for male one for female, and we can estimate the
9:25percentages of each. But when we look at the bar chart for height, the options
9:31just don't look as good. We can definitely do better. Let's run
9:36another analysis. Click on Analyze -> Descriptive Statistics -> Frequencies.
9:45This time, click on Charts, and then under Chart Type, click on Histogram. Let's also
9:53choose "Show normal curve on histogram." Notice that the chart values are now
9:58gray, because we don't need them. Click continue, but before clicking OK, let's do
10:04one more thing. Click on Statistics. Here we can choose other options like the
10:09mean, the standard deviatio,n the minimum, and maximum. We could also get the
10:14variance, the standard error of the mean, and the sum is good, too. As you see, we
10:21can pick as many of these options as we would like. If we change our mind, we can
10:25unselect them, too. Click continue and then OK. In the output
10:34window, we see all of the statistics that we asked for. For example, the average
10:39height was 65.8 inches. The tallest person? 70 inches tall. The shortest? 62
10:46inches tall. If we added up all of their heights, they would total 658 inches.
10:55But notice this first histogram for gender. It just doesn't look good, not like it
11:00did with the bar chart. The bars are connected, but gender is supposed to be
11:04discrete categories. We no longer see the labels for males and females. And the
11:10normal curve makes absolutely no sense. On the other hand, the histogram for
11:17height is much improved. The bars touch, indicating that the data are connected,
11:24and the superimposed normal curve makes sense with these data. We can see that
11:30the shape of the data match reasonably well with a normal distribution. The
11:37important thing to learn here is that you should choose the statistics and the
11:42graphs that are appropriate to your data. A nominal variable like gender should be
11:49reported with frequencies and a bar chart. Scale variables like height should
11:56be reported with a mean, standard deviation, and a histogram. We know that
12:03the average height for all participants is 65.8 inches, but
12:08what if we want to split that by males and females? Let me show you how. Click on
12:14Analyze, but instead of Descriptive Statistics, choose Compare Means and this
12:20first option, simply labeled, "Means." Here we have the options for dependent
12:27variables. "Layers" refers to the independent variable, or categorical
12:32variable. We did not really assign people to the condition called gender, so gender
12:38would really be what is called a "quasi independent variable." Still, we will use
12:44gender as our independent variable. We want to examine differences in height, so
12:51height will be the dependent variable. Now click OK. We can see the means and
12:59the standard deviations from males and females separately and together. There
13:06are 5 each for males and females, 10 total. We see that males were a few
13:12inches taller on the average than females.
13:15The total mean and standard deviation here are the same as the values that we
13:20got earlier using the Frequencies command.
13:25Overall frequency counts, charts, and descriptive statistics are a great way
13:30to take a peek at your data and see just what you have. It's a good idea to do
13:36this before running any other kind of analysis. In our next video, we will
13:41look a little bit more at these descriptive statistics and how to
13:46convert raw scores into z-scores. I'll see you then.