Chi Square Test Multiple Choice Questions
1. Which of the following is the explanatory variable in this study?
A. Exercise
B. Lung capacity
C. Smoking or not
D. Occupation
Answer: Occupation
2. Which of the following is a confounding variable in this study?
A. Exercise
B. Lung capacity
C. Smoking or not
D. Occupation
Answer: Exercise
3. A magazine printed a survey in its monthly issue and asked readers to fill it out and send it in. Over 1000 readers did so. This type of sample is called
A. a cluster sample
B. a self-selected sample
C. a stratified sample
D. a simple random sample
Answer: a self-selected sample
4. Which one of the following variables is not categorical?
A. Age of a person
B. Gender of a person: male or female
C. Choice on a test item: true or false
D. None of the above
Answer: Age of a person
5. Which of the following would indicate that a dataset is not bell-shaped?
A. The range is equal to 5 standard deviations
B. The range is larger than the interquartile range
C. The mean is much smaller than the median
D. None of the above
Answer: The mean is much smaller than the median
6. The value of a correlation is reported by a researcher to be r = −0.5. Which of the following statements is correct?
A. The x-variable explains −50% of the variability in the y-variable
B. The x-variable explains 50% of the variability in the y-variable
C. The x-variable explains −25% of the variability in the y-variable
D. The x-variable explains 25% of the variability in the y-variable
Answer: The x-variable explains 25% of the variability in the y-variable
7. Among people with age over 30, what’s the “risk” of always exceeding the speed limit?
A. 0.20
B. 0.40
C. 0.33
D. 0.50
Answer: 0.20
8. What is the approximate shape of the distribution?
A. Nearly symmetric
B. Skewed to the left
C. Skewed to the right
D. Bimodal (has more than one peak)
Answer: Skewed to the right
9. Which of the following is not a way that we could use central tendency?
A. make a guess for the value of an individual observation in our sample
B. determine how accurately we can guess the value of an individual observation in our sample
C. make a guess for the value of an observation that we will encounter in the future
D. describe a representative value for observations in our sample
Answer: determine how accurately we can guess the value of an individual observation in our sample
10. If we had a population of values and took many possible samples of some size (n) from that population, calculated the mean of each of these samples, and then took the standard deviation of these sample means, we would have:
A. calculated the population standard deviation
B. calculated the standard error of the mean
C. calculated the expected value of the mean
D. calculated the central limit theorem
Answer: calculated the standard error of the mean
11. A human resources manager is interested to know if the mean number of sick days taken per year differs between three specific job titles within a company. What test should the manager use to examine this question?
A. two-sample t-test for a difference in means
B. one-factor analysis of variance (ANOVA)
C. two-factor analysis of variance (ANOVA)
D. chi-square (Χ2) test for goodness-of-fit
Answer: one-factor analysis of variance (ANOVA)
12. A baker is interested in whether there is a difference in the mean ‘deliciousness’ ratings that customers give to a low-sugar option of a cookie and to a high-sugar option of a cookie, and whether the difference in mean rating for the low- and high-sugar options depends on the age group of the customer (child vs. adult). What test should the baker use to examine this question?
A. two-sample t-test for a difference in means
B. one-factor analysis of variance (ANOVA)
C. two-factor analysis of variance (ANOVA)
D. chi-square (Χ2) test for independence
Answer: two-factor analysis of variance (ANOVA)
13. A polling company finds that 60% of Democrats give a ‘favorable’ rating for a senator and 52% of Republicans give a ‘favorable’ rating for the same senator. What should the company use if they want to know if this is a ‘statistically significant’ finding?
A. binomial test for a single proportion
B. one-sample t-test for a single mean
C. z-test for a difference in two proportions
D. two-sample t-test for a difference in means
Answer: z-test for a difference in two proportions
14. A confidence interval describes what set of values?
A. a plausible set of values for an unknown population parameter
B. the set of values that we could have used for our null hypothesis that would have caused us to reject the null hypothesis
C. the set of values that we could have used for our null hypothesis that would have caused us to fail to reject the null hypothesis
D. a and c
Answer: a and c
15. For which scenario below would simulation be useful?
A. we aren’t sure if our sample data meets the assumptions for a theory-based test
B. we don’t know of a theory-based test for our sample statistic of interest
C. we are trying to do a quick, mental calculation to get a reasonable approximation
D. a and b
Answer: a and b
16. A pharmaceutical company is investigating a new drug (Drug A) as a potential replacement for an old drug (Drug B). They make several statements. Which of these statements most relates to the effect size of Drug A vs. Drug B.
A. patients using Drug A have a survival rate that is significantly higher than patients using Drug B
B. the patients in the study were aged 20 – 40 with early-stage disease progression and were randomly assigned by the experimenter to take one of two drugs
C. there were n = 100 patients in each group
D. on average, we would need to give 10 people Drug A to have one more survival than if we had given those 10 people Drug B
Answer: on average, we would need to give 10 people Drug A to have one more survival than if we had given those 10 people Drug B
17. In the same scenario as above, which of the statements most relates to whether we can determine that the relationship between drug and survival rate is a cause-and-effect relationship?
A. patients using Drug A have a survival rate that is significantly higher than patients using Drug B
B. the patients in the study were aged 20 – 40 with early-stage disease progression and were randomly assigned by the experimenter to take one of two drugs
C. there were n = 100 patients in each group
D. none of the above
Answer: the patients in the study were aged 20 – 40 with early-stage disease progression and were randomly assigned by the experimenter to take one of two drugs
18. Which is a question that we can answer directly with hypothesis testing?
A. what is the probability of observing our sample data by random chance, if some assumptions are true about the population?
B. does one of our variables of interest directly cause a change in another variable?
C. to what situations or populations can we generalize our sample data?
D. None of the above
Answer: what is the probability of observing our sample data by random chance, if some assumptions are true about the population?
19. In hypothesis testing, we use a to make an inference about a .
A. sample statistic; population parameter
B. sample statistic; sample statistic
C. population parameter; sample statistic
D. population parameter; population parameter
Answer: sample statistic; population parameter
20. Two dog owners are interested in whether their (shared) dog tends to run to a specific one of them first when they come home together. They walk through the door at the same time and record which person the dog runs to first. They repeat this procedure many times. What test should they use to ask if the dog runs to one of them more often than we would expect if the dog was choosing randomly between the owners?
A. binomial test for a single proportion
B. one-sample t-test for a single mean
C. z-test for a difference in two proportions
D. two-sample t-test for a difference in means
Answer: binomial test for a single proportion
21 A television enthusiast would like to know if the set of network television shows that were nominated for an Emmy were equally likely to have come from ABC, CBS, NBC, and FOX. What test asks whether the number of nominations from each network are consistent with a population distribution in which ¼ of the nominations were from each network.
A. two-sample t-test for a difference in means
B. one-factor analysis of variance (ANOVA)
C. two-factor analysis of variance (ANOVA)
D. chi-square (Χ2) test for goodness-of-fit
Answer: chi-square (Χ2) test for goodness-of-fit
22. Which does not happen as sample size (n) increases and as population or sample variance (σ2 or s2) decreases (assuming all else stays the same)?
A. the sample statistics that we would expect to observe by random chance become more clustered together
B. sample statistics become better estimates of population parameters
C. test statistics corresponding to our sample data (e.g., t, F, Χ2) increase
D. confidence intervals become wider
Answer: confidence intervals become wider
23. Which comparison of variability is accurate?
A. means of sets of observations are less variable than the individual observations
B. sums (totals) of sets of observations are less variable than the individual observations
C. differences between two means are less variable than the individual means
D. all of these comparisons are accurate
Answer: means of sets of observations are less variable than the individual observations
24. Simpson’s Paradox occurs when
A. No baseline risk is given, so it is not know whether or not a high relative risk has practical importance
B. A confounding variable rather than the explanatory variable is responsible for a change in the response variable
C. The direction of the relationship between two variables changes when the categories of a confounding variable are taken into account
D. None of the above
Answer: The direction of the relationship between two variables changes when the categories of a confounding variable are taken into account
25. A chi-square test involves a set of counts called “expected counts.” What are the expected counts?
A. Hypothetical counts that would occur of the alternative hypothesis were true
B. Hypothetical counts that would occur if the null hypothesis were true
C. The actual counts that did occur in the observed data
D. None of the above
Answer: Hypothetical counts that would occur if the null hypothesis were true
26. A chi-square test of the relationship between personal perception of emotional health and marital status led to rejection of the null hypothesis, indicating that there is a relationship between these two variables. One conclusion that can be drawn is:
A. Marriage leads to better emotional health
B. Better emotional health leads to marriage
C. The more emotionally healthy someone is, the more likely they are to be married
D. There are likely to be confounding variables related to both emotional health and marital status
Answer: There are likely to be confounding variables related to both emotional health and marital status
27. Most of the women in this sample felt that their actual weight was
A. about the same as their ideal weight
B. less than their ideal weight
C. greater than their ideal weight
D. no more than 2 pounds different from their ideal weight
Answer: greater than their ideal weight
28. The median of the distribution is approximately
A. −10 pounds
B. 10 pounds
C. 30 pounds
D. 50 pounds
Answer: 10 pounds
29. The newspaper also reported that “The number of children in the study who contracted asthma was relatively small, 265 of 3,535.” Which of the following is represented by 265/3535 = .075?
A. The overall risk of getting asthma for the children in this study
B. The baseline risk of getting asthma for the “non-athletic peers” in the study
C. The risk of getting asthma for children in the study who participated in sports
D. The relative risk of getting asthma for children who routinely participate in vigorous after-school sports on smoggy days and their non-athletic peers
Answer: The overall risk of getting asthma for the children in this study
30. Among people with age under 30 what are the odds that they always exceed the speed limit?
A. 1 to 2
B. 2 to 1
C. 1 to 1
D. 50%
Answer: 1 to 1
31. Of the following, which is the most important additional information that would be useful before making a decision about participation in school sports?
A. Where was the study conducted?
B. How many students in the study participated in after-school sports?
C. What is the baseline risk for getting asthma?
D. Who funded the study?
Answer: What is the baseline risk for getting asthma?
32. What is the relative risk of always exceeding the speed limit for people under 30 compared to people over 30?
A. 2.5
B. 0.4
C. 0.5
D. 30%
Answer: 2.5
33. Past data has shown that the regression line relating the final exam score and the midterm exam score for students who take statistics from a certain professor is: final exam = 50 + 0.5 × midterm One interpretation of the slope is
A. a student who scored 0 on the midterm would be predicted to score 50 on the final exam
B. a student who scored 0 on the final exam would be predicted to score 50 on the midterm exam
C. a student who scored 10 points higher than another student on the midterm would be predicted to score 5 points higher than the other student on the final exam
D. students only receive half as much credit (.5) for a correct answer on the final exam compared to a correct answer on the midterm exam
Answer: a student who scored 10 points higher than another student on the midterm would be predicted to score 5 points higher than the other student on the final exam
34. This is a randomized experiment rather than an observational study because:
A. Blood pressure was measured at the beginning and end of the study
B. The two groups were compared at the end of the study
C. The participants were randomly assigned to either walk or read, rather than choosing their own activity
D. A random sample of participants was used
Answer: The participants were randomly assigned to either walk or read, rather than choosing their own activity
35. Which of the following would be most likely to produce selection bias in a survey?
A. Using questions with biased wording
B. Only receiving responses from half of the people in the sample
C. Conducting interviews by telephone instead of in person
D. Using a random sample of students at a university to estimate the proportion of people who think the legal drinking age should be lowered
Answer: Using a random sample of students at a university to estimate the proportion of people who think the legal drinking age should be lowered
36. A polling agency conducted a survey of 100 doctors on the question “Are you willing to treat women patients with the recently approved pill RU-486”? The conservative margin of error associated with the 95% confidence interval for the percent who say ‘yes’ is
A. 50%
B. 10%
C. 5%
D. 2%
Answer: 10%
37. A list of 5 pulse rates is: 70, 64, 80, 74, 92. What is the median for this list?
A. 74
B. 76
C. 77
D. none of the above
Answer: 74
38. What is the effect of an outlier on the value of a correlation coefficient?
A. An outlier will always decrease a correlation coefficient
B. An outlier will always increase a correlation coefficient
C. An outlier might either decrease or increase a correlation coefficient, depending on where it is in relation to the other points
D. None of the above
Answer: An outlier might either decrease or increase a correlation coefficient, depending on where it is in relation to the other points
39. One use of a regression line is
A. to determine if any x-values are outliers
B. to determine if any y-values are outliers
C. to determine if a change in x causes a change in y
D. to estimate the change in y for a one-unit change in x
Answer: to estimate the change in y for a one-unit change in x
40. Pick the choice that best completes the following sentence. If a relationship between two variables is called statistically significant, it means the investigators think the variables are
A. related in the population represented by the sample
B. not related in the population represented by the sample
C. related in the sample due to chance alone
D. very important
Answer: related in the population represented by the sample
