| Course | MPH 550 Public Health Statistics (MPH/550) |
|---|---|
| Week | 5 |
| Paper type | Regression analysis paper |
| Length | about 1,205 words, 4 double-spaced pages plus title page and references |
| Format | APA 7 student paper |
| School | University of Phoenix |
| Program | MPH |
| Updated | September 2026 |
Free sample paper for MPH 550 Week 5
An R-Squared of 0.95 That Proves Less Than It Seems: Regressing County Diabetes on Physical Inactivity in Colorado
[Student Name]
University of Phoenix
MPH/550: Public Health Statistics
Week 5 Assignment
[Instructor Name]
[Date]
The health department, its analyst and the program proposal are composites written for a model paper; county estimates come from the CDC PLACES 2025 release, regression results were calculated from those estimates and other findings come from the sources cited.
A program manager at the county health department brought a chart to a planning meeting: Colorado counties with more physically inactive adults had more diabetes, and the points lay almost on a straight line. She proposed a county walking program and projected how much diabetes it would prevent using the line's slope. The analyst was asked to check the chart. This paper presents her correlation and regression analysis and what it can and cannot support.
The Data
Her source was the PLACES county file from CDC, age-adjusted and covering 2023, for all 64 Colorado counties: diagnosed diabetes as the outcome and physical inactivity, meaning adults who reported no exercise or other activity outside work during the previous 30 days, as the predictor (Centers for Disease Control and Prevention [CDC], 2025). The unit of analysis is the county.
The Scatterplot
The scatterplot showed a tight, upward, linear pattern. Boulder County sat at the lower left, with 12.2% inactivity, and Bent County at the upper right, with 28.9% inactivity and 13.1% diabetes. No curve or separate clusters appeared.
The Projection on the Table
The program manager's projection used the line directly. Pueblo County's inactivity estimate was 25.7%. If a walking program cut that to 20.7%, the slope of 0.365 implied diabetes would fall by about 1.8 points, which she translated into roughly 2,400 fewer adults with diabetes. The arithmetic was correct. The question was whether the slope could bear that interpretation.
Correlation
The Pearson correlation was 0.977, an extraordinarily strong positive linear association. For county-level health data, correlations this high are unusual and deserve suspicion before celebration.
Simple Linear Regression
The fitted line was: predicted diabetes equals 1.65 plus 0.365 times inactivity. The slope means that counties with one percentage point more inactive adults had, on average, 0.365 points more diabetes; its 95% confidence interval ran from 0.345 to 0.385. The intercept, 1.65, is the predicted diabetes at zero inactivity, far outside the data and not meaningful by itself.
R-Squared and Fit
The R-squared was 0.955: inactivity accounted for about 95% of the variation in diabetes across counties. The typical prediction error, the residual standard deviation, was 0.34 percentage points. The p-value for the slope was far below 0.001. By every standard summary, the model fits almost perfectly, which is exactly why the analyst kept checking. A fit this close in observational county data usually means the variables share a common source, not that one fully explains the other.
Residuals
Residuals showed no curve or widening pattern, so linearity and constant variance were reasonable. The largest positive residuals, counties with more diabetes than predicted, were Bent, Jefferson and Costilla, each about 0.9 points above the line. Pueblo County had the largest negative residual: predicted 11.0%, estimated 10.4%. The county the program targeted had less diabetes than its inactivity alone would predict.
Influential Points
The analyst checked whether a few counties drove the result. Removing Bent and Boulder, the two extremes, barely changed the slope, from 0.365 to 0.360, and no single county pulled the line noticeably on its own. The relationship is not an artifact of one or two unusual places.
Other Correlates
Diabetes correlated strongly with nearly everything related to disadvantage: 0.96 with lack of insurance, 0.90 with high blood pressure, 0.87 with obesity and 0.85 with smoking. It correlated weakly with depression, at minus 0.03. When one measure correlates above 0.85 with many others, a single-cause story becomes hard to defend.
Multiple Regression
The analyst added current smoking and lack of insurance as predictors. The inactivity coefficient fell from 0.365 to 0.213, with a standard error of 0.043; the insurance coefficient was 0.142, with a standard error of 0.032; smoking added little, 0.034 with a standard error of 0.033. R-squared rose only to 0.967.
Interpreting the Change
The inactivity slope fell because inactivity and lack of insurance travel together; their correlation was 0.95. Part of what looked like an inactivity effect reflected the broader disadvantage that both measure. With predictors this correlated, collinearity makes each coefficient unstable, and small changes in data or model could shift them considerably.
Why the Fit Is So Tight
Part of the answer lies in how the data were made. PLACES estimates come from multilevel models that combine survey responses with each county's demographic makeup, using the same census characteristics for every measure. CDC researchers showed that such model-based estimates track direct survey estimates closely at the county level, with correlations of 0.88 to 0.95 for the condition they studied (Zhang et al., 2014). Because each county's estimates share inputs, measures built from the same demographics will correlate more tightly than direct measurements would.
The Ecological Fallacy
Even if the correlation were fully real, it describes counties. In a classic demonstration using 1930 census data, the correlation between two characteristics across nine geographic divisions was 0.946, while the correlation for the same characteristics among individuals was only 0.203 (Robinson, 1950). A county-level slope cannot say how much a particular person's risk changes when they become active.
What Individual Studies Show
Individual-level research answers the program manager's real question more directly. A meta-analysis of 10 prospective cohort studies with 301,221 participants found a relative risk of type 2 diabetes of 0.69 for regular moderate physical activity compared with being sedentary, and 0.70 for regular brisk walking, with associations persisting after adjustment for body mass index (Jeon et al., 2007).
Advice to the Program Manager
The analyst supported the walking program but not the projection. The county regression cannot estimate cases prevented, because the slope reflects modeled data, shared disadvantage and county-level grouping. The cohort evidence supports physical activity as protective for individuals, and a program evaluation could measure participants' activity and later outcomes directly.
What a Better Analysis Would Use
The question of how much diabetes a walking program prevents needs individual data. The department's household survey, which records activity, diabetes, age, income and insurance for each respondent, supports a logistic regression at the person level. Better still, a program evaluation could compare participants with similar non-participants over several years. Either approach avoids the ecological fallacy and the shared modeling that inflate the county results.
Communicating the Finding
The analyst presented the chart with a caption the program manager agreed to: counties with more inactivity have more diabetes, but county patterns cannot estimate how many cases a program will prevent. She kept the scatterplot, because the pattern is real and persuasive, while removing the projection from the funding request.
Assumptions and Limits
The regression assumed independent observations, but neighboring counties share conditions, which could make standard errors too small. The county data are one year. Small counties' estimates rely heavily on the model. Age adjustment allows fair comparison but changes the values from the actual burden.
Conclusion
A correlation of 0.977 and an R-squared of 0.955 looked like strong evidence that inactivity drives diabetes across Colorado counties. Checking residuals, adding confounders and understanding how the estimates were built showed a more modest picture: a real association entangled with broader disadvantage and amplified by modeling. The ecological fallacy and individual cohort evidence placed the finding in context, supporting the program while retiring the projection.
References
Centers for Disease Control and Prevention. (2025). PLACES: Local data for better health, county data, 2025 release [Data set]. https://data.cdc.gov/d/swc5-untb
Jeon, C. Y., Lokken, R. P., Hu, F. B., & van Dam, R. M. (2007). Physical activity of moderate intensity and risk of type 2 diabetes: A systematic review. Diabetes Care, 30(3), 744-752. https://doi.org/10.2337/dc06-1842
Robinson, W. S. (1950). Ecological correlations and the behavior of individuals. American Sociological Review, 15(3), 351-357. https://doi.org/10.2307/2087176
Zhang, X., Holt, J. B., Lu, H., Wheaton, A. G., Ford, E. S., Greenlund, K. J., & Croft, J. B. (2014). Multilevel regression and poststratification for small-area estimation of population health outcomes: A case study of chronic obstructive pulmonary disease prevalence using the Behavioral Risk Factor Surveillance System. American Journal of Epidemiology, 179(8), 1025-1033. https://doi.org/10.1093/aje/kwu018
What the MPH 550 Week 5 instructions ask
The fifth MPH 550 assignment usually introduces correlation and regression. Prompts may ask students to describe the relationship between two variables with a scatterplot and correlation coefficient, fit a simple linear regression, interpret the slope, intercept and R-squared, add variables in a multiple regression, check assumptions such as linearity and residual patterns and explain why association does not establish causation. Some versions supply software output to interpret. Explain every coefficient in the units of the data. Strong papers interpret coefficients in plain words, examine residuals and influential points, discuss confounding and collinearity when variables are added and state the unit of analysis and what can be generalized.
How this MPH 550 Week 5 example is built
A program manager proposing a county walking program cites a striking chart linking inactivity and diabetes, which opens the paper. The analyst plots the 64 counties and calculates the correlation. A simple regression is fitted, and its slope, intercept, confidence interval and R-squared are interpreted. Residuals are examined, and Pueblo County emerges as the county with the lowest diabetes relative to its inactivity. Correlations with other measures are compared. A multiple regression adding smoking and insurance shows how the inactivity coefficient shrinks and how correlated predictors complicate interpretation. The ecological fallacy, modeled data and individual-level cohort evidence close the analysis, and the walking program's projection is replaced with a plan for direct evaluation.
MPH 550 Week 5 grading rubric: where the points go
The regression week is typically assessed on correct interpretation of coefficients, checked assumptions and careful reasoning about causation. Graders look for a scatterplot and correlation described, slope and intercept interpreted in context, R-squared explained, confidence intervals reported, residuals and outliers examined, multiple regression used to address confounding and limits such as the ecological fallacy, collinearity and data quality discussed. Individual-level research strengthens the interpretation. Explaining why a very high R-squared can mislead earns credit. Stating the unit of analysis clearly also earns marks. Readable tables, labeled output and APA citations close out the grade. Drafts that read a regression slope as proof of cause usually earn less.
MPH 550 Week 5 help: mistakes to avoid
Many MPH 550 Week 5 drafts report a strong correlation and conclude that one variable causes the other. Start with a scatterplot and look for curves, clusters and outliers. Interpret the slope in the data's units: how much the outcome changes for each one-unit change in the predictor. Explain R-squared as the share of variation explained, not as proof. Examine residuals to see which cases the model fits poorly and why. Add plausible confounders and watch how coefficients move. Ask what your unit of analysis is; county patterns may not hold for people. Finally, look for individual-level studies that test the relationship more directly, and compare their findings with yours. Where your model and those studies disagree, explain why.
Related MPH 550 sample papers
Other MPH 550 week samples
- MPH 550 Week 1: Descriptive Statistics
- MPH 550 Week 2: Probability and Sampling
- MPH 550 Week 3: Confidence Intervals and Tests
- MPH 550 Week 4: Comparing Groups
- MPH 550 Week 6: Survey Design
More MPH sample papers
- MPH 510 Week 5: Community Stewardship
- MPH 520 Week 5: Multilevel Intervention Design
- MPH 530 Week 5: Outbreak Investigation
- MPH 540 Week 5: Hierarchy of Controls
MPH 550 Week 5 questions, answered
What does MPH/550 Week 5 usually ask for?
The fifth statistics paper usually introduces correlation and regression, asking students to fit and interpret models, check assumptions and explain why association does not establish causation.
Where can I find a free MPH 550 Week 5 sample paper?
The county regression paper above is open to all readers free, with a note beside every coefficient. Send your data or output, and our first draft costs nothing.
What does R-squared mean?
The proportion of variation in the outcome that the model accounts for; a high value shows a close fit but does not show that the predictor causes the outcome.
What is the ecological fallacy?
The error of assuming that relationships observed between groups, such as counties, hold for individuals; a classic study found a group-level correlation of 0.946 where the individual correlation was 0.203.
Does physical activity lower diabetes risk for individuals?
A meta-analysis of 10 prospective cohorts with 301,221 participants found a relative risk of 0.69 for type 2 diabetes with regular moderate activity compared with being sedentary.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official University of Phoenix document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.
Request this one custom, free · All MPH 550 week samples · All courses