MPH 550 Week 4 Comparing Groups Example

Reviewed by Lenora Whitcombe, MSN, RN · University of Phoenix · Updated

This MPH 550 Week 4 example compares two groups of Colorado counties, the 16 with at least 50,000 residents and the 48 smaller ones, on five health measures from CDC PLACES. University of Phoenix MPH 550 covers analyzing health data and designing studies, and in week four MPH/550 students typically choose and run tests that compare groups, such as t-tests, rank tests, chi-square tests and ANOVA, and interpret them with effect sizes. The APA 7 paper finds adult smoking 2.05 points lower in larger counties, a difference that holds under a rank test and a multiple testing correction. Diabetes is 0.80 points lower with a p-value of 0.049, but that result does not survive a rank test. Research on effect size, unequal variances and multiple testing guides each choice.

CourseMPH 550 Public Health Statistics (MPH/550)
Week4
Paper typeGroup comparison paper
Lengthabout 1,178 words, 4 double-spaced pages plus title page and references
FormatAPA 7 student paper
SchoolUniversity of Phoenix
ProgramMPH
UpdatedSeptember 2026

Free sample paper for MPH 550 Week 4

1

Are Colorado's Smaller Counties Less Healthy? Comparing Two Groups of Counties With t-Tests, a Rank Test and Effect Sizes

[Student Name]

University of Phoenix

MPH/550: Public Health Statistics

Week 4 Assignment

[Instructor Name]

[Date]

The health department, its analyst and the state association request are composites written for a model paper; county estimates come from the CDC PLACES 2025 release, and other findings come from the sources cited.

What this part is doingThe title poses the association's question, so every test in the paper can be read as part of the answer.
2

The state association of local health officials asked the county's analyst to help answer a question for its legislative agenda: do Colorado's smaller, mostly rural counties carry a heavier chronic disease burden than larger ones? If so, the association would argue for funding weighted toward small counties. This paper compares the two groups and reports what the data support.

The Data and Groups

The analyst used CDC PLACES age-adjusted estimates for 2023 for all 64 Colorado counties (Centers for Disease Control and Prevention [CDC], 2025). She divided counties at 50,000 residents: 16 larger counties, which together hold about 90% of the state's population, and 48 smaller ones. The cutoff is simple and commonly used, though any cutoff is somewhat arbitrary.

Outcomes Chosen in Advance

Before running tests, she chose five outcomes: current smoking, diagnosed diabetes, physical inactivity, obesity and lack of health insurance among adults aged 18 to 64. Choosing in advance limits the temptation to search many outcomes and report only the ones that differ.

What this part is doingNaming the outcomes before testing is the first defense against false discoveries.
3

Looking Before Testing

The analyst first plotted each outcome for the two groups side by side. The plots showed that larger counties spread more widely on obesity and inactivity, running from Boulder's very low values to higher values in Pueblo and Weld, while smaller counties clustered more tightly. Diabetes showed a right tail among small plains counties. These pictures shaped the choice of test and warned that averages might hide variation within groups.

Choosing the Test

Each outcome is a continuous percentage, and there are two independent groups, pointing to a two-sample t-test. The classic Student's t-test assumes equal variances, but the groups differed in spread; for obesity, the standard deviation was 4.82 among larger counties and 3.34 among smaller ones, and the groups differ threefold in size.

Why Unequal Variances

A review of the question recommended the unequal variance, or Welch, t-test as a default whenever two groups are compared on a continuous outcome: it performs about as well as Student's test when variances are equal and far better when they are not, especially with unequal group sizes (Ruxton, 2006). The analyst used it for all five outcomes.

Results: Smoking

Adult smoking averaged 11.24% in larger counties and 13.30% in smaller ones, a difference of 2.05 points, with a 95% interval of 0.72 to 3.39 points; t was 3.13 with about 32 degrees of freedom and p was 0.004. Adults in smaller counties smoke at noticeably higher rates, a gap that survived every check.

Results: Diabetes

Diabetes averaged 7.98% in larger counties and 8.78% in smaller ones, a difference of 0.80 points, with an interval of 0.01 to 1.59 and p of 0.049.

Results: The Other Three

Physical inactivity was 2.07 points lower in larger counties, with an interval from 4.46 lower to 0.32 higher and p of 0.087. Obesity was 1.50 points lower, with an interval from 4.21 lower to 1.21 higher and p of 0.26. Lack of insurance was 1.39 points lower, with an interval from 3.46 lower to 0.68 higher and p of 0.18. All three point in the same direction, but each interval includes zero.

Checking Normality

The t-test assumes that group means are approximately normally distributed. With 48 smaller counties, the central limit theorem makes that reasonable even if individual values are skewed. With only 16 larger counties, skewness matters more, which is one reason the analyst added a rank test as a check rather than relying on the t-test alone.

Effect Sizes

A p-value shows whether a difference is distinguishable from zero, not how large it is. Effect sizes fill that gap, and reporting them helps readers judge practical importance, since large samples can make trivial differences significant and small samples can hide important ones (Sullivan & Feinn, 2012). Using Cohen's d, the smoking difference was 0.81 standard deviations, a large effect by common conventions; diabetes and inactivity were about 0.5, moderate; obesity was 0.40 and insurance 0.33, smaller.

A Robustness Check

Diabetes values were right-skewed, so the analyst ran a rank-based Mann-Whitney test, which compares the ordering of values rather than means. For smoking, the result held: medians of 10.7% and 13.9%, p of 0.007. For diabetes, the rank test gave p of 0.12. The diabetes finding depends on the method chosen, which signals fragility.

What this part is doingReporting a result that weakened under a second test shows the analyst did not shop for significance.
4

Multiple Testing

Five tests at the 0.05 level carry a much higher chance of at least one false positive than a single test. A review of adjustment methods explains that corrections are needed when a claim rests on any one of several tests being significant, while they matter less for a few prespecified primary questions interpreted individually (Bender & Lange, 2001). Because the association planned to cite whichever outcomes differed, the analyst applied a Bonferroni threshold of 0.01. Only smoking passed.

Interpreting the Diabetes Result

The diabetes p-value of 0.049 sits right at the conventional line, the interval nearly touches zero, the rank test does not confirm it and it fails the corrected threshold. The honest summary is that smaller counties may have modestly higher diabetes, but this comparison cannot establish it.

Other Tests for Other Questions

Different questions call for different tests. Comparing three groups, such as metropolitan, small city and rural counties, would use one-way ANOVA, followed by adjusted pairwise comparisons, or the Kruskal-Wallis test for skewed data. Comparing categories, such as the share of counties above a smoking threshold in each group, would call for chi-square, switching to Fisher's exact method if expected cell counts fall below five.

Power and Group Size

With only 16 larger counties, the comparison has limited power to detect moderate differences. The obesity and insurance results, with wide intervals, could reflect real differences the data were too thin to confirm. Adding more years would not help much, because PLACES estimates for adjacent years share the same model; a better route would be individual-level survey data with a rural and urban indicator.

Limits of the Comparison

The unit of analysis is the county, so results describe counties, not individuals. The estimates are modeled from survey data and demographic information, which may exaggerate or blur differences. A cutoff at 50,000 lumps resort counties with high incomes together with poorer farm counties. Age adjustment removes differences due to age, which is appropriate for comparison but hides the actual burden in older rural counties.

What the Association Can Say

The analyst advised the association to lead with smoking, where the difference is large, consistent across methods and robust to adjustment, and to describe the other measures as pointing in the same direction but not clearly different. That framing will hold up if legislators' staff check the numbers.

Conclusion

Comparing 16 larger and 48 smaller Colorado counties showed a clear, large difference in adult smoking and weaker, inconsistent differences in diabetes and other measures. Choosing a test suited to unequal variances, reporting effect sizes, checking results with a rank test and adjusting for multiple outcomes kept the analysis honest and gave the association claims it can defend.

5

References

Bender, R., & Lange, S. (2001). Adjusting for multiple testing: When and how? Journal of Clinical Epidemiology, 54(4), 343-349. https://doi.org/10.1016/S0895-4356(00)00314-0

Centers for Disease Control and Prevention. (2025). PLACES: Local data for better health, county data, 2025 release [Data set]. https://data.cdc.gov/d/swc5-untb

Ruxton, G. D. (2006). The unequal variance t-test is an underused alternative to Student's t-test and the Mann-Whitney U test. Behavioral Ecology, 17(4), 688-690. https://doi.org/10.1093/beheco/ark016

Sullivan, G. M., & Feinn, R. (2012). Using effect size, or why the P value is not enough. Journal of Graduate Medical Education, 4(3), 279-282. https://doi.org/10.4300/JGME-D-12-00156.1

What the MPH 550 Week 4 instructions ask

The fourth MPH 550 assignment commonly focuses on comparing groups. Prompts may ask students to choose a test suited to the outcome type and number of groups, check assumptions such as normality and equal variance, run a t-test, rank test, chi-square test or ANOVA, report effect sizes and confidence intervals and interpret results for a public health audience. Some versions supply a dataset and software output, while others ask students to explain how they would choose a test. Justify every test you pick. Strong papers state why a test fits, check assumptions before running it, report the size and direction of differences, consider multiple comparisons and note when results depend on the method chosen.

How this MPH 550 Week 4 example is built

A state association's question, whether rural and small-town counties carry a heavier chronic disease burden, opens the paper. Counties are grouped by population, and five outcomes are chosen before testing. Assumptions are checked, and the unequal variance t-test is justified. Results for smoking, diabetes, physical inactivity, obesity and lack of insurance are reported with differences, intervals, p-values and standardized effect sizes. A rank test checks robustness, and a multiple testing correction is applied. The fragile diabetes result is examined closely. Alternatives such as ANOVA and chi-square are explained for questions with more groups or categorical outcomes. Limits of grouping counties and using modeled estimates close the paper, along with advice on which claims the association can safely make.

MPH 550 Week 4 grading rubric: where the points go

The group comparison week is typically assessed on appropriate test selection, checked assumptions and honest interpretation. Graders look for tests matched to outcome type and groups, assumptions examined, differences reported with confidence intervals and effect sizes, p-values interpreted correctly, multiple comparisons considered and conclusions stated in plain language. Sources on test choice and effect size strengthen the reasoning. Showing that a result holds, or fails, under a second method earns credit. Stating the unit of analysis and the limits of grouped data also earns marks. Clean tables and properly formatted references round out the evaluation. Drafts that run many tests and report only the significant ones usually score lower.

MPH 550 Week 4 help: mistakes to avoid

MPH 550 Week 4 drafts often pick a test by habit and report whichever p-values fall below 0.05. Start by naming the outcome type and the number of groups: two means suggest a t-test, three or more suggest ANOVA, counts or categories suggest chi-square. Check whether variances look equal and whether distributions are skewed; if not, use the unequal variance version or a rank test. Report the difference and its interval, then a standardized effect size so readers can judge importance. If you test several outcomes, say so and adjust. Finally, check whether a borderline result changes with a different reasonable method, and tell your reader if it does.

Related MPH 550 sample papers

Other MPH 550 week samples

More MPH sample papers

MPH 550 Week 4 questions, answered

What does MPH/550 Week 4 usually ask for?

The fourth statistics paper commonly focuses on comparing groups with suitable tests, such as t-tests, rank tests, chi-square or ANOVA, reporting effect sizes and interpreting results.

Where can I find a free MPH 550 Week 4 sample paper?

You can read the county comparison paper above for free, and notes explain each test choice. Share your data or output, and the first paper we draft for you is free.

Why use the unequal variance t-test?

It does not assume the two groups have the same variance, performs well when they do and protects against errors when they do not, which makes it a sensible default for comparing two means.

What is an effect size?

A measure of the magnitude of a difference or association, such as Cohen's d, which tells readers how large a difference is, something a p-value alone cannot show.

Why adjust for multiple testing?

Each additional test adds a chance of a false positive; with five tests at the 0.05 level, the chance of at least one false alarm rises well above 5% unless the threshold is adjusted.

Write yours, or have the desk draft it

This paper is an original model document written by our desk, not a submitted student paper and not an official University of Phoenix document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.