MPH 550 Week 1 Descriptive Statistics and Data Quality Example

Reviewed by Lenora Whitcombe, MSN, RN · University of Phoenix · Updated

This MPH 550 Week 1 example summarizes a real public health dataset with descriptive statistics and then asks how far its numbers can be trusted, using CDC PLACES estimates of adult diabetes and obesity for all 64 Colorado counties. University of Phoenix MPH 550 develops the statistical reasoning behind public health decisions, and in its first week MPH/550 students typically describe data with measures of center, spread and shape and assess data quality. The APA 7 paper reports age-adjusted diabetes prevalence averaging 8.57% across counties, with a median of 8.2% and a range from 6.0% in Douglas County to 13.1% in Bent County. Pueblo County sits at 10.4%. Research on how these small-area estimates are modeled from BRFSS survey data, and on the survey's validity, frames the data quality section.

CourseMPH 550 Public Health Statistics (MPH/550)
Week1
Paper typeDescriptive statistics paper
Lengthabout 1,192 words, 4 double-spaced pages plus title page and references
FormatAPA 7 student paper
SchoolUniversity of Phoenix
ProgramMPH
UpdatedSeptember 2026

Free sample paper for MPH 550 Week 1

1

Sixty-Four Counties, One Question: Describing Diabetes and Obesity Across Colorado and Judging Whether the Numbers Can Be Trusted

[Student Name]

University of Phoenix

MPH/550: Public Health Statistics

Week 1 Assignment

[Instructor Name]

[Date]

The health department, its analyst and the commissioners are composites written for a model paper; county estimates come from the CDC PLACES 2025 release, and other findings come from the sources cited.

What this part is doingThe title states the scale of the dataset and the two tasks, describing and judging.
2

At a spring meeting, a county commissioner asked the health department's analyst a direct question: is the county really worse than the rest of Colorado on diabetes, or does it just feel that way? The analyst answered with data. This paper describes how she summarized adult diabetes and obesity across all 64 Colorado counties using descriptive statistics, and how she judged whether the numbers deserved trust.

The Data

The analyst used the county file of CDC's PLACES project, 2025 release, which reports estimates for 2023 (Centers for Disease Control and Prevention [CDC], 2025). For each county, it gives the estimated percentage of adults with diagnosed diabetes and with obesity, both crude and age-adjusted, along with a 95% confidence interval and the county's population. She used age-adjusted values, because Colorado counties differ widely in age, and diabetes rises with age.

What this part is doingChoosing age-adjusted values is the first analytic decision, and naming the reason shows judgment.
3

Variable Types

The unit of analysis is the county, not the person. County name is a nominal variable. Prevalence of diabetes and obesity are continuous variables expressed as percentages. Population is a count. Recognizing that each row summarizes thousands of people, or in some counties fewer than a thousand, matters for everything that follows. The county populations themselves are highly unequal: the median county has about 15,100 residents, 27 counties have fewer than 10,000 and the largest has more than 700,000. That imbalance shapes both the averages and the reliability of each estimate.

Center: Diabetes

Across the 64 counties, age-adjusted diabetes prevalence averaged 8.57%. The median was 8.2%, meaning half of counties fell below that value and half above. The mean sits above the median because a few counties have much higher values that pull the average upward.

Spread: Diabetes

The standard deviation was 1.58 percentage points. The range ran from 6.0% in Douglas County to 13.1% in Bent County, a spread of 7.1 points. The middle half of counties fell between 7.2% and 9.6%, an interquartile range of 2.4 points. The highest county's diabetes rate was more than double the lowest's.

Center and Spread: Obesity

Obesity prevalence averaged 27.07%, with a median of 27.1% and a standard deviation of 3.78 points. The lowest value was 17.4% in Boulder County and the highest 35.1% in Otero County. Because the mean and median nearly match, the obesity distribution is roughly symmetric.

Shape

The diabetes distribution is right-skewed: most counties cluster between 7% and 10%, with a tail of higher values in the southeastern plains. A histogram would show a peak near 8% and a thin right tail. When a tail like this is present, the analyst prefers to lead with the median and the quartiles, reporting the mean alongside for comparison.

Outliers

The analyst used z-scores to check for outliers. Bent County's diabetes value of 13.1% lies 13.1 minus 8.57, divided by 1.58, or about 2.9 standard deviations above the mean. Under the interquartile range rule, the upper fence is 9.6 + 1.5 times 2.4, or 13.2%, so Bent sits just inside it. She treated Bent as unusual but plausible, not an error, since neighboring plains counties also run high.

Where Pueblo County Falls

Pueblo County's age-adjusted diabetes estimate was 10.4%, with a confidence interval of 8.9% to 12.0%. That places it above the third quartile; 54 of the 64 counties had lower estimates, and Pueblo sits about 1.2 standard deviations above the county mean. Its obesity estimate of 32.3% ranked higher still, with 57 counties below it. Note, though, that Pueblo's confidence interval overlaps the values of many counties near the middle of the distribution, so its exact rank is uncertain.

What this part is doingPlacing the county within the distribution answers the commissioner's question with a percentile, not an impression.
4

Counting Counties or Counting People

A county mean treats Hinsdale County, with 765 residents, the same as El Paso County, with 744,215. Weighted by population, mean diabetes prevalence across Colorado was 8.04%, lower than the 8.57% simple mean, because large Front Range counties have lower rates. The analyst reported both and explained the difference: the simple mean describes the typical county; the weighted mean describes the typical Coloradan.

Summary Table

The analyst's table listed, for diabetes and obesity, the number of counties, mean, standard deviation, median, quartiles, minimum and maximum, with Pueblo County's value and percentile in a final row. A footnote stated that values were age-adjusted estimates for 2023.

Data Quality: How the Numbers Were Made

The commissioner's trust depends on how the data were produced. PLACES estimates are not direct counts. The Behavioral Risk Factor Surveillance System, a telephone survey, interviews too few people in most counties to estimate prevalence directly, so CDC uses multilevel regression and poststratification, combining survey responses with census population counts to model local values.

Evidence on the Model

When CDC researchers developed this approach for chronic obstructive pulmonary disease, model-based estimates agreed closely with direct survey estimates where both were available; the correlation was 0.99 at the state level and 0.88 to 0.95 at the county level (Zhang et al., 2014).

A Clue in the Confidence Intervals

The confidence intervals revealed the modeling. Hinsdale County, with 765 residents, had a diabetes interval 2.3 points wide; El Paso County, nearly a thousand times larger, had an interval 2.4 points wide. With direct survey data, the smaller county's interval would be far wider. Similar widths signal that estimates lean on the model, so small counties' values partly reflect their demographic makeup rather than local measurements.

Self-Reported Data

The underlying survey relies on self-report. A systematic review found BRFSS prevalence estimates comparable to other self-reported national surveys but less similar to surveys that include physical measurements, with strong validity evidence for some topics and little for others (Pierannunzi et al., 2013). Diagnosed diabetes also misses people with undiagnosed disease, so true prevalence is higher.

What the Analyst Told the Commissioner

The analyst's answer was careful: by these estimates, the county ranks in the top quarter of Colorado counties for adult diabetes and obesity, about 1.2 standard deviations above the county average. The estimates are modeled rather than measured locally, and diagnosed diabetes undercounts the true burden. The pattern is strong enough to guide planning but not precise enough to rank counties separated by fractions of a point.

Checking the Data Before Summarizing

Before calculating anything, the analyst checked the file. Every county had one age-adjusted value for each measure, no values were missing and all fell between 0 and 100. She confirmed that the county names matched Colorado's 64 counties and that the year was the same for all rows. Such checks are routine, but a single duplicated or mislabeled row can distort a small dataset.

Limits

The analysis describes counties, not people, and says nothing about which residents have diabetes. One year of data cannot show trends. The age-adjusted values allow fair comparison but do not show the actual share of adults living with the disease locally.

Conclusion

Descriptive statistics turned 64 county estimates into an answer: diabetes prevalence is right-skewed, centered near 8%, with the county in the upper quarter. Checking the shape, outliers and weighting kept the summary honest. Examining how PLACES estimates are modeled, and what self-report can miss, showed how much confidence the numbers deserve.

5

References

Centers for Disease Control and Prevention. (2025). PLACES: Local data for better health, county data, 2025 release [Data set]. https://data.cdc.gov/d/swc5-untb

Pierannunzi, C., Hu, S. S., & Balluz, L. (2013). A systematic review of publications assessing reliability and validity of the Behavioral Risk Factor Surveillance System (BRFSS), 2004-2011. BMC Medical Research Methodology, 13, Article 49. https://doi.org/10.1186/1471-2288-13-49

Zhang, X., Holt, J. B., Lu, H., Wheaton, A. G., Ford, E. S., Greenlund, K. J., & Croft, J. B. (2014). Multilevel regression and poststratification for small-area estimation of population health outcomes: A case study of chronic obstructive pulmonary disease prevalence using the Behavioral Risk Factor Surveillance System. American Journal of Epidemiology, 179(8), 1025-1033. https://doi.org/10.1093/aje/kwu018

What the MPH 550 Week 1 instructions ask

The opening MPH 550 assignment generally centers on summarizing a health dataset and judging its quality. Prompts may ask students to identify variable types, calculate or interpret the mean, median, standard deviation, range and percentiles, describe a distribution's shape, flag outliers, present results in a table or chart and discuss where the data came from, how they were collected and what errors they may contain. Some versions supply a dataset, while others ask students to find public data. Cite the dataset as a reference. Strong papers choose statistics that suit the data's shape, interpret each number in plain words and treat data quality as a real question rather than a closing caveat.

How this MPH 550 Week 1 example is built

A county commissioner asking whether the county is really worse than the rest of Colorado on diabetes opens the paper. The analyst pulls PLACES estimates for all 64 counties and identifies the variables and their types. Measures of center and spread are calculated for diabetes and obesity, and the difference between mean and median reveals a right skew. Bent County is examined as a possible outlier using a z-score. A population-weighted mean is compared with the simple county mean. The data quality section explains that the estimates are modeled from survey responses, which shapes their confidence intervals, and reviews evidence on self-reported survey data. The analyst's plain-language answer to the commissioner closes the paper.

MPH 550 Week 1 grading rubric: where the points go

The descriptive statistics week is typically assessed on correct calculation, sound choice of summary measures and honest discussion of data quality. Graders look for variables classified by type, center and spread reported with suitable measures, distribution shape described, outliers identified with a stated method, results shown in a clear table and interpreted in plain words and a discussion of how the data were produced and what limits them. Citing the dataset and methodological sources strengthens the paper. Distinguishing between a county average and a population average earns credit. Tables formatted clearly and APA citations complete the grade. Drafts that report numbers without interpretation, or skip data quality, tend to earn less.

MPH 550 Week 1 help: mistakes to avoid

MPH 550 Week 1 papers often paste a column of statistics without saying what any of them mean. For each number, add a sentence a commissioner could follow: the typical county, how much counties differ, which are unusual. Check the shape before choosing measures; when the mean and median differ, say why and which one fits. Use a stated rule for outliers, such as a z-score or the interquartile range, rather than intuition. Ask whether each unit should count equally or be weighted by population. Then investigate the data: who collected them, how, from whom and whether the numbers are measured directly or modeled. That investigation is often the most useful part of the paper.

Related MPH 550 sample papers

Other MPH 550 week samples

More MPH sample papers

MPH 550 Week 1 questions, answered

What does MPH/550 Week 1 usually ask for?

The opening statistics paper generally centers on describing a health dataset with measures of center, spread and shape, presenting results clearly and judging the data's quality.

Where can I find a free MPH 550 Week 1 sample paper?

Read the Colorado diabetes paper above without paying; margin notes explain each statistic. Send your dataset or prompt, and your first paper is on us.

When should you use the median instead of the mean?

If a distribution has a long tail or a few extreme values, the median better represents a typical value, because extreme values pull the mean toward them.

What are PLACES estimates?

County and local estimates of health measures published by CDC, produced by modeling Behavioral Risk Factor Surveillance System responses together with census population data rather than by surveying each area directly.

Is self-reported survey data reliable?

A systematic review found BRFSS prevalence estimates comparable to other self-report surveys but less similar to surveys that include physical measurements, with strong validity evidence for some topics and little for others.

Write yours, or have the desk draft it

This paper is an original model document written by our desk, not a submitted student paper and not an official University of Phoenix document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.