A Significant t-Test and a Nonsignificant Chi-Square From the Same Coaching Pilot: Comparing Two Groups of Patients With Diabetes and Reading P Values, Confidence Intervals and Effect Sizes Correctly
[Student Name]
University of Phoenix
DNP/701: Biostatistics and Epidemiology
Week 5 Assignment
[Instructor Name]
[Date]
The clinic and its figures are a composite written for a model paper.
Our rural clinic network piloted nurse telephone coaching for adults with diabetes and an A1c above 8%. Sixty patients received biweekly coaching calls for six months, and 60 similar patients from another site received usual care. I compared the two groups on two outcomes: change in A1c and the proportion reaching an A1c below 8%. This paper presents the hypotheses, the tests, the results and their interpretation.
Hypotheses
For A1c change, the null hypothesis states that coached and uncoached patients share the same average change, and the two-sided alternative states that their averages are unequal. For goal attainment, the null hypothesis states that equal shares of each group reach an A1c under 8%, and the alternative states that the shares are unequal. I used two-sided tests at an alpha of .05, set before analysis.
Choosing the Tests
A1c change is a continuous variable, and its distribution in each group was roughly symmetric, so an independent-samples t-test is appropriate. Because the group variances were not identical, I used Welch's version, which does not assume equal variances. Reaching goal is a yes-or-no outcome, so a chi-square test of independence compares the two proportions.
A1c Change Results
The coaching group's mean A1c fell by 0.9 percentage points, standard deviation 1.2. The usual care group's fell by 0.4, standard deviation 1.3. The difference in mean change is 0.5 points. Its standard error comes from adding each group's variance divided by its size, 1.44 over 60 and 1.69 over 60, and taking the square root, which gives about 0.23. The t statistic is 0.5 divided by 0.23, or about 2.19, with a p value of about .03. The 95% confidence interval for the difference is approximately 0.05 to 0.95 percentage points.
Goal Attainment Results
In the coaching group, 27 of 60 patients, or 45%, reached an A1c below 8%. In usual care, 18 of 60, or 30%, did. Expected counts under the null hypothesis are 22.5 reaching goal and 37.5 not reaching goal in each group. The chi-square statistic is 2.88 with one degree of freedom, and the p value is about .09. The gap between the two groups is 15 points, and its 95% interval runs from roughly minus 2 to plus 32.
One test crossed .05 and one did not, but both confidence intervals leaned the same way; the data told one story, not two.
What the P Values Mean
Writing for the national statistics society, Wasserstein and Lazar (2016) set out six principles. Among them: a p value shows how poorly the data fit a specified statistical model; it is not the chance that a hypothesis is correct or that the findings were produced by luck alone; decisions should not hinge on whether it crosses a cutoff; and it says nothing about how large or important an effect is.
Avoiding Common Misreadings
Greenland et al. (2016) catalog common misinterpretations. A p value of .03 does not mean there is a 3% chance that coaching has no effect. A p value of .09 does not show that coaching has no effect on goal attainment. And a significant result in one test and a nonsignificant result in another do not show that the two outcomes differ in response to coaching. They recommend interpreting confidence intervals as ranges of effect sizes reasonably compatible with the data.
Reading the Intervals
The A1c interval, 0.05 to 0.95, suggests that coaching's effect could be very small or nearly a full percentage point. The goal attainment interval, minus 2 to plus 32 points, includes no difference but extends to a large benefit. Both intervals favor coaching. The goal attainment test lacked precision because proportions from 60 patients per group carry wide uncertainty.
Effect Size
Sullivan and Feinn (2012) argue that the p value alone is not enough because it depends on sample size, and that effect sizes, such as Cohen's d, describe the magnitude of a difference independent of sample size. Using a pooled standard deviation of about 1.25, Cohen's d for A1c change is 0.5 divided by 1.25, or 0.40, a small-to-moderate effect by conventional benchmarks. For goal attainment, the risk difference of 15 points corresponds to a number needed to coach of about 7.
One-Sided or Two-Sided
A one-sided test would have assumed that coaching could only help. Because coaching could in principle harm, for example by prompting overly aggressive insulin changes that cause hypoglycemia, a two-sided test was the honest choice, set before looking at the data.
Multiple Comparisons
I tested two outcomes. Testing many outcomes increases the chance that at least one crosses .05 by chance. With two prespecified outcomes, the concern is modest, but I report both results rather than only the significant one, which is the more important protection against misleading conclusions.
A Nonparametric Check
As a sensitivity analysis, I ran the Wilcoxon rank-sum test on A1c change, which does not assume normality. It gave a similar result, p = .04, which supports the t-test conclusion.
Assumptions and Limits
The groups came from different sites and were not randomized, so differences between sites may explain part of the result. The t-test assumes independent observations and reasonably normal distributions; both held approximately. The sample was small, which widens intervals and reduces power.
Power
A post hoc power calculation is not informative once results are known. More useful is planning: to detect a 15-point difference in goal attainment with 80% power, each group would need roughly 160 patients. Our pilot was too small to test that outcome precisely.
Interpreting for the Clinic
My summary for our director: "Coaching was associated with about a half-point greater A1c reduction, with a plausible range from very small to nearly one point, and more patients reaching goal, though that difference was less certain. The evidence favors coaching and supports a larger, better-designed evaluation."
Practical Versus Statistical Thinking
Our director does not need to know whether a p value crossed .05. She needs to know whether coaching probably helps, how much and how sure we are. Framing results as effect estimates with ranges answers those questions directly and avoids the trap of treating .05 as a line between truth and falsehood.
Software
I ran the analysis in R using the t.test function with the Welch correction and the chisq.test function without continuity correction, and I checked the results against hand calculations.
Conclusion
A Welch t-test showed a statistically significant difference in A1c change, and a chi-square test did not reach significance for goal attainment. Following the American Statistical Association statement and guidance on misinterpretation, these results do not conflict: both estimates favor coaching, and the difference in significance reflects precision, not a difference in effect. Confidence intervals and effect sizes give the clinic a more honest picture than p values alone.
References
Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337-350. https://doi.org/10.1007/s10654-016-0149-3
Sullivan, G. M., & Feinn, R. (2012). Using effect size, or why the P value is not enough. Journal of Graduate Medical Education, 4(3), 279-282. https://doi.org/10.4300/JGME-D-12-00156.1
Wasserstein, R. L., & Lazar, N. A. (2016). The ASA statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129-133. https://doi.org/10.1080/00031305.2016.1154108
How this DNP 701 Week 5 example is structured
The DNP/701 Week 5 work usually addresses hypothesis testing for comparing groups. This paper states the hypotheses, chooses tests that fit each outcome, shows the calculations and then interprets the results with confidence intervals and effect sizes rather than p values alone. Students search this week as DNP 701 Week 5, DNP701 Wk 5 or DNP/701 Wk 5; all three are the same assignment.
DNP/701 Week 5 questions, answered
What does DNP/701 Week 5 usually ask for?
Many sections ask students to select and conduct statistical tests comparing groups, such as t-tests and chi-square tests, and to interpret the results for practice.
What does a p value actually mean?
It is the probability of obtaining data at least as extreme as those observed if the null hypothesis and all other model assumptions were true; it is not the probability that the null hypothesis is true or that the result occurred by chance.
Does a nonsignificant result mean there is no difference?
No; it means the data are compatible with no difference, but they may also be compatible with a meaningful difference, especially in a small study, and the confidence interval shows that range.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official University of Phoenix document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.