PSY 315 Week 3 Hypothesis Testing and t Tests Example

Reviewed by Queenie Halstead, MA · University of Phoenix · Updated

This PSY 315 Week 3 example runs a full hypothesis test from start to finish, writing the hypotheses, choosing among one-sample, independent and paired t tests and reporting effect sizes and confidence intervals alongside p values. University of Phoenix PSY 315 introduces inferential testing in Week 3, and PSY/315 trains psychology students to follow the steps of a significance test, avoid Type I and Type II errors and explain what a significant finding can and cannot tell a reader. The sample stays with the same Phoenix class survey used in Weeks 1 and 2. It compares exam scores of short and longer sleepers, tests the class average against a college benchmark and checks stress before and after a brief breathing exercise, then discusses what the numbers mean in context.

CoursePSY 315 Statistical Reasoning in Psychology (PSY/315)
Week3
Paper typeHypothesis testing report
Lengthabout 1,063 words, 4 double-spaced pages plus title page and references
FormatAPA 7 student paper
SchoolUniversity of Phoenix
ProgramBS in Psychology
UpdatedOctober 2026

Free sample paper for PSY 315 Week 3

1

Do Short Sleepers Score Lower? Hypothesis Testing and Three t Tests on Class Survey Data

[Student Name]

University of Phoenix

PSY/315: Statistical Reasoning in Psychology

Week 3 Assignment

[Instructor Name]

[Date]

The survey, the students and all numbers are composites written for a model paper; statistical guidance comes from the sources listed.

What this part is doingThe title poses the research question as the hypothesis test will answer it.
2

Description in Week 1 and probability in Week 2 set up the central task of inferential statistics: deciding whether a pattern in a sample is likely to reflect something real in the population or could easily have arisen by chance. This report applies hypothesis testing to three questions from the composite class survey and reports each result with its effect size and confidence interval.

The Logic of a Significance Test

Every significance test starts from a null hypothesis, typically the claim of zero difference or zero effect, set against an alternative claiming some effect exists. The researcher sets an alpha level, commonly .05, collects data and computes a test statistic. If the result would be very unlikely under the null hypothesis, with p below alpha, the null is rejected. Two mistakes are possible. A Type I error rejects a true null hypothesis, a false alarm. A Type II error fails to reject a false one, a miss. Lowering alpha reduces false alarms but raises the risk of misses unless the sample is larger.

Question One: Sleep and Exam Scores

The survey split students into 42 who slept fewer than six hours and 78 who slept six or more. Under the null, short and longer sleepers share one population mean; under the alternative, their means are not the same. Because the groups contain different students, an independent-samples t test fits.

Short sleepers averaged 72.10 (SD = 11.80), and longer sleepers averaged 78.60 (SD = 10.30). The pooled standard deviation was 10.84, and the standard error of the difference was 2.08. The test gave t(118) = 3.13, p = .002. The difference of 6.50 points had a 95 percent confidence interval from 2.39 to 10.61, and Cohen's d was 0.60, a medium effect.

What this part is doingReporting the interval alongside p shows the plausible size of the gap, not just that one exists.
3

Interpreted carefully, short sleepers scored lower, and the gap is large enough to matter for a letter grade. The design does not show that short sleep causes lower scores. Students who work night jobs, care for children or feel more stressed might both sleep less and study less.

Checking Assumptions

The independent t test assumes roughly normal scores within each group and similar variances. Histograms of exam scores in both groups were close to bell-shaped, and the two standard deviations, 11.80 and 10.30, differ by much less than a factor of two. With groups of this size, the test is robust to modest departures. Rerunning the comparison with Welch's correction for unequal spread gave nearly the same result.

Question Two: The Class Against a Benchmark

The department reports that students across the college average 74 on this common exam. A one-sample t test asks whether this class differs. With a mean of 76.30, a standard deviation of 11.20 and 120 students, the standard error is 1.02, and t(119) = 2.25, p = .026. The effect size, d = 0.21, is small. With a large sample, even a modest difference reaches significance, which is why the effect size matters: the class did slightly better than the college as a whole, not dramatically better.

A p value of .026 says the class probably differs from the benchmark; an effect size of 0.21 says the difference is modest.

Question Three: Stress Before and After a Breathing Exercise

In one section, 25 students rated their stress before and after a five-minute guided breathing exercise at the start of class. Because each student gave two ratings, a paired t test compares the differences within each person. Ratings fell from an average of 6.40 to 5.60, a mean change of 0.80 with a standard deviation of 1.50. The test gave t(24) = 2.67, p = .013, with an effect size of 0.53 based on the differences.

Two cautions apply. Stress ratings are ordinal, so a Wilcoxon signed-rank test would be a reasonable check; it gave the same conclusion. And without a comparison group that sat quietly for five minutes, the drop might reflect settling into class rather than the breathing itself.

What p Values Can and Cannot Say

Cohen (1994) criticized the habit of treating p below .05 as the goal of research, pointing out that a small p does not tell a researcher how likely the null is to be correct and that rejecting a null of exactly zero says almost nothing about whether an effect is big enough to care about. He urged researchers to report effect sizes and confidence intervals. Wasserstein and Lazar (2016), writing for the American Statistical Association, set out six principles, among them that p by itself cannot convey effect size or practical weight and that crossing one cutoff such as .05 is a poor basis for a scientific or policy decision.

Lakens (2013) offered a practical guide to computing and reporting effect sizes for t tests and analysis of variance, noting that effect sizes allow results to be compared across studies and combined in meta-analyses. The three tests above follow that advice.

What this part is doingPairing each p value with an effect size applies the critiques rather than merely citing them.
4

Errors in This Context

A Type I error in question one would mean concluding that short sleepers score lower when, in the population, they do not; the college might launch a sleep campaign aimed at a problem that does not exist. A Type II error in question three would mean missing a real benefit of the breathing exercise because 25 students were too few to detect it, and instructors might abandon a helpful routine.

Writing the Results Section

A results paragraph for question one might read: students who slept fewer than six hours scored lower on the exam (M = 72.10, SD = 11.80) than students who slept six hours or more (M = 78.60, SD = 10.30), t(118) = 3.13, p = .002, d = 0.60, 95% CI [2.39, 10.61]. Each statistic is italicized in the final paper, means and standard deviations come first and the sentence states the direction of the difference before any number, so a reader learns the finding even if they skip the statistics.

Conclusion

Three t tests matched to three designs show that short sleepers in this survey scored about six and a half points lower, that the class modestly outperformed the college benchmark and that stress ratings dropped after a short breathing exercise. Reporting effect sizes and confidence intervals with each p value, checking assumptions and resisting causal claims make these findings informative rather than merely significant.

5

References

Cohen, J. (1994). The earth is round (p < .05). American Psychologist, 49(12), 997-1003. https://doi.org/10.1037/0003-066X.49.12.997

Lakens, D. (2013). Calculating and reporting effect sizes to facilitate cumulative science: A practical primer for t-tests and ANOVAs. Frontiers in Psychology, 4, Article 863. https://doi.org/10.3389/fpsyg.2013.00863

Wasserstein, R. L., & Lazar, N. A. (2016). The ASA statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129-133. https://doi.org/10.1080/00031305.2016.1154108

What the PSY 315 Week 3 instructions ask

Week 3 assignments in PSY 315 usually ask students to conduct and interpret hypothesis tests. Typical requirements include stating null and alternative hypotheses, choosing the correct t test for the design, setting an alpha level, calculating or running the test in software, reporting the result in APA style and explaining Type I and Type II errors. Many versions now expect effect sizes and confidence intervals as well as p values. Some provide a data file and ask for output; others pose word problems. Match the test to the design, check assumptions, report every statistic with its degrees of freedom and write a conclusion in plain language that a reader without statistics training could follow. Cite the textbook and an outside source.

How this PSY 315 Week 3 example is built

Our worked report tests three questions with the Week 1 survey. Short sleepers averaged 72.1 on the exam and longer sleepers 78.6, and an independent-samples t test finds the gap significant with a medium effect size and a confidence interval from about 2 to 11 points. The class mean of 76.3 is compared with a college benchmark of 74 using a one-sample t test, which finds a small but significant difference. A paired t test shows stress ratings dropped after a five-minute breathing exercise. A classic critique of significance testing and a professional statement on p values shape how each result is interpreted, and a primer on effect sizes guides the reporting.

PSY 315 Week 3 grading rubric: where the points go

Hypothesis testing reports are usually graded on correct choice of test, accurate calculations and sound interpretation. Instructors look for hypotheses written in symbols and words, for the right degrees of freedom and for results reported in APA format, such as t(118) = 3.13, p = .002. Credit goes to effect sizes and confidence intervals, to checks on assumptions such as normality and equal variances and to conclusions that avoid claiming proof or causation from a survey. Graders also reward a clear explanation of what a Type I or Type II error would mean in the specific study, not just a textbook definition, and a short comment on practical importance.

PSY 315 Week 3 help: mistakes to avoid

Students often choose an independent test when the same people were measured twice, which calls for a paired test, or the reverse. Ask whether each person contributes one score or two. Another frequent error is saying a nonsignificant result proves no difference exists, when it may reflect a small sample. Some reports state that p is the probability the null hypothesis is true, a common misreading. Others report p values alone without effect sizes, or claim that short sleep causes lower grades from survey data. Identify the design, report the full set of statistics and interpret them cautiously in context. A tutor can help you match each research question to the right t test.

Related PSY 315 sample papers

Other PSY 315 week samples

More BS in Psychology sample papers

PSY 315 Week 3 questions, answered

What does PSY 315 Week 3 usually cover?

It usually covers the steps of hypothesis testing, Type I and Type II errors and one-sample, independent and paired t tests.

Where can I find a free PSY 315 Week 3 sample paper?

Scroll up for the complete PSY 315 Week 3 report running three t tests on survey data, free to read.

When do I use a paired t test?

When the same people are measured twice, or when each person in one group is matched with a person in the other.

What does a p value mean?

The probability of a result at least as extreme as the one observed if the null hypothesis were true, not the chance the null is true.

Why report an effect size?

Because significance depends on sample size, while an effect size shows how large the difference is in practical terms.

Write yours, or have the desk draft it

This paper is an original model document written by our desk, not a submitted student paper and not an official University of Phoenix document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.