PSY 315 Week 4 Analysis of Variance Example

Reviewed by Queenie Halstead, MA · University of Phoenix · Updated

This PSY 315 Week 4 example uses analysis of variance to compare three groups at once, partitioning variability, computing the F ratio, following up with post hoc tests and reporting eta squared. University of Phoenix PSY 315 turns to ANOVA in Week 4, and through PSY/315 psychology students learn why running many t tests inflates false alarms, how between-group and within-group variance form the F ratio and how to tell which groups differ after a significant result. The sample uses the composite community college survey from earlier weeks, where three sections in Phoenix prepared for the same exam with lecture review, practice quizzes or a study guide. It runs the ANOVA, applies Tukey's test, adds a second, nonsignificant ANOVA on caffeine groups and explains what each result supports.

CoursePSY 315 Statistical Reasoning in Psychology (PSY/315)
Week4
Paper typeAnalysis of variance report
Lengthabout 1,020 words, 4 double-spaced pages plus title page and references
FormatAPA 7 student paper
SchoolUniversity of Phoenix
ProgramBS in Psychology
UpdatedOctober 2026

Free sample paper for PSY 315 Week 4

1

Lecture Review, Practice Quizzes or Study Guides? A One-Way ANOVA on Three Class Sections

[Student Name]

University of Phoenix

PSY/315: Statistical Reasoning in Psychology

Week 4 Assignment

[Instructor Name]

[Date]

The sections, the students and all numbers are composites written for a model paper; statistical guidance and research findings come from the sources listed.

What this part is doingThe title lists the three conditions the analysis will compare.
2

The t tests in Week 3 compared two means at a time. Many psychology questions involve three or more groups, and analysis of variance answers them with a single test. This report applies a one-way ANOVA to three class sections that prepared differently for the same exam, follows up the result with post hoc comparisons and reports a second analysis that did not reach significance.

The Question and the Data

Three sections of one introductory course at a Phoenix, Arizona, two-year college sat the same exam. Each instructor used a different review method in the week before: Section A reviewed through lecture, Section B took three short practice quizzes with feedback and Section C received a detailed study guide. Each section had 40 students who completed the class survey, for 120 in all. The question is whether mean exam scores differ across the three review methods.

Why One Test Instead of Three

Comparing three groups in pairs would need three t tests. If each uses an alpha of .05, the chance of at least one false alarm across the set climbs to about .14. Bender and Lange (2001) explained that when several comparisons address one overall question, the error rate for the family of tests should be controlled, and they described when adjustments are needed and how to apply them. ANOVA checks every mean in a single step at an alpha of .05, and follow-up comparisons then hold the error rate down across the pairs.

Hypotheses

Under the null hypothesis, all three review methods would produce the same population mean; the alternative holds that one or more of them would not. ANOVA does not say which one; that is the job of the follow-up tests.

Partitioning the Variability

Exam scores across all 120 students had a mean of 76.3 and a total sum of squares of 14,927.4. ANOVA divides that total into two parts. The between-groups sum of squares reflects how far each section's mean lies from the grand mean, weighted by section size: 1,589.6. The within-groups sum of squares reflects how much students vary inside their own section: 13,337.8. Together they add back to the total.

What this part is doingShowing that the two sums of squares add to the total lets a reader check the arithmetic.
3

The Summary Table

Dividing each sum of squares by its degrees of freedom gives a mean square. Between groups, with 2 degrees of freedom, the mean square is 794.8. Within groups, with 117 degrees of freedom, it is 114.0. The F ratio is 794.8 divided by 114.0, or 6.97. With 2 and 117 degrees of freedom, p = .001, so the null hypothesis is rejected.

In APA style: exam scores differed across review methods, F(2, 117) = 6.97, p = .001, eta squared = .11.

Effect Size

Eta squared, found by dividing the review-method share of variability by the total, shows that review method accounts for about 11 percent of the variability in exam scores, a medium effect by common guidelines. Olejnik and Algina (2003) pointed out that eta squared and related measures can be inflated in small samples or distorted by design features and recommended generalized versions so effects can be compared fairly across different designs. Omega squared, a less biased estimate, gives a slightly smaller value here, about .09.

Which Sections Differ?

Tukey's honestly significant difference test compares every pair while holding the familywise error rate at .05. With the mean square within of 114.0 and 40 students per group, the smallest difference that counts as significant is about 5.7 points. The quiz section, at 80.9, scored 8.9 points above the lecture section at 72.0, a significant difference. The quiz and study guide sections differed by 4.9 points and the study guide and lecture sections by 4.0, neither reaching the threshold.

ANOVA said the three sections were not all alike; Tukey's test showed that only one pair, quizzes against lecture, truly stood apart.

A Plausible Explanation

Roediger and Karpicke (2006) had students either restudy prose passages or take practice recall tests on them. Restudying helped more on a test given minutes later, but practice testing produced much better retention after two days and one week. Practice quizzes may have given Section B the benefit of retrieval practice. The ANOVA cannot confirm this, however, because sections were intact classes rather than randomly assigned groups. Different instructors, class times and student backgrounds could also explain the gap.

What this part is doingLinking the result to retrieval research offers a reason while the design note keeps the claim modest.
4

Checking Assumptions

ANOVA assumes independent observations, roughly normal scores in each group and similar variances. The three standard deviations were close, between 10.1 and 11.3, and Levene's test was not significant. Histograms in each section were near bell-shaped. With equal group sizes, ANOVA is fairly robust to modest violations.

A Nonsignificant Comparison

A second ANOVA compared exam scores across caffeine groups: no caffeine (20 students, mean 77.5), one or two cups (68 students, mean 77.8) and three or more (32 students, mean 72.4). The result was F(2, 117) = 2.74, p = .068, eta squared = .04. Because the overall test was not significant, no post hoc tests were run. This does not show that caffeine is unrelated to performance; heavy users scored about five points lower, and a larger sample might detect a real but modest effect.

Extending to a Two-Way Design

A natural next study would cross review method with sleep, creating six cells: short and longer sleepers within each section. A two-way ANOVA would then test three questions at once: whether review method matters, whether sleep matters and whether the benefit of practice quizzes depends on how much students sleep. That last question, the interaction, is often the most interesting. If quizzes helped well-rested students but did little for exhausted ones, the advice to the college would change, and a one-way analysis would never reveal it.

Conclusion

A one-way ANOVA showed that exam scores differed across three review methods, with practice quizzes producing the clearest advantage over lecture review. Eta squared placed the effect in the medium range, Tukey's test located the difference and research on retrieval practice offers a reason. The nonsignificant caffeine analysis and the lack of random assignment are reminders to read both results with care.

5

References

Bender, R., & Lange, S. (2001). Adjusting for multiple testing: When and how? Journal of Clinical Epidemiology, 54(4), 343-349. https://doi.org/10.1016/S0895-4356(00)00314-0

Olejnik, S., & Algina, J. (2003). Generalized eta and omega squared statistics: Measures of effect size for some common research designs. Psychological Methods, 8(4), 434-447. https://doi.org/10.1037/1082-989X.8.4.434

Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249-255. https://doi.org/10.1111/j.1467-9280.2006.01693.x

What the PSY 315 Week 4 instructions ask

The fourth PSY 315 assignment usually asks students to conduct and interpret a one-way analysis of variance. Common requirements include stating hypotheses for three or more group means, explaining between-group and within-group variability, completing or producing an ANOVA summary table, reporting the F ratio with degrees of freedom and p value, calculating an effect size and running post hoc comparisons when the result is significant. Some versions also introduce two-way designs and interactions. Explain why ANOVA is preferred over several t tests, check assumptions such as equal variances and interpret the follow-up tests in words. Report results in APA style and cite the textbook and one additional source.

How this PSY 315 Week 4 example is built

Our worked report compares exam scores for three sections of 40 students. The practice quiz section averaged 80.9, the study guide section 76.0 and the lecture review section 72.0. The ANOVA gives F(2, 117) = 6.97, p = .001, with eta squared of .11, and Tukey's test shows that only the quiz and lecture sections differ significantly. Research on test-enhanced learning offers a plausible reason the quiz group did best. A second ANOVA comparing caffeine groups is not significant, and the paper explains why that does not prove caffeine is irrelevant. Sources on effect size measures and on correcting for multiple comparisons guide the analysis and its limits.

PSY 315 Week 4 grading rubric: where the points go

ANOVA reports are usually graded on correct hypotheses, an accurate summary table, appropriate follow-up tests and clear interpretation. Instructors look for sums of squares that add up, degrees of freedom computed correctly and an F ratio reported in APA style. Credit goes to an effect size such as eta squared, to post hoc tests used only after a significant overall result and to conclusions that respect the design, such as noting that intact class sections were not randomly assigned. A short explanation of why ANOVA controls the familywise error rate earns credit, as does a figure showing group means with error bars and a sentence on practical meaning.

PSY 315 Week 4 help: mistakes to avoid

A frequent error is running three separate t tests instead of one ANOVA, which raises the chance of a false alarm across the set. Another is stopping at a significant F without testing which groups differ, or running post hoc tests after a nonsignificant F. Students also mislabel the summary table, putting the within-group mean square in the numerator, or forget that F is always positive. Some reports call a nonsignificant result proof of no difference. Others claim the teaching method caused the gap when sections were not assigned at random. Build the table step by step, report an effect size and follow up only when the omnibus test allows. A tutor can help you check each row of an ANOVA table.

Related PSY 315 sample papers

Other PSY 315 week samples

More BS in Psychology sample papers

PSY 315 Week 4 questions, answered

What does PSY 315 Week 4 usually cover?

It usually covers one-way analysis of variance, the F ratio, effect sizes and post hoc comparisons, sometimes with two-way designs.

Where can I find a free PSY 315 Week 4 sample paper?

Above is a free, complete PSY 315 Week 4 ANOVA report comparing three class sections.

Why not just run several t tests?

Each test carries its own risk of a false alarm, so running many raises the overall chance of at least one Type I error.

What does the F ratio compare?

Variability between group means against variability within groups; a large ratio suggests real differences among the means.

When do you run a post hoc test?

After a significant ANOVA, to find out which specific pairs of groups differ while controlling the overall error rate.

Write yours, or have the desk draft it

This paper is an original model document written by our desk, not a submitted student paper and not an official University of Phoenix document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.