PSYCH 655 Week 2 Reliability Example

Reviewed by Queenie Halstead, MA · University of Phoenix · Updated

This PSYCH 655 Week 2 example examines reliability, the consistency of scores, by studying a vocational rehabilitation office's hands-on work-sample battery: whether clients score the same on retest, whether two evaluators agree and what internal consistency statistics can and cannot show. Reliability is the whole of Week 2 in University of Phoenix PSYCH 655; for PSYCH/655, MS in Psychology students explain classical test theory, compute and interpret reliability coefficients, use the standard error of measurement and choose the right reliability evidence for a test's use. The study is run by a composite vocational evaluator in Albuquerque whose office relies on work samples for training recommendations. She draws on the original paper defining coefficient alpha, a critique arguing for alternatives to alpha and a guideline for selecting and reporting intraclass correlations.

CoursePSYCH 655 Psychometrics (PSYCH/655)
Week2
Paper typeReliability analysis
Lengthabout 1,164 words, 4 double-spaced pages plus title page and references
FormatAPA 7 student paper
SchoolUniversity of Phoenix
ProgramMS in Psychology
UpdatedOctober 2026

Free sample paper for PSYCH 655 Week 2

1

Would Two Evaluators Give the Same Score? Reliability Evidence for a Vocational Work-Sample Battery

[Student Name]

University of Phoenix

PSYCH/655: Psychometrics

Week 2 Assignment

[Instructor Name]

[Date]

The vocational rehabilitation office, its work samples and the reliability data are composites written for a model paper; research findings come from the sources listed.

What this part is doingThe title asks the reliability question that matters most for a rater-scored test.
2

A score that changes depending on the day of testing or the person scoring it cannot support important decisions. Reliability is the consistency of scores, and different kinds of reliability evidence answer different questions. This paper examines the reliability of a work-sample battery used at the vocational rehabilitation office where I work.

The Battery

Our office in Albuquerque uses a six-station work-sample battery to help clients and counselors choose training paths. Stations include small-parts assembly, sorting mail by zip code, measuring and cutting materials to specification, operating a cash register simulation, assembling a simple circuit board and following a set of written multistep instructions to complete a task. Each station yields a score combining speed and accuracy, and scores are combined into a profile. Recommendations based on the battery affect whether clients are funded for particular training programs, so the stakes are meaningful.

True Scores and Error

Classical test theory proposes that any observed score equals a true score plus error. The true score is the average a person would obtain over infinite independent testings; error includes anything that makes an observed score differ from it, such as fatigue, distraction, luck in guessing or differences between scorers. Reliability is the proportion of observed score variance due to true score variance, ranging from zero to one. No single reliability coefficient captures all sources of error, so test users must choose evidence that fits their purpose.

Internal Consistency and Its Limits

Cronbach (1951) introduced coefficient alpha as a general formula for the internal consistency of a test, showing that it equals the mean of all possible split-half coefficients and that it reflects how strongly items or parts covary. Alpha became the most widely reported reliability statistic in psychology.

McNeish (2018) argued that alpha is widely misused because it rests on assumptions rarely met in practice, especially that every item relates equally strongly to the underlying construct, called tau equivalence, and that errors are uncorrelated. When these assumptions fail, alpha can underestimate or overestimate reliability. McNeish recommended alternatives such as omega coefficients computed from factor models and urged researchers to choose reliability estimates that fit their data.

What this part is doingExplaining alpha's assumptions prepares the reader for why it is not the main evidence here.
3

Our battery's six stations measure different skills by design, so high internal consistency across stations is neither expected nor desirable; alpha across stations was .58, which reflects the battery's breadth rather than poor measurement. Internal consistency within each station matters more, but the most important evidence for our purposes is stability over time and agreement between raters.

Test-Retest Reliability

Thirty clients completed the full battery twice, two weeks apart, with no training between sessions. The correlation between the two sets of total scores was .84, and station correlations ranged from .71 for circuit board assembly to .88 for sorting. Scores rose slightly on retest, by about four percent on average, a practice effect, particularly on the timed stations. A two-week interval balances memory and practice effects against real change in skills.

Interrater Reliability

Several stations require evaluator judgment. Koo and Li (2016) provided a guideline for selecting and reporting intraclass correlation coefficients, explaining that there are several forms depending on whether raters are a random sample or fixed, whether a single rater or the mean of raters is used and whether absolute agreement or consistency is of interest. They suggested that values below .5 indicate poor reliability, .5 to .75 moderate, .75 to .9 good and above .9 excellent, and urged researchers to report which form they used.

Two evaluators independently scored twenty clients. Using a two-way random effects model for absolute agreement and single raters, interrater reliability was .91 for timed stations, where scoring depended mainly on a stopwatch and error counts, but only .62 for the written instructions station, where evaluators judged whether each step was completed correctly.

On the stopwatch stations, two evaluators agreed almost perfectly; on the station that needed judgment, they agreed only moderately.

Uncertainty Around One Score

The standard error of measurement estimates how much an individual's observed score would vary around their true score. It equals the standard deviation times the square root of one minus the reliability. For the battery total, with a standard deviation of 12 points and test-retest reliability of .84, the standard error is about 4.8 points. A client scoring 70 has a 95 percent confidence band of roughly 61 to 79. For the written instructions station, lower reliability produces a much wider band, which counselors should understand before treating small score differences as meaningful.

What this part is doingTurning reliability into a confidence band shows counselors what the numbers mean for one client.
4

How Reliable Is Reliable Enough?

The level of reliability needed depends on the stakes. For research comparing group averages, coefficients around .70 may be acceptable. For decisions about individuals, such as whether a client is funded for a specific training program, much higher reliability is expected, often .85 to .90 or above for key scores, because error in an individual score directly affects that person. Our battery total, at .84, approaches that level; the written instructions station, at .62 for interrater agreement, falls well short and should not drive decisions on its own.

Reliability Is Not Validity

High reliability shows that scores are consistent, not that they measure what the office intends or predict success in training. A station could be scored with perfect agreement and still have little to do with the jobs clients pursue. Validity evidence, the subject of next week, must be examined separately. Reliability sets an upper limit on validity: a test cannot correlate with an outcome more strongly than its scores correlate with themselves.

Sources of Error

The main sources of error were practice effects on timed stations and evaluator judgment on the written instructions station, where the rubric left room for interpretation about partially completed steps. Fatigue may also affect clients with certain disabilities in later stations.

Clients With Disabilities

Our clients have a wide range of disabilities, which raises reliability questions of its own. A client with chronic pain may perform differently on a morning and an afternoon, and a client with a seizure disorder may have good and bad weeks. For such clients, a single administration may not represent typical performance, and evaluators may need to schedule testing across two sessions or note conditions that could have depressed scores. These considerations do not lower the reliability coefficient for the group, but they matter for the meaning of an individual's score.

Improvements

I recommended rewriting the written instructions station's scoring rubric with specific examples of complete, partial and incorrect steps, training evaluators with practice videos until agreement reaches at least .80, rotating the order of stations to spread fatigue effects and reporting confidence bands with every score.

Conclusion

Reliability evidence must match how scores are used. For our work-sample battery, test-retest and interrater reliability matter more than alpha. The battery is stable over time and reliably scored on timed stations, while the judgment-based station needs a clearer rubric. Confidence bands make the remaining uncertainty visible to the people who rely on the scores.

5

References

Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297-334. https://doi.org/10.1007/BF02310555

Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155-163. https://doi.org/10.1016/j.jcm.2016.02.012

McNeish, D. (2018). Thanks coefficient alpha, we'll take it from here. Psychological Methods, 23(3), 412-433. https://doi.org/10.1037/met0000144

What the PSYCH 655 Week 2 instructions ask

The second week of PSYCH 655 usually focuses on reliability, the first property any score needs before it can mean anything. Students are typically asked to explain true score and error in classical test theory, describe test-retest, alternate forms, internal consistency and interrater reliability, interpret coefficients and the standard error of measurement and consider factors that raise or lower reliability. Some versions supply data to analyze. Match each kind of reliability to the decision at hand, calculate and explain at least one coefficient and a confidence band around an individual score, discuss sources of error specific to the test and recommend improvements. Support the analysis with measurement research cited in APA style.

How this PSYCH 655 Week 2 example is built

Hana Kobayashi, the evaluator who writes this worked paper, investigates the reliability of her office's six-station work-sample battery, which includes small-parts assembly, sorting, measuring and following written instructions. Thirty clients completed the battery twice, two weeks apart, and two evaluators scored twenty clients independently. The original alpha paper explains internal consistency as average item agreement. A critique shows that alpha assumes conditions rarely met and recommends alternatives. An intraclass correlation guideline shows how to measure rater agreement properly. Test-retest reliability was .84 overall, interrater agreement was .91 for timed stations but .62 for an instruction-following station and Hana uses the standard error to explain uncertainty.

PSYCH 655 Week 2 grading rubric: where the points go

Reliability papers earn credit for accurate concepts, correct calculations and interpretations that connect each number to a real decision about a person. Faculty look for true score and error to be explained, for the type of reliability to match the test's use, for coefficients to be interpreted with standards for the stakes involved and for the standard error of measurement to be used to build confidence bands around individual scores. Credit goes to identifying specific sources of error, such as rater judgment or practice effects, and recommending improvements. Reporting coefficient alpha as the only evidence for a performance test loses points. Lay out every calculation step and list sources in APA form.

PSYCH 655 Week 2 help: mistakes to avoid

In this unit, reliability papers often report a single coefficient without saying which kind it is or why it fits the test's use. Another frequent error is treating coefficient alpha as proof that a test measures one thing, or as the right statistic for rater-scored performance tasks. Some students ignore the standard error of measurement and interpret individual scores as exact. Others forget practice effects on retest. Identify the decisions the scores support, choose and calculate the matching reliability evidence, build confidence intervals around scores and pinpoint where error comes from. A tutor can help you compute a standard error of measurement and explain a confidence band in words a counselor would use.

Related PSYCH 655 sample papers

Other PSYCH 655 week samples

More MS in Psychology sample papers

PSYCH 655 Week 2 questions, answered

What does PSYCH 655 Week 2 usually cover?

Reliability of test scores, including classical test theory, types of reliability, coefficients and the standard error of measurement.

Where can I find a free PSYCH 655 Week 2 sample paper?

The full PSYCH 655 Week 2 reliability study of a vocational work-sample battery is above, free to read.

What is the standard error of measurement?

An estimate of how much an individual's observed score would vary around their true score, used to build confidence bands.

Is coefficient alpha always the right reliability measure?

No; alpha reflects internal consistency under strict assumptions, and other evidence, such as test-retest or interrater reliability, may fit better.

What is interrater reliability?

The degree to which different raters give consistent scores to the same performances.

Write yours, or have the desk draft it

This paper is an original model document written by our desk, not a submitted student paper and not an official University of Phoenix document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.