PSYCH 647 Week 2 Performance Criteria and Measurement Example

Reviewed by Queenie Halstead, MA · University of Phoenix · Updated

This PSYCH 647 Week 2 example works through the criterion problem, deciding how to measure each performance dimension so the numbers are relevant, reliable and fair, by weighing objective records against supervisor ratings for water utility field technicians. Week 2 of University of Phoenix PSYCH 647 examines criteria and measurement, and the PSYCH/647 paper has MS in Psychology students explain criterion relevance, deficiency and contamination, compare objective and judgmental measures and plan for reliability. The plan is written by a composite HR analyst at an Albuquerque water authority who has the six dimensions from last week and a pile of work-order data. She relies on a history of the criterion problem, a meta-analysis on whether objective and subjective measures can substitute for each other and a meta-analysis of how consistently raters judge job performance.

CoursePSYCH 647 Human Performance, Assessment, and Feedback (PSYCH/647)
Week2
Paper typeCriterion measurement paper
Lengthabout 1,160 words, 4 double-spaced pages plus title page and references
FormatAPA 7 student paper
SchoolUniversity of Phoenix
ProgramMS in Psychology
UpdatedOctober 2026

Free sample paper for PSYCH 647 Week 2

1

Counts, Ratings or Both? Choosing Performance Criteria and Measures for Water Utility Field Technicians

[Student Name]

University of Phoenix

PSYCH/647: Human Performance, Assessment, and Feedback

Week 2 Assignment

[Instructor Name]

[Date]

The water utility, its records and the measurement plan are composites written for a model paper; research findings come from the sources listed.

What this part is doingThe title states the choice the plan must make for every dimension.
2

Defining performance is the first step; measuring it is the second, and it is harder. Every measure captures only part of what matters and some of what does not. This paper plans how to measure the six performance dimensions defined last week for field technicians at a water utility.

The Dimensions

The Mesa del Sol Water Authority's new appraisal system for 140 field technicians covers technical proficiency, safety, customer communication, adaptivity, proactivity and teamwork and citizenship, with counterproductive incidents recorded separately. The question now is how to measure each.

The Criterion Problem

Austin and Villanova (1992) traced the history of the criterion problem in industrial psychology from 1917 to 1992. They described the gap between the ultimate criterion, the full conceptual meaning of success in a job, and actual criteria, the measures organizations use, which are always imperfect. Actual criteria can be deficient, missing parts of the ultimate criterion, and contaminated, including variance unrelated to it. The authors argued that criterion development had received far less attention than predictor development and called for careful, construct-oriented work on what criteria mean.

For our technicians, the ultimate criterion is something like "keeps safe water flowing to customers while protecting themselves and the public and supporting the utility's mission." No single number captures it.

What this part is doingIntroducing the ultimate criterion gives a standard against which each actual measure can be judged.
3

Objective Measures Available

The utility's work-order system records time to restore service after main breaks, repeat visits to the same address within thirty days, documentation errors flagged by the records team and meter reading accuracy. The safety office records incidents and near misses. These measures are cheap and appear objective, but each carries contamination. Repair time depends on pipe depth, soil, traffic and the size of the break. Repeat visits depend partly on aging infrastructure. Safety incidents are rare and partly random.

Can Objective and Subjective Measures Substitute?

Bommer et al. (1995) meta-analyzed studies that measured employee performance both objectively, such as sales or production, and subjectively, through ratings. The two types correlated positively but only moderately, around .39 after corrections, and the authors concluded that they should not be used interchangeably. The correlation was higher when objective and subjective measures tapped the same content, such as ratings of quantity compared with production counts. Each type captures aspects of performance the other misses.

For the utility, this means repair times cannot replace supervisors' judgment of technical skill, and ratings cannot replace documentation error counts. Both are needed, matched to the dimensions they best capture.

How Reliable Are Ratings?

Viswesvaran et al. (1996) meta-analyzed the reliability of job performance ratings. The average interrater reliability of supervisory ratings of overall performance was about .52, meaning two supervisors rating the same employee agreed only moderately. Peer ratings were less reliable still, around .42. Reliability within a single rater over time was higher than agreement between raters, suggesting that each rater brings an idiosyncratic perspective.

These findings caution against relying on one supervisor. Our technicians work in crews of three or four, often out of their supervisor's sight, which makes a single supervisor's view even narrower.

Two supervisors watching the same technician agree only about half as well as most people assume.

Who Sees What

Different observers see different slices of a technician's work. Supervisors manage several crews and visit job sites for perhaps an hour a week per crew, so they see results and occasional moments of work. Lead technicians work beside crew members every day and see safety habits, problem solving and how people treat one another, but they are peers and may be reluctant to rate harshly. Customers see only the few minutes when a technician explains why the water is off. The records team sees paperwork, not fieldwork. Matching each dimension to the observers best placed to see it is the practical core of reducing deficiency: no single observer sees everything, but together they see most of what matters.

Rating Formats

Behaviorally anchored rating scales give raters concrete descriptions at several points on each dimension, such as, for safety, "skips trench shoring on short jobs" at the low end and "checks shoring and atmosphere before every entry and corrects others' lapses" at the high end. Anchors drawn from the critical incidents collected last week make ratings more consistent and give technicians clearer feedback than numbers alone.

The Measurement Plan

Technical proficiency: supervisor and lead technician ratings on behaviorally anchored scales, combined with rework rates adjusted for district pipe age and documentation error counts. Safety: ratings by supervisors and the safety officer, informed by monthly unannounced job-site observations using a checklist, with incidents recorded as context rather than scored directly. Customer communication: a short text survey sent to customers after service visits, asking about clarity and courtesy, combined with supervisor ratings. Adaptivity and proactivity: ratings by supervisors and lead technicians, with documented examples such as reported hazards or suggestions adopted. Teamwork and citizenship: ratings by supervisors and peers on the crew, using anchored scales.

What this part is doingAssigning each measure to the dimension it fits best follows from the finding that measures are not interchangeable.
4

Adjusting for Context

To reduce contamination from district conditions, objective measures will be compared within districts rather than across the whole utility, and rework rates will be adjusted for pipe age using the utility's asset database. Technicians assigned to older districts will not be penalized for the condition of the pipes.

Weighting the Sources

Combining sources raises a further question: how much weight each should carry. For technical proficiency, the lead technician's daily view and the adjusted rework rate will count most, with the supervisor's rating as a check. For customer communication, customer surveys will count most once enough responses accumulate, since customers are the people the dimension serves. Weights will be published to technicians in advance and reviewed after the first year, when the reliability data show which sources agree and which stand apart.

Checking Reliability

In the first year, two supervisors will independently rate a sample of thirty technicians they both observe, allowing us to estimate interrater agreement. Where agreement is low on a dimension, we will revise anchors and provide additional rater training.

What the Plan Still Misses

Even with multiple sources, the plan may be deficient. It captures little of how technicians handle rare emergencies, such as a major break during a monsoon storm, which may be the most important moments of the job. Supervisors will be asked to document critical incidents throughout the year, both positive and negative, so these moments inform ratings.

Acceptance and Cost

The plan requires more work than the old form: unannounced observations, peer ratings and customer surveys. Technicians and the union reviewed it and supported multiple sources because they reduce dependence on a single supervisor. Customer surveys cost about $3,000 a year through the utility's existing text platform.

Conclusion

No measure captures performance fully. Research on the criterion problem, the limited interchangeability of objective and subjective measures and the modest reliability of single raters points to a plan that matches measures to dimensions, uses multiple sources, adjusts for context and checks reliability.

5

References

Austin, J. T., & Villanova, P. (1992). The criterion problem: 1917-1992. Journal of Applied Psychology, 77(6), 836-874. https://doi.org/10.1037/0021-9010.77.6.836

Bommer, W. H., Johnson, J. L., Rich, G. A., Podsakoff, P. M., & MacKenzie, S. B. (1995). On the interchangeability of objective and subjective measures of employee performance: A meta-analysis. Personnel Psychology, 48(3), 587-605. https://doi.org/10.1111/j.1744-6570.1995.tb01772.x

Viswesvaran, C., Ones, D. S., & Schmidt, F. L. (1996). Comparative analysis of the reliability of job performance ratings. Journal of Applied Psychology, 81(5), 557-574. https://doi.org/10.1037/0021-9010.81.5.557

What the PSYCH 647 Week 2 instructions ask

The second week of PSYCH 647 typically asks students to select and evaluate measures of job performance for the dimensions defined in the first week. Prompts commonly include the criterion problem, ultimate and actual criteria, relevance, deficiency and contamination, objective production and personnel data, judgmental ratings, multiple raters and reliability, along with practical issues such as cost and acceptance. Some versions provide a job and ask for a measurement plan. Take each dimension in turn, propose at least one measure, explain what it captures and misses, consider who is best placed to observe the behavior and describe how reliability will be checked. Cite journal studies in APA style.

How this PSYCH 647 Week 2 example is built

With six dimensions defined, the HR analyst Grace Okonkwo now works out how to measure the six technician dimensions defined in Week 1. Work-order systems record repair times, rework and documentation errors; supervisors observe safety and teamwork; customers can rate communication. A history of the criterion problem warns that every measure is a partial window. A meta-analysis shows that objective and subjective measures correlate only moderately and cannot simply substitute. A meta-analysis of rater consistency shows that a single supervisor's ratings are only moderately reliable. Grace builds a plan combining records, structured ratings from supervisors and lead technicians, customer surveys and checks for contamination.

PSYCH 647 Week 2 grading rubric: where the points go

Criterion papers in this course earn credit for applying measurement concepts to each dimension in turn and for a plan that balances relevance, reliability and practicality. Instructors look for criterion relevance, deficiency and contamination to be explained with examples, for objective and judgmental measures to be compared honestly, for rater reliability to be addressed through multiple sources or training and for fairness to be considered. Credit goes to matching each measure to the dimension it best captures and to explaining what the overall system will miss. A single measure for all dimensions earns little, however convenient it looks. Clear tables described in text and references formatted in APA style complete the plan.

PSYCH 647 Week 2 help: mistakes to avoid

On this topic, students often assume that objective numbers are automatically better than ratings, overlooking contamination from factors outside a person's control, such as equipment, territory or luck. Others assume that one supervisor's rating captures performance, ignoring evidence that single raters agree only moderately with other raters. Some papers list measures without saying which dimension each addresses. Others ignore cost and acceptance, proposing measurement systems no organization would sustain past the first year. Go dimension by dimension, name what each measure includes and excludes, use multiple sources where possible and plan a reliability check. A tutor can help you build a measurement matrix linking each dimension to its sources and to the contamination each source carries.

Related PSYCH 647 sample papers

Other PSYCH 647 week samples

More MS in Psychology sample papers

PSYCH 647 Week 2 questions, answered

What does PSYCH 647 Week 2 usually cover?

Performance criteria and measurement, including the criterion problem, objective versus judgmental measures and reliability.

Where can I find a free PSYCH 647 Week 2 sample paper?

The full PSYCH 647 Week 2 measurement plan for water utility field technicians is posted above, free.

What is criterion contamination?

When a performance measure includes factors unrelated to the person's performance, such as equipment age or district conditions.

Are objective performance measures better than ratings?

Not necessarily; they correlate only moderately with ratings, and each captures different aspects of performance.

How reliable are supervisor ratings?

Research suggests ratings by a single supervisor are only moderately reliable, so multiple raters or structured methods help.

Write yours, or have the desk draft it

This paper is an original model document written by our desk, not a submitted student paper and not an official University of Phoenix document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.