| Course | PSYCH 647 Human Performance, Assessment, and Feedback (PSYCH/647) |
|---|---|
| Week | 3 |
| Paper type | Appraisal methods and rater error analysis |
| Length | about 1,177 words, 4 double-spaced pages plus title page and references |
| Format | APA 7 student paper |
| School | University of Phoenix |
| Program | MS in Psychology |
| Updated | October 2026 |
Free sample paper for PSYCH 647 Week 3
Why Every Technician Got a Four: Rating Formats, Rater Errors and Rater Training for a Water Utility's Appraisals
[Student Name]
University of Phoenix
PSYCH/647: Human Performance, Assessment, and Feedback
Week 3 Assignment
[Instructor Name]
[Date]
The water utility, its supervisors and rating data are composites written for a model paper; research findings come from the sources listed.
Performance appraisal depends on human judgment, and human judgment has predictable limits. Some rating problems come from how people process information; others come from what raters are trying to achieve. This paper analyzes rating patterns at a water utility and recommends formats and training to improve them.
The Rating Pattern
At the Mesa del Sol Water Authority in Albuquerque, supervisors rated about 140 field technicians each year on an eight-item graphic scale from one to five. Across five years, eighty-one percent of all item ratings were four, twelve percent were five, six percent were three and fewer than one percent were one or two. Within individual technicians, ratings rarely varied across items: a technician rated four on attitude was almost always rated four on quality of work, dependability and initiative. Technicians whom coworkers described as the best and the weakest on their crews often received identical ratings.
Rater Errors in the Data
The data show several familiar patterns. Leniency is the tendency to rate everyone higher than their performance warrants; here, almost no one received below four. Range restriction appears in the narrow spread, which makes ratings useless for distinguishing performance. Halo, the tendency to let an overall impression shape ratings on distinct dimensions, appears in the near-identical ratings across items. These patterns mean that the ratings carried little information about who was performing well or poorly on what.
Why Raters Do It
Interviews with six supervisors revealed reasons beyond cognitive error. Ratings below three required a written improvement plan and often led to union grievances. Merit raises were tied to ratings, and supervisors did not want to cost good workers money because of a few weak areas. Several said that everyone works hard in a difficult job and that low ratings hurt morale. One said simply, "If I give a two, I'm in grievance hearings for a month. Fours keep the peace."
Murphy (2008) explained the weak relationship between job performance and ratings of job performance by distinguishing raters' ability to evaluate accurately from their motivation to report evaluations accurately. Many rating problems, Murphy argued, stem less from inaccurate perception than from raters' goals: avoiding conflict, maintaining relationships, motivating employees or protecting themselves. Interventions that improve raters' skills without addressing these goals will have limited effects.
How Much of a Rating Is the Rater?
Scullen et al. (2000) analyzed multisource ratings of managers and partitioned variance into components. Each rater's personal leanings explained a bigger slice of the scores than the manager being scored did. In other words, a rating said more about who did the rating than about who was rated. The finding supports using multiple raters and focusing training on shared standards.
At the utility, each technician was rated by one supervisor, so each rating carried that supervisor's idiosyncrasies in full.
A four told you less about a technician than it told you about which supervisor had signed the form.
Other Errors to Watch
Two further errors matter for the new system. Recency gives the last few weeks before the appraisal too much weight; a technician who handled a dramatic main break in October may outshine one who worked steadily all year. A rater may also warm to technicians who remind them of themselves, the same hometown, the same fishing weekends, the same blunt way of talking. Both can be reduced by keeping brief notes of critical incidents throughout the year and by rating against behavioral anchors rather than impressions.
Comparing Rating Formats
Graphic rating scales, like the utility's old form, are simple but vague. Anchored scales print a short description of real behavior beside several points on each dimension. They cut guesswork and tell employees what a higher score looks like, but they take weeks of incident collection to build, and raters sometimes hunt for an anchor that matches exactly instead of judging the closest fit. Behavioral observation scales ask how often employees perform specific behaviors. Forced distribution requires raters to place set percentages of employees in each category, eliminating leniency but creating competition and resentment, especially in small teams where true performance may not follow a curve. Management by objectives focuses on goal achievement, which suits some roles but can reward outcomes outside a person's control. Narrative methods offer rich detail but are hard to compare.
For technicians in small crews whose work depends on conditions, forced distribution would be unfair and harmful to teamwork. Behaviorally anchored scales built from the critical incidents collected earlier fit the job best.
Rater Training
Woehr and Huffcutt (1994) meta-analyzed rater training studies and compared four approaches. Rater error training, which teaches raters to avoid errors like halo and leniency, reduced those errors but sometimes reduced accuracy too. Performance dimension training, which familiarizes raters with dimensions, and behavioral observation training, which improves observation and recording, had moderate effects. Frame-of-reference training, which builds a shared understanding of performance standards through definitions, examples of different performance levels, practice ratings and feedback, produced the largest improvements in rating accuracy.
The finding suggests that telling supervisors to avoid giving everyone a four is less useful than teaching them what a two, a three and a five look like for each dimension.
Recommendations
First, adopt the behaviorally anchored scales developed from technicians' critical incidents for each of the six dimensions. Second, provide frame-of-reference training for all supervisors and lead technicians: a half-day session with written examples and short videos of technicians performing at different levels, practice ratings and discussion until raters align. Third, address rater motivation by separating developmental feedback conversations from pay decisions, holding them at different times of year, replacing the requirement for a full improvement plan with a brief documented conversation for ratings of two on a single dimension and working with the union to agree that ratings on individual dimensions are coaching information rather than grounds for discipline. Fourth, use multiple raters, combining supervisor and lead technician ratings as planned last week. Fifth, give supervisors feedback on their own rating distributions compared with peers.
Calibration Meetings
Supervisors will also meet once a year in calibration sessions, reviewing anonymized examples and their ratings together to align standards across districts. Research suggests such discussions help reduce idiosyncratic rater effects, though they can also introduce group pressure, so a facilitator will keep the focus on behavioral evidence.
Measuring Improvement
Success will be judged on several indicators rather than one.
After the first cycle, we will examine the spread of ratings, the correlation between dimensions within technicians, agreement between supervisors and lead technicians and technicians' perceptions of fairness.
Limits
Better ratings may produce more low ratings and initial resistance. The union agreement is essential, and some leniency may persist where supervisors value harmony above accuracy.
Conclusion
Ratings at the utility clustered at four because of rater errors and, more importantly, because the system rewarded leniency. Research on idiosyncratic rater effects, rater motivation and training points to anchored scales, frame-of-reference training, multiple raters and changes to the incentives that made fours the safest answer.
References
Murphy, K. R. (2008). Explaining the weak relationship between job performance and ratings of job performance. Industrial and Organizational Psychology, 1(2), 148-160. https://doi.org/10.1111/j.1754-9434.2008.00030.x
Scullen, S. E., Mount, M. K., & Goff, M. (2000). Understanding the latent structure of job performance ratings. Journal of Applied Psychology, 85(6), 956-970. https://doi.org/10.1037/0021-9010.85.6.956
Woehr, D. J., & Huffcutt, A. I. (1994). Rater training for performance appraisal: A quantitative review. Journal of Occupational and Organizational Psychology, 67(3), 189-205. https://doi.org/10.1111/j.2044-8325.1994.tb00562.x
What the PSYCH 647 Week 3 instructions ask
Week 3 of PSYCH 647 usually focuses on how performance is appraised and why ratings go wrong. Prompts may include graphic rating scales, behaviorally anchored and behavioral observation scales, forced distribution and ranking, management by objectives, narrative methods, rater errors such as leniency, severity, central tendency, halo and recency, rater motivation and politics and rater training approaches including frame-of-reference training. Some versions supply rating data to diagnose. Explain each format's strengths and weaknesses for the job, show which errors appear in the data, distinguish errors of judgment from deliberate distortion and recommend formats and training supported by research. Support each recommendation with journal findings in APA form.
How this PSYCH 647 Week 3 example is built
Five years of rating records are the starting point; the analyst, Grace Okonkwo, reviews five years of ratings at the Mesa del Sol Water Authority, where eighty-one percent of technicians received four out of five on every item. She identifies leniency, halo and range restriction, and interviews show that supervisors inflated ratings to avoid grievances and protect raises. A study of rating variance found that idiosyncratic rater effects accounted for more variance than the person being rated. A meta-analysis of rater training found frame-of-reference training most effective for accuracy. An analysis of weak rating-performance links points to rater motivation. Grace pairs behaviorally anchored scales with frame-of-reference training and changes that reduce incentives to inflate.
PSYCH 647 Week 3 grading rubric: where the points go
Appraisal method papers in this week earn credit for accurate description of rating formats and rater errors, sound diagnosis of real rating data and recommendations supported by research. Faculty look for rating formats to be compared on what they require of raters and what they offer employees, for rater errors to be defined and identified in evidence, for the role of rater motivation and organizational context to be recognized and for training approaches to be evaluated by their effects. Credit goes to understanding that errors are not only cognitive slips but also strategic choices. A plan that addresses both ability and motivation scores highest, with references in APA style.
PSYCH 647 Week 3 help: mistakes to avoid
Students often define halo, leniency and central tendency without showing them in actual ratings, which leaves the analysis abstract. Another common error is assuming rater training alone will fix ratings, when raters inflate deliberately to avoid conflict or protect employees' pay. Some papers recommend forced distribution without considering its effects on teamwork and morale in small units. Others treat behaviorally anchored scales as a cure-all, though anchors cannot stop a rater who has reasons to inflate. Show the errors in data, explain why raters make them, match formats to the job and pair training with changes to rater incentives. A tutor can help you calculate simple statistics, such as the share of ratings at each scale point, to show rating patterns.
Related PSYCH 647 sample papers
Other PSYCH 647 week samples
- PSYCH 647 Week 1: Defining Job Performance
- PSYCH 647 Week 2: Criteria and Measurement
- PSYCH 647 Week 4: Giving and Receiving Feedback
- PSYCH 647 Week 5: Performance Management Systems
- PSYCH 647 Week 6: Evaluating an Appraisal Program
More MS in Psychology sample papers
- PSYCH 639 Week 3: Ethics in Testing and Selection
- PSYCH 642 Week 3: Interviews and Assessment Centers
- PSYCH 644 Week 3: Memory Processes and Models
- PSYCH 645 Week 3: Cognitive-Behavioral Approaches
PSYCH 647 Week 3 questions, answered
What does PSYCH 647 Week 3 usually cover?
Performance appraisal methods, common rater errors and rater training approaches.
Where can I find a free PSYCH 647 Week 3 sample paper?
The full PSYCH 647 Week 3 analysis of why every water utility technician got a four is available above, free.
What is the halo effect in performance appraisal?
When a rater's overall impression of an employee colors ratings on separate dimensions.
What is frame-of-reference training?
Rater training that builds a shared understanding of performance dimensions and standards using examples and practice with feedback.
Why do supervisors inflate ratings?
Often to avoid conflict, protect employees' pay or promotions, maintain morale or avoid extra documentation required for low ratings.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official University of Phoenix document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.
Request this one custom, free · All PSYCH 647 week samples · All courses