PSYCH 642 Week 2 Selection Tests and Their Validity Example

Reviewed by Queenie Halstead, MA · University of Phoenix · Updated

This PSYCH 642 Week 2 example examines selection tests and their validity by checking whether two tests an aircraft maintenance company already uses, a mechanical comprehension test and a hands-on troubleshooting exercise, actually predict how new mechanics perform, and what the answer is worth in dollars. Week 2 of University of Phoenix PSYCH 642 turns to tests and their validity, and for PSYCH/642 MS in Psychology students explain the kinds of validity evidence, design or interpret a criterion-related study, handle problems such as range restriction and estimate utility. The study is run by a composite talent acquisition manager in Mesa with eighteen months of hiring and audit data. It leans on a classic review of predictor validity and utility, a critical review of how criterion-related validation should be done and a review of ability and personality as predictors.

CoursePSYCH 642 Personnel Psychology (PSYCH/642)
Week2
Paper typeTest validation paper
Lengthabout 1,194 words, 4 double-spaced pages plus title page and references
FormatAPA 7 student paper
SchoolUniversity of Phoenix
ProgramMS in Psychology
UpdatedOctober 2026

Free sample paper for PSYCH 642 Week 2

1

Do Our Mechanic Tests Predict Who Passes Inspection? A Local Validation Study and Utility Estimate at an Aircraft Maintenance Shop

[Student Name]

University of Phoenix

PSYCH/642: Personnel Psychology

Week 2 Assignment

[Instructor Name]

[Date]

The maintenance company, its tests and the validation figures are composites written for a model paper; research findings come from the sources listed.

What this part is doingThe title asks the question a validation study is designed to answer: do scores predict the outcome that matters?
2

An employer may use a test for years without knowing whether it predicts anything. Personnel psychology offers methods for finding out and for estimating what a valid test is worth. This paper applies those methods to the two tests my company uses to hire aircraft mechanics.

The Tests

Copperline Aero Services requires all mechanic applicants to hold FAA airframe and powerplant certification. Beyond that, applicants take two tests. The first is a commercial mechanical comprehension test, thirty-five multiple-choice items about gears, levers, pressure and electrical circuits. The second is a forty-five-minute hands-on troubleshooting exercise developed in house: applicants diagnose a planted fault in a hydraulic test rig and a wiring harness while a lead mechanic scores their approach and outcome on a checklist. Applicants must pass both, but no one has examined whether either predicts performance.

Kinds of Validity Evidence

Validity refers to how well evidence and theory support the use of scores for a purpose. Content evidence is about overlap with the job itself; here the troubleshooting exercise scores well, because tracing a fault on a hydraulic rig is what mechanics do on the hangar floor. Criterion-related evidence asks whether scores relate to measures of job performance. Construct evidence asks whether a test measures the attribute it claims to measure. For an employer, criterion-related evidence answers the most practical question: do higher scorers become better mechanics?

What a Century of Research Suggests

Hunter and Hunter (1984) reviewed the validity of many predictors of job performance and concluded that general cognitive ability tests predicted performance across jobs, with higher validity for more complex jobs, and that work samples and some other predictors also showed strong validity. They also argued that substituting less valid predictors for valid ones carries large economic costs, introducing utility estimates into a wide audience. Later reanalyses have revised some of their figures downward, but the core idea that valid selection pays off has held.

What this part is doingStarting from the research base sets expectations before the local numbers appear.
3

The Local Study

I conducted a predictive validation study using the eighty-four mechanics hired over the past eighteen months who had completed at least nine months of work. Test scores were recorded at hiring but did not vary much in their influence on decisions beyond the pass mark. The study used two criteria. First, quality findings: the number of significant discrepancies found by quality inspectors in each mechanic's completed tasks, adjusted for hours worked. Second, supervisor ratings on a structured form covering technical accuracy, documentation, safety and teamwork, completed by each mechanic's lead after training on the form.

Scores on the troubleshooting exercise correlated .38 with fewer quality findings and .33 with supervisor ratings. Scores on the mechanical comprehension test correlated .21 with fewer findings and .18 with ratings. With eighty-four people, the troubleshooting correlations are statistically reliable, while the comprehension test's correlations are smaller and less certain, with confidence intervals that include values near zero.

The forty-five minutes at the hydraulic rig told us more about future inspection results than thirty-five multiple-choice items did.

Problems in Local Validation

A critical review of validation practice (Van Iddekinge & Ployhart, 2008) offered recommendations that fit our situation closely. The authors emphasized the importance of criterion quality, since contaminated or deficient criteria distort validity estimates; the effects of range restriction, which occurs when only applicants who passed a test are available for study and which typically lowers observed correlations; the need for adequate sample sizes; and the value of examining validity across subgroups. They encouraged practitioners to report both observed and corrected correlations and to explain their corrections.

Our study faces each issue. Range restriction is present because only applicants who passed both tests were hired. Using the applicant pool's score spread, a standard correction raises the troubleshooting exercise's correlation with quality findings from .38 to about .45 and the comprehension test's from .21 to about .27. Supervisor ratings may be contaminated by leads' impressions of mechanics they like, which is why quality findings, recorded by independent inspectors, are the stronger criterion.

Fairness Across Groups

Among the eighty-four hires, numbers of women and of some ethnic groups were too small for separate validity estimates. Pass rates for applicants over two years showed that 61 percent of women and 70 percent of men passed the troubleshooting exercise, a ratio above the four-fifths guideline, while the comprehension test showed a larger gap. These results support keeping the exercise and reviewing the comprehension test.

Adding a Measure of Conscientiousness

Schmitt (2014) reviewed research on cognitive ability and personality as predictors of performance at work. Ability predicted task performance well across jobs, while personality traits, especially conscientiousness, added prediction for outcomes such as rule following, safety and counterproductive behavior. Combining measures often improved overall prediction because they captured different aspects of performance. Much of Schmitt's review stressed that measures should be matched to the outcomes an organization cares about.

For aircraft mechanics, careful documentation and strict adherence to procedures are central to safety. A brief conscientiousness measure could add prediction for these behaviors, which our troubleshooting exercise does not directly capture. I recommended a pilot in which applicants complete such a measure without it affecting decisions, followed by a validation study.

What this part is doingLinking the next predictor to the job's safety demands keeps the recommendation grounded.
4

What the Exercise Is Worth

Utility analysis estimates the dollar value of a selection procedure. Using a common approach, value depends on the number hired, the validity of the procedure, the spread of performance in dollar terms and the average score of those selected compared with all applicants, minus testing costs. Using conservative assumptions, thirty hires a year, a corrected validity of .40, a performance spread estimated at forty percent of the average mechanic's $74,000 salary and a modest selection ratio, the troubleshooting exercise yields an estimated benefit of several hundred thousand dollars a year in avoided rework and inspection delays, far exceeding its cost of about $600 per applicant. These figures are estimates, but even if they are halved, the exercise pays for itself.

Concurrent Versus Predictive Designs

An alternative design would have tested current mechanics and compared their scores with their present performance, a concurrent study. It is faster, since data can be collected in weeks, but current employees differ from applicants: they have learned on the job, they may be less motivated to do well on a test that cannot affect them and those who performed poorly may already have left. Our predictive design avoids some of these problems because scores were collected at hiring, though it took eighteen months to accumulate enough performance data.

Recommendations

Keep the troubleshooting exercise and expand its rating checklist so different leads score consistently. Retain the comprehension test only as a minimum screen for now, pending review of its adverse impact and modest validity. Pilot a conscientiousness measure. Repeat the validation study every two years as more mechanics are hired.

Conclusion

A local validation study showed that the hands-on troubleshooting exercise predicts the quality of mechanics' work, while the comprehension test adds less. Research on predictor validity, validation practice and the combination of ability and personality supports keeping the exercise, reviewing the written test and exploring a conscientiousness measure for safety-related behavior.

5

References

Hunter, J. E., & Hunter, R. F. (1984). Validity and utility of alternative predictors of job performance. Psychological Bulletin, 96(1), 72-98. https://doi.org/10.1037/0033-2909.96.1.72

Schmitt, N. (2014). Personality and cognitive ability as predictors of effective performance at work. Annual Review of Organizational Psychology and Organizational Behavior, 1, 45-65. https://doi.org/10.1146/annurev-orgpsych-031413-091255

Van Iddekinge, C. H., & Ployhart, R. E. (2008). Developments in the criterion-related validation of selection procedures: A critical review and recommendations for practice. Personnel Psychology, 61(4), 871-925. https://doi.org/10.1111/j.1744-6570.2008.00133.x

What the PSYCH 642 Week 2 instructions ask

The second week of PSYCH 642 typically examines employment tests and the evidence an employer needs before relying on them. Prompts may cover content, criterion-related and construct validity evidence, concurrent and predictive designs, reliability, range restriction and criterion problems, cognitive ability, personality, job knowledge and work sample tests, adverse impact and utility analysis. Many assignments ask students to evaluate the tests an employer uses or to plan a validation study. Choose criteria that reflect real performance, explain the design and its weaknesses, report correlations with their sample sizes and confidence, correct where appropriate, consider fairness and translate results into practical value. List journal research in APA style.

How this PSYCH 642 Week 2 example is built

In this sample, Lena Park sets out to learn whether Copperline Aero Services' two hiring steps predict new mechanics' performance. Eighty-four mechanics hired over eighteen months took a mechanical comprehension test and a hands-on troubleshooting exercise. She compares scores with first-year inspection findings and supervisor ratings on a structured form. The troubleshooting exercise correlates .38 with fewer inspection findings; the comprehension test correlates .21. A classic review of predictor validity frames these results, and a critical review of validation practice guides her handling of range restriction and criterion quality. A review of ability and personality suggests adding a conscientiousness measure. A utility estimate shows the exercise is worth its cost.

PSYCH 642 Week 2 grading rubric: where the points go

Validation papers earn the most when the design is sound, the criteria reflect performance that matters to the employer and the results are interpreted with appropriate caution. Graders look for validity to be treated as evidence for an inference, for criterion-related designs and their limits to be explained, for correlations to be reported with sample sizes and appropriate corrections, for fairness to be checked across groups and for utility to be estimated with stated assumptions. Credit goes to recognizing criterion contamination and deficiency and to practical recommendations. Overstating small-sample results or ignoring range restriction loses points. Tables described clearly in text and APA references finish a strong paper.

PSYCH 642 Week 2 help: mistakes to avoid

In this unit, validation papers often report a correlation without saying how many people it rests on or what the criterion measured, which leaves readers unable to judge it. Another common slip is ignoring range restriction, since only hired applicants have performance data, which tends to shrink observed correlations. Some students treat supervisor ratings as a perfect criterion, overlooking bias and leniency. Others compute utility with unrealistic assumptions. Describe the sample, the criterion and the design in plain terms, report and correct correlations carefully, check subgroup differences and present utility with conservative inputs. A tutor can help you work through a range restriction correction with your own numbers and explain the result in a sentence a manager would follow.

Related PSYCH 642 sample papers

Other PSYCH 642 week samples

More MS in Psychology sample papers

PSYCH 642 Week 2 questions, answered

What does PSYCH 642 Week 2 usually cover?

Employment tests, types of validity evidence, criterion-related validation studies, reliability, adverse impact and utility analysis.

Where can I find a free PSYCH 642 Week 2 sample paper?

Above is the complete PSYCH 642 Week 2 local validation study of an aircraft maintenance shop's mechanic tests, free to read.

What is criterion-related validity?

Evidence that test scores relate to an important outcome, such as job performance, measured at the same time or later.

What is range restriction?

A narrowing of score ranges, as when only high scorers are hired, which tends to reduce observed correlations.

What is utility analysis?

A method for estimating the economic value of a selection procedure based on its validity, the spread of performance and the number hired.

Write yours, or have the desk draft it

This paper is an original model document written by our desk, not a submitted student paper and not an official University of Phoenix document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.