| Course | PSYCH 655 Psychometrics (PSYCH/655) |
|---|---|
| Week | 1 |
| Paper type | Test principles and norms paper |
| Length | about 1,172 words, 4 double-spaced pages plus title page and references |
| Format | APA 7 student paper |
| School | University of Phoenix |
| Program | MS in Psychology |
| Updated | October 2026 |
Free sample paper for PSYCH 655 Week 1
A Raw Score of 38 Compared With Whom? Standardization, Norms and Old Norm Tables in a Vocational Rehabilitation Office
[Student Name]
University of Phoenix
PSYCH/655: Psychometrics
Week 1 Assignment
[Instructor Name]
[Date]
The vocational rehabilitation office, its evaluator, clients and test scores are composites written for a model paper; testing standards and research findings come from the sources listed.
A test score on its own means little. Its meaning comes from how the test was given and from the group it is compared against. This paper explains standardization and norms through a case from the vocational rehabilitation office where I work, where the same client's performance looked very different depending on which norm table was used.
The Case
I am a vocational evaluator at a state vocational rehabilitation office in Albuquerque, New Mexico, helping people with disabilities prepare for and find work. Mr. Begay, fifty-one, worked for twenty-two years as a heavy equipment mechanic at a coal mine near Farmington until the mine closed and a back injury limited his lifting. He hopes to retrain as a maintenance technician at a hospital or school district. As part of his evaluation, I gave him a timed mechanical reasoning test with sixty items about gears, pulleys, levers and electrical circuits.
His raw score was 38. On the publisher's current general adult norms, published in 2019, that score corresponds to the 52nd percentile. A colleague, however, pointed to an older norm table in our office binder, labeled "skilled trades applicants," on which 38 falls at the 18th percentile. Mr. Begay's counselor wanted to know which was right, since one suggested strong aptitude and the other suggested he might struggle in training.
What a Test Is
A psychological test is a standardized procedure for sampling behavior and describing it with scores or categories. Standardization means that the test is administered, scored and interpreted in the same way for everyone, with the same items, instructions, time limits and scoring rules. Without standardization, differences in scores could reflect differences in procedure rather than in the people tested.
Under the joint testing standards (American Educational Research Association et al., 2014), anyone giving a test should follow standardized procedures unless an accommodation or modification is justified, should document any departures and should interpret scores in light of the evidence for the intended use.
Raw Scores and Derived Scores
A raw score, such as 38 correct out of 60, has no meaning by itself; it depends on the test's difficulty. Derived scores place the raw score in relation to a norm group. To get a z score, subtract the norm group's average from the raw score and divide by the group's standard deviation; the result says how many spreads above or below average the person sits. On the 2019 general norms, the mean raw score is 37.4 and the standard deviation 8.6, so Mr. Begay's z is about 0.07, essentially average. A T score rescales z to a mean of 50 and standard deviation of 10, giving about 51. A standard score with a mean of 100 and standard deviation of 15 gives about 101. A percentile rank tells what percentage of the norm group scored lower, here about 52.
On the older skilled trades table, the mean was 45.0 and the standard deviation 7.5. Mr. Begay's z there is about minus 0.93, a percentile near 18.
Which Norm Group Fits?
Norms are only as useful as their match to the person and the decision. The 2019 general adult norms were based on a national sample of about 2,400 adults stratified by age, sex, education, race and ethnicity and region, matching recent census figures. The older skilled trades table came from a 1990s sample of 310 applicants to apprenticeship programs at three manufacturing companies in the Midwest, mostly men in their twenties who had passed a screening interview. That group differs from Mr. Begay in age, region, era and selection: it consisted of people already screened for mechanical interest.
The question also matters. If the decision is whether Mr. Begay's mechanical reasoning is typical for adults, the general norms fit. If it were whether he would rank highly among applicants to a competitive apprenticeship, a current applicant norm would be relevant, but the old table is too small, too dated and too different to serve.
Mr. Begay did not change between the two tables; only the people he was compared with did.
Why Old Norms Mislead
Average performance on many ability tests has changed over time. Trahan et al. (2014) meta-analyzed studies comparing scores on newer and older norms of intelligence tests and found that scores rose by about 2.3 points per decade on average, consistent with the Flynn effect, though gains varied across tests and periods. Using outdated norms can inflate or deflate scores relative to current populations, which has serious consequences when scores determine eligibility for services or legal decisions.
Mechanical reasoning tests are not intelligence tests, but the same logic applies: norms collected thirty years ago describe a different population, with different schooling and different exposure to machines and technology.
How Good Are Psychological Tests?
Meyer et al. (2001) reviewed evidence on the validity of psychological tests and compared it with the validity of medical tests. Many psychological tests showed validity coefficients comparable to those of widely used medical tests, and they argued that multimethod assessment, combining tests with interviews and other sources, offers more complete information than any single method. They also emphasized that test results must be interpreted in context by trained professionals.
Departures From Standard Administration
Mr. Begay asked to stand during part of the test because sitting for long periods aggravates his back. Standing does not change the test's content or timing, so it is an accommodation that preserves standardization. Had he needed extra time, the score would be harder to interpret against norms collected under standard timing, and I would have documented the change.
Criterion-Referenced Interpretation
Norms answer how a person compares with others. Some decisions call for a different question: can the person do what a job requires? Criterion-referenced interpretation compares performance with a defined standard, such as a passing level set by experts for safe work. For Mr. Begay's retraining goal, a criterion-referenced work sample showing he could diagnose and repair a faulty pump assembly was arguably more relevant than any comparison with other adults. The two kinds of interpretation complement each other: norms describe his standing; criteria describe his readiness.
What I Reported
I reported Mr. Begay's mechanical reasoning as average for adults on current national norms, explained the norm group and noted that the older table was not appropriate. I combined the result with a hands-on work sample and his work history, which together suggested strong practical mechanical skills. His counselor approved enrollment in a building maintenance certificate program.
Office Practice
The episode led to a small but lasting change in how the office handles norms.
The office removed the old table from its binder and adopted a policy that every reported score name its norm group and publication year.
Conclusion
A score's meaning depends on standardized administration and an appropriate, current norm group. Mr. Begay's average performance on current norms and low performance on an old, mismatched table illustrate why test users must examine norms, not just numbers.
References
American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
Meyer, G. J., Finn, S. E., Eyde, L. D., Kay, G. G., Moreland, K. L., Dies, R. R., Eisman, E. J., Kubiszyn, T. W., & Reed, G. M. (2001). Psychological testing and psychological assessment: A review of evidence and issues. American Psychologist, 56(2), 128-165. https://doi.org/10.1037/0003-066X.56.2.128
Trahan, L. H., Stuebing, K. K., Fletcher, J. M., & Hiscock, M. (2014). The Flynn effect: A meta-analysis. Psychological Bulletin, 140(5), 1332-1360. https://doi.org/10.1037/a0037173
What the PSYCH 655 Week 1 instructions ask
Week 1 of PSYCH 655 usually introduces the foundations of psychological testing that every later week builds on. Prompts commonly ask students to define tests and assessment, describe how tests are standardized, explain norm-referenced and criterion-referenced interpretation, convert among raw scores, z scores, T scores, percentiles and standard scores and judge the adequacy of a norm sample. Some versions provide a score report to interpret. Show the steps of score conversion with numbers, ask whether the norm group matches the person tested in age, education, language and era, explain what changes when administration departs from the standard and cite the professional standards and research in APA style.
How this PSYCH 655 Week 1 example is built
Hana Kobayashi, the vocational evaluator writing this sample, gives a timed test to a fifty-one-year-old former mine mechanic, Mr. Begay, measuring his mechanical reasoning. His raw score of 38 converts to the 52nd percentile on the publisher's general adult norms from 2019 but the 18th percentile on an older table some colleagues still use for "skilled trades applicants." The joint testing standards explain that norms must fit the intended use and population. A review of testing evidence shows that well-normed tests yield validity comparable to many medical tests. A meta-analysis of rising scores shows why old norms can mislead. Hana explains standardization, score conversions and which norms to use.
PSYCH 655 Week 1 grading rubric: where the points go
In this opening week, testing foundations papers are graded on accurate explanation of standardization and norms, correct score conversions and sound judgment about which norms fit. Instructors look for raw scores to be distinguished from derived scores, for calculations to be shown and correct, for norm sample characteristics such as size, representativeness and date to be evaluated and for departures from standard administration to be considered. Credit goes to connecting principles to a real decision and to citing the professional standards. Treating a percentile as a percent correct, or ignoring the norm group entirely, loses points. Show the arithmetic, and cite sources in APA style.
PSYCH 655 Week 1 help: mistakes to avoid
A frequent error is confusing percentile ranks with percent correct, or reporting a score without naming the norm group it was compared against. Another is assuming that any published norm table fits any person, when age, education, language, culture and the year of norming can change a score's meaning. Some students state conversion formulas without working an example. Others overlook nonstandard administration, such as reading items aloud or extending time. Work through a conversion step by step, describe the norm group and how well it matches the person and explain what the score does and does not mean. A tutor can help you convert raw scores into several derived scores with a worked example and explain each one in a sentence.
Related PSYCH 655 sample papers
Other PSYCH 655 week samples
- PSYCH 655 Week 2: Reliability
- PSYCH 655 Week 3: Validity Evidence
- PSYCH 655 Week 4: Intelligence and Achievement Tests
- PSYCH 655 Week 5: Personality and Career Assessment
- PSYCH 655 Week 6: Test Fairness and Ethics
More MS in Psychology sample papers
- PSYCH 645 Week 1: Psychodynamic and Attachment
- PSYCH 647 Week 1: Defining Job Performance
- PSYCH 650 Week 1: Classification and the DSM
- PSYCH 658 Week 1: Need and Content Theories
PSYCH 655 Week 1 questions, answered
What does PSYCH 655 Week 1 usually cover?
Foundations of psychological testing, including standardization, norms, score types and norm-referenced interpretation.
Where can I find a free PSYCH 655 Week 1 sample paper?
The complete PSYCH 655 Week 1 paper on one mechanic's raw score of 38 judged against two norm groups is above, free.
What is a norm group?
The sample of people whose scores define what is typical on a test, against which an individual's score is compared.
Is a percentile the same as percent correct?
No; a percentile tells what percentage of the norm group scored below a person, not how many items were answered correctly.
Why do old test norms matter?
Average performance on many tests has shifted over decades, so outdated norms can make scores look higher or lower than they are.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official University of Phoenix document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.
Request this one custom, free · All PSYCH 655 week samples · All courses