NSG/533 Week 4: Reliability, Validity and Item Analysis, sample paper

Reviewed by Lenora Whitcombe, MSN, RN · University of Phoenix

This page holds a complete NSG/533 Week 4 sample paper on reliability, validity and item analysis, in true APA form. It reports the statistics from the first administration of a revised 50-item unit exam in a composite maternal-newborn course, interprets the reliability coefficient, the difficulty and discrimination of individual items and distractor performance, gathers validity evidence and decides item by item what to keep, revise or remove.

1

What the Numbers Said About the New Exam: Reliability, Validity Evidence and an Item Analysis of a Revised Maternal-Newborn Unit Test

[Student Name]

University of Phoenix

NSG/533: Educational Assessment and Evaluation

Week 4 Assignment

[Instructor Name]

[Date]

The school, the course, the exam and its statistics are a composite written for a model paper.

What this part is doingThe title promises numbers and what they said. The reader expects interpretation and decisions, not only statistics.
2

In Week 3, I rewrote recall items from the maternal-newborn course exams as scenario-based items requiring application and analysis. The first revised unit exam, 50 items with 34 at the apply level or higher, has now been given to 56 students. This paper interprets its statistics and gathers evidence on whether its scores can be trusted to mean what we intend.

Overall Results

The mean score was 76.4%, lower than the previous exam's 84%, with a standard deviation of 8.2 percentage points and a range of 54% to 94%. The lower mean was expected, since the items now require more than recall, and the wider spread suggests the exam separates students more than before.

Reliability

The Kuder-Richardson Formula 20 coefficient, a measure of internal consistency for items scored right or wrong, was 0.78. For a classroom test used for grading, values of about 0.70 or higher are generally considered acceptable, and values in the 0.80s are desirable for decisions with higher stakes (Oermann & Gaberson, 2021). The coefficient indicates that the items measure a consistent body of knowledge and skill reasonably well. The standard error of measurement, about 3.8 percentage points, means a student's true score could plausibly lie a few points above or below the observed score, which matters for students near the passing cutoff of 75%.

Item Difficulty

Item difficulty, the proportion of students who answered correctly, ranged from 0.21 to 0.98. Most items, 38 of 50, fell between 0.40 and 0.90, a useful range for discriminating among students. Six items were very easy, above 0.90; four of these were recall items retained to cover objective 1, which is expected. Six items were very difficult, below 0.40, and needed a closer look, since an item can be hard because it tests difficult content or because it is flawed.

What this part is doingEach statistic is defined in plain terms, interpreted against a standard and connected to a practical consequence, such as the meaning of the standard error for students near the cutoff.
3

Item Discrimination

The discrimination index for each item compares the proportion of high scorers and low scorers who answered correctly; the point-biserial correlation gives similar information. Forty-one items had point-biserial values of 0.20 or higher, showing that they separated stronger from weaker students. Six items had low values between 0.00 and 0.19. Three items had negative values, meaning that low scorers answered them correctly more often than high scorers. A negative discrimination index is the item analysis's clearest warning; it usually means the item is miskeyed, ambiguous or rewards the wrong reasoning.

Looking Closely at Problem Items

Item 17, difficulty 0.29, point-biserial minus 0.12. A prioritization item about four newborns. On review, option C, which most high scorers chose, was also defensible: a newborn with a glucose of 38 mg/dL was as urgent as the keyed answer, a newborn with grunting. The item had two defensible answers. Decision: remove from scoring for this administration, credit all students, and revise so only one option is an emergency.

Item 23, difficulty 0.33, point-biserial minus 0.05. A medication item on methylergonovine. The stem did not state the patient's blood pressure, which is essential because the drug is contraindicated with hypertension; strong students hesitated to choose it without that information. Decision: revise by adding the blood pressure to the stem.

Item 31, difficulty 0.36, point-biserial 0.34. A hard but well-performing unfolding case item on recognizing hemorrhage. Decision: keep; difficulty reflects challenging content that strong students mastered.

Item 42, difficulty 0.62, point-biserial minus 0.08. Distractor analysis showed that high scorers were split between the key and option D. On review, the key was wrong; option D was correct. Decision: rekey, rescore and review the answer key process.

Distractor Analysis

Distractors should attract some students, mostly low scorers. Across the exam, 11 distractors were chosen by no student, which means they did no work and made those items easier, in effect three-option items. They will be replaced with plausible options based on common student errors seen in clinical and in written work.

Why the Flaws Matter

The flaws found here are typical. When researchers reviewed ten end-of-course nursing exams, nearly half of all items contained at least one item-writing flaw, and removing the flawed items changed which students passed and which scored highly (Tarrant & Ware, 2008). The miskeyed item and the item with two defensible answers on this exam would have lowered the scores of the students who reasoned most carefully, which is the opposite of what an exam should do.

Validity Evidence

Reliability alone does not show that the exam measures what we intend. Downing (2003) describes validity as the degree to which evidence supports the intended interpretation of scores, drawing on several sources. Four kinds of evidence were gathered.

Content evidence: the exam blueprint from Week 2 shows that items cover objectives 1, 2 and 3 in proportion to their weight, and two faculty independently classified each item's cognitive level, agreeing on 90%.

Response process evidence: five students completed a think-aloud session on eight items after the exam. On the revised items, their spoken reasoning matched the intended thinking, recognizing cues and prioritizing, which supports the claim that items measure judgment rather than test-taking tricks. On item 23, students said they "needed the blood pressure," confirming the flaw.

Relationship to other variables: exam scores correlated moderately, r = 0.52, with scores on the simulation rubric for the same objectives, which is expected if both measure related skills by different methods.

Consequences: the lower mean led to more students near the cutoff, and three students scored below 75%. The consequence is appropriate if the exam is valid, but it makes careful review of flawed items essential before grades are final.

Limits of the Statistics

With 56 students, item statistics are somewhat unstable; a single student's answers can shift a discrimination index noticeably. Decisions about items with borderline values will therefore wait for a second administration, while clear problems, such as negative discrimination or a wrong key, are acted on now. Statistics are also only one input: an item with good numbers can still be clinically outdated, so every item is reviewed for currency each year regardless of how it performs, and any item that no longer matches current practice guidance is retired even if its statistics are strong.

Decisions and Next Steps

Three items will be revised, one rekeyed and scored correctly for all students, one removed from scoring with credit given, 11 distractors replaced and the answer key process changed to require a second faculty review before each exam. The item analysis will be repeated after the next administration.

Conclusion

The revised exam is reasonably reliable, separates stronger from weaker students better than the old exam and has supporting evidence that it measures the application and analysis the course objectives require. The item analysis found a handful of flawed items, including a miskeyed item and one with two defensible answers, which were corrected before grades were final. Week 5 will move from a single exam to evaluating the whole course.

What this part is doingThe conclusion summarizes what the statistics support and what was corrected. Every source cited in the paper appears in the reference list.
4

References

Downing, S. M. (2003). Validity: On the meaningful interpretation of assessment data. Medical Education, 37(9), 830-837. https://doi.org/10.1046/j.1365-2923.2003.01594.x

Oermann, M. H., & Gaberson, K. B. (2021). Evaluation and testing in nursing education (6th ed.). Springer Publishing.

Tarrant, M., & Ware, J. (2008). Impact of item-writing flaws in multiple-choice questions on student achievement in high-stakes nursing assessments. Medical Education, 42(2), 198-206. https://doi.org/10.1111/j.1365-2923.2007.02957.x

How this NSG 533 Week 4 example is structured

The NSG/533 description stresses that nurse educators count on reliable and valid information. This paper reads real kinds of exam statistics the way faculty receive them from scoring software, interprets each in plain language, weighs validity evidence from several sources and ends with decisions, because an item analysis matters only if it changes the exam. Students search this week as NSG 533 Week 4, NSG533 Wk 4 or NSG/533 Wk 4; all three are the same assignment.

NSG/533 Week 4 questions, answered

What does NSG/533 Week 4 usually ask for?

The course description stresses reliable and valid evaluation tools. Many sections ask students to examine the reliability and validity of an assessment and interpret an item analysis.

What is a good item discrimination index?

Positive values show that high scorers answered the item correctly more often than low scorers. Values of about 0.30 or higher are generally considered good; values near zero or negative suggest a problem with the item.

What does KR-20 measure?

The internal consistency of a test scored right or wrong, an estimate of how consistently the items measure the same body of knowledge. Values of 0.70 or higher are usually considered acceptable for classroom tests.

Write yours, or have the desk draft it

This paper is an original model document written by our desk, not a submitted student paper and not an official University of Phoenix document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.