Grading a Student or Judging a Course? Assessment, Evaluation and the Revised Taxonomy in an Accelerated BSN Maternal-Newborn Course
[Student Name]
University of Phoenix
NSG/533: Educational Assessment and Evaluation
Week 1 Assignment
[Instructor Name]
[Date]
The school, the course and its data are a composite written for a model paper.
I lead the maternal-newborn course in an accelerated second-degree BSN program, where students who already hold a bachelor's degree in another field complete nursing school in sixteen months. The course runs for eight weeks and enrolls 56 students each cohort. Last year, students' average score on the course exams was 84%, yet on the standardized maternal-newborn content exam given at the end of the program, the cohort scored below the national mean, and several clinical instructors said students could recite normal values but struggled to decide what to do when a postpartum patient's condition changed. This paper uses the foundations of assessment and evaluation to understand that mismatch.
Assessment and Evaluation
The two terms are often used interchangeably, but they serve different purposes. Assessment is the process of gathering information about what students know and can do, through tests, observations, papers and questions in class. Evaluation is the process of making a judgment based on that information, such as assigning a grade, deciding that a student has met a clinical competency or deciding that a course needs revision (Oermann & Gaberson, 2021). A quiz used only to show students what they need to review is assessment; the same quiz counted toward a grade becomes part of evaluation.
The distinction matters for my course because the same data serve both purposes. Exam scores assign grades to individual students, and, added together, they are also part of how the program judges whether the course is working. If the exams measure the wrong thing, both judgments are wrong.
Formative and Summative Purposes
Formative assessment happens during learning and is meant to improve it: practice questions, feedback on a care plan draft, a debriefing after simulation. Summative evaluation happens at the end of a unit or course and records achievement: a final exam, a final clinical evaluation. Oermann and Gaberson (2021) recommend that formative assessment be frequent and low stakes, so that students can learn from mistakes before they are graded. In my course, almost all assessment is summative: two exams, a final and a pass-or-fail clinical evaluation. Students receive little feedback until they are graded.
Validity and Reliability in Brief
Two properties determine whether assessment data can be trusted. Validity refers to whether the interpretation of scores is supported, whether a test actually measures what we claim it measures; Downing (2003) describes validity as a property of the meaning we give to scores rather than of the test itself, supported by several kinds of evidence. Reliability refers to consistency, whether a student would receive a similar score on another occasion or from another grader. A test can be reliable without being valid: our exams produce consistent scores, but if they measure recall when we claim they measure clinical judgment, the scores are consistently misleading.
The Revised Taxonomy
The revised taxonomy of the cognitive domain describes six levels of thinking, arranged from simpler to more complex: remember, understand, apply, analyze, evaluate and create (Anderson & Krathwohl, 2001). It also distinguishes kinds of knowledge, from factual to conceptual, procedural and metacognitive. Its value for a nursing faculty member is that it turns the vague complaint that students cannot think into a measurable question: at what level are we asking them to think?
Classifying the Course Objectives
The course has six objectives. I classified each by its verb and its demand on students. One objective, identifying normal physiological changes of pregnancy and the postpartum period, is at the remember and understand levels. Three objectives are at the apply and analyze levels: prioritizing nursing care for women and newborns with common complications, interpreting assessment findings to recognize deterioration and applying evidence to teaching families. One is at evaluate, judging the effectiveness of nursing interventions. One, planning culturally responsive care, is at create. Five of six objectives therefore require thinking above the understand level.
Classifying the Test Items
I then classified all 150 items on last year's two unit exams and final exam. Two faculty classified the items independently and agreed on 91% of them, resolving the rest by discussion. Of the 150 items, 98, or 65%, were at the remember or understand level, such as identifying the normal range for fundal height or the timing of the newborn's first bath. Thirty-nine, or 26%, were at the apply level, and only 13, 9%, required analysis or evaluation, such as deciding which of four postpartum patients to see first.
The Mismatch
The comparison shows a clear mismatch. Five of six course objectives require application or higher, but nearly two-thirds of the test items measure remembering and understanding. Students can earn an 84% average by knowing facts, while the objectives, the clinical instructors and the standardized exam all expect them to use those facts to make decisions. The high course scores are reliable, since students would likely score similarly on a similar test, but they do not support the interpretation that students have met the course objectives, which is a validity problem.
Other Sources of Evidence
The mismatch is supported by other data. Clinical instructors' comments on final evaluations mentioned difficulty prioritizing or recognizing changes in 19 of 56 students. On the standardized content exam, the cohort's lowest scores were in the subscales that test clinical judgment and prioritization. Students in the course evaluation rated exams as "fair" but said practice questions from the licensure review books "felt like a different subject."
Why the Mismatch Happened
The mismatch is not the result of carelessness. Recall items are easier and faster to write, the item bank we inherited was written years ago for a different course, and faculty rarely receive training in writing higher-level items. The accelerated schedule also leaves little time between exams to revise them. Naming these causes matters, because the remedy must address them, not only the items.
A Note on the Other Domains
The revised taxonomy addresses the cognitive domain, but the course also has psychomotor and affective aims, such as performing a newborn assessment correctly and respecting families' cultural practices. These cannot be measured by written tests at all, and the current course measures them only through a pass-or-fail clinical form. Any redesign must include methods suited to these domains, which Week 2 will address.
What the Rest of the Course Will Address
The following weeks will apply the course's content to this problem: matching assessment methods to objectives in Week 2, writing items and rubrics at the right level in Week 3, examining reliability, validity and item analysis in Week 4, evaluating the course with multiple data sources in Week 5 and communicating results to stakeholders in Week 6.
Conclusion
Assessment gathers information and evaluation makes judgments with it; both depend on tools that measure what they claim to measure. Applying the revised taxonomy to the maternal-newborn course showed that its objectives ask students to apply and analyze while most of its test items ask them to remember. The course's high scores are consistent but do not show that students can do what the course promises, which is where improvement must begin.
References
Anderson, L. W., & Krathwohl, D. R. (Eds.). (2001). A taxonomy for learning, teaching, and assessing: A revision of Bloom's taxonomy of educational objectives. Longman.
Downing, S. M. (2003). Validity: On the meaningful interpretation of assessment data. Medical Education, 37(9), 830-837. https://doi.org/10.1046/j.1365-2923.2003.01594.x
Oermann, M. H., & Gaberson, K. B. (2021). Evaluation and testing in nursing education (6th ed.). Springer Publishing.
How this NSG 533 Week 1 example is structured
The NSG/533 description centers on reliable, valid tools aligned to an educational taxonomy and on using assessment and evaluation data to improve programs. This paper defines the key terms, then applies them to a real kind of course, classifying objectives and test items by taxonomy level so the gap between what is taught and what is measured becomes visible. Students search this week as NSG 533 Week 1, NSG533 Wk 1 or NSG/533 Wk 1; all three are the same assignment.
NSG/533 Week 1 questions, answered
What does NSG/533 Week 1 usually ask for?
The course description centers on reliable, valid evaluation tools aligned to an educational taxonomy. Many sections begin by asking students to distinguish assessment from evaluation and to explain how a taxonomy guides the design of assessments.
What is the difference between assessment and evaluation?
Assessment gathers information about learning, often to guide teaching and give feedback. Evaluation uses information to make a judgment, such as a grade or a decision about whether a course or program is working.
Why align test items with a taxonomy?
So that each item measures learning at the level the objective requires. If objectives ask students to analyze and apply, tests that ask only for recall measure the wrong thing.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official University of Phoenix document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.