NSG/533 Week 5: Evaluating a Course With Quantitative and Qualitative Data, sample paper

Reviewed by Lenora Whitcombe, MSN, RN · University of Phoenix

This page holds a complete NSG/533 Week 5 sample course evaluation, in true APA form. It evaluates the first cohort of a composite accelerated BSN maternal-newborn course after its assessments were redesigned, using quantitative data from exams, simulation, clinical evaluations and a standardized content exam and qualitative data from student focus groups and instructor comments, and judges which changes worked and which did not.

1

One Cohort Later: Evaluating a Redesigned Maternal-Newborn Course With Exam, Simulation, Clinical, Standardized Test and Student Voice Data

[Student Name]

University of Phoenix

NSG/533: Educational Assessment and Evaluation

Week 5 Assignment

[Instructor Name]

[Date]

The school, the course and all data are a composite written for a model paper.

What this part is doingThe title names the time frame and every data source. The reader expects a balanced judgment, not only good news.
2

Over the past four weeks, I redesigned the assessments in the maternal-newborn course of our accelerated BSN program: exams shifted toward application and analysis, a graded simulation with a clinical judgment rubric and written assignments with analytic rubrics were added and weekly formative practice was introduced. The first cohort under the redesign has now finished. This paper evaluates the course using several kinds of data.

Evaluation Questions

The evaluation asks four questions. Did students' performance on higher-level thinking improve? Do the new assessments give consistent, trustworthy information? How did students experience the changes? What should change next?

Data Sources

Quantitative data: exam statistics for both cohorts; simulation rubric scores; clinical evaluation outcomes; and the cohort's scores on the standardized maternal-newborn content exam taken at the end of the program. Qualitative data: two student focus groups of eight students each, led by a faculty member from another course; open comments from the course evaluation; and written comments from clinical instructors.

Results: Exams

The redesigned unit exams had a mean of 77.8% across the course, compared with 84.1% for the previous cohort on the old exams. Reliability coefficients ranged from 0.76 to 0.81, compared with 0.72 to 0.75 on the old exams. On a set of 12 scenario-based items that appeared on both cohorts' final exams, the redesigned cohort answered 68% correctly, compared with 59% for the previous cohort. The lower overall mean reflects harder items, while the shared items suggest improved higher-level performance.

What this part is doingResults are reported with comparisons to the previous cohort and with an explanation of why a lower mean does not mean worse learning. The shared-items comparison is the fairest test.
3

Results: Simulation and Clinical

On the postpartum hemorrhage simulation, scored with the clinical judgment rubric Lasater (2007) developed, adapted for this scenario, on four dimensions scored 1 to 4, the cohort's median scores were 3 for noticing, 3 for interpreting, 3 for responding and 2 for reflecting. Two raters scored 14 students independently; they agreed exactly on 79% of dimension scores and within one level on all. Clinical instructors noted difficulty with prioritizing or recognizing changes in 9 of 55 students, compared with 19 of 56 in the previous cohort.

Results: Standardized Content Exam

The redesigned cohort's mean on the standardized maternal-newborn exam was at the national mean, compared with below the mean for the previous cohort, and scores on the clinical judgment subscale rose the most. Because the exam is given at the end of the program, other courses may have contributed.

Results: Student Voice

Focus group notes were analyzed by two faculty for themes. Students said the new exams felt harder but fairer, because they "looked like what happens on the unit." Three themes appeared in both groups. First, weekly practice questions with immediate explanations were the most valued change. Second, the simulation was stressful but taught them more than any lecture, especially the debriefing. Third, the written case analysis felt disconnected from the exams, and several students did not understand how the rubric's levels differed. Course evaluation comments echoed these themes: 41 of 52 comments on assessments were positive, and the most common criticism concerned the rubric.

Results: Instructor Voice

Clinical instructors wrote that students asked better questions and recognized changes in patients sooner. Two instructors noted that students still struggled to reflect on their own performance, matching the low simulation score on reflecting.

Were the Comparisons Fair?

Comparing two cohorts invites the question of whether they were alike. The two cohorts had similar mean prior grade point averages and similar proportions of students with health care experience, and the same faculty taught both. The shared items were identical on both finals. These similarities do not rule out every difference, but they make it less likely that the improvement came from a stronger group of students rather than from the redesign. Validity arguments rest on accumulated evidence of this kind rather than on a single number (Downing, 2003).

Interpretation

The data answer the evaluation questions with reasonable confidence. Higher-level performance appears to have improved: shared exam items, fewer clinical concerns and the standardized exam all point the same way, and the students' own account supports it. The new assessments appear reasonably reliable, with better internal consistency and acceptable agreement between simulation raters. Students experienced the changes as harder and fairer. The weakest areas are reflection and the written case analysis rubric.

The evidence has limits. One cohort cannot rule out differences between groups, such as differences in prior education. The standardized exam reflects the whole program. Focus group participants volunteered and may differ from other students. Oermann and Gaberson (2021) emphasize that evaluation decisions should rest on several sources of evidence and be revisited as data accumulate, and the conclusions here are offered in that spirit.

What the Data Could Not Show

The data could not show whether students will perform better as nurses, which is the course's real purpose. Clinical instructor comments and the standardized exam are the closest measures available within the program. A follow-up survey of graduates' employers, discussed in Week 6, may add evidence later.

Decisions

Keep the exam redesign, the weekly practice and the simulation. Revise the case analysis rubric by adding an annotated example at each performance level and a class session in which students score a sample paper, since students did not understand the levels. Strengthen reflection by adding a guided reflection prompt after simulation and clinical days, using the reflection aspect of clinical judgment as a structure. Repeat the evaluation with the next cohort, including the shared exam items.

Cost of the Redesign

The evaluation should also count costs. Faculty spent about 40 additional hours on item writing and review, simulation required eight hours of simulation center time and two trained raters, and grading the written assignments added about 25 hours per cohort. These costs are modest compared with the benefits found, but they must be planned for, and the communication in Week 6 will present them to leadership alongside the results.

Unexpected Findings

Two findings were unexpected. Students who scored lowest on the first unit exam improved more by the final than in the previous cohort, which suggests that the weekly practice questions helped weaker students most. And simulation raters disagreed most often on the responding dimension, which suggests that the rubric's description of that level needs clearer examples before the next cohort. Both findings will be checked again with the next cohort before any change to the rubric or the practice questions is made permanent, since a single cohort can produce patterns that do not repeat.

Conclusion

The first cohort after the redesign showed improved performance on higher-level items, fewer clinical concerns and a standardized exam score at the national mean, with students describing the assessments as harder but fairer. Reflection and the case analysis rubric need work. Combining numbers with students' and instructors' words produced a clearer and more trustworthy judgment than either kind of data alone. Week 6 will plan how to communicate these results to stakeholders.

What this part is doingThe conclusion states what improved, what needs work and why mixed data made the judgment stronger. Every source cited in the paper appears in the reference list.
4

References

Downing, S. M. (2003). Validity: On the meaningful interpretation of assessment data. Medical Education, 37(9), 830-837. https://doi.org/10.1046/j.1365-2923.2003.01594.x

Lasater, K. (2007). Clinical judgment development: Using simulation to create an assessment rubric. Journal of Nursing Education, 46(11), 496-503. https://doi.org/10.3928/01484834-20071101-04

Oermann, M. H., & Gaberson, K. B. (2021). Evaluation and testing in nursing education (6th ed.). Springer Publishing.

How this NSG 533 Week 5 example is structured

The NSG/533 description calls for gathering qualitative and quantitative data that tell the educator how well learners are performing and how effective the program is. This paper states the evaluation questions first, reports each data source with its limits, compares the redesigned cohort with the previous one and ends with judgments and the changes they lead to. Students search this week as NSG 533 Week 5, NSG533 Wk 5 or NSG/533 Wk 5; all three are the same assignment.

NSG/533 Week 5 questions, answered

What does NSG/533 Week 5 usually ask for?

The course description calls for using qualitative and quantitative data to evaluate educational effectiveness. Many sections ask students to evaluate a course or program with multiple sources of data and draw conclusions.

Why use both quantitative and qualitative data?

Numbers show whether outcomes changed and by how much; students' and instructors' words help explain why. Together they give a fuller and more trustworthy picture.

Can one cohort show that a course change worked?

One cohort gives an early signal but cannot rule out differences between cohorts. Conclusions should be stated cautiously and checked with later cohorts.

Write yours, or have the desk draft it

This paper is an original model document written by our desk, not a submitted student paper and not an official University of Phoenix document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.