Unknown for One Patient in Three: Assessing the Quality of Race, Ethnicity and Language Data in a Community Health Center's Record Before Using It to Measure Disparities
[Student Name]
University of Phoenix
DNP/715: Information Systems and Health Care Delivery Technology
Week 3 Assignment
[Instructor Name]
[Date]
The health center and figures are composites written for a model paper.
I am a DNP student and nurse practitioner at a community health center serving about 14,000 adults. My project will address blood pressure control, and our board wants to know whether control differs by race, ethnicity and preferred language. Before I produce those comparisons, I need to know whether the record's demographic data can bear the weight. This paper assesses their quality.
Why Assess Before Analyzing
Weiskopf and Weng (2013) reviewed studies of electronic health record data quality assessment for research reuse and named five qualities that usable data must have: they must be present, accurate, consistent across sources, believable and up to date. They also identified seven categories of assessment methods, including comparison with gold standards, data element agreement, data source agreement and element presence. They emphasized that data quality is task dependent: data adequate for one purpose may be inadequate for another.
A Harmonized Framework
Kahn et al. (2016) harmonized data quality terminology into three categories, conformance, completeness and plausibility, and two assessment contexts: verification, checking data against organizational expectations, and validation, checking against an external gold standard. I use Kahn et al.'s categories for my checks and Weiskopf and Weng's dimensions to interpret them.
What Other Systems Found
Polubriaginof et al. (2019) examined race and ethnicity data in large national databases and a New York City health system's record and found race or ethnicity unknown for 25% of patients in the national data and 57% in the health system. When patients recorded their own race and ethnicity, 86% provided meaningful information, and 66% reported information that differed from what the record contained. Patient self-recording substantially improved data quality.
A disparity measured with data that are missing for a third of patients is not a finding; it is a guess about the two thirds who were counted.
Completeness
Of 14,112 adult patients seen in the past year, race was recorded as unknown, declined or blank for 4,590, or 32.5%. Ethnicity was missing for 21.8%. Preferred language was missing for only 3.1%, because front-desk staff must enter it to schedule interpreters. Completeness varied by site: our newest site had 51% missing race, compared with 19% at our oldest.
Conformance
Our record allows free-text entries in the race field at one site, producing 214 different spellings and categories, some not matching federal reporting categories. These fail conformance checks and cannot be grouped reliably.
Plausibility
Plausibility checks found 312 patients whose race changed between visits in ways unlikely to reflect self-identification changes, such as alternating between two categories at each visit. This suggests that staff sometimes enter race by observation rather than by asking.
Correctness
To validate correctness, I compared the record with patient self-report for a random sample of 150 patients who completed a brief tablet survey in the waiting room. Race matched for 71%. Most mismatches involved patients who identified as multiracial or Hispanic, echoing Polubriaginof et al. (2019).
Currency
Race and ethnicity are relatively stable, but language preference can change. Language had not been updated in more than five years for 38% of patients.
Implications for My Analysis
With a third of race data missing and nonrandom by site, and 29% of recorded race disagreeing with self-report in my sample, any comparison of blood pressure control by race would be unreliable. If missing data are concentrated among patients with worse control, which is plausible at our newest site, disparities could be understated or overstated.
Missing Is Not Random
Missing race data were not spread evenly. They were more common among patients seen once, patients registered through the urgent care entrance and patients whose preferred language was not English. If these groups also differ in blood pressure control, excluding them would bias any comparison, which is why completeness must improve before conclusions are drawn.
Improvement Plan
First, move collection to patient self-report using a tablet at check-in, with categories consistent with federal standards and options for multiple races. Second, remove free-text entry. Third, train front-desk staff to explain why the questions are asked and never to assign race by observation. Fourth, prompt updates of language preference annually.
Aligning With Reporting Standards
Federal reporting for health centers uses defined race and ethnicity categories, and our free-text entries cannot be mapped to them. Aligning our fields with those categories, while allowing multiple selections and a write-in for detail, satisfies reporting and respects how patients describe themselves.
Interim Analysis Approach
Until data improve, I will report blood pressure control by language, whose data are more complete, and by race only for sites with less than 15% missing, with explicit notes on the limitations.
Language as a Model
Language data are nearly complete because they serve an immediate operational purpose: scheduling interpreters. When staff see a direct use for a data element, they collect it reliably. Explaining to staff how race and ethnicity data will be used to improve care, and sharing results with them, may improve completeness in the same way.
Measuring Data Quality Over Time
I will track completeness, conformance and agreement with self-report monthly for six months after the change, using the same checks, and I will share site-level results with each site's staff.
Who Enters the Data
At our sites, race is entered by front-desk staff during registration, often at a busy window with a line of patients. Staff told me they sometimes skip the question or guess to save time and avoid awkwardness. Data quality problems often begin in workflows designed for speed rather than accuracy, and fixing them requires attention to the people who collect data, not only the fields they fill.
Using Multiple Data Sources
Weiskopf and Weng (2013) describe data source agreement as one assessment method. Our health center also collects race on a separate sliding-fee application. Comparing the two sources for 500 patients found agreement in 78%, and the application, completed by patients themselves, was more often complete. Linking the two sources could improve completeness immediately while self-report at check-in is introduced.
Reporting Limitations Transparently
When I present blood pressure control by group, I will include a table of data completeness by site and group, so the board can see how much confidence each comparison deserves. Transparent reporting of data quality is part of honest analysis.
Ethics and Trust
Some patients decline to share race because of past discrimination. Explaining the purpose, to find and fix unequal care, respects their choice while encouraging participation.
Conclusion
Before measuring blood pressure disparities, I assessed the quality of race, ethnicity and language data using two data quality frameworks. Race was missing for a third of patients, varied by site, included nonconforming free text and disagreed with self-report in nearly a third of those sampled. These findings make a race-based analysis unreliable now, and an improvement plan based on patient self-report, supported by published evidence, will make it possible later.
References
Kahn, M. G., Callahan, T. J., Barnard, J., Bauck, A. E., Brown, J., Davidson, B. N., Estiri, H., Goerg, C., Holve, E., Johnson, S. G., Liaw, S.-T., Hamilton-Lopez, M., Meeker, D., Ong, T. C., Ryan, P., Shang, N., Weiskopf, N. G., Weng, C., Zozus, M. N., & Schilling, L. (2016). A harmonized data quality assessment terminology and framework for the secondary use of electronic health record data. eGEMs, 4(1), Article 1244. https://doi.org/10.13063/2327-9214.1244
Polubriaginof, F. C. G., Ryan, P., Salmasian, H., Shapiro, A. W., Perotte, A., Safford, M. M., Hripcsak, G., Smith, S., Tatonetti, N. P., & Vawdrey, D. K. (2019). Challenges with quality of race and ethnicity data in observational databases. Journal of the American Medical Informatics Association, 26(8-9), 730-736. https://doi.org/10.1093/jamia/ocz113
Weiskopf, N. G., & Weng, C. (2013). Methods and dimensions of electronic health record data quality assessment: Enabling reuse for clinical research. Journal of the American Medical Informatics Association, 20(1), 144-151. https://doi.org/10.1136/amiajnl-2011-000681
How this DNP 715 Week 3 example is structured
The DNP/715 Week 3 work usually focuses on electronic health records and data quality. This paper evaluates specific data elements against named quality dimensions, quantifies each problem and shows how data quality determines whether a planned analysis is valid. Students search this week as DNP 715 Week 3, DNP715 Wk 3 or DNP/715 Wk 3; all three are the same assignment.
DNP/715 Week 3 questions, answered
What does DNP/715 Week 3 usually ask for?
Many sections ask students to examine electronic health record data quality, including completeness, accuracy and consistency, and its implications for practice and projects.
What are the dimensions of EHR data quality?
A widely cited review identified five: completeness, correctness, concordance, plausibility and currency; a later harmonized framework groups quality checks into conformance, completeness and plausibility.
Why does race and ethnicity data quality matter?
These data are needed to measure and reduce health disparities; if they are missing or wrong, disparities can be hidden or misstated, and patient self-report generally improves their quality.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official University of Phoenix document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.