An administrative dataset is reusable only relative to a proposed research use. Researchers should not ask whether the dataset is simply “good,” but whether its population, measurements, coding, provenance, processing, and access conditions are sufficiently understood and accurate for the intended question.[1][2] This framework separates what data producers should document from what analysts should independently verify, while retaining the principal cautions about coverage, coding changes, missingness, and restricted access.
| Dimension | Producers should document | Analysts should verify |
|---|---|---|
| Conformance | Expected formats, structures, value ranges, identifiers, permissible codes, and known departures from them. | Whether records and fields conform to the stated structural and coding rules, and whether nonconforming values affect the proposed analysis. Conformance is one of the core data-quality categories identified for secondary use.[3] |
| Completeness | Required fields, periods or sites with missing data, missingness mechanisms where known, and whether absence of a record means absence of the event. | Amounts and patterns of missing variables, observations, and linkages; whether missingness could bias inclusion, exposure, outcome, or confounder measurement.[4][5] |
| Plausibility and validity | Data-entry practices, care setting, who entered the data, staff training, and systematic procedures used during collection. | Whether values are clinically or administratively plausible and whether the measures are valid for the intended construct. Large volume does not guarantee accuracy, especially where billing-related codes may be inaccurate or strategically applied.[6][7][8][9] |
| Fitness for purpose | The original operational purpose, collection context, intended population, and known limitations. | Whether the dataset actually measures what the research question requires, rather than assuming that routine availability implies suitability.[10][11] |
A reusable dataset needs an explicit account of where it came from and how its population was formed. Producers should state the database name and type, geographic region, time frame, care or administrative setting, collection period, original purpose, and the relationship between the database population and the underlying source population. A database label alone does not explain what the data contain or how they were generated.[12][13]
For comparisons across places or time, producers should record changes in eligibility rules, participating sites, population composition, clinical or administrative practice, software, and coding conventions. Analysts should test whether observed differences could instead reflect changes in which patients were included, which tests were performed, or how diagnoses and procedures were coded. Differences between hospitals and populations can alter both testing and diagnostic algorithms.[14][15]
Coverage must be assessed separately from nominal database size. Producers should describe inclusion and exclusion criteria, selection codes and algorithms, linkage success, and the stages by which the study population was drawn from the initial database. Analysts should compare that population with the relevant source population and investigate whether missing or unlinked individuals create selection bias or limit generalisability.[16][17]
Producers should provide the definitions, code lists, algorithms, linkage rules, and ascertainment procedures used to identify people, exposures, outcomes, confounders, and effect modifiers. Where possible, they should report comparisons with a reference standard using measures such as sensitivity, specificity, predictive values, or kappa statistics. Analysts should check that the validation population and setting are sufficiently similar to their own, because validation evidence from another population or database may not transfer directly.[18][19][20]
Verification and validation are different forms of evidence. Verification asks whether data agree with an organisation’s own records; validation compares them with an accepted gold standard. Analysts should not treat internal record checks as equivalent to gold-standard validation.[21][22]
For processing and reproducibility, producers should document cleaning, linkage techniques, linkage-quality assessment, selection steps, code lists, algorithms, and any survey wording or derived variables, subject to legal and contractual limits. Analysts should reproduce the permitted processing where possible and inspect flow diagrams showing linked and unlinked records, exclusions, and the final analytic population.[23][24]
Producers should maintain a change history for classification systems, code definitions, software, and coding guidance. Analysts should look for transitions such as ICD-9 to ICD-10, upcoding, opportunistic coding, stigma-related undercoding, and provider incentives. These mechanisms can change ascertainment even when the underlying condition or activity has not changed.[25][26]
Reusability also depends on whether the data and methods can actually be inspected. Producers should state what investigators could access, which fields or records were withheld, and whether licensing, proprietary algorithms, copyright, intellectual-property rules, or other laws restrict publication of code lists or analytic methods. Analysts should determine whether those restrictions prevent replication, constrain the sample, or alter the feasible study design.[27][28]
The minimum standard for reuse is an auditable chain from origin to analysis: researchers should be able to determine how the population was covered, how variables were recorded and coded, how records were cleaned and linked, what validation supports the measures, what changed over time, and what access restrictions remain. The final decision should weigh conformance, completeness, plausibility, comparability, coverage, ascertainment, and reproducibility against the specific research purpose. Even after adjustment, missing variables and unmeasured confounding remain limitations because analysts can only control for information present in the dataset.[48][49]
Get more accurate answers with Super Pandi, upload files, personalized discovery feed, save searches and contribute to the PandiPedia.
Let's look at alternatives: