88

How should researchers assess whether an administrative dataset is reusable?. Organize the report around quality dimensions, comparability, metadata and provenance, validation, processing, and governance. Include a practical assessment framework that distinguishes what producers should document from what analysts should verify, while preserving the sources' cautions about coverage, coding changes, and access restrictions.

Assessing Whether an Administrative Dataset Is Reusable

An administrative dataset is reusable only relative to a proposed research use. Researchers should not ask whether the dataset is simply “good,” but whether its population, measurements, coding, provenance, processing, and access conditions are sufficiently understood and accurate for the intended question.[1][2] This framework separates what data producers should document from what analysts should independently verify, while retaining the principal cautions about coverage, coding changes, missingness, and restricted access.

Quality dimensions and the producer-analyst division of responsibility

DimensionProducers should documentAnalysts should verify
ConformanceExpected formats, structures, value ranges, identifiers, permissible codes, and known departures from them.Whether records and fields conform to the stated structural and coding rules, and whether nonconforming values affect the proposed analysis. Conformance is one of the core data-quality categories identified for secondary use.[3]
CompletenessRequired fields, periods or sites with missing data, missingness mechanisms where known, and whether absence of a record means absence of the event.Amounts and patterns of missing variables, observations, and linkages; whether missingness could bias inclusion, exposure, outcome, or confounder measurement.[4][5]
Plausibility and validityData-entry practices, care setting, who entered the data, staff training, and systematic procedures used during collection.Whether values are clinically or administratively plausible and whether the measures are valid for the intended construct. Large volume does not guarantee accuracy, especially where billing-related codes may be inaccurate or strategically applied.[6][7][8][9]
Fitness for purposeThe original operational purpose, collection context, intended population, and known limitations.Whether the dataset actually measures what the research question requires, rather than assuming that routine availability implies suitability.[10][11]

Comparability, coverage, metadata, and provenance

A reusable dataset needs an explicit account of where it came from and how its population was formed. Producers should state the database name and type, geographic region, time frame, care or administrative setting, collection period, original purpose, and the relationship between the database population and the underlying source population. A database label alone does not explain what the data contain or how they were generated.[12][13]

For comparisons across places or time, producers should record changes in eligibility rules, participating sites, population composition, clinical or administrative practice, software, and coding conventions. Analysts should test whether observed differences could instead reflect changes in which patients were included, which tests were performed, or how diagnoses and procedures were coded. Differences between hospitals and populations can alter both testing and diagnostic algorithms.[14][15]

Coverage must be assessed separately from nominal database size. Producers should describe inclusion and exclusion criteria, selection codes and algorithms, linkage success, and the stages by which the study population was drawn from the initial database. Analysts should compare that population with the relevant source population and investigate whether missing or unlinked individuals create selection bias or limit generalisability.[16][17]

Validation, ascertainment, and processing

Producers should provide the definitions, code lists, algorithms, linkage rules, and ascertainment procedures used to identify people, exposures, outcomes, confounders, and effect modifiers. Where possible, they should report comparisons with a reference standard using measures such as sensitivity, specificity, predictive values, or kappa statistics. Analysts should check that the validation population and setting are sufficiently similar to their own, because validation evidence from another population or database may not transfer directly.[18][19][20]

Verification and validation are different forms of evidence. Verification asks whether data agree with an organisation’s own records; validation compares them with an accepted gold standard. Analysts should not treat internal record checks as equivalent to gold-standard validation.[21][22]

For processing and reproducibility, producers should document cleaning, linkage techniques, linkage-quality assessment, selection steps, code lists, algorithms, and any survey wording or derived variables, subject to legal and contractual limits. Analysts should reproduce the permitted processing where possible and inspect flow diagrams showing linked and unlinked records, exclusions, and the final analytic population.[23][24]

Coding changes, incentives, and access restrictions

Producers should maintain a change history for classification systems, code definitions, software, and coding guidance. Analysts should look for transitions such as ICD-9 to ICD-10, upcoding, opportunistic coding, stigma-related undercoding, and provider incentives. These mechanisms can change ascertainment even when the underlying condition or activity has not changed.[25][26]

Reusability also depends on whether the data and methods can actually be inspected. Producers should state what investigators could access, which fields or records were withheld, and whether licensing, proprietary algorithms, copyright, intellectual-property rules, or other laws restrict publication of code lists or analytic methods. Analysts should determine whether those restrictions prevent replication, constrain the sample, or alter the feasible study design.[27][28]

Practical assessment workflow

  1. Define the use case. Specify the target population, time period, exposures, outcomes, comparison groups, and decisions the data will support. Judge fitness against this use, not against a universal quality label.[29][30]
  2. Establish provenance. Obtain documentation of origin, purpose, setting, geography, time frame, source population, eligibility, and data-generation processes.[31][32]
  3. Profile quality. Measure conformance, completeness, and plausibility, then assess whether data-entry practices and missingness threaten the intended analysis.[33][34][35]
  4. Test comparability and coverage. Identify population, site, practice, software, eligibility, and coding changes; quantify or describe linkage loss and exclusions; and assess generalisability.[36][37][38]
  5. Assess ascertainment. Review definitions, algorithms, code lists, linkage rules, and validation evidence, distinguishing verification from gold-standard validation.[39][40][41]
  6. Reconstruct processing. Follow the cleaning, selection, linkage, and derivation steps and inspect the flow from source records to the analytic cohort.[42][43]
  7. Audit change and access risk. Record coding and software changes, incentives, restrictions, unavailable fields, and limits on reproducing methods or code lists.[44][45]
  8. Make a use-specific decision. Classify the dataset as reusable, reusable with qualification, or unsuitable for the proposed use, explaining residual risks and any protocol deviations rather than hiding them.[46][47]

Conclusion

The minimum standard for reuse is an auditable chain from origin to analysis: researchers should be able to determine how the population was covered, how variables were recorded and coded, how records were cleaned and linked, what validation supports the measures, what changed over time, and what access restrictions remain. The final decision should weigh conformance, completeness, plausibility, comparability, coverage, ascertainment, and reproducibility against the specific research purpose. Even after adjustment, missing variables and unmeasured confounding remain limitations because analysts can only control for information present in the dataset.[48][49]