100

Open Data Quality Requires Looking Inside the File

A dataset can be published, licensed, machine-readable, and highly rated by a portal, yet still fail real-world reuse. Portal and metadata scores indicate access and discovery, not the quality of the data itself.[1][2]

đŸ§” 1/5

Start with the reuse case and legal gate: verify a licence, avoid terms that prohibit commercial exploitation, and check whether the data is available in standard formats through APIs, web services, or downloads.[3]

đŸ§” 2/5

Then inspect the schema against a domain standard. Check accurate feature names, required fields, semantic meaning, and completeness. Test actual values for consistent types, and measure missingness from nulls and empty strings.[4]

đŸ§” 3/5

Do not stop at structure. Profile values for inconsistencies and anomalies, including unusual observations or possible sensor faults. The evidence supports content-level checks, but not a single universal anomaly score or dedicated anomaly dimension.[5]

đŸ§” 4/5

Report separate results for format, schema accuracy, schema completeness, data-type consistency, and data completeness. Track updates, document reuse, and monitor progress, while remembering: licensing and portals enable reuse; only direct content assessment shows whether data is fit for it.[6]

đŸ§” 5/5