Exposing the Organisational Damage Caused by Dark Data and how to Prevent it

Authors

  • Graham Chant Stewart Barr & Associates

DOI:

https://doi.org/10.34190/eckm.27.1.4720

Keywords:

dark, data, cataloguing, lifecycle, management, governance

Abstract

Dark data refers to data that organisations routinely collect, process, and store but do not use for decision-making, analysis, or value creation. Although often overlooked, dark data has become a significant organisational concern due to the volume, velocity, and heterogeneity of contemporary information ecosystems. Its presence creates four interrelated clusters of organisational problems. First, it undermines strategic and operational decision-making by obscuring potentially valuable insights and distorting evidence bases. Second, it creates operational inefficiencies, including duplicated data, uncontrolled storage growth, and technical debt. Third, dark data introduces compliance, security, and ethical risks because organisations may unknowingly retain sensitive, regulated, or high-risk information without adequate oversight. Finally, it reflects and reinforces organisational and cultural barriers, such as weak stewardship, siloed practices, and low data literacy, which prevent coordinated and responsible data use. This study investigates how these problems can be reduced or minimised through coordinated operational activities rather than isolated, technology-centric interventions. A comprehensive literature review methodology was employed, using structured search, screening, and thematic synthesis across peer-reviewed publications and scholarly frameworks in data governance, information management, organisational risk, and knowledge management. This approach ensured systematic coverage of relevant domains and enabled the integration of theoretical and empirical perspectives. The review findings demonstrate that dark data cannot be eliminated entirely, but its organisational impact can be substantially mitigated by four categories of coordinated action: (1) data discovery and cataloguing to improve visibility and accountability; (2) lifecycle control and data-quality management to reduce redundancy and unmanaged growth; (3) organisational enablement to strengthen stewardship, literacy, and incentives; and (4) governance-by-design architectural practices that embed rules, controls, and monitoring directly into data platforms and pipelines. Together, these activities provide a coherent framework for reducing dark data and structuring future empirical research on dark data governance.

Downloads

Published

2026-08-25