It’s one thing to lose a computer file or for that matter any hard-copy documents related to one’s personal affairs. It’s another thing to lose files and data which other people depend on. I’m reminded of the fires that destroyed the warehouses of Universal Studios in the San Fernando Valley in 2008. According to news reports, original master recordings of over 800 recording artists were destroyed in whole or part, almost a complete who’s who of American popular music. The fire reportedly destroyed recording data of rock-and-rollers like Buddy Holly, country icons like Dolly Parton, jazz virtuosos like Aretha Franklin, blues giants like Muddy Waters, among many others. Clearly the first and foremost principles of data management deserve to be data preservation and data integrity.
Unfortunately the mindset needed for assuring the satisfactory exercise of these principles has often been adopted momentarily at best. The destruction of many libraries throughout the ages suggest we have been doomed to repeating the same mistake, again and again. Perhaps one of the many curses of the human condition is the tendency to substitute aspiration for principle.

source: Global Datavault, https://www.globaldatavault.com/blog/information-destruction-history/, accessed 03/09/2021
In terms of research itself, research practices tend to be most effective when information is available and easy to retrieve. While the Internet has generally opened up previously closed avenues to knowledge and information, one of the drawbacks of using the Internet as a tool for research is the over abundance of data storage locations. When research data is stored in more than a couple of locations, data becomes difficult to keep track of and use effectively. Critical evidence can be lost not because the data has been erased but because its location is no longer known or insufficient cataloging occurred. Using only one tool for research such as Zotero, in which data is stored and centralized in on place and cataloging can be semi-automated, mitigates the risk of loosing track of references, internet resources, and other research upon which one’s current research depends.
A second problem involves the double-edged nature of digital media. As easy as it is to duplicate digital data, it can be just as easy to delete. Depending on the time table of the research, internet pages can suddenly disappear without a trace, leaving the visitor watching in dismay a 404-page-not-found animation. Depending on the value of the missing information, a search in the Wayback Machine may or may not yield the version of the page initially accessed. To some extent the ephemerality of the web poses a serious risk to the quality of research. As a result, an additional evaluation of internet data comes into play in terms of determining the need for redundant data preservation by downloading the web page.
While automated redundancy more broadly has gotten better since the early days of the Internet, we are still not at that point when all applications automagically save all significant versions into a 100% redundant versioning system. Total and automated redundancy goes against the right to be forgotten, the right to be anonymous, and the right not to be tied to a moment in the past. Control over the degree and nature of redundancy to some extent offers freedoms at the cost of the discipline and responsibility to judiciously save a version.
As research goes beyond the work of a single individual to encompass group collaboration, effective data management takes on an even greater importance. Without clear roles for managing data, the likelihood of encountering problems will persist. Rock-solid technologies may be in place, but if the responsibilities for the management of data are not clearly defined and assigned, state-of-the-art storage technologies will not by themselves prevent data loss or the reduction of data integrity.


