Personal blog – Data management

This past week the team met with Lara Alonzo from coronastories.world . She alogn with one of her partners, discussed with us among other things, data!  They mentioned constructing their own site and obtaining the project’s main content; videos of people across the world answering two questions relating to their lives during Covid19, by solicitation. They obtained the 30-40 seconds videos through email, social media, even Whatsapp. However, they had to compile each video, catergorize them, and edit them by adding lower thirds graphics and subtitles. The participants consent to exposing their information by submitting their videos.

This conversation got me thinking about the Covid 19 Archive ‘s own data. Similarly, our team will need to collect the media, alogn with consent forms for the participants, and possibly, the school information. This sensative data needs to be organized, and most importantly protected. This week I created a second draft of a consent form that can also act as a way to collect media. I did this on a shared Google Form. What’s great about this form, if we manage to work out the signature component, is how it creates an esay to use categorized spreadsheet.

As the media producer, I will be responsible to collect,and organize raw media from our participant. This one needs to be stored locally, on my computer, to be edited, and exported to the proper format for our plaform and possibly youtube as another hosting site. However, I will be providing the team with all of the raw, drafts, and final products. Therefore, the media data, will be stored and organized on at least three places: locally on my computer, on a server for the team to access, and lastly, the final product, will be stored on our site’s storage and maybe on the Youtube server.

Other possible pieces contructed for promotion and marketing purposes, will be store on my peresonal Canva account, where they can be easily edited, and shared. However,  the final verisons of these these can also be downloaded and stored.

Personal Blog #3: Data is Everything, Everything is Data

forms of data

Last week’s presentation by Stephen Zweibel changed my perception of data. Previously (and naively), I considered data to be computer-generated but actually it also includes non-digital items. From his slide “Forms of Data”, it appears that everything is data so therefore, data is everything! Which makes sense – without data, our projects would not exist. Even with our Rebus project, we will be dealing with digital files and code, etc. but even the digital files are based off of actual objects, data in the form of pieces of paper that were printed. 

I also started to think about big vs. small data. Our project is using small data, perhaps even tiny data, for scholarly purposes. On the other side of the spectrum is big data, big tech, AI, IoT- types of data, analytics and data used for profit and marketing. What about medium data, is there such a thing?

So, a DMP is essential. Our process to follow the checklist was straightforward. Although we “checked” everything off, there are still a lot of unknowns at this point. But, the DMP is now firmly implanted in my head and going forward, obtaining and preserving data will be a constant concern.  If my team finds that we’re unable to manage the data, that’s a major hurdle to overcome. I can already see how all the files in Google Drive can become uncontrollable if we don’t stay organized! Meanwhile, who knows how Google is profiting from the data we’re storing on their servers.

Rachel’s Public Journal – Data Management+

Steven’s DMP checklist was very logical and it helped our team talk through our current conceptualization of the project. I posit that one of the biggest unseen benefits to completing the data management plan early in the project (though really, we’re in the thick of it now — I think it will feel like we are just getting started until we’re racing to the finish line, if it’s like any other of my projects during the pandemic) was how it crystallized some of the artifacts around our site. 

I think the reason that the data document was so logical for us to conceptualize was that an archive —or, as I am inclined to believe – what we may have, to Lisa Rhody’s point in her talk a couple weeks ago, a ‘collection” 🤔 — may have more concrete ideas about what data will be used due to the nature of its thing-ness. Nevertheless, our project still has ambiguities that were fun and helpful to sort out with the team using the data management checklist and provided examples. 

I think my biggest non-DMP challenge overall has been tamping down my control freak nature, which is great for solo projects and horrible for group work — lessons learned as a young undergrad — and letting the team as a whole guide the project with their interests, ideas, and areas of expertise. It’s been great so far, and I’m learning so much. 

Additionally this week we identified some areas for interaction within the WordPress platform which will be both the stretch goal and a fun technical challenge for me. No spoilers, but hopefully it will allow users to interact with the collected rebus puzzles that already have solutions. I did a bit of outreach, and also “solved” a French rebus critical of Napoleon’s efforts while away from France with my mom, who normally likes to Zoom in and do a Sunday crossword with me. I may just extend this to my nascent network of relevant rebus riddlers, and see if we can “swarm on the (rebus) problem,” one of my favorite agile practices, for fun and project progress.

Lisa’s Public Journal – Week Six — Data Management

Our team has had several meetings where we have touched on different aspects of data management.  We are using GitHUB for development, which has good version controls for the codebase.  That system is self-contained and doesn’t require that we create a system from whole cloth.  However, we have agreed to a version control system for file names which is limited to letters, numbers, and the underscore symbol and which includes the creation date.

A bigger challenge is how to collaborate on content creation asynchronously.  We have settled on using a shared Google spreadsheet, with each tab being a different aspect of the website: Splash page, timeline page, about page, et cetera.  We have also agreed on using Tidy Data* standards for the tracking of all assets.  The Tidy Data standard, when applied to a spreadsheet, uses the structure below:

  1. Every column is a variable.
  2. Every row is an observation.
  3. Every cell is a single value.

We expect to start populating the spreadsheet soon so that our developer has some content to push into our site’s wireframe.  Our expectation is to have the bulk of our research completed by 15 April.  While the spreadsheet will host information about the different data sources, assets we create will be hosted on Wikimedia Commons if a picture or on Soundcloud if an audio file.  In this way, we can extend the life of the assets without the worry of having to pay for their storage.

Our team leader is hosting all our shared Google files on their personal Google account.  Archives are being saved to the Library on our team Commons site.  It will be interesting to see how collaborating in a shared spreadsheet works over time. My biggest fear is that we inadvertently lose information.  However, my hope is that by agreeing to conform to these standards at the beginning of the project, we will avoid our developer losing time managing the uploads during the production process.

*SOURCE: https://vita.had.co.nz/papers/tidy-data.html and https://vita.had.co.nz/papers/tidy-data.pdf.

Bianca’s weekly post: 4×6 nostalgia

Reflections on how (if) data management can improve my research practices:

Although our group (ReadingRebus) chose to manage our project data through a combination of Google Folders/Docs and a Google Sheet tracking each team member’s contribution on a timeline, I thought I’d explore further a more task-designated tool that I came across when we were still in the project planning stage.

Trello’s design appealed to me because it looked as if it might replicate my ancient dissertation-research logging practices: writing each quotation or idea on a 4×6 index card, with a bibliographic shorthand in one corner and the topic/chapter designation in the other.  I shuffled them endlessly as I reorganized my argument and its supporting materials, bound them with elastic bands, and stored them in shoe boxes.

And then I threw all the cards away when the dissertation was done.  For some time, I regretted that I hadn’t used pre-punched catalogue cards so I could claim a discarded library file cabinet to house them, forever, hauling the wall-sized furniture around the world with my increasing boxes of books and dead-tree teaching files . . . . (The cards would still be valueless but the cabinets have appreciated considerably.)

Optimistic for a lighter equivalent, I sketched out two projects on Trello, one where I created my own board, another where I used a template.  The former is for a manuscript that I’m revising, but where some of the material will need to be broken up into smaller sections or go into a different chapter.  I can see how the column-of-post-its design could help me imagine different sequences and move the bits and pieces accordingly.  The latter is a new project, based off a short presentation that I both need to expand and to write up in a different form.  Again, having a preexisting structure (in this case a sequence of slides), makes it easy to envision the project as a growing series of equal components or steps.

However, trying to use Trello in the early stages of a project, where I don’t know where I’m heading, may result in one endless data column that can only be sorted after I start understanding my claims.  Before that I just need somewhere to dump things and perhaps a DM program will just require unnecessary preemptive organization.  For example, the Trello templates are a wash.  There are far too many options, most of which are meant for non-academic projects or daily tasks.  And even if one opts for a very simple template, as I did for my second project, one wastes time eliminating features that one doesn’t want or need (inspirational photos of up-tilted faces for the “TODO” (sic) column or images of human cartwheels for “DONE”) .

None of this addresses the real challenge: envisioning the “deliverables” of a scholarly research project and approximating how long it will take to find the data to support one’s hazy vision of a possible conclusion or to replace it with something more plausible.

If a project starts with a question (i.e. “Might early modern women have become literate by sewing letters rather than writing them?”), one might be able to chart the areas where one should look and assign them to a timeline, but one could hardly work backwards from an anticipated answer (beyond ‘Yes/No”) through the stages of the project to a fixed start date.  A program like Zotero, that uses data that won’t get tossed, in a preexisting structure (i.e. a standardized bibliography of works consulted), might be a better tether for the amorphous beast of potentially useful/useless matter that is raw research.

Hence the short answer for me is that I’m not sure data management–at least as found in conventional data management programs–can improve my research practices.  However it certainly should help shape that accumulated research into a legible form by a fixed date–say–into a conference paper due March 15th?!  We’ll see . . . .

Freedom Of Speech: Data Management Plan

1. Data Collection

  • What data will you collect or create?
    • Our metadata about court cases, and our initial filtering, comes from the Washington University in St. Louis (WUSTL) law school database, also known as The Supreme Court Database (SCD). The dataset is comprised of two .csv files that we have downloaded which cover: (1) cases up until 1945; and then (2) post-1945. These two .csv files have been stitched together in a chronological manner. Each case then has the actual, court-published opinion texts as well, scraped from the Justia website and attached to the dataset.
  • How will the data be collected or created?
    • After having downloaded our data from the SCD, the next step is to filter by First Amendment. The download produced a larger dataset than we are actually interested in, so the next step is to filter for cases that address the Freedom of Speech specifically. These cases will be taken from the book Landmark Supreme Court Cases: The Most Influential Decisions of the Supreme Court of the United States (Vol. 3. 2nd ed.) by Richard A. Leiter and Roy M. Mersky. In its complete form, the data will be a JSON file with nested features.
  • Is it possible to regenerate the data? What are the implications for your research if the data are lost or become unusable later?
    • It is possible to regenerate our data by recreating the steps mentioned above and rescraping Justia for the opinion texts. We have backups of all of the processes we have created in order to get our final, usable dataset.
  • What are the tools or software you will be using to create/process/analyze/visualize the data?
    • We will be using a Jupyter Notebook to create, process, and analyze the data. Some visualizations will be sketched out using a mixture of matplotlib/seaborn for initial topic modeling/text analysis, and then reimagined in d3.js.

2. Documentation and Metadata

  • What documentation and metadata will accompany the data?
    • We will use README.md files to document our data’s features and the processes we work with. These processes include: combining SCD’s codebook with the codebook we create for the features we generate in the process of building out the existing database; processing documentation of cleaning, scraping, and filtering; and linking to SCD, Justia, and the textbook we’re using for analysis purposes (Landmark Supreme Court Cases). 
    • In order to ensure good project and data documentation, we will be taking notes throughout the creation, cleaning, and processing phases and then compiling those notes into README files.
    • Each of us, with respect to our roles in the project, will be responsible for an aspect of data management.
    • We will use categorical descriptive terminology in order to name our directories and files, and because our data is already mostly created, we will be following the standards of the already-created datasets.

3. Ethics and Legal Compliance

  • How will you manage any ethical issues?
    • Managing ethical issues largely entails deferring to experts when possible—especially through Landmark Supreme Court Cases—rather than deciding ourselves what “counts” as a landmark case. We will also aim to be transparent about where our data is from and any decisions we have made in curating it. 
  • How will you manage copyright and Intellectual Property Rights (IP/IPR) issues?
    • The SCD allows us to download and transform the data, with attribution. Regarding the Landmark Cases textbook, we will be reaching out to the publisher in order to request permission to use the content therein.

4. Storage and Backup

  • How will the data be stored and backed up during the research?
    • To use Stephen Zweibel’s language, we will use Github for “far”, flash drives for “near,” and our computers for “here.” Doing this should allow for sufficient storage. Regarding backup, each of us will back the data up locally, as well as on flash drives just in case.
  • How will you manage access and security?
    • Our Github repository/organization is closed to the public in terms of who can edit and manage the repository, so only our group can control it. The repository can still be forked, but the root data will not be mutated.

5. Selection and Preservation

  • Which data are of long-term value and should be retained, shared, and/or preserved?
    • After cleaning, processing, and finalizing the dataset, we will be keeping it in our Github repository. We will be keeping it there for the foreseeable future.
  • What is the long-term preservation plan for the dataset?
    • Github will also be used for long-term preservation.

6. Data Sharing

  • How will you share the data?
    • We will make the data immediately available upon upload, via the Freedom of Speech* site. The data itself will be a public, raw dataset available through Github.
  • Are any restrictions on data sharing required?
    • Our case data is in the public domain, but we may have to restrict our data if we are given a limited license by the publishers of the Landmark Cases textbook.
    • **need to check copyright**
  • Who is your possible audience? Who may use the data now, or later?
    • Our possible audience would be: any audience interested in First Amendment rights; law students who specialize in Constitutional Law; and members of the general public who want to know more about the fundamentals of legal studies; we also hope to reach a more “casual” or amorphous audience on social media platforms like Twitter, where Freedom of Speech is a relevant/hot topic.
  • What tools/software are required to access your data?
    • No special tools or software are required to access the data; if the user has a web browser, they can access it through our Github page, which will open the data in a new window.

7. Responsibilities and Resources

  • Who will be responsible for data management?
    • Each member of our team will exercise due diligence in implementation of the standards outlined above.
  • What resources will you require to deliver your plan?
    • We have all the resources we need in order to deliver our plan: Discord for communication, Jupyter Notebooks + Python for code and processing, Figma for design ideation and UX/UI wireframing/prototyping, Observable Notebooks for prototyping visualizations in d3, and Github for storage and hosting.

Data Redundancy, Data Management Roles, and the DMP

It’s one thing to lose a computer file or for that matter any hard-copy documents related to one’s personal affairs.  It’s another thing to lose files and data which other people depend on.  I’m reminded of the fires that destroyed the warehouses of Universal Studios in the San Fernando Valley in 2008. According to news reports, original master recordings of over 800 recording artists were destroyed in whole or part, almost a complete who’s who of American popular music.  The fire reportedly destroyed recording data of rock-and-rollers like Buddy Holly, country icons like Dolly Parton, jazz virtuosos like Aretha Franklin, blues giants like Muddy Waters, among many others.  Clearly the first and foremost principles of data management deserve to be data preservation and data integrity.

Unfortunately the mindset needed for assuring the satisfactory exercise of these principles has often been adopted momentarily at best.  The destruction of many libraries throughout the ages suggest we have been doomed to repeating the same mistake, again and again.  Perhaps one of the many curses of the human condition is the tendency to substitute aspiration for principle.

information destruction infographic

source: Global Datavault, https://www.globaldatavault.com/blog/information-destruction-history/, accessed 03/09/2021

click image to enlarge

In terms of research itself, research practices tend to be most effective when information is available and easy to retrieve.  While the Internet has generally opened up previously closed avenues to knowledge and information, one of the drawbacks of using the Internet as a tool for research is the over abundance of data storage locations.  When research data is stored in more than a couple of locations, data becomes difficult to keep track of and use effectively.  Critical evidence can be lost not because the data has been erased but because its location is no longer known or insufficient cataloging occurred.  Using only one tool for research such as Zotero, in which data is stored and centralized in on place and cataloging can be semi-automated, mitigates the risk of loosing track of references, internet resources, and other research upon which one’s current research depends.

A second problem involves the double-edged nature of digital media.  As easy as it is to duplicate digital data, it can be just as easy to delete.  Depending on the time table of the research, internet pages can suddenly disappear without a trace, leaving the visitor watching in dismay a 404-page-not-found animation.  Depending on the value of the missing information, a search in the Wayback Machine may or may not yield the version of the page initially accessed.  To some extent the ephemerality of the web poses a serious risk to the quality of research.  As a result, an additional evaluation of internet data comes into play in terms of determining the need for redundant data preservation by downloading the web page.

While automated redundancy more broadly has gotten better since the early days of the Internet, we are still not at that point when all applications automagically save all significant versions into a 100% redundant versioning system.  Total and automated redundancy goes against the right to be forgotten, the right to be anonymous, and the right not to be tied to a moment in the past.  Control over the degree and nature of redundancy to some extent offers freedoms at the cost of the discipline and responsibility to judiciously save a version.

As research goes beyond the work of a single individual to encompass group collaboration, effective data management takes on an even greater importance.  Without clear roles for  managing data, the likelihood of encountering problems will persist.  Rock-solid technologies may be in place, but if the responsibilities for the management of data are not clearly defined and assigned, state-of-the-art storage technologies will not by themselves prevent data loss or the reduction of data integrity.

Protecting Our Students Through Careful Data Collection

Because my group will be handling the sensitive personal information of students under 18, thinking about the safe collection, maintenance, and storage of our data has been top of mind. We want to be smart about what information we collect from students during their submissions and not just collect personal information for the sake of it. After chatting with Rebecca Banchik last week, we were relieved to hear that we wouldn’t need to go through the IRB process. The meeting was not only productive in relieving us of the potential IRB complications and delays, but also because Rebecca was able to help us work through what kind of information we would be collecting and how we would display it.

We knew that we would require parental approval in addition to the student consent, and Rebecca confirmed that. However she explained that we have a few options for displaying student information on the site. We could give students the option to use a pseudonym or only display their first name rather than first and last names. We could also keep the geographic location very broad at a state or country level, rather than zoom into a smaller school district. She explained that we want to follow the HIPAA guidelines: any details that would fall under that HIPAA umbrella (zip code, last name, date of birth) are off limits. We don’t want anyone to be able to plug a few of these personal identifiers into a search engine and find the student. This level of flexibility in how we present the student’s info as well as clear language on exactly how we will be using their info will need to be presented to the student and parent.

Needless to say, the consent/submission form has been our biggest hurdle to work through so far. Maggi set us up with a wonderful draft that she based on her own professional experience working with obtaining permission. We made some tweaks, including adding in our questions/prompts to spark student inspiration, and are using this in our first phase of outreach this week. We decided to use Google Forms to gather the student info which will then be automatically populated into a spreadsheet for us. This makes it easy to manage and reference the gold source list of entries all in one place. Phil brought up the excellent point that we’ll need to be able to migrate this list once the project moves beyond this semester’s prototype phase. Learning about the care needed here has been enlightening, and it’s certainly made me a better researcher.

In addition to this front end data collection, we have a lot of behind the scenes data decisions to make in regards to our categorization and metadata practices. As someone who spends professional time to come up with the best labels for audience and product tags in the content I manage, this kind of organization work really excites me! Consistency will be key here. We’ll need to come up with a clear labeling system for each submission type (video, poem, audio interview) along with other identifiers like geography (state country) and student grade level. That way entries are searchable not only by us, but also by the user via our project’s search bar.

Leslie pointing at organized piles of binders

We’ve made really fantastic progress so far with all of these careful data considerations, and I feel good knowing my team is up for handling the next sticky decision to work out.

Personal Blog: The First DMP Is the Deepest

Our deliverables these past two weeks have been very challenging to me. I really didn’t want to deliver them. But having gone through the process of creating a Work Plan and a Data Management Plan (DMP), I get it. It’s so much more fun having ideas and having your head in the clouds, but, again, I get it.

In my other class, we’ve been doing related work in reviewing other digital projects. To do so, we’ve been following a template put forth by Miriam Posner (see her video How Did They Make That), in which she asks: (1) What are the sources (or data)? (2) What did they do to them/how was the data processed? And (3) how is the project presented? I’m very glad to have watched this video before working on the DMP for this class because it helped me understand how vast a category “data” really is and to start thinking about it more expansively–beyond numbers and calculations.

Having now completed my first DMP, I have to believe future ones will be easier, or if not easier then at least feel more approachable. And I am optimistic this process will change how I write a proposal in the first place. With a better sense of the data I want/am able to collect, I think I will be able to start with stronger research questions.

It was also challenging thinking about how our group data will be stored and maintained–servers seem so far away and their capacity limitless. But I’m learning this is very far from true. So how long will our data be stored? We wrote 3 to 5 years, but in my heart I wrote for-ev-er.

Gif of Officer Saying For-Ev-Er

Elena’s Journal – Week 6 (DMP Considerations)

Compared to writing the Work Plan, working on the DMP was relatively easy. Maybe it’s because the Work Plan was a pain to write…or maybe it’s because sitting down and talking about data feels more concrete than trying to figure out the steps of an ever-changing project.

Stephen Zweibel’s presentation was illuminating: it helped me and my team figure out what constitutes our data and how we want to treat collect, process, and preserve it. The actual process of writing our DMP happened two days later when my team and I met over Zoom and had a very productive (and nerdy) conversation about our data. It got pretty intense, but in a good way: Montage and I found ourselves discussing every possibility and ended up writing a detailed plan that considers our Fridge Data, our Collaborators’ personal information, and the crowdsourced materials that we will collect in the Archive. Each of these kinds of data comes with its own advantages and challenges, so we had to consider what to do with each of them and the DMP was three times as long as we had envisioned in the beginning. Luckily Jean stepped in at the end and edited the text so that it was legible for humans!

Now that we have a Work Plan, a Project Timeline, and a DMP I feel more ready than ever to tackle every step of this project and to help my team reach its goals. Moreover, writing this DMP gave me a more structured thought process when considering research data and archival materials. The questions we had to answer guided me in thinking about the short- and long-term preservation of data, which is something that I had always intuitively considered: now I finally have the vocabulary to talk about it. Another benefit of working on this DMP was that I realized it isn’t that hard to do it if you know your project well enough. In the future, I will probably apply for grants for DH projects, and it is good to know that I can draft a DMP over the course of an afternoon, with the help of my team members.