The George Washington University Research Data Management Survey
Report date: May 19, 2026
The Research Data Management Task Force
Current members:
- Emily Blumenthal, Libraries and Academic Innovation
- Clark Gaylord, Research Technology Services, GW Information Technology
- Brie McDonald, Himmelfarb Health Sciences Library
- Brittany Smith, Himmelfarb Health Sciences Library
- Jennifer Strickland, Office of the Vice Provost for Research
- Lowell Williams, GW Information Technology
Former members:
- Tima Atie, Data Governance, GW Information Technology
- Zoë Hammatt, Office of the Vice Provost for Research
- Kelly O’Keefe, Office of the Vice Provost for Research
- Lydia Payne-Johnson, Data Governance, GW Information Technology
Contents
- Executive Summary
- Advancing Research Data Stewardship at GW: Context and Institutional Response
- Methods: The GW Research Data Management Survey
- Results: Research Data Management Landscape at GW
- Discussion
- Recommendations
- References
- Appendices
- Appendix A: The Research Data Management Task Force Charge
- Appendix B: GW Research Data Survey, Phase I: Data Access, Storage, and Preservation Needs
- Appendix C: Respondents and Responses by GW School and Department
- Appendix D: Survey Response Distributions by School and Question
- Appendix E: GW Resources for Research Data Management
Executive Summary
As part of ongoing efforts to better understand and support research data practices at GW, the Research Data Management Task Force conducted a university-wide survey to assess how researchers at the George Washington University manage, store, and preserve their data. The survey gathered both quantitative and qualitative input from a diverse group of researchers across disciplines.
The findings highlight a gap between the varied and complex research data produced at GW and the current mix of systems used to manage, store, and preserve those data. While practices vary across individuals and disciplines, many researchers rely on decentralized or external solutions, particularly when working across institutions. At the same time, there is strong interest in a GW-supported research data repository or metadata catalog, suggesting a shared recognition of the potential value of more coordinated infrastructure. Adoption, however, will depend on addressing concerns related to sustainability, institutional reliability, and long-term access.
The results also point to an opportunity to improve institutional visibility into research data. Limited insight into where data are stored and how they are managed may present challenges for compliance, particularly for regulated or export-controlled data. Responses further indicate a need for clearer guidance around data ownership, classification, and stewardship, as well as expanded training to support consistent and responsible data practices across the research lifecycle.
Taken together, these findings suggest a phased and collaborative approach to strengthening research data stewardship at GW. Near-term efforts should focus on expanding researcher education, clarifying institutional roles and expectations, and engaging schools and units (such as through focus groups coordinated by Associate Deans for Research) to refine data management and infrastructure requirements and to ensure alignment with disciplinary needs.
Over the longer term, a structured evaluation of research data repository and metadata catalog options will be critical to identifying sustainable, scalable solutions.
Advancing this work directly supports GW’s strategic priorities, particularly Priority One: Generate Scholarship with Impact and Priority Three: Strengthen Our Foundation for Excellence, by enhancing the visibility, integrity, and long-term impact of GW research while building the infrastructure and institutional capacity necessary to sustain a growing and increasingly complex research enterprise.
Advancing Research Data Stewardship at GW: Context and Institutional Response
In recent years, expectations for research data management, sharing, and long-term stewardship have increased substantially. Federal agencies and global research organizations increasingly require that research data be treated as a durable, accessible scholarly output. Policies such as the memorandum from the White House Office of Science and Technology Policy (2022) and updates to the National Institutes of Health Public Access Policy (2025) reflect a broader shift toward open science, requiring that data be discoverable, well-documented, and preserved for reuse. At the same time, recent disruptions affecting federal data infrastructure have underscored the vulnerability of critical datasets and the growing importance of institutional stewardship.
In response to these developments, the George Washington University (GW) convened a Research Data Management (RDM) Task Force, following participation in the summer 2024 Summit for Academic Institutional Readiness in Data Sharing (STAIRS). This group brought together representatives from Libraries and Academic Innovation, Himmelfarb Health Sciences Library, Research Technology Services, Information Security, Data Governance, and the Office of the Vice Provost for Research to establish a baseline understanding of research data practices and assess institutional needs. The RDM Task Force was tasked with assessing data infrastructure to help identify pathways to fulfilling the objectives of the GW Strategic Framework (2025). For example, expanding GW’s global research footprint (Priority One) is increasingly dependent on accessible, well-managed data. By identifying existing gaps in infrastructure, the Task Force provides information necessary to support the systems, expertise, and capacity for innovation called for in Strengthening Our Foundation for Excellence (Priority Three).
For research-intensive universities, these developments represent both a compliance requirement and a strategic opportunity. Evidence from the scholarly literature demonstrates that accessible and well-curated research data are associated with increased citation rates and broader research impact (Colavizza et al., 2020; Piwowar & Vision, 2013). The work of the Task Force is a critical first step in ensuring that GW can leverage these practices to enhance research visibility, support reproducibility, and secure the long-term impact of its scholarship.
As the Task Force began its work, the broader data environment continued to evolve. Emerging threats to federal datasets and infrastructure highlighted the fragility of research data and reinforced the need for coordinated institutional action. In response, the Task Force temporarily expanded its focus to support data preservation efforts and community education. This work demonstrated both the value of cross-campus collaboration and the risks associated with gaps in infrastructure.
Following this initial phase, the Task Force reconvened with updated membership and a refined scope focused on the distinct policy, compliance, and infrastructure needs of scholarly output. With this renewed focus, the Task Force returned to its core objective: to assess current practices, identify gaps, and develop evidence-based recommendations regarding institutional support for research data management, including the potential need for an institutional research data repository or catalog (see Appendix A).
Research data repositories and data catalogs are widely adopted mechanisms for enabling systematic data stewardship, discovery, and long-term preservation. A research data repository provides a secure institutional platform to store, manage, and share datasets (GFZ German Research Centre For Geosciences et al., 2013; NNLM, 2026b; OpenAIRE, n.d.), while a data catalog serves as a metadata-driven index that enhances discovery of data regardless of where it is stored (NNLM, 2026a). These systems are complementary and together form the foundation of a coordinated institutional data ecosystem (Cox et al., 2017).
Across peer institutions, investment in such infrastructure is now standard. Among Association of American Universities (AAU) institutions, approximately 80% have implemented either a research data repository, a data catalog, or both (Blumenthal, 2026). These systems provide several key benefits:
- Funder Compliance: Institutional data infrastructure allows researchers to comply with funder data management requirements. A robust research data management ecosystem, characterized by a data repository or data catalog, can support research data being discoverable and accessible (White House Office of Science and Technology Policy (OSTP), 2022).
- FAIR Principles: Data stored in a data catalog or repository enables researchers to meet FAIR principles (Findable, Accessible, Interoperable, Reusable), which are becoming the global standard for quality research data management (Rehnert & Takors, 2022).
- Data Security & Inventory: Institutional data infrastructure creates a necessary inventory of GW’s research data assets, allowing the institution to understand its risk exposure and meet mandates for data security, even if the data are hosted elsewhere.
- Enhance Research Visibility and Impact: Consistent metadata creation and preservation workflows lead to more discoverable research. This can increase citations of GW publications and foster new collaborations (Colavizza et al., 2020; Dellureficio, 2021).
- Long-Term Preservation: Institutional data catalogs/repositories support the long-term integrity and accessibility of valuable research outputs that often form the basis of future grant proposals and scholarly output (Narlock et al., 2024).
Beyond supporting individual researchers, these investments generate broader institutional advantages. Coordinated data infrastructure strengthens research integrity and reproducibility, reduces the risk of data loss or noncompliance, and enables more efficient use of research investments through data reuse (Piwowar & Vision, 2013). It also builds institutional capacity by centralizing expertise in data stewardship and fostering collaboration across libraries, information technology, and research offices, directly supporting GW’s goal of strengthening the systems and partnerships that underpin research excellence (Akers & Doty, 2013; Cox et al., 2017).
Together, these trends and institutional priorities underscore the need for a coordinated, evidence-based approach to research data stewardship at GW. As the first step in this process, the Task Force sought to understand how GW researchers are currently managing, storing, and sharing their data; what challenges they face; and what types of institutional support would be most valuable.
To address these questions, the Task Force developed and administered a university-wide survey focused on research data access, storage, and preservation. The findings from this survey provide the foundation for the analysis and recommendations that follow.
Methods: The GW Research Data Management Survey
The RDM Task Force conducted a formal assessment of institutional research data practices between October 2025 and February 2026. The primary objective was to understand the types of digital research data being generated, current storage practices, and the potential value of a centralized institutional data repository or metadata catalog.
The survey was designed to capture the research data lifecycle across four key areas. First, participants identified the types of digital research products they generate, including structured datasets, text-based materials, software code, multimedia, and lab documentation. Second, the survey assessed data classification based on security and privacy requirements, using the Institutional Data Classification Guide to categorize data as Public (freely available), Restricted (confidential/internal use only based on business need), or Regulated (legally protected with strict controls on use). Third, respondents described their current storage and preservation practices, including the use of institutional platforms, discipline-specific repositories, and local hardware. Finally, the survey explored the potential utility of a centrally maintained research data repository or metadata catalog and included an open-ended section for participants to share questions, concerns, or support needs related to data sharing and preservation (see Appendix B for the full survey).
The survey was promoted through several high-visibility channels, including the university-wide Infomail and newsletters from Himmelfarb Health Sciences Library, Libraries and Academic Innovation, and the Office of the Vice Provost for Research. To broaden reach across disciplines, Associate Deans of Research (ADRs) conducted targeted outreach within their respective schools and units. This approach yielded 121 respondents representing seven schools and more than 30 departments across the university (see Appendix C for a complete breakdown of respondents by school and departmental affiliation).
Figure 1. Number of Respondents by GW School
Respondent distribution across GW schools, with the largest representation from the Columbian College of Arts and Sciences and notable participation across multiple schools, alongside a portion of respondents who did not indicate an affiliation.
Results: Research Data Management Landscape at GW
Data Diversity and Classification
Researchers at GW produce a wide range of digital materials. Over half of respondents reported creating multiple types of content – including documentation, multimedia, structured datasets, and text-based materials – with an average of 2.68 data categories per person.
Nearly half of respondents also work with multiple data classification types simultaneously. While most primarily handle public data, a substantial portion manages sensitive information; 77 respondents explicitly reported working with regulated or restricted data. Qualitative responses highlight the specific needs of this group:
- “Security for protected data (make sure not searchable)”
- “I need an easily searchable database that can then be deidentified for sharing.”
Figure 2. Digital Research Products Produced
Types of digital research products created by respondents, highlighting a wide range of outputs, including structured datasets, text-based materials, multimedia, code, and documentation.
Figure 3. Data Classification
Distribution of how respondents classify their data, with most identifying their data as public or restricted and fewer reporting regulated or other classifications.
Current Storage and Preservation Practices
Data storage practices of GW researchers are highly fragmented. Most researchers use a combination of storage solutions, while a small subset (10 respondents) relies exclusively on personal platforms or local hard drives. Three respondents initially reported that their data was not preserved at all.
This fragmentation is often driven by the need for stability outside institutional systems (n=6), particularly for multi-institutional projects:
- “It was pretty hard to get clear answers to questions about where we could store documents and data in an ongoing way throughout the data collection process across multiple universities...”
- “I used to use my university repository at my last university but switched to using Zenodo instead since it is not uncommon for faculty to move universities over time...”
Figure 4. Current Data Storage
Where respondents currently store or preserve their data, showing a mix of institutional, personal, and repository-based solutions, with relatively few indicating data are not preserved or shared.
Institutional Demand for a Data Repository or Catalog
Support for a centralized GW Data Repository or Metadata Catalog is strong. A large majority of respondents (87%; 105 of 121) indicated they would use such a service.
Figure 5. GW Data Repository or Metadata Catalog Uses and Perceptions
Ways respondents currently use — or would consider using — the GW data repository, with strong interest in long-term preservation, access, and data management support services.
Demand by School and Discipline
Interest is consistent across disciplines, with several schools reporting 100% interest in using an institutional data repository for long-term preservation:
- School of Medicine and Health Sciences: 18/18 (100%)
- School of Engineering and Applied Science: 5/5 (100%)
- Law School: 3/3 (100%)
- School of Nursing: 2/2 (100%)
Other schools also demonstrated strong interest:
- Columbian College of Arts and Sciences: 34/46 (74%)
- Milken Institute School of Public Health: 5/7 (71%)
See Appendix D for a more detailed breakdown of responses to the survey questions by the respondent’s GW school affiliation.
Primary Use Cases and Priorities
Respondents anticipate using an institutional data repository for multiple purposes, selecting an average of 2.5 functional needs. Among the 24 respondents who identified a single primary use, priorities were:
- Long-term preservation: 19 respondents
- Access and discovery: 3 respondents
- Data description: 1 respondent
- Other uses: 1 respondent
Barriers to Adoption and “Would Not Use” Responses
Only 13% of respondents (16 of 121) indicated they would not use an institutional data repository. Qualitative responses (n=6) and general concerns (n=12) reveal three primary barriers:
- Disciplinary Preference: “I am happy with the other data repositories that I use, and would feel like my data [are] less discoverable in a GW-specific data repository.”
- Institutional Trust: “In the 13 years I have been at GW, the stability and turnover in policy and implementation of GW managed IT resources has been ‘turbulent’ at best. I would not trust critical data to this internal GW organizational unit.”
- Technical Limitations: “The [current] environment... does not allow storing information in several file formats, including djvu. That means that I was unable to fulfil my NFS (sic) grant commitment...”
Infrastructure Needs and Educational Gaps
Respondents identified both technical and educational needs for successful implementation. Twelve respondents provided a “wish list” focused on system performance:
- “Data uploading and downloading speeds are currently very slow, which significantly affects efficiency when managing and sharing large research datasets.”
Additionally, seven respondents emphasized the need for training and guidance:
- “Is there training available for new faculty on database best practices and tips & tricks for using GW-sanctioned data tools?”
- “I’d like clearer guidance on best practices for long-term storage and accessibility, especially for large or multimedia datasets.”
Discussion
General Observations
Adoption of a GW research data repository or metadata catalog appears to be driven primarily by current storage practices and data sensitivity. Interest is consistently high across user groups, including those already using institutional platforms and those relying on siloed or personal storage. This suggests that a centralized GW solution could serve as a unifying, user-friendly hub to reduce fragmentation.
However, data sensitivity introduces an important nuance. While demand is strong among researchers working with public and restricted data, those handling regulated data may require additional assurances related to encryption, compliance, and security. Addressing these concerns will be critical to building institutional trust.
Observations by School
Cross-school comparisons were limited due to small sample sizes in some schools (see Appendix C for response rates and breakdown by survey question). However, responses generally aligned with overall positive sentiment toward a repository or catalog. Schools such as the Law School, School of Engineering and Applied Science, and School of Medicine and Health Sciences skewed toward support, while responses from the Milken Institute School of Public Health and Columbian College of Arts and Sciences were more evenly split between support and non-use.
Concerns most widely shared across schools related to security/data integrity (12 responses) and institutional boundaries (6 responses), with some overlap between these categories (3 respondents). These concerns often focused on long-term preservation, protection from internal or external interference, and access to data after leaving GW. They also intersected with questions about institutional capacity, including staffing and infrastructure, and whether supporting external repositories might be more effective.
Technical requirements were among the most frequently cited needs (12 responses), particularly among respondents from CCAS and SMHS. This may indicate the need for targeted outreach to better understand and support discipline-specific requirements. Although fiscal concerns were rare, both instances were raised by GWSPH faculty and focused on potential costs to PIs and the level of institutional financial support.
Other Correlations
Some additional patterns emerged across responses. Researchers currently using external repositories were more likely to express concerns about access after leaving GW and challenges with cross-institutional collaboration. Similarly, respondents who indicated they would not use a GW solution disproportionately cited concerns about institutional boundaries, suggesting this as a key area for improving adoption.
Finally, use case preferences may shape attitudes toward adoption. Respondents interested in metadata catalog functions (e.g., access and discovery, description, inventory tracking) were slightly more likely to express reluctance to use an institutional platform, while those prioritizing repository functions (e.g., long-term preservation, security monitoring) were more likely to express support. This suggests that additional education on the value of an institutional metadata catalog might lead to increased adoption by users who are currently unclear about how it might address their expressed needs.
Conclusions
The findings from the GW Research Data Management Survey highlight both strong demand for coordinated institutional support and a gap between the diverse and complex data produced by researchers at GW and the current mix of systems used to manage, store, and preserve those data. While researchers are working across multiple data types and classifications, existing practices are often decentralized and vary widely across individuals and projects.
At the same time, there is clear interest in a GW-supported research data repository or metadata catalog, suggesting a shared recognition of the value of more coordinated institutional support. However, responses also indicate that adoption will depend on addressing key concerns related to sustainability, cost, institutional boundaries, and long-term reliability. Patterns in the responses further suggest that trust and clarity will be critical factors in future implementation. Researchers currently using external repositories emphasized the importance of maintaining long-term access to their data, while others expressed uncertainty about institutional capacity or continuity, pointing to the need for transparent planning and communication.
A metadata catalog deserves particular attention in this context. Unlike a repository, which addresses the actual storage of research data, a catalog creates institutional visibility into the research data landscape regardless of storage location. This makes the catalog the foundational investment, enabling compliance monitoring, supporting data governance, and positioning GW to build additional infrastructure incrementally and in response to demonstrated needs. The survey findings, especially the noted fragmentation of current storage practices and challenges working with protected data, underscore the value of this approach.
Data sensitivity adds an additional layer of complexity. While many researchers appear ready to engage with institutional solutions, those working with regulated or export-controlled data require clear assurances regarding security, compliance, and access controls. It is also important to note that response rates varied across schools, and individuals with strong perspectives on data management may have been more likely to participate. While this may influence some findings, the consistency of key themes across responses suggests that the results reflect meaningful institutional patterns.
Taken together, these findings point to a clear opportunity for GW to strengthen its research data infrastructure in ways that are responsive to researcher needs while aligned with institutional priorities. Advancing this work represents a strategic opportunity to enhance scholarly impact, strengthen the university’s research enterprise, and build the institutional foundation necessary to support long-term research excellence.
Recommendations
Consistent with GW’s strategic framework, particularly Generate Scholarship with Impact (Priority One) and Strengthen Our Foundation for Excellence (Priority Three), the following recommendations are designed to address our findings through a coordinated, phased, and sustainable approach to research data stewardship at GW.
Conduct Targeted School- and Discipline-Level Engagement
While interest in institutional data infrastructure is widespread, needs and concerns vary across disciplines, data types, and compliance environments. Targeted engagement will be essential to ensure that solutions are responsive, scalable, and broadly adopted.
Recommendation
Conduct school- and unit-level focus groups, coordinated through Associate Deans for Research (or equivalent leadership), to:
- Validate survey findings at the local level
- Identify discipline-specific requirements (e.g., sensitive data, large-scale computing, external repositories)
- Surface barriers to adoption, including concerns about sustainability, trust, and workflow integration
Strategic Alignment
This approach supports Priority One by enabling interdisciplinary and domain-specific research needs to be addressed effectively, and Priority Three by strengthening institutional responsiveness and collaboration.
Develop a Strategic Roadmap for Research Data Infrastructure
Survey results indicate a clear need for coordinated, institutionally supported infrastructure, including potential investment in a research data repository, data catalog, or hybrid model. However, long-term success will depend on thoughtful planning, sustainable resourcing, and alignment with existing systems and services.
Recommendation
Develop a phased institutional roadmap that:
- Prioritizes metadata catalog infrastructure as a foundational layer from which repository and storage decisions follow
- Evaluates repository, catalog, and hybrid infrastructure models
- Assesses integration with existing platforms (e.g., Box, Research NAS, external repositories, compliance systems)
- Identifies governance, staffing, and funding models
- Establishes sustainability and service ownership across units
Strategic Alignment
This approach directly advances Priority Three by building the systems and infrastructure necessary to support scholarly excellence and innovation, while enabling Priority One through improved research outputs and impact.
Strengthen Institutional Capacity for Data Stewardship and Support
Effective research data management requires not only infrastructure but also coordinated expertise, training, and support across the research lifecycle. Survey findings suggest variability in researcher knowledge of best practices, funder requirements, and available services.
Recommendation
Expand and coordinate research data support services by:
- Enhancing training on data management planning, FAIR principles, and data sharing requirements
- Increasing visibility and accessibility of existing services (e.g., Libraries, IT, Office of Research)
- Building cross-unit expertise through shared staffing models or communities of practice
- Embedding data support into existing research workflows (e.g., grant development, IRB processes)
Strategic Alignment
This approach supports Priority Three by strengthening institutional expertise and service delivery, and Priority One by enabling high-quality, reproducible, and impactful research.
Conclusion
Together, these recommendations position GW to build a coordinated, sustainable, and strategically aligned approach to research data stewardship, one that strengthens the university’s research enterprise, enhances scholarly impact, and supports its long-term institutional goals.
References
- Akers, K. G., & Doty, J. (2013). Disciplinary differences in faculty research data management practices and perspectives. International Journal of Digital Curation, 8(2), 5–26. https://doi.org/10.2218/ijdc.v8i2.263
- Blumenthal, E. (2026). Research Data Repositories and Metadata Catalogs. The George Washington University. https://doi.org/10.17605/OSF.IO/54KFX
- Colavizza, G., Hrynaszkiewicz, I., Staden, I., Whitaker, K., & McGillivray, B. (2020). The citation advantage of linking publications to research data. PLOS ONE, 15(4), e0230416. https://doi.org/10.1371/journal.pone.0230416
- Cox, A. M., Kennan, M. A., Lyon, L., & Pinfield, S. (2017). Developments in research data management in academic libraries: Towards an understanding of research data service maturity. Journal of the Association for Information Science and Technology, 68(9), 2182–2200. https://doi.org/10.1002/asi.23781
- Dellureficio, A. (2021). Data Catalogs, Metadata, and Data Discovery: How data catalogs extend and enhance the IR ecosystem. Medical Institutional Repositories in Libraries (MIRL). https://hsrc.himmelfarb.gwu.edu/mirl/2021/program/26
- GFZ German Research Centre For Geosciences, Humboldt-Universität Zu Berlin, Germany Karlsruhe Institute Of Technology (KIT), Purdue University Libraries, Bertelmann, R., Buys, M., Cousijn, H., Dierolf, U., Elger, K., Fenner, M., Ferguson, L. M., Fritze, F., Fuchs, C., Goebelbecker, H.-J., Gundlach, J., Kindling, M., Kloska, G., Klump, J., Kramer, C., … van de Sandt, S. (2013). Registry of Research Data Repositories. DataCite. Registry of Research Data Repositories. https://doi.org/10.17616/R3D
- GW Strategic Framework | The George Washington University. (2025, October 21). GW Strategic Framework. https://strategicframework.gwu.edu/
- Narlock, M. R., Calvert, S., Taylor, S., Marquez, R. P., & Parkman, A. (2024). Knowledge Infrastructures Are Growing Up: The Case for Institutional (Data) Repositories 10 Years After the Holdren Memo. Data Science Journal, 23(46). https://doi.org/10.5334/dsj-2024-046
- NIH Public Access Policy Overview | Grants & Funding. (2025, June 23). https://grants.nih.gov/policy-and-compliance/policy-topics/public-access/nih-public-access-policy-overview
- NNLM. (2026a, February 20). Data Catalog. NNLM. https://www.nnlm.gov/resources/data-glossary/data-catalog
- NNLM. (2026b, February 20). Repository. NNLM. https://www.nnlm.gov/resources/data-glossary/repository
- OpenAIRE. (n.d.). What are repositories? OpenAIRE. Retrieved March 20, 2026, from https://www.openaire.eu/where-can-i-read-more-about-fp7
- Piwowar, H. A., & Vision, T. J. (2013). Data reuse and the open data citation advantage. PeerJ, 1, e175. https://doi.org/10.7717/peerj.175
- Rehnert, M., & Takors, R. (2022). FAIR research data management as community approach in bioengineering. Engineering in Life Sciences, 23(1), e2200005. https://doi.org/10.1002/elsc.202200005
- Summit for Academic Institutional Readiness in Data Sharing (STAIRS). (2024, August 5). Data Curation Network. https://datacuration.network/summit-for-academic-institutional-readiness-in-data-sharing-stairs/
- White House Office of Science and Technology Policy (OSTP). (2022). Desirable Characteristics of Data Repositories for Federally Funded Research. Executive Office of the President of the United States. https://doi.org/10.5479/10088/113528
Appendices
Appendix A: The Research Data Management Task Force Charge
The revised RDM Task Force Charge is as follows:
- Building on the success of GW’s participation in the STAIRS Summit, in Academic Year 2024-25, a Research Data Management Task Force will work to engage with GW faculty and researchers in order to establish a baseline understanding of current practices related to the management, preservation, and open sharing of research data at GW.
- Create an engagement plan to survey GW faculty and researchers to gather information including:
- What are faculty/researchers doing with their research data?
- How are they making those decisions?
- What are their needs and pain points?
- What is their understanding of funder requirements and compliance?
- Identify colleagues to help implement this plan and assist with data collection and analysis.
- Answer the question and make a recommendation about whether GW needs an Institutional Data Repository (IDR).
- Document and communicate findings and recommendations to stakeholders.
Appendix B: GW Research Data Survey, Phase I: Data Access, Storage, and Preservation Needs
The GW Research Data Survey, Phase I: Data Access, Storage, and Preservation Needs aims to assess faculty data needs, acquisition behaviors, and storage practices, with particular attention to the vulnerability of research data. Specifically, the survey will identify the types of data faculty require, determine whether they have successfully accessed and downloaded these datasets, document where and how they store data, and evaluate whether these datasets are federally maintained or at risk due to accessibility, security, or preservation concerns.
Additionally, this survey will inform efforts to centralize and coordinate data preservation at George Washington University, ensuring long-term stability and increasing access to critical research data.
This survey consists of 13 questions and should take approximately 10-15 minutes to complete.
How Your Responses Will Be Used
The information gathered will be used to prioritize data backup efforts and determine how best to support long-term storage and accessibility of research data. Identifying information will not be collected unless you choose to opt in by sharing your email address. Only aggregate data will be shared outside of the members of the Research Data Management Task Force.
The Research Data Management Task Force
The Research Data Management Task Force is a collaboration between Libraries and Academic Innovation, Himmelfarb Health Sciences Library, GW Information Technology, and the Office of the Vice Provost for Research.
- What types of digital research data or products do you generate in your work? These might include anything you create, collect, or analyze as part of your research.
(Select all that apply and/or describe in your own words)- Structured datasets (e.g., spreadsheets, tables, survey results, including de-identified EMR files; genomic libraries, etc.)
- Text-based materials (e.g., interview transcripts, field notes, open-ended survey responses, including surveys in Qualtrics or REDCap)
- Images or multimedia (e.g., photographs, audio or video recordings, medical imagery, including NI Core, Flow Core, etc.)
- Software/code/scripts (e.g., analysis scripts, custom tools, GitHub projects, including NSQIP outcome data, analysis of medicare data, etc.)
- Models or simulations (e.g., statistical models, computational simulations, genomic pipelines)
- Documentation (e.g., codebooks, lab protocols, README files, LabArchives)
- Other (please describe)
- What is the data classification, based on security or privacy, of the digital research data you collect?
- Public – information intended for open access with no legal restrictions (e.g., public directories, deidentified survey responses, synthetic data for teaching)
- Restricted – data that must be protected due to institutional policy, contractual obligations, or privacy concerns (e.g., interview transcripts with potentially identifiable information, lab data under embargo until publication)
- Regulated – data protected by law or regulation, requiring strict controls (e.g., identifiable health data (HIPAA), student academic records (FERPA), biometric or genetic data, export-controlled technical specifications)
- Other (please describe)
- Where do you currently store your digital research data for sharing and/or long-term preservation?
(Select all that apply)- Institutional platform (e.g., GW Box, GW Drive, Research NAS)
- Discipline-specific repository (e.g., Dataverse, GenBank, ICPSR)
- Federal data repository (e.g., CDC WONDER, NIH dbGaP)
- General-purpose repository (e.g., Dryad, Figshare, OSF, QDR, Zenodo)
- Other repository mandated by funder/grant requirements, please describe
- Personal platform (e.g., your own Box, Drive, Dropbox)
- Local computer or hard drive
- Not currently preserved or shared
- Other (please describe)
- How would you use a GW research data repository or metadata catalog? (Not sure what terms like "data repository" or "metadata catalog" mean, or why this matters? See our FAQs for a quick explanation. Select all that apply and/or describe in your own words)
- Long-term preservation – I want to securely store my data for future access
- Inventory tracking – I need a central place to record where my data are housed (even if stored elsewhere)
- Access and discovery – I want myself and others to be able to find and request access to datasets (when appropriate)
- Describing my data – I want to add context like keywords, methods, or formats to make my data easier to understand
- Monitoring data security – I want to keep track of where my data are stored across platforms to help ensure appropriate protections are in place
- I would not use a GW research data repository
- Other: ___________
- What are your questions, concerns, or support needs related to research data preservation and sharing?
- (Optional) Would you like us to follow up with you? If yes, please provide your email address and a member of the RDM Task Force will follow up with you within 3 business days
(Refer to the Data Classification Guide for detailed definitions and examples. Select all that apply)
Appendix C: Respondents and Responses by GW School and Department
| School | Department | Number of Respondents |
|---|---|---|
| Columbian College of Arts and Sciences | 46 | |
| Art History | 1 | |
| Biological Sciences | 2 | |
| Chemistry | 4 | |
| Corcoran School of the Arts and Design | 3 | |
| Economics | 2 | |
| Mathematics | 1 | |
| Organizational Sciences and Communication | 1 | |
| Physics | 6 | |
| Psychological and Brain Sciences | 3 | |
| Public Policy and Public Administration | 1 | |
| Sociology | 1 | |
| Statistics | 1 | |
| No Department Selected | 20 | |
| Graduate School of Education and Human Development | 7 | |
| Counseling and Human Development | 2 | |
| Educational Leadership | 2 | |
| Human and Organizational Learning | 1 | |
| Other | 1 | |
| No Department Selected | 1 | |
| Law School | 3 | |
| Law | 3 | |
| Milken Institute School of Public Health | 7 | |
| Biostatistics | 1 | |
| Biostatistics & Bioinformatics | 1 | |
| Environmental & Occupational Health | 1 | |
| Exercise and Nutrition | 1 | |
| Global Health | 1 | |
| Prevention and Community Health | 1 | |
| No Department Selected | 1 | |
| School of Engineering and Applied Science | 5 | |
| Computer Science | 1 | |
| Engineering Management and Systems Engineering | 2 | |
| Other | 2 | |
| School of Medicine and Health Sciences | 18 | |
| Biomedical Laboratory Sciences | 3 | |
| Clinical Research and Leadership | 2 | |
| Health, Human Function, and Rehabilitation Sciences | 5 | |
| Other | 1 | |
| Physician Assistant Studies | 2 | |
| No Department Selected | 5 | |
| School of Nursing | 2 | |
| School of Nursing | 1 | |
| No Department Selected | 1 | |
| Other | 1 | |
| Other | 1 | |
| No School Selected | 32 | |
| No Department Selected | 32 | |
| Grand Total | 121 | |
Table 1. The number of respondents by their self-identified GW School and Departmental affiliation.
Figure 6. Proportion of “Yes” Responses by School and Question
Heatmap illustrating the share of respondents within each school who selected “Yes” for each survey question. Proportions are calculated relative to the total number of respondents from each school, enabling meaningful comparison despite differences in response volume.
Appendix D: Survey Response Distributions by School and Question
The following visualizations illustrate survey responses categorized by GW school affiliation. As this was an optional field, participants who chose not to self-identify are grouped under "No Response." Data are presented for the Columbian College of Arts and Sciences (CCAS), the Graduate School of Education and Human Development (GSEHD), the Milken Institute School of Public Health (GWSPH), the Law School (Law), the School of Nursing (Nursing), other affiliation not listed (Other), the School of Engineering and Applied Science (SEAS), and the School of Medicine and Health Sciences (SMHS).
Figure D1. Digital Research Products Produced by School
Bar charts illustrating the types of digital research products generated within each school, including datasets, text-based materials, code, multimedia, and documentation.
Figure D2. Data Classification by School
Bar charts showing how respondents within each school classify their data (Public, Restricted, Regulated, or Other).
Figure D3. Current Data Storage Practices by School
Bar charts showing how respondents within each school currently store and preserve their data, including use of institutional platforms, repositories, and local storage.
Figure D4. Perceived Uses of the GW Data Repository by School
Bar charts summarizing how respondents across schools report current or potential uses of a centralized institutional repository, including preservation, access, and data management support.
Appendix E: GW Resources for Research Data Management
In addition to the recommendations outlined in this report, the RDM Task Force identified a range of existing resources that support research data management practices across the institution. These materials include guides, training resources, service information, templates, and related documentation developed by campus units and partner organizations.
- LAI Data Services Resources
- Himmelfarb NIH Data Management & Sharing Plan (DMSP) Research Guide
- OVPR's Research Commons
- GW Data Classification Guide
The RDM Task Force recommends maintaining and expanding this resource collection as services and institutional needs evolve.