Paleo Data Portal Orientation
This page is intended to help data providers 1) determine whether the {% include resource_link filename=’pdp.yml’ %} will be an appropriate tool for managing and sharing data associated with their collection and 2) understand key information for leveraging the portal’s built-in tools for data mobilization.
Introduction
Thank you for your interest in using the Paleo Data Portal (PDP)A Symbiota portal that supports the management and sharing of data associated with paleontological specimens that are made available for research via permanent repositories. The scope of the PDP is limited to extinct organisms and their traces (i.e., fossils). Geological samples, archaeological and anthropological materials, as well as neontological specimen data fall outside this scope and should not be cataloged in this portal., a Symbiota-based data portal for managing and publishing fossil specimen data. This documentation is intended to help data providers 1) determine whether this resource will be an appropriate tool for managing and sharing data associated with their collection and 2) understand key information for leveraging the portal’s built-in tools for data mobilization.
Portal Scope
The Paleo Data Portal’s scope is described on the home page, https://paleo.symbiota.org. The portal is intended for managing and sharing specimen data from fossil collections that can be used for research and are held in public trust, for example, by a university collection or nonprofit museum. The portal’s scope is limited to extinct organisms and their traces (i.e., fossils) and collections that intend to use the portal for active data management. The Paleo Data Portal is not intended for use by teaching collections, collections that are not publicly accessible, geological samples that do not contain fossils, archaeological or anthropological materials, or neontological (extant) specimen data.
Portal Sustainability
The Paleo Data Portal’s underlying code and infrastructure are maintained by the Symbiota Support Hub (SSHSymbiota Support Hub. Group based at the University of Kansas that maintains the most commonly used version of Symbiota, as well as a number of Symbiota-based portals.) at the University of Kansas. As a result, the SSH incurs costs to make this resource available—for example, by responding to Help Desk tickets, developing and hosting webinars, improving the Symbiota code and fixing bugs, running data backups, and keeping the portal and its physical server secure. Therefore, while it is not required, data providers are strongly encouraged to budget support for the Symbiota Support Hub in funding proposals and/or your annual operating budget in a capacity is feasible for your collection. Please write to the SSH’s Help Desk if you would like to contribute funds toward routine portal maintenance costs via the service center, KU Symbiota, or in a forthcoming proposal. The SSH also accepts donations.
User Support
Technical
The Paleo Data Portal runs on Symbiota, which is open source software used for creating themed data portals to manage and share biodiversity specimen data. Symbiota is based on the Darwin Core (DwC) data standard to enable easy data sharing with other Darwin CoreBiodiversity data standard maintained by TDWG, with terms for sharing species occurrence and specimen data.-aligned databases, like the Global Biodiversity Information FacilityInternational network and data infrastructure providing open access to biodiversity data, including fossils., known as “GBIF”. The portal itself is hosted and maintained by a team at the University of Kansas, the Symbiota Support Hub (aka “SSH”). The SSH develops the Symbiota code, maintains the software’s official user documentation, “Symbiota Docs”, and keeps the Paleo Data Portal’s software up-to-date (along with many other SSH-hosted Symbiota portals). The SSH also administers the portal’s underlying “backend” database and maintains its physical server at the University of Kansas. If you share data hosted on SSH-managed infrastructure, you should review the SSH’s Data Sharing Policy.
Community
The goals of the Paleo Data Portal overlap with several related projects.
The Paleo Data Portal is part of a larger initiative working to improve cyberinfrastructure for fossil collections and paleontological research. In 2023, this project acquired funding from the US National Science Foundation, which enabled new software enhancements for Symbiota and provisioned for the establishment of the Paleo Data Portal and associated user documentation. This project is working to address long outstanding issues, such as improving digitally available taxonomy for extinct organisms and the alignment of existing data standards with the needs of fossil collections and researchers. The Paleo Data Portal integrates with this project in several ways, for example, by helping hone use cases demonstrating how the data needs in the domain of paleontology differ from that of extant biological research and collections.
The aforementioned initiative dovetails with the goals of the Paleo Data Working Group (aka “PDWGPaleo Data Working Group. Community of practice centered around collections-based paleo and informatics professionals”), which is a community of practice for paleontological collections and informatics professionals who aim to develop and promote best practices for managing and digitizing fossil specimens. If you join the Paleo Data Portal as a data provider, participation in PDWG is not mandatory but it is strongly encouraged. While the Symbiota Support Hub can help with technical issues, PDWG members can provide discipline-specific support to Paleo Data Portal users. For example, if you are unsure how to digitize something in your fossil collection, you can ask PDWG for assistance via one of its communication channels, such as Slack or attending a biweekly “Happy HourName of the regular meetings hosted by the Paleo Data Working Group” meeting.
If you join the Paleo Data Portal as a data provider, you will be invited to PDWG’s Slack space and Google Group. Please email PaleoDataWG@gmail.com if you have a preferred Google-enabled email that differs from the address associated with your user account in the portal.
Contacts
Inquiries about the Paleo Data Portal can be directed to:
- The Symbiota Support Hub’s Help Desk: help@symbiota.org
- The Paleo Data Portal’s Steering Committee: PaleoDataWG@gmail.com
The Symbiota Support Hub can assist with technical tasks, such as offering advice related to image hosting, “backend” database data manipulation, software bugs, and portal performance problems, among many other issues. If you have more paleo/discipline-specific questions, such as how to represent fossil specimen data in the portal (e.g. “How do I catalog this fossil?”), please contact the Paleo Data Working Group by attending a biweekly meeting, posting a message in our shared Slack space, or by emailing the address listed above.
User Documentation
To learn how to use the Paleo Data Portal and maximize the impact of your digitization efforts, data providers should become familiar with the resources made available for this purpose. Additional information to help you get started using the portal is outlined below.
Symbiota Docs
“Symbiota Docs” is the official user documentation for Symbiota maintained by the Symbiota Support Hub. This resource includes written instructions and recordings to help you learn how to navigate most features in your portal. The SSH also hosts regular webinars for user training and networking purposes. Once your collection is set up in the Paleo Data Portal, you will occasionally receive email notifications about these events from the SSH.
Paleo Data Knowledge Hub
Additional discipline/paleo-specific documentation has been created for the Paleo Data Portal and these resources are integrated into the Paleo Data Knowledge Hub. Links to this information are available in the portal’s main navigation menu for ease of reference. The Symbiota “quick start” tutorial and data management how-to guide are essential resources for learning how to navigate the portal and catalog/digitally represent your fossil specimen data. An additional guide has been developed to help you think through possible workflows for digitizing your collection.
Collection Metadata & Contacts
Data providers should keep contacts for their collection up-to-date in the portal (instructions). While this information is helpful for researchers, it is equally helpful to the Symbiota Support Hub and the portal’s Steering Committee; we may use this information to send you important updates about the portal, for instance, if it will be temporarily offline during routine server maintenance. New data providers are advised to add the following email addresses to your safe senders/contact lists to ensure related messages are not directed to spam: help@symbiota.org | hub@symbiota.org | PaleoDataWG@gmail.com. The SSH’s “Communications FAQ” can be found here.
GRSciColl Registration
Data providers are encouraged to additionally register their collection’s metadataInformation describing data, e.g., who collected it, where, when, how. with the Global Registry of Scientific Collections (aka “GRSciCollGlobal Registry of Scientific Collections. Worldwide catalogue of scientific collections supported by a community comprised of national editors and representatives from registered institutions who keep the information up to date.”). This can happen before, during, or after your collection begins using the Paleo Data Portal. Conveniently, you can register new GRSciColl records and suggest changes to existing records associated with your institution and collection without a website log in (instructions). Once your collection is established in the Paleo Data Portal, its unique URLUniform Resource Locator. A type of URI that specifies the location of a resource on the internet by describing its primary access mechanism. E.g. https://… in the portal can be added to its corresponding GRSciColl record (example).
Highlighted Features & Related Services
Cataloging Tools
Once you are oriented to the portal’s essential features using the “quick start” tutorial, you should review the data management how-to guide. This page will help you learn to catalog your specimen data, either by formatting and importing a spreadsheet of existing data, or by creating catalog records directly in the portal. If you maintain a spreadsheet of existing catalog records, this page contains tips and links to external resources to help you prepare your data for import, as well as a template file (and example file) that can be used for this purpose. This page also includes a number of example records to help you learn how Symbiota can be used to approach some frequently encountered cataloging scenarios in fossil collections—for example, how to represent part/counterpart relationships (USNMP7427 & USNMP7428, USNMP42726), multitaxon slabs with cataloged (USNMPAL449450) and uncataloged (USNMPAL566311) taxa, and associating cataloged specimens with digitally available literature (USNMV4735).
Data Redaction
By default, catalog records in Symbiota portals are publicly visible; however, because fossil collecting locations sometimes require data redaction, the Paleo Data Portal includes features to obscure location-related data on individual catalog records. In the Paleo Data Portal it is the data provider’s responsibility to redact sensitive specimen data from public access. You can apply security settings to individual catalog records directly in the portal, or by including a “recordSecurity” field when you prepare your spreadsheet of catalog records for import. More information about how these features work can be found in Symbiota Docs.
Data Publishing
Once a collection maintains catalog records in the Paleo Data Portal, sending a copy of this data to the GBIF data portal is encouraged. Doing so will increase the visibility and impact of your data mobilization efforts (examples). If applicable, data that are redacted in Symbiota will remain obscured in GBIFGlobal Biodiversity Information Facility. International network and data infrastructure providing open access to biodiversity data, including fossils. unless the data provider opts to change these settings in Symbiota at the time of data publication. Establishing the data publishing “pipeline” is typically straightforward, and once it’s set up, you can easily refresh your GBIF dataset directly from the Paleo Data Portal. The Symbiota Support Hub can assist with configuring data publishing for your collection in the Paleo Data Portal (recommended: CC PaleoDataWG@gmail.com).
Media Hosting
Accurately representing the physical state and historical/curatorial context of fossils can be challenging to convey with textual data alone; therefore, associating catalog records with images of the fossil specimens and/or their associated labels is recommended when feasible. Symbiota-based portals, including the Paleo Data Portal, support the association of media (images, audio) with specimen data (instructions). If you intend to associate images with your catalog records, these files can be hosted externally (e.g. self-hosted by your university) or by the Symbiota Support Hub. Because many universities and museums maintain platforms with image-hosting abilities, investigating your options before proceeding is advised, for example, by contacting your institution’s library and archives. Self-hosting images can afford greater flexibility in terms of file size and resolution that can be shared and displayed in the portal. In general, you should aim to provide images that clearly illustrate what the cataloged fossil material represents and, if applicable, with legible text.
If you opt to have the Symbiota Support Hub host your images, be aware that this service may result in fees due to SSH personnel time associated with media server management and the facilitation of batch image ingestion. The SSH’s media hosting rates are posted here. The SSH only hosts “web ready” media (<10 MB/each + JPEG format preferred) and does not offer “archival” media storage. If you would like to know more about the SSH’s image hosting services, please write to the SSH’s Help Desk.
Taxonomy & Identifications/Determinations
Taxonomic Thesaurus
Every Symbiota portal contains a “Taxonomic Thesaurus” that serves as the centralized source of taxonomy for all catalog records and associated resources in that portal. The Taxonomic Thesaurus in the Paleo Data Portal is not and should not be considered a taxonomic authority, as is true for most Symbiota portals; rather, it is primarily a tool for data discovery. The Taxonomic Thesaurus is used to automatically associate the scientific names assigned to cataloged specimens with higher taxonomy (e.g. Kingdom, Phylum, Order, etc.). For example, if you were to catalog a specimen/specimen lot as “Trilobita”, the Taxonomic Thesaurus will automatically associate this catalog record with Kingdom = “Animalia” and Phylum = “Arthropoda” when the catalog record is searched on or its associated data are exported from the portal. Taxonomy in the Paleo Data Portal can be viewed using the Taxonomy Explorer and Tree Viewer.
Identifications/Determinations
As you begin to use the Paleo Data Portal, you may notice that not all scientific names applied to your catalog records will correspond to higher taxonomy in the portal, and in some cases, the higher taxonomy may not align with your preferred higher taxonomy. (Again, this is because the portal’s Thesaurus is not a taxonomic authorityThe author(s) who formally published the scientific name of a taxon, often included after the species name..) For this reason, data providers are encouraged to include override values for familyThe full scientific name of the family in which the dwc:Taxon is classified. and scientificNameAuthorshipThe authorship information for the dwc:scientificName formatted according to the conventions of the applicable dwc:nomenclaturalCode. (when applicable/known/feasible) when cataloging fossil specimens directly in the portal or importing this information using a spreadsheet to increase the discoverability of these catalog records. Other scenarios frequently encountered in fossil collections are the identification of specimens to higher taxon level (i.e., “Trilobita” instead of a species-level taxon), sometimes with the inclusion of identification qualifiersNotation that modifies a taxon name to reflect uncertainty (e.g. cf., aff., ?)., like “?”, “aff.”, “cf.” (e.g. “Trilobita?”). In order for your catalog records to properly link to the Taxonomic Thesaurus, the qualifier should be recorded in identificationQualifierA brief phrase or a standard term (“cf.”, “aff.”) to express the determiner’s doubts about the dwc:Identification. and rather than appended to the taxon name specified in scientificNameThe full scientific name, with authorship and date information if known. When forming part of a dwc:Identification, this should be the name in lowest level taxonomic rank that can be determined. This term should not contain identification qualifications, which should instead be supplied in the dwc:identificationQualifier term..
Paleo Ecosystem Context
The Paleo Data Portal is part of a larger initiative to address the inadequacies of existing cyberinfrastructure for fossil collections and research, which includes improving digitally available taxonomy for extinct organisms. For this reason, all higher taxonomy in the Paleo Data Portal is sourced from ChecklistBank’s Catalog of Life (COL) “Extended Release” taxonomic checklist, which is a taxonomic compilation that draws from sources such as the World Register of Marine SpeciesAuthoritative taxonomy database for global marine species. (WoRMSWorld Register of Marine Species. Authoritative taxonomy database for global marine species.), the Interim Register of Marine and Nonmarine Genera (IRMNG), and the Paleobiology DatabaseGlobal, community-driven database of fossil occurrence records supported by literature. (PBDBPaleobiology Database. Global, community-driven database of fossil occurrence records supported by literature.). By using the ChecklistBank COLCatalogue of Life. A resource designed to compile and share the most complete authoritative list of the world’s species, as maintained by hundreds of global taxonomists. checklist to populate the portal’s Taxonomic Thesaurus, the Paleo Data Portal will help document shortcomings of existing taxonomic resources, allowing the Paleo Data Working Group to make recommendations for improvement to cyberinfrastructure maintainers.
This a resource with a tooltip: Bauer et al. (2022)The purpose of this document is to provide a framework for how to mobilize information via Wikidata about people working in and/or associated with scientific collections. Building on previous Wikidata documentation produced by Siobhan Leachman (2020, https://doi.org/10.5281/zenodo.4724139) participants of the Using Wikidata to Capture and Share Information about People in Paleontology workshop (held March 29-31, 2022) created this framework to formalize and share practical knowledge gained from the workshop.. It should display as a link to an external resource with a tooltip that appears on hover. Tooltips displayed using this widget are based on the annotation field. If that field contains any formatting code, it may break the tooltip. The first appearance of the word tooltip should be underlined and include a tooltip but no link.
This paragraph includes a term defined manually in the glossaries/custom.yml file, Internationalized Resource IdentifierVariation of a URI that expands the characterset..
This is a paragraph including Darwin CoreBiodiversity data standard maintained by TDWG, with terms for sharing species occurrence and specimen data. terms, like geodeticDatumThe ellipsoid, geodetic datum, or spatial reference system (SRS) upon which the geographic coordinates given in dwc:decimalLatitude and dwc:decimalLongitude are based. and maximumDistanceAboveSurfaceInMetersThe greater distance in a range of distance from a reference surface in the vertical direction, in meters. Use positive values for locations above the surface, negative values for locations below. If depth measures are given, the reference surface is the location given by the depth, otherwise the reference surface is the location given by the elevation.. Those terms should display as links to the Darwin Core Quick Reference Guide and should show a definition on hover. Only the first appearance of a term on each page should include the tooltip, so geodeticDatum in this sentence should appear as plain text. Note that terms do not use a widget or any other syntax. The script that builds the page identifies them automatically.
Here are some items that should not include tooltips:
-
dwc:institutionID
- dwc:collectionID
- dwc:datasetID
But these should:
- institutionIDAn identifier for the institution having custody of the object(s) or information referred to in the record.
- collectionIDAn identifier for the collection or dataset from which the record was derived.
- datasetIDAn identifier for the set of data. May be a global unique identifier or an identifier specific to a collection or institution.
Notices
Created with {: .notice}
Created with {: .notice--info}
Created with {: .notice--warning}
Created with {: .notice--danger}
Created with {: .notice--tip}
Related content
- Develop a digitization workflow using Symbiota (How to guide): This guide is intended to help data providers develop digitization workflows for fossil specimens using Symbiota.
- Manage data about fossil specimens using Symbiota (How to guide): This guide is intended to complement introductory documentation for Symbiota data providers, as well as the official Symbiota user documentation, Symbiota Docs. Symbiota Docs provides general guidance for working in Symbiota-based data portals and should be referenced for basic functions and workflows. This manual expands on this resource to provide discipline-specific information for fossil collections.
- Start using Symbiota (Tutorial): This tutorial is designed orient new and prospective data providers to using Symbiota, a tool for managing and publishing fossil specimen data.
- Symbiota Paleo Data Portal (Highlighted resource): This is a landing page that describes what the Symbiota Paleo Data Portal is and why it is important in the context of paleo data. You can dive deeper via the links to related resources aggregated here.
Browse additional related content, including PDWG Happy Hours and links out to external resources, via topic: symbiota
Related content
- Georeference collection localities (How to guide): This page contains general information about community practices for georeferencing collection localities, and also aggregates links to additional resources with more specific information.
- Manage data about geography (How to guide): This page contains general information about community practices for managing data about geography, and also aggregates links to additional resources with more specific information.
- Manage data about sensitive localities (How to guide): This page contains general information about community practices for managing data about sensitive localities, and also aggregates links to additional resources with more specific information.
- Manage data about shared localities (How to guide): This page contains general information about community practices for managing data about shared localities, and also aggregates links to additional resources with more specific information.
Browse additional related content, including PDWG Happy Hours and links out to external resources, via topic: geography and georeference
Related content
- Data wrangling (Explanation): This page explains what it means to wrangle data by transforming and mapping it from one form into another with the intent of making the data more appropriate for a specific use. This page also links out to tools used for data wrangling, and resources for learning how to wrangle data.
- Georeference collection localities (How to guide): This page contains general information about community practices for georeferencing collection localities, and also aggregates links to additional resources with more specific information.
- Manage data about geography (How to guide): This page contains general information about community practices for managing data about geography, and also aggregates links to additional resources with more specific information.
- Manage data about sensitive localities (How to guide): This page contains general information about community practices for managing data about sensitive localities, and also aggregates links to additional resources with more specific information.
- Manage data about shared localities (How to guide): This page contains general information about community practices for managing data about shared localities, and also aggregates links to additional resources with more specific information.
- Wrangle data in OpenRefine (How to guide): This page contains information about how to manipulate and transform data using the OpenRefine software. It also links out to tutorials and additional resources.
Browse additional related content, including PDWG Happy Hours and links out to external resources, via topic: geography, georeference, and data wrangling