Paleo Data Portal orientation
The content on this page has not been finalized. Contributors can mark a page as complete and remove this warning by adding status: published to the front matter in the Markdown source file.
This page is intended to help data providers 1) determine whether the Paleo Data Portal will be an appropriate tool for managing and sharing data associated with their collection and 2) understand key information for leveraging the portal’s built-in tools for data mobilization. If you are not already familiar with the portal, an introductory explanation can be found here.
Portal scope & guidelines
If you are interested in using the Paleo Data Portal (PDP)A Symbiota portal that supports the management and sharing of data associated with paleontological specimens that are made available for research via permanent repositories. The scope of the PDP is limited to extinct organisms and their traces (i.e., fossils). Geological samples, archaeological and anthropological materials, as well as neontological specimen data fall outside this scope and should not be cataloged in this portal. to manage and share fossil specimen data, first consider the portal’s scope and community guidelines. The portal’s scope is stated on its home page and below:
This portal supports the management and sharing of data associated with paleontological specimens that are made available for research via permanent repositories. Its scope is limited to extinct organisms and their traces (i.e., fossils). Geological samples, archaeological and anthropological materials, as well as neontological specimen data fall outside this scope and should not be cataloged in this portal.
Further, the portal is intended for actively managing and sharing specimen data from fossil collections that can be used for research and are held in public trust, for example, by a university collection or nonprofit museum. It is not intended for use by teaching collections or private collections. Publicly accessible research collections that have not directly benefited from the US National Digitization effort or do not have access to secure cyberinfrastructure to maintain their specimen data are especially encouraged to participate.
If the Paleo Data Portal does not meet the needs of your collection, other collection management systemsSoftware for curating, describing, and managing specimens. Also known as a Collection Information System (CIS). (CMSCollection Management System. Software for curating, describing, and managing specimens. Also known as a Collection Information System (CIS).) may be more suitable. Krimmel (2022)Materials for a lesson developed as part of a course, Introduction to Biodiversity Specimen Digitization, offered by the iDigBio Digitization Academy with the goal of introducing the creation of digital data about biodiversity specimens to those who are just beginning this activity. provides an overview of different CMS options and what to consider when selecting one.
Portal sustainability
The Paleo Data Portal’s underlying code and infrastructure are maintained by the Symbiota Support Hub, aka “SSHSymbiota Support Hub. Group based at the University of Kansas that maintains the most commonly used version of Symbiota, as well as a number of Symbiota-based portals.”, at the University of Kansas. The SSH incurs costs to make this resource available—for example, by responding to Help Desk tickets, developing and hosting webinars, maintaining Symbiota DocsA central repository for documentation regarding Symbiota-based data portals. This site is maintained by the Symbiota Support Hub, but all Symbiota users may contribute. Information on Symbiota Docs complements documentation on the Paleo Data Knowledge Hub about using Symbiota for fossil specimens., improving the Symbiota code and fixing bugs, running data backups, and keeping the portal and its physical server secure. The SSH also keeps the Paleo Data Portal’s software up-to-date, along with many other SSH-hosted portals in the Symbiota Ecosystem. Therefore, while it is not required, data providers are strongly encouraged to budget support for the Symbiota Support Hub in funding proposals and/or your annual operating budget in a capacity is feasible for your collection. Please contact the SSH if you would like to contribute funds toward portal hosting and maintenance. The SSH also accepts donations; if you donate, you can optionally indicate that your are a Paleo Data Portal community member.
User documentation
To learn how to use the Paleo Data Portal and maximize the impact of your digitization efforts, data providers should become familiar with the resources made available for this purpose. Additional information to help you get started using the portal is outlined below.
Symbiota Docs
Symbiota DocsA central repository for documentation regarding Symbiota-based data portals. This site is maintained by the Symbiota Support Hub, but all Symbiota users may contribute. Information on Symbiota Docs complements documentation on the Paleo Data Knowledge Hub about using Symbiota for fossil specimens. is the official user documentation for Symbiota maintained by the Symbiota Support Hub. This resource includes written instructions and recordings to help you learn how to navigate most features in your portal. The SSH also hosts regular webinars for user training and networking purposes. Once your collection is set up in the Paleo Data Portal, you will occasionally receive email notifications about these events from the SSH.
Paleo Data Knowledge Hub
Additional discipline/paleo-specific documentation has been created for the Paleo Data Portal and these resources are integrated into this website. Links this information can be found in the portal’s main navigation menu for ease of reference. The Symbiota “quick start” tutorial and data management how-to guide are essential resources for learning how to navigate the portal and catalog/digitally represent your fossil specimen data. An additional guide has been developed to help you think through possible workflows for digitizing your collection.
User support
Technical
The Paleo Data Portal is part of the global SymbiotaSymbiota is an adaptable, customizable software that enables self-defined and self-governed communities of practice to create collaborative biodiversity data portals. Data in web-based Symbiota portals can be instantly searchable via public search features, and there are many tools to connect these data to other collections and to aggregators (e.g., iDigBio, GBIF). Symbiota portals include many tools for efficient digitization, data import, and data mobilization, as well as duplicate matching tools (across and within collections), CSV and Darwin Core Archive import and export options, and linking tools to external sources or specimen records. Ecosystem, which is largely maintained by Symbiota Support Hub, as described above. The SSH administers the Paleo Data Portal’s underlying “backend” database and maintains its physical server at the University of Kansas. In addition to portal’s specific scope and guidelines, data providers should also be aware that the SSH maintains its own Data Sharing Policy.
The Symbiota Support Hub can assist with technical tasks, such as offering advice related to image hosting, “backend” database data manipulation, software bugs, and portal performance problems, among many other issues. You can contact the SSH’s Help Desk by writing to help@symbiota.org or submitting a ticket through this interface.
Community
The goals of the Paleo Data Portal overlap with several related projects.
The Paleo Data Portal is part of a larger initiative working to improve cyberinfrastructure for fossil collections and paleontological research. In 2023, this project acquired funding from the US National Science Foundation, which enabled new software enhancements for Symbiota and provisioned for the establishment of the Paleo Data Portal and associated user documentation. This project is working to address long outstanding issues, such as improving digitally available taxonomy for extinct organisms and the alignment of existing data standards with the needs of fossil collections and researchers. The Paleo Data Portal integrates with this project in several ways, for example, by helping hone use cases demonstrating how the data needs in the domain of paleontology differ from that of extant biological research and collections.
The aforementioned initiative dovetails with the goals of the Paleo Data Working Group (aka “PDWGPaleo Data Working Group. Community of practice centered around collections-based paleo and informatics professionals”), which is a community of practice for paleontological collections and informatics professionals who aim to develop and promote best practices for managing and digitizing fossil specimens. If you join the Paleo Data Portal as a data provider, participation in PDWG is not mandatory but it is strongly encouraged. If you join the Paleo Data Portal as a data provider, you will be invited to PDWG’s Slack space and Google Group.
While the Symbiota Support Hub can help with technical issues, PDWG members can provide discipline-specific support to Paleo Data Portal users. For example, if you are unsure how to digitize something in your fossil collection, you can ask PDWG for assistance via one of its communication channels, such as Slack, attending a biweekly “Happy HourName of the regular meetings hosted by the Paleo Data Working Group” meeting, or contacting PaleoDataWG@gmail.com.
Collection metadata & contacts
Data providers should keep contacts for their collection up-to-date in the portal (instructions). While this information is helpful for researchers, it is equally helpful to the Symbiota Support Hub and the portal’s Steering Committee; we may use this information to send you important updates about the portal, for instance, if it will be temporarily offline for routine server maintenance. New data providers are advised to add the following email addresses to your safe senders/contact lists to ensure related messages are not directed to spam: help@symbiota.org | hub@symbiota.org | PaleoDataWG@gmail.com. The SSH’s “Communications FAQ” can be found here.
GRSciColl registration
Data providers are encouraged to additionally register their collection’s metadataInformation describing data, e.g., who collected it, where, when, how. with the Global Registry of Scientific Collections. This can happen before, during, or after your collection begins using the Paleo Data Portal. Conveniently, you can register new GRSciCollGlobal Registry of Scientific Collections. Worldwide catalogue of scientific collections supported by a community comprised of national editors and representatives from registered institutions who keep the information up to date. records and suggest changes to existing records associated with your institution and collection without a website log in (instructions). Once your collection is established in the Paleo Data Portal, its unique URLUniform Resource Locator. A type of URI that specifies the location of a resource on the internet by describing its primary access mechanism. E.g. https://… in the portal can be added to its corresponding GRSciColl record (example).
Highlighted features & related services
Cataloging tools
Once you are oriented to the portal’s essential features using the “quick start” tutorial, you should review the data management how-to guide. This documentation will help you learn to catalog your specimen data, either by formatting and importing a spreadsheet of existing data, or by creating catalog records directly in the portal. If you maintain a spreadsheet of existing catalog records, how-to guide contains tips and links to external resources to help you prepare your data for import, as well as a template file (and example file) that can be used for this purpose. It also includes a number of example records to help you approach frequently encountered cataloging scenarios in fossil collections—for example, how to represent part/counterpart relationships (USNMP7427 & USNMP7428, USNMP42726), multitaxon slabs with cataloged (USNMPAL449450) and uncataloged (USNMPAL566311) taxa, and associating cataloged specimens with digitally available literature (USNMV4735).
Data redaction
By default, catalog records in Symbiota portals are publicly visible; however, because fossil collecting locations sometimes require data redaction, the Paleo Data Portal includes features to obscure location-related data on individual catalog records. In the Paleo Data Portal it is the data provider’s responsibility to redact sensitive specimen data from public access. You can apply security settings to individual catalog records directly in the portal, or by including a “recordSecurity” field when you prepare your spreadsheet of catalog records for import. More information about how these features work can be found in Symbiota Docs.
Data publishing
Once a collection maintains catalog records in the Paleo Data Portal, sending a copy of this data to the GBIFGlobal Biodiversity Information Facility. International network and data infrastructure providing open access to biodiversity data, including fossils. data portal is encouraged (examples). Doing so will increase the visibility and impact of your data mobilization efforts. If applicable, data that are redacted in Symbiota will remain obscured in GBIF unless the data provider opts to change these settings in Symbiota at the time of data publication. Establishing the data publishing “pipeline” is typically straightforward, and once it’s set up, you can easily refresh your GBIF dataset directly from the Paleo Data Portal (instructions). Contact the SSH (CC: PaleoDataWG@gmail.com) for assistance with configuring data publishing for your collection in the Paleo Data Portal.
Media hosting
Accurately representing the physical state and historical/curatorial context of fossils can be challenging to convey with textual data alone; therefore, associating catalog records with images of the fossil specimens and/or their associated labels is recommended when feasible. Symbiota-based portals, including the Paleo Data Portal, support the association of media (images, audio) with specimen data (instructions). If you intend to associate images with your catalog records, these files can be hosted externally (e.g. self-hosted by your university) or by the Symbiota Support Hub. They can also be added to catalog records one-by-one or in bulk. Because many universities and museums maintain platforms with image-hosting abilities, investigating your options before proceeding is advised, for example, by contacting your institution’s library and archives. Self-hosting images can afford greater flexibility in terms of file size and resolution that can be shared and displayed in the portal. In general, you should aim to provide images that clearly illustrate what the cataloged fossil material represents and, if applicable, with legible text.
If you opt to have the Symbiota Support Hub host your images, be aware that this service may result in fees due to SSH personnel time associated with media server management and the facilitation of batch image ingestion. The SSH’s media hosting rates are posted here. The SSH only hosts “web ready” media (<10 MB/file + JPEG format preferred) and does not offer “archival” media storage. If you would like to know more about the SSH’s image hosting services, please contact the SSH.
Taxonomy & identifications/determinations
Taxonomic thesaurus
Every Symbiota portal contains a “Taxonomic Thesaurus” that serves as the centralized source of taxonomy for all catalog records and associated resources in that portal. The Taxonomic Thesaurus in the Paleo Data Portal is not and should not be considered a taxonomic authority, as is true for most Symbiota portals; rather, it is primarily a tool for data discovery. The Taxonomic Thesaurus is used to automatically associate the scientific names assigned to cataloged specimens with higher taxonomy (e.g. Kingdom, Phylum, Order, etc.). For example, if you were to catalog a specimen/specimen lot as “Trilobita”, the Taxonomic Thesaurus will automatically associate this catalog record with kingdomThe full scientific name of the kingdom in which the dwc:Taxon is classified. = “Animalia” and phylumThe full scientific name of the phylum or division in which the dwc:Taxon is classified. = “Arthropoda” when the catalog record is searched on or its associated data are exported from the portal. Taxonomy in the Paleo Data Portal can be viewed using the Taxonomy Explorer and Tree Viewer.
Identifications/determinations
As you begin to use the Paleo Data Portal, you may notice that not all scientific names applied to your catalog records will correspond to higher taxonomy in the portal, and in some cases, the higher taxonomy may not align with your preferred higher taxonomy. (Again, this is because the portal’s Thesaurus is not a taxonomic authorityThe author(s) who formally published the scientific name of a taxon, often included after the species name..) For this reason, data providers are encouraged to include override values for familyThe full scientific name of the family in which the dwc:Taxon is classified. and scientificNameAuthorshipThe authorship information for the dwc:scientificName formatted according to the conventions of the applicable dwc:nomenclaturalCode. (when applicable/known/feasible) when cataloging fossil specimens directly in the portal or importing this information using a spreadsheet to increase the discoverability of these catalog records. Other scenarios frequently encountered in fossil collections include the identification of specimens to a higher taxon level (i.e., “Trilobita” instead of a species-level taxon), sometimes with the inclusion of identification qualifiersNotation that modifies a taxon name to reflect uncertainty (e.g. cf., aff., ?). (e.g. “Trilobita?”). In order for your catalog records to properly link to the portal’s Taxonomic Thesaurus, the qualifier should be recorded in identificationQualifierA brief phrase or a standard term (“cf.”, “aff.”) to express the determiner’s doubts about the dwc:Identification. rather than appended to the taxon name specified in scientificNameThe full scientific name, with authorship and date information if known. When forming part of a dwc:Identification, this should be the name in lowest level taxonomic rank that can be determined. This term should not contain identification qualifications, which should instead be supplied in the dwc:identificationQualifier term..
Paleo Data Ecosystem context
The Paleo Data Portal is part of a larger initiative to address the inadequacies of existing cyberinfrastructure for fossil collections and research, which includes improving digitally available taxonomy for extinct organisms. For this reason, all higher taxonomy in the Paleo Data Portal is sourced from ChecklistBank’s Taxonomic resource with highly variable coverage for fossils. Catalog of Life (COLCatalogue of Life. A resource designed to compile and share the most complete authoritative list of the world’s species, as maintained by hundreds of global taxonomists.) “Extended Release” taxonomic checklist, which is a taxonomic compilation that draws from sources such as the World Register of Marine SpeciesAuthoritative taxonomy database for global marine species. (WoRMSWorld Register of Marine Species. Authoritative taxonomy database for global marine species.), the Interim Register of Marine and Nonmarine Generaa provisional compilation of genus and species names covering marine and non-marine or fossil and recent taxa (IRMNGa provisional compilation of genus and species names covering marine and non-marine or fossil and recent taxa), and the Paleobiology DatabaseGlobal, community-driven database of fossil occurrence records supported by literature. (PBDBPaleobiology Database. Global, community-driven database of fossil occurrence records supported by literature.). By using the ChecklistBank COL checklist to populate the portal’s Taxonomic Thesaurus, the Paleo Data Portal will help document shortcomings of existing taxonomic resources, allowing the Paleo Data Working Group to make recommendations for improvement to cyberinfrastructure maintainers.
Related content
- Develop a digitization workflow using Symbiota (How to guide): This guide is intended to help data providers develop digitization workflows for fossil specimens using Symbiota.
- Manage data about fossil specimens using Symbiota (How to guide): This guide is intended to complement introductory documentation for Symbiota data providers, as well as the official Symbiota user documentation, Symbiota Docs. Symbiota Docs provides general guidance for working in Symbiota-based data portals and should be referenced for basic functions and workflows. This manual expands on this resource to provide discipline-specific information for fossil collections.
- Start using Symbiota (Tutorial): This tutorial is designed orient new and prospective data providers to using Symbiota, a tool for managing and publishing fossil specimen data.
- Symbiota Paleo Data Portal (Highlighted resource): This is a landing page that describes what the Symbiota Paleo Data Portal is and why it is important in the context of paleo data. You can dive deeper via the links to related resources aggregated here.
Browse additional related content, including PDWG Happy Hours and links out to external resources, via topic: symbiota