Enabling Easy Data Discoverability through Metadata Validation

pexels-jsme-mila-523821574-29372694

Background

The drug discovery lifecycle involves several stakeholders who generate, handle, transfer, and analyse data. Often, each stakeholder follows different data classification and organisation practices, which affects data reusability and makes it difficult for users retrieving data in downstream analyses. 

Quark addresses this issue in data management by enforcing metadata validation. In this blog, we explore how Quark uses metadata to ensure frictionless data discoverability.

Data Discoverability Challenges

Metadata is information about a dataset that defines and structures it, making it discoverable and useful for downstream analysis.

For example, the title and headers of this article are its metadata. Data parameters such as gender, tumour status, genomic variant type, patient ID and sample ID function as metadata, since they structure sample data and make it discoverable.

However, many data custodians are faced with the challenge of consolidating inconsistent data capture (storage) and data retrieval (query) terms. A researcher who has no prior knowledge of the metadata terms that were used to record patient information will face two primary difficulties in discovering data: 

  • Unconsolidated terms: Patient gender may be recorded as chromosomal information (XX or XY) instead of ‘female’ or ‘male’. A scientist who is specifically searching for female patients for their study may not be able to retrieve and reuse this data for their analysis. It won’t be discoverable since the filtering term (female) is incompatible with the stored data parameter (XX) on file.  
  • Irreconcilable pipeline differences: Multi-omics datasets capture a wide range of data. Some of this information is uniform across datasets (sample ID, patient ID, tumour status, gender) whereas others are specific to the data-type and pipeline used to analyse them (like different pipelines for genomics, transcriptomics, or proteomics). Unique data fields makes it difficult for data custodians to unify data retrieval and reusability.  

Ensuring consistent metadata is thus a challenge that faces data custodians. Imposing a uniform rule and validation standard across different datasets is unfeasible since data parameters vary between workflows. Also, a uniform standard restricts the reusability of otherwise high-quality datasets.   

Further complicating the issue is the exponential data volumes being generated in the current omics era. Disparate and segregated data generation makes it difficult to enforce a consistent data organisation standard that would facilitate searchability. 

Quark addresses the issue of data discoverability by managing an exhaustive metadata index or metadata catalogue, which can be enforced by the project’s administrator.

Quark Improves Sample Findability by Enabling an Exhaustive Metadata Index

Quark tailors metadata for every secondary analysis pipeline by curating an exhaustive list of data parameters. Metadata standards are enforced without imposing on the user by enabling Global-level and Pipeline-level metadata validation.

Enforcing Global Metadata Consistency 

Global metadata are the default parameters that apply to all pipelines on the platform. Unless otherwise specified, any pipeline that runs on the Quark platform will automatically validate sample information based on the preset global metadata.

Project administrators can view and add new global parameters on Quark by following these steps: 

  • Sign-in: to the Admin portal and select Metadata from the left Navigator Pane. 
  • Enable Metadata Validation: Administrators have the option to enable or disable metadata validation through the toggle on the top-right. 
Metadata Window on Quark
  • Set Default Parameters: Default Global Metadata features set on Quark are shown below, and include age, sex, sample_ID, patient_ID. Each metadata property can be made mandatory or optional. Mandatory properties must be added to the samplesheets by users while running any secondary analysis pipeline on Quark. Optional properties are validated only if present in the samplesheet.

Default Global Metadata Features on Quark

  • Consolidate Data using Synonyms: Data retrieval terms are consolidated using the ‘Synonym’ field—for example, the synonyms provided for ‘patient’ is a string of alternative terms such as ‘subject’ that are often used to query samples while building cohorts. Synonyms ensure that different terms used by users to describe the same property can be stored in a single database field for ease-of-retrieval. 

Enforcing Pipeline Metadata Consistency

Global metadata definitions do not always meet the requirements of every pipeline. Some pipelines need unique metadata indices that have to remain consistent. For example, pipelines that predict protein-folding (AlphaFold) need protein sequence inputs rather than age or gender inputs (global metadata properties).

Pipeline metadata validation allows administrators to customise such requirements and override global metadata properties. 

 Customised Pipeline Metadata Features on Quark

Project administrators can “Add New” Properties, select the Pipeline they wish to customize for metadata, and fill the prompted fields to enforce the requirements of a pipeline run to ensure reproducibility across the project and institution. 

Examples of Using Metadata for Simplified Data Analysis

Metadata has various applications in both secondary and tertiary analytics, where data needs to be readily discoverable and reusable. 

Secondary data analysis: With metadata validation, Quark users are immediately alerted about errors in their sample sheets—like missing data columns.

On Quark, Project administrators can mark certain fields, such as “Patients,” as mandatory for a particular pipeline to launch. When a sample is missing a mandatory data column, Quark immediately alerts the user, along with prompts to fix the errors. 

The figure below provides an example of how Metadata Validation precludes erroneous pipeline runs, saving ineffective cost and time expenditures. Quark helps Project Administrators enforce uniform standards from an early-stage (at secondary data analysis), thus streamlining data pipelines to run successfully. 

Metadata Validation on Quark: Pre-Run Error Message Alerts Users About Issues in their Sample Data

Tertiary data analysis: On Quark, all data analysis results are unified onto a single tab. Bench-scientists and bioinformaticians can query and filter their samples for cohort comparisons, using Quark’s exhaustive metadata index. 

For example, Quark’s intuitive interface allows users to directly filter sample results, based on specific genotypic or phenotypes, such as gender (Figure below).

Filtering and Building Cohorts based on Gender

Quark’s metadata thus makes it possible to query, filter and discover data between and across multiple analyses. Quark simplifies the user’s experience of finding data through the maintenance of an updated, comprehensive and exhaustive index of metadata parameters.

Conclusion

Ensuring metadata consistency is important for ensuring data reusability and discovery. By enforcing metadata consistency, administrators consolidate data retrieval terms and facilitate easier:

  • Data querying
  • Cohort creation
  • Cohort comparison
  • Data visualisation.

Quark enables an exhaustive metadata catalogue that ensures users have a seamless experience in retrieving or reusing data for their project needs. Metadata validation unifies how data is stored across pipelines on the platform, enforcing data reproducibility at the institutional level. 

Quark’s two-pronged global and pipeline-level metadata consistency enforcement enables project administrators to customise data standards. Administrators can enforce or disable these standards as required. 

Request a demo to learn more about Quark.

Leave a Reply

Discover more from Quark Bioinformatics Platform

Subscribe now to keep reading and get access to the full archive.

Continue reading