Background
Omics technologies have enlarged the scope and scale of early exploratory insights, and are increasingly used to accelerate biomarker validation. However, their rising popularity is undercut by the high computational costs associated with multi-dimensional data analysis, presenting a cost-exclusionary barrier to scaling drug discovery pipelines.
Even with powerful infrastructure like on-premise high performance computing (HPC), scientists are challenged to scale their workflows on-demand, leading to under-utilised resources. Additionally, powerful compute services require technical expertise beyond programming to deploy pipelines at scale, further delaying bench-scientists from independently scaling their own workflows.
Scaling Bioinformatics Workflows: Importance and Challenges
Early exploratory insights often require additional sample data for validation, which challenges existing infrastructure.
For example, a lab that has an on-premise HPC may not be able to scale their workflow in a timely, cost-effective manner for omics technologies like single cell RNA Sequencing (scRNA Seq.), which generate large data volumes. When scaled to larger sample sizes, the associated computational requirements multiply, and scaling cannot be done cost-effectively. The exclusionary cost of high-throughput runs thus raises the barrier of entry for bench-scientists for accessing high-compute infrastructure.
Additionally, any modification requires IT intervention for correct deployment of the pipeline on an HPC cluster. This is because a job submission on an HPC requires:
- ordering multiple computing tasks correctly,
- aligning tasks to the right queue with appropriate CPU and memory assigned,
- prioritising tasks based on run-time and queue availability, and
- re-running unsuccessful tasks.
A limited number of users can simultaneously deploy their pipelines on an on-premise HPC, regardless of how powerful it may be, causing delays and backlogs as projects queue indefinitely.
Finally, project workflows are not always linear; a pipeline may need a few reruns to adjust parameters or tools used for analysis, which challenges the design of a scalable workflow.
Other barriers to scalability include irreproducible workflows, segregated data sources, lack of computational expertise (e.g., knowledge of HPC, cloud, AWS Batch, etc.), and limited flexibility in existing systems.
Quark overcomes these barriers by leveraging cloud services such as Amazon Web Services (AWS) to enable scientists scale their own workflows cost-effectively.
Quark: Delivering Cloud-Enabled Scalable Workflows
Cloud services enable users to provision and scale computing resources on-demand, based on their workload. Cloud also facilitates scientific collaborations regardless of geographical location. However, a challenge in integrating with cloud services is that it may require prior knowledge in compute modalities like HPC, AWS Batch, Kubernetes, etc.
As illustrated earlier, the complexity of deploying a pipeline in each environment is time-consuming, and requires IT/DevOps expertise. Quark overcomes this entry-barrier for bench-scientists and democratises access to powerful compute resources, through:
- an intuitive user-interface;
- enabling versatile workflow auto-scaling, and;
- providing cost-effective instances.
Quark’s Intuitive Interface Simplifies Accessibility to Advanced Compute Modalities
Though cloud delivers affordable compute resources at scale, complex modalities like HPC require high skilled expertise to optimally deploy bioinformatics pipelines on them. Quark lowers this barrier-for-entry for bench-scientists through a single-click interface.
Our clients’ testimonials highlight the intuitive nature of the Quark interface–
Quark has allowed our scientists to discover pipelines and run them in a self-service manner with no IT intervention. This has shortened the time for results. Truly game changing for research.
Director of Computational Engineering, Leading Fortune 500 BioTech
Quark’s intuitive and simple interface to run pipelines allows me to get my research done faster. I’m able to run 1000s of pipelines every week without having to dabble scripts or reach out to IT.
Director of BioInformatics, Fortune 500 BioTech
Thus, bench-scientists can access powerful computational resources and scale their workflows hassle-free, without needing IT expertise.
Quark enhances the user’s experience by integrating services like AWS Healthomics, simplifying access to various resources without needing prior knowledge of Healthomics workflows or cloud services.
Quark Enables Versatile Auto-Scaling
Quark leverages cloud infrastructure to auto-scale workflows, through:
- On-demand provisioning: Bioinformatics workflows are dynamic. When a study is updated with new samples, or an early data-insight requires validation, Quark enables scientists to avail computing resources on-demand without needing to wait on tickets or queues.
- Multiple compute modalities: Quark leverages different compute modalities like HPC, AWS Batch, and Kubernetes, providing users with the optimal resources needed for their workflows. For example, containerised pipelines with long runtimes are best managed and run with Kubernetes, which can auto-scale based on configured metrics. For batch jobs, Quark leverages AWS Batch, which can seamlessly integrate with other AWS Cloud Services.
Therefore, Quark automates workflow scaling and simplifies the user experience so that they do not have to configure their infrastructure for every pipeline run, or make calculations to manage compute costs.
Additionally, users can easily choose between CPUs or GPUs for their workflows. Programs for protein structure prediction, or omics workflows like NVIDIA Parabricks, typically need higher-end GPU instances which are available on Quark.
Quark Enables Cost-Effective Scaling
With Quark, scientists can take advantage of both on-demand compute resources and cost-effective Spot Instances. Spot Instances allow users to access unused cloud resources, such as AWS Elastic Compute Cloud (EC2), at significantly reduced prices. This enables users to run their analytics efficiently while minimising costs.
An advantage of deploying bioinformatics pipelines on cloud services is that cloud providers continually release Instance types based on new higher-performance chips. These new Instances match the same price of older chips, incrementally improving the cost/performance ratio over time.
Other aspects of managing powerful compute-infrastructure, like obtaining security and compliance certifications, is in part taken care of by cloud providers.
Therefore, users do not need to invest in on-premise infrastructure and can directly commission cloud services through Quark without undue wait-times.
Conclusion
Drug discovery pipelines often stall due to issues in scalability when there’s limited availability of compute infrastructure. Since omics data is multi-dimensional, it requires sophisticated and expensive computational infrastructure for data analysis, which raises the barrier of entry for bench-scientists. Even with on-premise HPC, resources are often under-utilised since scientists cannot scale their workflows on-demand when needed– unlike cloud computing, where workflows can be scaled up and down on-demand, thereby optimising costs.
A platform that adapts to rapidly changing technologies is crucial for a future-ready and scalable workflow. The compute infrastructure should be able to accommodate advancements in omics technologies that amass large volumes of multi-dimensional data. Integrating with cloud services helps democratise access to compute resources, allowing users to scale their workflows in an affordable, efficient manner.
Quark simplifies access to cloud-enabled services and delivers cost-effective, yet powerful computing solutions tailored to a project’s needs. Bench-scientists can leverage different types of compute modalities to get actionable insights faster, thus overcoming computational and expensive entry-barriers for scaling workflows.
Request a demo to learn more about Quark.