An automated infrastructure to support high-throughput bioinformatics

Articolo

Data di Pubblicazione:

2014

Abstract:

The number of domains affected by the big data phenomenon is constantly increasing, both in science and industry, with high-throughput DNA sequencers being among the most massive data producers. Building analysis frameworks that can keep up with such a high production rate, however, is only part of the problem: current challenges include dealing with articulated data repositories where objects are connected by multiple relationships, managing complex processing pipelines where each step depends on a large number of configuration parameters and ensuring reproducibility, error control and usability by non-technical staff. Here we describe an automated infrastructure built to address the above issues in the context of the analysis of the data produced by the CRS4 next-generation sequencing facility. The system integrates open source tools, either written by us or publicly available, into a framework that can handle the whole data transformation process, from raw sequencer output to primary analysis results.

Tipologia CRIS:

01.01 Articolo in rivista

Keywords:

Bioinformatics; MapReduce; NGS

Elenco autori:

Angius, Andrea

Autori di Ateneo:

ANGIUS ANDREA

Link alla scheda completa:

https://iris.cnr.it/handle/20.500.14243/245211

Dati Generali

URL

http://www.scopus.com/inward/record.url?eid=2-s2.0-84908632088&partnerID=q2rCbXpz

An automated infrastructure to support high-throughput bioinformatics

Angius, Andrea

Dati Generali

URL