Cited 0 times in Scipus Cited Count

SEQprocess: a modularized and customizable pipeline framework for NGS processing in R package

DC Field Value Language
dc.contributor.authorJoo, T-
dc.contributor.authorChoi, JH-
dc.contributor.authorLee, JH-
dc.contributor.authorPark, SE-
dc.contributor.authorJeon, Y-
dc.contributor.authorJung, SH-
dc.contributor.authorWoo, HG-
dc.date.accessioned2020-10-21T07:21:09Z-
dc.date.available2020-10-21T07:21:09Z-
dc.date.issued2019-
dc.identifier.urihttp://repository.ajou.ac.kr/handle/201003/18867-
dc.description.abstractBACKGROUNDS: Next-Generation Sequencing (NGS) is now widely used in biomedical research for various applications. Processing of NGS data requires multiple programs and customization of the processing pipelines according to the data platforms. However, rapid progress of the NGS applications and processing methods urgently require prompt update of the pipelines. Recent clinical applications of NGS technology such as cell-free DNA, cancer panel, or exosomal RNA sequencing data also require appropriate customization of the processing pipelines. Here, we developed SEQprocess, a highly extendable framework that can provide standard as well as customized pipelines for NGS data processing.
RESULTS: SEQprocess was implemented in an R package with fully modularized steps for data processing that can be easily customized. Currently, six pre-customized pipelines are provided that can be easily executed by non-experts such as biomedical scientists, including the National Cancer Institute's (NCI) Genomic Data Commons (GDC) pipelines as well as the popularly used pipelines for variant calling (e.g., GATK) and estimation of allele frequency, RNA abundance (e.g., TopHat2/Cufflink), or DNA copy numbers (e.g., Sequenza). In addition, optimized pipelines for the clinical sequencing from cell-free DNA or miR-Seq are also provided. The processed data were transformed into R package-compatible data type 'ExpressionSet' or 'SummarizedExperiment', which could facilitate subsequent data analysis within R environment. Finally, an automated report summarizing the processing steps are also provided to ensure reproducibility of the NGS data analysis.
CONCLUSION: SEQprocess provides a highly extendable and R compatible framework that can manage customized and reproducible pipelines for handling multiple legacy NGS processing tools.
-
dc.language.isoen-
dc.subject.MESHData Analysis-
dc.subject.MESHHigh-Throughput Nucleotide Sequencing-
dc.subject.MESHHumans-
dc.subject.MESHReproducibility of Results-
dc.subject.MESHSoftware-
dc.subject.MESHWorkflow-
dc.titleSEQprocess: a modularized and customizable pipeline framework for NGS processing in R package-
dc.typeArticle-
dc.identifier.pmid30786880-
dc.identifier.urlhttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC6383233/-
dc.subject.keywordNext generation sequencing-
dc.subject.keywordPipeline-
dc.subject.keywordPreprocessing-
dc.subject.keywordRNA sequencing-
dc.subject.keywordWhole exome sequencing-
dc.contributor.affiliatedAuthor최, 지혜-
dc.contributor.affiliatedAuthor우, 현구-
dc.type.localJournal Papers-
dc.identifier.doi10.1186/s12859-019-2676-x-
dc.citation.titleBMC bioinformatics-
dc.citation.volume20-
dc.citation.date2019-
dc.citation.startPage90-
dc.citation.endPage90-
dc.identifier.bibliographicCitationBMC bioinformatics, 20. : 90-90, 2019-
dc.identifier.eissn1471-2105-
dc.relation.journalidJ014712105-
Appears in Collections:
Journal Papers > School of Medicine / Graduate School of Medicine > Physiology
Files in This Item:
30786880.pdfDownload

qrcode

해당 아이템을 이메일로 공유하기 원하시면 인증을 거치시기 바랍니다.

Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.

Browse