The storage and data transferring of large genome data are becoming important concerns for biomedical researchers. We present a novel multi-reference based genome compression method with a hierachical structure. Our approach works for the de facto standard alignment format (i.e., BAM) compression that is the pressing need at present. We align new sequences to a reference sequence using SOAP3, a GPU-based aligning software, and summarize mapping properties and information for exact mapped reads. To increase the exact aligning rate, we also realign the approximately mapped and unmapped reads by changing the reference sequence or shortening the read length. Meanwhile, we further the study using “lossy” quality values through k-means clustering scheme and find its minute effect on downstream applications. The proposed method has achieved compression ratios from 0.5 to 0.65, which corresponds to space savings of 35%-50%, on experimental datasets.

Features

  • Efficient
  • Fast
  • Promising

Project Samples

Project Activity

See All Activity >

Categories

Bio-Informatics

License

Academic Free License (AFL)

Follow Hierachical_DNAcoder

Hierachical_DNAcoder Web Site

Other Useful Business Software
Contractor Foreman is the most affordable all-in-one construction management software for contractors and is trusted by contractors in more than 75 countries. Icon
Contractor Foreman is the most affordable all-in-one construction management software for contractors and is trusted by contractors in more than 75 countries.

For Residential, Commercial and Public Works Contractors

Starting at $49/m for the WHOLE company, Contractor Foreman is the most affordable all-in-one construction management system for contractors. Our customers in 75+ countries and industry awards back it up. And it's all backed by a 100 day guarantee.
Learn More
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of Hierachical_DNAcoder!

Additional Project Details

Operating Systems

Linux

Languages

English

Intended Audience

Engineering, Information Technology, Science/Research

Programming Language

C++

Related Categories

C++ Bio-Informatics Software

Registered

2013-06-18