谷歌浏览器插件
订阅小程序
在清言上使用

A Novel Deep Contrastive Convolutional Autoencoder Based Binning Approach for Taxonomic Independent Metagenomics Data

Sharanbasappa D. Madival, Girish Kumar Jha,Dwijesh Chandra Mishra, Sunil Kumar,Neeraj Budhlakoti,Anu Sharma,Krishna Kumar Chaturvedi, S. Kabilan,Mohammad Samir Farooqi, Sudhir Srivastava

Journal of Plant Biochemistry and Biotechnology(2024)

引用 0|浏览0
暂无评分
摘要
In this study, we present an innovative binning approach for metagenomics data that combines Natural Language Processing (NLP) with a Deep Contrastive Convolutional Autoencoder (DCAE). We used NLP for feature extraction, specifically focusing on Tetra-nucleotide frequency (TNF) through CountVec and (Term Frequency -Inverse Document Frequency) TF-IDF, further enriched by integrating GC-Content into their respective feature matrices. The DCAE, equipped with advanced convolutional layers and a contrastive loss function, excels at capturing intricate patterns in the data, providing a sophisticated representation for binning. By applying k-means clustering to the latent representations obtained from the DCAE, our approach consistently achieves impressive results. To assess the performance of our method, we utilized three standard benchmark metagenomics datasets: 10s, 25s, and Sharon datasets. Across all datasets, we observed Silhouette Indices exceeding 0.6 and Rand Indices surpassing 0.8, demonstrating the superior performance of our proposed method. Compared to existing methodologies, our approach not only surpasses the Rand Index and Silhouette Index of current unsupervised methods but also performs on par with semi-supervised methods across datasets. This underscores the effectiveness and versatility of our approach in metagenomics analysis.
更多
查看译文
关键词
Binning,Deep contrastive convolutional autoencoder,Natural language processing,Metagenomics,Tetra nucleotide frequency
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要