Multilingual translation for zero-shot biomedical classification using BioTranslator

Nature Communications(2023)

引用 2|浏览69
暂无评分
摘要
Existing annotation paradigms rely on controlled vocabularies, where each data instance is classified into one term from a predefined set of controlled vocabularies. This paradigm restricts the analysis to concepts that are known and well-characterized. Here, we present the novel multilingual translation method BioTranslator to address this problem. BioTranslator takes a user-written textual description of a new concept and then translates this description to a non-text biological data instance. The key idea of BioTranslator is to develop a multilingual translation framework, where multiple modalities of biological data are all translated to text. We demonstrate how BioTranslator enables the identification of novel cell types using only a textual description and how BioTranslator can be further generalized to protein function prediction and drug target identification. Our tool frees scientists from limiting their analyses within predefined controlled vocabularies, enabling them to interact with biological data using free text. Here, the authors develop the cross-modal translation method BioTranslator to translate the textual description to non-text biological data. This approach frees scientists from limiting their analysis within predefined controlled vocabularies.
更多
查看译文
关键词
Classification and taxonomy,Computational models,Data mining,Science,Humanities and Social Sciences,multidisciplinary
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要