Database

Using Ontology Fingerprints to disambiguate gene name entities in the biomedical literature

Chen, G., Zhao, J., Cohen, T., Tao, C., Sun, J., Xu, H., Bernstam, E. V., Lawson, A., Zeng, J., Johnson, A. M., Holla, V., Bailey, A. M., Lara-Guerra, H., Litzenburger, B., Meric-Bernstam, F., Jim Zheng, W..

Ambiguous gene names in the biomedical literature are a barrier to accurate information extraction. To overcome this hurdle, we generated Ontology Fingerprints for selected genes that are relevant for personalized cancer therapy. These Ontology Fingerprints were used to evaluate the association between genes and biomedical literature to disambiguate gene names. We obtained 93.6% precision for the test gene set and 80.4% for the area under a receiver-operating characteristics curve for gene and article association. The core algorithm was implemented using a graphics processing unit-based MapReduce framework to handle big data and to improve performance. We conclude that Ontology Fingerprints can help disambiguate gene names mentioned in text and analyse the association between genes and articles.

Database URL: http://www.ontologyfingerprint.org