NCBI
PubMed
A service of the
U.S. National Library of Medicine
and the
National Institutes of Health
My NCBI
[Sign In]
[Register]
All Databases
PubMed
Nucleotide
Protein
Genome
Structure
OMIM
PMC
Journals
Books
Search
Database name
PubMed
Protein
Nucleotide
GSS
EST
Structure
Genome
Books
CancerChromosomes
Conserved Domains
dbGaP
3D Domains
Gene
Genome Project
GENSAT
GEO Profiles
GEO DataSets
HomoloGene
Journals
MeSH
NCBI Web Site
NLM Catalog
OMIA
OMIM
PMC
PopSet
Probe
Protein Clusters
PubChem BioAssay
PubChem Compound
PubChem Substance
SNP
SRA
Taxonomy
ToolKit
ToolKitAll
UniGene
UniSTS
for
Search term
Go
Clear
Advanced Search
Limits
Preview/Index
History
Clipboard
Details
Your browser version may not work well with NCBI's Web applications. More information
here...
Display
Summary
Brief
Abstract
AbstractPlus
Citation
MEDLINE
XML
UI List
LinkOut
ASN.1
Related Articles
Cited in Books
CancerChrom Links
Domain Links
3D Domain Links
dbGaP Links
GEO DataSet Links
Gene Links
Gene (OMIM) Links
Gene (GeneRIF) Links
Genome Links
Project Links
GENSAT Links
GEO Profile Links
HomoloGene Links
Nucleotide Links
Nucleotide (RefSeq) Links
Nucleotide (Weighted) Links
EST Links
EST (RefSeq) Links
GSS Links
GSS (RefSeq) Links
OMIA Links
OMIM (calculated) Links
OMIM (cited) Links
BioAssay Links
Compound Links
Compound (MeSH Keyword)
Compound (Publisher) Links
Substance Links
Substance (MeSH Keyword)
Substance (Publisher) Links
PMC Links
Cited in PMC
PopSet Links
Probe Links
Protein Links
Protein (RefSeq) Links
Protein (Weighted) Links
Protein Cluster Links
Cited Articles
SNP Links
SNP (Cited)
Structure Links
Taxonomy via GenBank
UniGene Links
UniSTS Links
Show
5
10
20
50
100
200
500
Sort By
Pub Date
First Author
Last Author
Journal
Title
Send to
Text
File
Printer
Clipboard
Collections
E-mail
Order
All: 1
Review: 0
Click to change filter selection through MyNCBI.
1:
Bioinformatics.
2005 Jan 15;21(2):248-56. Epub 2004 Aug 27.
Related Articles
,
Links
Gene name ambiguity of eukaryotic nomenclatures.
Chen L
,
Liu H
,
Friedman C
.
Department of BioMedical Informatics, Columbia University New York, NY 10032, USA. lifeng.chen@dbmi.columbia.edu
MOTIVATION: With more and more scientific literature published online, the effective management and reuse of this knowledge has become problematic. Natural language processing (NLP) may be a potential solution by extracting, structuring and organizing biomedical information in online literature in a timely manner. One essential task is to recognize and identify genomic entities in text. 'Recognition' can be accomplished using pattern matching and machine learning. But for 'identification' these techniques are not adequate. In order to identify genomic entities, NLP needs a comprehensive resource that specifies and classifies genomic entities as they occur in text and that associates them with normalized terms and also unique identifiers so that the extracted entities are well defined. Online organism databases are an excellent resource to create such a lexical resource. However, gene name ambiguity is a serious problem because it affects the appropriate identification of gene entities. In this paper, we explore the extent of the problem and suggest ways to address it. RESULTS: We obtained gene information from 21 organisms and quantified naming ambiguities within species, across species, with English words and with medical terms. When the case (of letters) was retained, official symbols displayed negligible intra-species ambiguity (0.02%) and modest ambiguities with general English words (0.57%) and medical terms (1.01%). In contrast, the across-species ambiguity was high (14.20%). The inclusion of gene synonyms increased intra-species ambiguity substantially and full names contributed greatly to gene-medical-term ambiguity. A comprehensive lexical resource that covers gene information for the 21 organisms was then created and used to identify gene names by using a straightforward string matching program to process 45,000 abstracts associated with the mouse model organism while ignoring case and gene names that were also English words. We found that 85.1% of correctly retrieved mouse genes were ambiguous with other gene names. When gene names that were also English words were included, 233% additional 'gene' instances were retrieved, most of which were false positives. We also found that authors prefer to use synonyms (74.7%) to official symbols (17.7%) or full names (7.6%) in their publications. CONTACT: lifeng.chen@dbmi.columbia.edu
Publication Types:
Comparative Study
Evaluation Studies
Research Support, Non-U.S. Gov't
Research Support, U.S. Gov't, Non-P.H.S.
Validation Studies
PMID: 15333458 [PubMed - indexed for MEDLINE]
Display
Summary
Brief
Abstract
AbstractPlus
Citation
MEDLINE
XML
UI List
LinkOut
ASN.1
Related Articles
Cited in Books
CancerChrom Links
Domain Links
3D Domain Links
dbGaP Links
GEO DataSet Links
Gene Links
Gene (OMIM) Links
Gene (GeneRIF) Links
Genome Links
Project Links
GENSAT Links
GEO Profile Links
HomoloGene Links
Nucleotide Links
Nucleotide (RefSeq) Links
Nucleotide (Weighted) Links
EST Links
EST (RefSeq) Links
GSS Links
GSS (RefSeq) Links
OMIA Links
OMIM (calculated) Links
OMIM (cited) Links
BioAssay Links
Compound Links
Compound (MeSH Keyword)
Compound (Publisher) Links
Substance Links
Substance (MeSH Keyword)
Substance (Publisher) Links
PMC Links
Cited in PMC
PopSet Links
Probe Links
Protein Links
Protein (RefSeq) Links
Protein (Weighted) Links
Protein Cluster Links
Cited Articles
SNP Links
SNP (Cited)
Structure Links
Taxonomy via GenBank
UniGene Links
UniSTS Links
Show
5
10
20
50
100
200
500
Sort By
Pub Date
First Author
Last Author
Journal
Title
Send to
Text
File
Printer
Clipboard
Collections
E-mail
Order
About Entrez
Text Version
Entrez PubMed
Overview
Help
|
FAQ
Tutorials
New/Noteworthy
E-Utilities
PubMed Services
Journals Database
MeSH Database
Single Citation Matcher
Batch Citation Matcher
Clinical Queries
Special Queries
LinkOut
My NCBI
Related Resources
Order Documents
NLM Mobile
NLM Catalog
NLM Gateway
TOXNET
Consumer Health
Clinical Alerts
ClinicalTrials.gov
PubMed Central
Write to the Help Desk
NCBI
|
NLM
|
NIH
Department of Health & Human Services
Privacy Statement
|
Freedom of Information Act
|
Disclaimer