Scalable feature selection, classification and signature generation for organizing large text databases into hierarchical topic taxonomies.

Chakrabarti, Soumen [y otros]

Scalable feature selection, classification and signature generation for organizing large text databases into hierarchical topic taxonomies.

We explore how to organize large text databases hierarchically by topic to aid better searching, browsing and filtering. Many corpora, such as internet directories, digital libraries, and patent databases are manually organized into topic hierarchies, also called taxonomies. Similar to indices for relational data, taxonomies make search and access more efficient. However, the exponential growth in the volume of on-line textual information makes it nearly impossible to maintain such taxonomic organization for large, fast-changing corpora by hand.


TAXONOMIES
ORGANIZING DATABASES

H004.65 VER