Repository | Book | Chapter

176380

Knowledge discovery from semistructured texts

Hiroshi Sakamoto Hiroki Arimura Setsuo Arikawa

pp. 586-599

Abstract

This paper surveys our recent results on the knowledge discovery from semistructured texts, which contain heterogeneous structures represented by labeled trees. The aim of our study is to extract useful information from documents on the Web. First, we present the theoretical results on learning rewriting rules between labeled trees. Second, we apply our method to the learning HTML trees in the framework of the wrapper induction. We also examine our algorithms for real world HTML documents and present the results.

Publication details

Published in:

Arikawa Setsuo, Shinohara Ayumi (2002) Progress in discovery science: final report of the Japanese discovery science project. Dordrecht, Springer.

Pages: 586-599

DOI: 10.1007/3-540-45884-0_45

Full citation:

Sakamoto Hiroshi, Arimura Hiroki, Arikawa Setsuo (2002) „Knowledge discovery from semistructured texts“, In: S. Arikawa & A. Shinohara (eds.), Progress in discovery science, Dordrecht, Springer, 586–599.