Skip to Main Content
The vision information of Web page is applied for information extraction, which avoids using the sophisticate natural language processing technology. This paper combines the natural language processing technology with vision character of HTML page in the application of information extraction for Web page, we carried out relevant research. We propose a Web Page Information extraction algorithm based on vision character, we use the vision character rule of web page, in respect of the detailed problem of coarse-grained web page segmentation and the restructure problem of the smallest web page segmentation, we analyze the vision character of page block and finally accurate determine the topic data region. After using the information extraction technology of web page, it reduces the information block of web page content and thus reduces the cost of index generating, and also increases the hit rate of search engine.