An Efficient Approach for Web Indexing of Big Data through Hyperlinks in Web Crawling

Joint Authors

Devi, R. Suganya
Manjula, D.
Siddharth, R. K.

Source

The Scientific World Journal

Issue

Vol. 2015, Issue 2015 (31 Dec. 2015), pp.1-9, 9 p.

Publisher

Hindawi Publishing Corporation

Publication Date

2015-06-07

Country of Publication

Egypt

No. of Pages

9

Main Subjects

Medicine
Information Technology and Computer Science

Abstract EN

Web Crawling has acquired tremendous significance in recent times and it is aptly associated with the substantial development of the World Wide Web.

Web Search Engines face new challenges due to the availability of vast amounts of web documents, thus making the retrieved results less applicable to the analysers.

However, recently, Web Crawling solely focuses on obtaining the links of the corresponding documents.

Today, there exist various algorithms and software which are used to crawl links from the web which has to be further processed for future use, thereby increasing the overload of the analyser.

This paper concentrates on crawling the links and retrieving all information associated with them to facilitate easy processing for other uses.

In this paper, firstly the links are crawled from the specified uniform resource locator (URL) using a modified version of Depth First Search Algorithm which allows for complete hierarchical scanning of corresponding web links.

The links are then accessed via the source code and its metadata such as title, keywords, and description are extracted.

This content is very essential for any type of analyser work to be carried on the Big Data obtained as a result of Web Crawling.

American Psychological Association (APA)

Devi, R. Suganya& Manjula, D.& Siddharth, R. K.. 2015. An Efficient Approach for Web Indexing of Big Data through Hyperlinks in Web Crawling. The Scientific World Journal،Vol. 2015, no. 2015, pp.1-9.
https://search.emarefa.net/detail/BIM-1079073

Modern Language Association (MLA)

Devi, R. Suganya…[et al.]. An Efficient Approach for Web Indexing of Big Data through Hyperlinks in Web Crawling. The Scientific World Journal No. 2015 (2015), pp.1-9.
https://search.emarefa.net/detail/BIM-1079073

American Medical Association (AMA)

Devi, R. Suganya& Manjula, D.& Siddharth, R. K.. An Efficient Approach for Web Indexing of Big Data through Hyperlinks in Web Crawling. The Scientific World Journal. 2015. Vol. 2015, no. 2015, pp.1-9.
https://search.emarefa.net/detail/BIM-1079073

Data Type

Journal Articles

Language

English

Notes

Includes bibliographical references

Record ID

BIM-1079073