Skip navigation

Learning URI selection criteria to improve the crawling of linked open data

Learning URI selection criteria to improve the crawling of linked open data

Huang, Hai ORCID: 0000-0003-1412-0567 and Gandon, Fabien (2019) Learning URI selection criteria to improve the crawling of linked open data. In: The Semantic Web: 16th International Conference, ESWC 2019, Portorož, Slovenia, June 2–6, 2019, Proceedings. Lecture Notes in Computer Science, 11503 . Springer, Cham, Switzerland, pp. 194-208. ISBN 978-3030213473 ISSN 0302-9743 (Print), 1611-3349 (Online) (doi:https://doi.org/10.1007/978-3-030-21348-0_13)

Full text not available from this repository. (Request a copy)

Abstract

As the Web of Linked Open Data is growing the problem of crawling that cloud becomes increasingly important. Unlike normal Web crawlers, a Linked Data crawler performs a selection to focus on collecting linked RDF (including RDFa) data on the Web. From the perspectives of throughput and coverage, given a newly discovered and targeted URI, the key issue of Linked Data crawlers is to decide whether this URI is likely to dereference into an RDF data source and therefore it is worth downloading the representation it points to. Current solutions adopt heuristic rules to filter irrelevant URIs. Unfortunately, when the heuristics are too restrictive this hampers the coverage of crawling. In this paper, we propose and compare approaches to learn strategies for crawling Linked Data on the Web by predicting whether a newly discovered URI will lead to an RDF data source or not. We detail the features used in predicting the relevance and the methods we evaluated including a promising adaptation of FTRL-proximal online learning algorithm. We compare several options through extensive experiments including existing crawlers as baseline methods to evaluate their efficacy.

Item Type: Conference Proceedings
Title of Proceedings: The Semantic Web: 16th International Conference, ESWC 2019, Portorož, Slovenia, June 2–6, 2019, Proceedings
Uncontrolled Keywords: Linked Open Data, Web crawling
Subjects: Q Science > QA Mathematics > QA75 Electronic computers. Computer science
Faculty / Department / Research Group: Faculty of Liberal Arts & Sciences
Faculty of Liberal Arts & Sciences > School of Computing & Mathematical Sciences (CAM)
Last Modified: 15 Jan 2021 09:31
Selected for GREAT 2016: None
Selected for GREAT 2017: None
Selected for GREAT 2018: None
Selected for GREAT 2019: None
Selected for REF2021: None
URI: http://gala.gre.ac.uk/id/eprint/30839

Actions (login required)

View Item View Item