Compacting Frequent Star Patterns in RDF Graphs

Farah Karim; Maria-Esther Vidal; Sören Auer

doi:10.48550/arXiv.2003.05238

Details

Originalsprache	Englisch
Seiten (von - bis)	561-585
Seitenumfang	25
Fachzeitschrift	Journal of Intelligent Information Systems
Jahrgang	55
Ausgabenummer	3
Frühes Online-Datum	15 Apr. 2020
Publikationsstatus	Veröffentlicht - Dez. 2020

Abstract

Knowledge graphs have become a popular formalism for representing entities and their properties using a graph data model, e.g., the Resource Description Framework (RDF). An RDF graph comprises entities of the same type connected to objects or other entities using labeled edges annotated with properties. RDF graphs usually contain entities that share the same objects in a certain group of properties, i.e., they match star patterns composed of these properties and objects. In case the number of these entities or properties in these star patterns is large, the size of the RDF graph and query processing are negatively impacted; we refer these star patterns as frequent star patterns. We address the problem of identifying frequent star patterns in RDF graphs and devise the concept of factorized RDF graphs, which denote compact representations of RDF graphs where the number of frequent star patterns is minimized. We also develop computational methods to identify frequent star patterns and generate a factorized RDF graph, where compact RDF molecules replace frequent star patterns. A compact RDF molecule of a frequent star pattern denotes an RDF subgraph that instantiates the corresponding star pattern. Instead of having all the entities matching the original frequent star pattern, a surrogate entity is added and related to the properties of the frequent star pattern; it is linked to the entities that originally match the frequent star pattern. Since the edges between the entities and the objects in the frequent star pattern are replaced by edges between these entities and the surrogate entity of the compact RDF molecule, the size of the RDF graph is reduced. We evaluate the performance of our factorization techniques on several RDF graph benchmarks and compare with a baseline built on top gSpan, a state-of-the-art algorithm to detect frequent patterns. The outcomes evidence the efficiency of proposed approach and show that our techniques are able to reduce execution time of the baseline approach in at least three orders of magnitude. Additionally, RDF graph size can be reduced by up to 66.56% while data represented in the original RDF graph is preserved.

ASJC Scopus Sachgebiete

Informatik (insg.)
Software
Informatik (insg.)
Artificial intelligence
Informatik (insg.)
Information systems
Informatik (insg.)
Hardware und Architektur
Informatik (insg.)
Computernetzwerke und -kommunikation

Zitieren

Compacting Frequent Star Patterns in RDF Graphs. / Karim, Farah; Vidal, Maria-Esther; Auer, Sören.
in: Journal of Intelligent Information Systems, Jahrgang 55, Nr. 3, 12.2020, S. 561-585.

Publikation: Beitrag in Fachzeitschrift › Artikel › Forschung

Karim, F, Vidal, M-E & Auer, S 2020, 'Compacting Frequent Star Patterns in RDF Graphs', Journal of Intelligent Information Systems, Jg. 55, Nr. 3, S. 561-585. https://doi.org/10.48550/arXiv.2003.05238, https://doi.org/10.1007/s10844-020-00595-9

Karim, F., Vidal, M.-E., & Auer, S. (2020). Compacting Frequent Star Patterns in RDF Graphs. Journal of Intelligent Information Systems, 55(3), 561-585. https://doi.org/10.48550/arXiv.2003.05238, https://doi.org/10.1007/s10844-020-00595-9

Karim F, Vidal ME, Auer S. Compacting Frequent Star Patterns in RDF Graphs. Journal of Intelligent Information Systems. 2020 Dez;55(3):561-585. Epub 2020 Apr 15. doi: 10.48550/arXiv.2003.05238, 10.1007/s10844-020-00595-9

Karim, Farah ; Vidal, Maria-Esther ; Auer, Sören. / Compacting Frequent Star Patterns in RDF Graphs. in: Journal of Intelligent Information Systems. 2020 ; Jahrgang 55, Nr. 3. S. 561-585.

Download

@article{010be1fd82ee420fb925744c34b32530,

title = "Compacting Frequent Star Patterns in RDF Graphs",

abstract = "Knowledge graphs have become a popular formalism for representing entities and their properties using a graph data model, e.g., the Resource Description Framework (RDF). An RDF graph comprises entities of the same type connected to objects or other entities using labeled edges annotated with properties. RDF graphs usually contain entities that share the same objects in a certain group of properties, i.e., they match star patterns composed of these properties and objects. In case the number of these entities or properties in these star patterns is large, the size of the RDF graph and query processing are negatively impacted; we refer these star patterns as frequent star patterns. We address the problem of identifying frequent star patterns in RDF graphs and devise the concept of factorized RDF graphs, which denote compact representations of RDF graphs where the number of frequent star patterns is minimized. We also develop computational methods to identify frequent star patterns and generate a factorized RDF graph, where compact RDF molecules replace frequent star patterns. A compact RDF molecule of a frequent star pattern denotes an RDF subgraph that instantiates the corresponding star pattern. Instead of having all the entities matching the original frequent star pattern, a surrogate entity is added and related to the properties of the frequent star pattern; it is linked to the entities that originally match the frequent star pattern. Since the edges between the entities and the objects in the frequent star pattern are replaced by edges between these entities and the surrogate entity of the compact RDF molecule, the size of the RDF graph is reduced. We evaluate the performance of our factorization techniques on several RDF graph benchmarks and compare with a baseline built on top gSpan, a state-of-the-art algorithm to detect frequent patterns. The outcomes evidence the efficiency of proposed approach and show that our techniques are able to reduce execution time of the baseline approach in at least three orders of magnitude. Additionally, RDF graph size can be reduced by up to 66.56% while data represented in the original RDF graph is preserved.",

keywords = "cs.DB, Semantic Web, Linked data, Knowledge graph, RDF compaction",

author = "Farah Karim and Maria-Esther Vidal and S{\"o}ren Auer",

note = "Funding information: Farah Karim is supported by the German Academic Exchange Service (DAAD); this work is partially funded by the EU H2020 project IASiS (GA No.727658).",

year = "2020",

month = dec,

doi = "10.48550/arXiv.2003.05238",

language = "English",

volume = "55",

pages = "561--585",

journal = "Journal of Intelligent Information Systems",

issn = "0925-9902",

publisher = "Springer Netherlands",

number = "3",

}

Download

TY - JOUR

T1 - Compacting Frequent Star Patterns in RDF Graphs

AU - Karim, Farah

AU - Vidal, Maria-Esther

AU - Auer, Sören

N1 - Funding information: Farah Karim is supported by the German Academic Exchange Service (DAAD); this work is partially funded by the EU H2020 project IASiS (GA No.727658).

PY - 2020/12

Y1 - 2020/12

N2 - Knowledge graphs have become a popular formalism for representing entities and their properties using a graph data model, e.g., the Resource Description Framework (RDF). An RDF graph comprises entities of the same type connected to objects or other entities using labeled edges annotated with properties. RDF graphs usually contain entities that share the same objects in a certain group of properties, i.e., they match star patterns composed of these properties and objects. In case the number of these entities or properties in these star patterns is large, the size of the RDF graph and query processing are negatively impacted; we refer these star patterns as frequent star patterns. We address the problem of identifying frequent star patterns in RDF graphs and devise the concept of factorized RDF graphs, which denote compact representations of RDF graphs where the number of frequent star patterns is minimized. We also develop computational methods to identify frequent star patterns and generate a factorized RDF graph, where compact RDF molecules replace frequent star patterns. A compact RDF molecule of a frequent star pattern denotes an RDF subgraph that instantiates the corresponding star pattern. Instead of having all the entities matching the original frequent star pattern, a surrogate entity is added and related to the properties of the frequent star pattern; it is linked to the entities that originally match the frequent star pattern. Since the edges between the entities and the objects in the frequent star pattern are replaced by edges between these entities and the surrogate entity of the compact RDF molecule, the size of the RDF graph is reduced. We evaluate the performance of our factorization techniques on several RDF graph benchmarks and compare with a baseline built on top gSpan, a state-of-the-art algorithm to detect frequent patterns. The outcomes evidence the efficiency of proposed approach and show that our techniques are able to reduce execution time of the baseline approach in at least three orders of magnitude. Additionally, RDF graph size can be reduced by up to 66.56% while data represented in the original RDF graph is preserved.

AB - Knowledge graphs have become a popular formalism for representing entities and their properties using a graph data model, e.g., the Resource Description Framework (RDF). An RDF graph comprises entities of the same type connected to objects or other entities using labeled edges annotated with properties. RDF graphs usually contain entities that share the same objects in a certain group of properties, i.e., they match star patterns composed of these properties and objects. In case the number of these entities or properties in these star patterns is large, the size of the RDF graph and query processing are negatively impacted; we refer these star patterns as frequent star patterns. We address the problem of identifying frequent star patterns in RDF graphs and devise the concept of factorized RDF graphs, which denote compact representations of RDF graphs where the number of frequent star patterns is minimized. We also develop computational methods to identify frequent star patterns and generate a factorized RDF graph, where compact RDF molecules replace frequent star patterns. A compact RDF molecule of a frequent star pattern denotes an RDF subgraph that instantiates the corresponding star pattern. Instead of having all the entities matching the original frequent star pattern, a surrogate entity is added and related to the properties of the frequent star pattern; it is linked to the entities that originally match the frequent star pattern. Since the edges between the entities and the objects in the frequent star pattern are replaced by edges between these entities and the surrogate entity of the compact RDF molecule, the size of the RDF graph is reduced. We evaluate the performance of our factorization techniques on several RDF graph benchmarks and compare with a baseline built on top gSpan, a state-of-the-art algorithm to detect frequent patterns. The outcomes evidence the efficiency of proposed approach and show that our techniques are able to reduce execution time of the baseline approach in at least three orders of magnitude. Additionally, RDF graph size can be reduced by up to 66.56% while data represented in the original RDF graph is preserved.

KW - cs.DB

KW - Semantic Web

KW - Linked data

KW - Knowledge graph

KW - RDF compaction

UR - http://www.scopus.com/inward/record.url?scp=85083792835&partnerID=8YFLogxK

U2 - 10.48550/arXiv.2003.05238

DO - 10.48550/arXiv.2003.05238

M3 - Article

VL - 55

SP - 561

EP - 585

JO - Journal of Intelligent Information Systems

JF - Journal of Intelligent Information Systems

SN - 0925-9902

IS - 3

ER -

Research@Leibniz University

Compacting Frequent Star Patterns in RDF Graphs

Autoren

Organisationseinheiten

Externe Organisationen

Details

Abstract

ASJC Scopus Sachgebiete

Zitieren

Von denselben Autoren

DataDesc: A framework for creating and sharing technical metadata for research software interfaces

Federated Querying of Scholarly Communication Infrastructures

A Reputation System for Scientific Contributions Based on a Token Economy

SWARM-SLR: Streamlined Workflow Automation for Machine-Actionable Systematic Literature Reviews

Effective Context Selection in LLM-Based Leaderboard Generation: An Empirical Study