Estudo comparativo de métodos de embeddings de grafo para tarefa de agrupamento com foco na análise de robustez ao ruído
Carregando...
Arquivos
Data
Autores
Título da Revista
ISSN da Revista
Título de Volume
Editor
Universidade Federal de São Carlos
Resumo
The production of information is currently at its highest level in history, growing year after year with no signs of slowing down in the near future. This massive volume of data originates from a wide variety of sources and formats, ranging from sensor readings collected by Internet of Things (IoT) devices, such as smartwatches, to video sharing on social media platforms. Together, these data constitute the massive body of information and the field of study known as Big Data, which can be characterized by the 5 Vs: volume, velocity, variety, veracity, and value, summarizing the richness of this virtually unlimited source of data. However, not only has data production increased, but the quality and complexity of these data have also evolved, with tables, images, and networks becoming increasingly larger and more detailed. Although these data hold great potential, allowing valuable information and insights to be extracted, their raw form is not suitable for direct use by most Machine Learning algorithms. Therefore, preprocessing is required to refine the data, thereby improving the performance of these models. Among the tasks in this field, clustering aims to identify patterns in data in order to partition them into groups based on their similarity, without relying on external labels that define such a partition. When considered in the context of graphs, clustering can be translated into the task of community detection, in which, besides the possible features associated with the nodes, the connections between them are also taken into account. Another approach to this problem is graph embedding, which transforms graph nodes into numerical vectors in a latent space while preserving notions of proximity from the original space. As a result, clustering and community detection become equivalent tasks. The embedding process can be applied to different types of data and various data components in order to provide a unified vector representation. With this in mind, the objective of this work is to evaluate the performance of different graph embedding and clustering models on a collection of graphs for the task of grouping nodes, with community detection serving as the analogous problem in the graph domain. In addition, robustness analyses were conducted to evaluate the resilience of these models to the presence of noise in graphs, since real-world data frequently contain noisy information that may remain untreated and negatively affect the quality of downstream tasks. The experimental results demonstrate that each graph scenario and combination of embedding and clustering models exhibits distinct behaviors and performance characteristics, requiring different hyperparameter configurations and presenting different levels of robustness to noise. Although promising results were obtained in some cases, the practical viability of this approach requires further investigation before it can be conclusively confirmed or ruled out.
Descrição
Palavras-chave
Citação
UENO, Igor Kenji Kawai. Estudo comparativo de métodos de embeddings de grafo para tarefa de agrupamento com foco na análise de robustez ao ruído. 2026. Trabalho de Conclusão de Curso (Graduação em Engenharia de Computação) – Universidade Federal de São Carlos, Campus São Carlos, 2026. Disponível em: https://repositorio.ufscar.br/handle/20.500.14289/24430.