• português (Brasil)
    • English
    • español
  • English 
    • português (Brasil)
    • English
    • español
  • Login
About
  • Policies
  • Instructions to authors
  • Contact
    • Policies
    • Instructions to authors
    • Contact
View Item 
  •   Home
  • Centro de Ciências Exatas e de Tecnologia - CCET
  • Programas de Pós-Graduação
  • Ciência da Computação - PPGCC
  • Teses e dissertações
  • View Item
  •   Home
  • Centro de Ciências Exatas e de Tecnologia - CCET
  • Programas de Pós-Graduação
  • Ciência da Computação - PPGCC
  • Teses e dissertações
  • View Item
JavaScript is disabled for your browser. Some features of this site may not work without it.

Browse

All of DSpaceCommunities & CollectionsBy Issue DateAuthorsAdvisorTitlesSubjectsCNPq SubjectsGraduate ProgramDocument TypeThis CollectionBy Issue DateAuthorsAdvisorTitlesSubjectsCNPq SubjectsGraduate ProgramDocument Type

My Account

Login

Extração automática de relações semânticas a partir de textos escritos em português do Brasil

Thumbnail
View/Open
5456.pdf (1.808Mb)
Date
2013-07-11
Author
Taba, Leonardo Sameshima
Metadata
Show full item record
Abstract
Information extraction (IE) is one of the many applications in Natural Language Processing (NLP); it focuses on processing texts in order to retrieve specific information about a certain entity or concept. One of its subtasks is the automatic extraction of semantic relations between terms, which is very useful in the construction and improvement of linguistic resources such as ontologies and lexical bases. Moreover, there s a rising demand for semantic knowledge, as many computational NLP systems need that information in their processing. Applications such as information retrieval from web documents and automatic translation to other languages could benefit from that kind of knowledge. However, there aren t sufficient human resources to produce that knowledge at the same rate of its demand. Aiming to solve that semantic data scarcity problem, this work investigates how binary semantic relations can be automatically extracted from Brazilian Portuguese texts. These relations are based on Minsky s (1986) theory and are used to represent common sense knowledge in the Open Mind Common Sense no Brasil (OMCS-Br) project developed at LIA (Laboratório de Interação Avanc¸ada), partner of LaLiC (Laborat´orio de Lingu´ıstica Computacional), where this research was conducted, both in Universidade Federal de São Carlos (UFSCar). The first strategies for this task were based on searching textual patterns in texts, where a certain textual expression indicates that there is a specific relation between two terms in a sentence. This approach has high precision but low recall, which led to the research of methods that use machine learning as their main model, encompassing techniques such as probabilistic and statistical classifiers and also kernel methods, which currently figure among the state of the art. Therefore, this work investigates, implements and evaluates some of these techniques in order to determine how and to which extent they can be applied to the automatic extraction of binary semantic relations in Portuguese texts. In that way, this work is an important step in the advancement of the state of the art in information extraction for the Portuguese language, which still lacks resources in the semantic area, and also advances the Portuguese language NLP scenario as a whole.
URI
https://repositorio.ufscar.br/handle/ufscar/543
Collections
  • Teses e dissertações

UFSCar
Universidade Federal de São Carlos - UFSCar
Send Feedback

UFSCar

IBICT
 

 


UFSCar
Universidade Federal de São Carlos - UFSCar
Send Feedback

UFSCar

IBICT