Quotation Extraction for Portuguese

Lieferzeit: Lieferbar innerhalb 14 Tagen

35,90 

ISBN: 3659288152
ISBN 13: 9783659288159
Autor: Ducca Fernandes, William Paulo/Milidiú, Ruy Luiz
Verlag: LAP LAMBERT Academic Publishing
Umfang: 64 S.
Erscheinungsdatum: 09.01.2018
Auflage: 1/2018
Format: 0.5 x 22 x 15
Gewicht: 113 g
Produktform: Kartoniert
Einband: Kartoniert
Artikelnummer: 3467522 Kategorie:

Beschreibung

Quotation Extraction consists of identifying quotations from a text and associating them to their authors. In this work, we present a Quotation Extraction system for Portuguese. Quotation Extraction has been previously approached using dierent techniques and for several languages. Our proposal diers from previous work since we use Machine Learning to automatically build specialized rules instead of human-derived rules. Machine Learning models usually present stronger generalization power compared to human-derived models. In addition, we are able to easily adapt our model to other languages, needing only a list of verbs of speech for a given language. The previously proposed systems would probably need a rule set adaptation to correctly classify the quotations, which would be time consuming. We tackle the Quotation Extraction task using one model for the Entropy Guided Transformation Learning algorithm and another one for the Structured Perceptron algorithm. In order to train and evaluate the system, we have build the GloboQuotes corpus, with news extracted from the globo.com portal.

Autorenporträt

Graduated in 2008 from the Universidade Federal de Juiz de Fora (UFJF) in Computer Science. Has a Masters in Informatics from the Pontifícia Universidade Católica do Rio de Janeiro (PUC-Rio). Nowadays is PhD student in Informatics at PUC-Rio. His research focuses on Natural Language Processing, Information Extraction and Machine Learning.

Herstellerkennzeichnung:


OmniScriptum SRL
Str. Armeneasca 28/1, office 1
2012 Chisinau
MD

E-Mail: info@omniscriptum.com

Das könnte Ihnen auch gefallen …