Hostname: page-component-5d84bcc8dc-lb775 Total loading time: 0 Render date: 2026-08-19T18:24:51.219Z Has data issue: false hasContentIssue false

Weighted finite-state transducers for normalization of historical texts

Published online by Cambridge University Press:  01 April 2019

Izaskun Etxeberria*
Affiliation:
IXA Group, University of the Basque Country, Donostia-San Sebastián, Spain
Iñaki Alegria
Affiliation:
IXA Group, University of the Basque Country, Donostia-San Sebastián, Spain
Larraitz Uria
Affiliation:
IXA Group, University of the Basque Country, Donostia-San Sebastián, Spain
*
*Corresponding author. Email: izaskun.etxeberria@ehu.eus

Abstract

This paper presents a study about methods for normalization of historical texts. The aim of these methods is learning relations between historical and contemporary word forms. We have compiled training and test corpora for different languages and scenarios, and we have tried to read the results related to the features of the corpora and languages. Our proposed method, based on weighted finite-state transducers, is compared to previously published ones. Our method learns to map phonological changes using a noisy channel model; it is a simple solution that can use a limited amount of supervision in order to achieve adequate performance. The compiled corpora are ready to be used for other researchers in order to compare results. Concerning the amount of supervision for the task, we investigate how the size of training corpus affects the results and identify some interesting factors to anticipate the difficulty of the task.

Information

Type
Article
Copyright
© Cambridge University Press 2019 

Access options

Get access to the full version of this content by using one of the access options below. (Log in options will check for institutional or personal access. Content may require purchase if you do not have access.)

Article purchase

Temporarily unavailable