Hostname: page-component-89b8bd64d-r6c6k Total loading time: 0 Render date: 2026-05-08T09:47:45.990Z Has data issue: false hasContentIssue false

Improving shift-reduce constituency parsing with large-scale unlabeled data

Published online by Cambridge University Press:  19 June 2013

MUHUA ZHU
Affiliation:
Natural Language Processing Lab, Northeastern University, Shenyang 110819, China e-mails: zhumuhua@gmail.com, zhujingbo@mail.neu.edu.cn, wanghuizhen@mail.neu.edu.cn
JINGBO ZHU
Affiliation:
Natural Language Processing Lab, Northeastern University, Shenyang 110819, China e-mails: zhumuhua@gmail.com, zhujingbo@mail.neu.edu.cn, wanghuizhen@mail.neu.edu.cn
HUIZHEN WANG*
Affiliation:
Natural Language Processing Lab, Northeastern University, Shenyang 110819, China e-mails: zhumuhua@gmail.com, zhujingbo@mail.neu.edu.cn, wanghuizhen@mail.neu.edu.cn
*
Corresponding author.

Abstract

Shift-reduce parsing has been studied extensively for diverse grammars due to the simplicity and running efficiency. However, in the field of constituency parsing, shift-reduce parsers lag behind state-of-the-art parsers. In this paper we propose a semi-supervised approach for advancing shift-reduce constituency parsing. First, we apply the uptraining approach (Petrov, S. et al. 2010. In Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing (EMNLP), Cambridge, MA, USA, pp. 705–713) to improve part-of-speech taggers to provide better part-of-speech tags to subsequent shift-reduce parsers. Second, we enhance shift-reduce parsing models with novel features that are defined on lexical dependency information. Both stages depend on the use of large-scale unlabeled data. Experimental results show that the approach achieves overall improvements of 1.5 percent and 2.1 percent on English and Chinese data respectively. Moreover, the final parsing accuracies reach 90.9 percent and 82.2 percent respectively, which are comparable with the accuracy of state-of-the-art parsers.

Information

Type
Articles
Copyright
Copyright © Cambridge University Press 2013 

Access options

Get access to the full version of this content by using one of the access options below. (Log in options will check for institutional or personal access. Content may require purchase if you do not have access.)

Article purchase

Temporarily unavailable