Font Size: a A A

A multi-modular approach to model selection in statistical natural language processing

Posted on:2003-08-28Degree:Ph.DType:Thesis
University:The University of ChicagoCandidate:Higgins, Derrick CharlesFull Text:PDF
GTID:2468390011479268Subject:Language
Abstract/Summary:
The effectiveness of statistical methods of inferring a probabilistic language model from observed data is limited by the set of models which they may select from the hypothesis space. This dissertation will argue that the models typically used in computational linguistics, consisting of monolithic representations on a single linguistic level, could be improved upon by incorporating other aspects of the full linguistic structure. For example, parsing models ought to benefit from the availability of semantic information about a sentence.; This study will investigate under what circumstances it is possible to identify the best set of components which comprise a grammar in a linguistic model which is multi-modular. The framework for this investigation is statistical language modeling, which has proven an effective way of solving problems in applied natural language processing, and is becoming an ever more popular approach to the problem of inferring grammars on the basis of language data.; We present paradigm cases of computational models using information from multiple modules of linguistic structure. In the first experiment, we investigate how a multi-modular account of quantifier scope preferences can be constructed on the basis of tagged data from a treebank, so that the scope module interacts with an independent SPSG “treebank grammar” of the traditional sort. In the second experiment, we apply the same methodology to an investigation of the phenomenon known to theoretical linguists as A-bar dependencies . In the final experiment, we turn to unsupervised learning methods in constructing a model of the induction of morphological categories which is comprised of multiple interacting components.; These examples are intended to provide evidence of the feasibility of integrating multi-modular theories of grammar with the tools of statistical language modeling, and suggest ways in which corpus-driven training of these models can provide an empirically grounded way of addressing linguistic questions.
Keywords/Search Tags:Model, Language, Statistical, Multi-modular, Linguistic
Related items