Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
In this paper, we describe our hybrid approach to two key NLP technologies: biomedical named entity recognition (Bio-NER) and (Bio-SRL). In Bio-NER, our system successfully integrates linguistic features into the CRF framework. In addition, we employ web lexicons and template-based post-processing to further boost its performance. Through these broad linguistic features and the nature of CRF, our system outperforms state-of-the-art machine-learning-based systems, especially in the recognition of protein names (F=78.5%). In Bio-SRL, first, we construct a proposition bank on top of the popular biomedical GENIA treebank following the PropBank annotation scheme. We only annotate the predicate-argument structures (PAS's) of thirty frequently used biomedical verbs (predicates) and their corresponding arguments. Second, we use our proposition bank to train a biomedical SRL system, which uses a maximum entropy (ME) machine-learning model. Thirdly, we automatically generate argument-type templates, which can be used to improve classification of biomedical argument roles. Our experimental results show that a newswire English SRL system that achieves an F-score of 86.29% in the newswire English domain can maintain an F-score of 64.64% when ported to the biomedical domain. By using our annotated biomedical corpus, we can increase that F-score by 22.9%. Adding automatically generated template features further increases overall F-score by 0.47% and adjunct (AM) F-score by 1.57%, respectively.
This paper explores several important issues in developing syntactically annotated Korean corpora for higher-level language processing, including semantic-discourse parsing, question-answering, machine translation, information retrieval, etc.In particular, we compare the Penn Korean Treebank (PKT) and the Korean Treebank of the 21st Century Sejong Project (ST) and discuss four critical issues in syntactic annotation.We argue for the use of more sophisticated morphosyntactic information, and based on our comparative study, we propose revisions in the syntactic annotation schemes of the existing Korean Treebanks in order to improve the quality of annotated corpora and their usability both for conducting theoretical research and for developing computational tools.The results of our comparative study reveal four significant issues in syntactic annotations: the syntactic analysis of verbal complexes, the hierarchical structure of noun phrases, the representation of traces, and the marking of zero elements.These factors may trigger erroneous syntactic representations for certain linguistic phenomena, and they may increase difficulties in data search and lessen reliability in computational processing.Thus, evaluating and improving the syntactic annotation of Treebanks is an important task for aspects of both theoretical and computational linguistics.
This study uses a spatial logit model to evaluate the statistical effect of conditions of communities on municipal bond ratings. It finds that private (non-farm) earnings in the community positively explain bond ratings with statistical significance, while earnings from personal transfers negatively affect ratings. Own source of revenues of local governments (local taxes) increase ratings and inter-governmental revenues (transfers to local governments) negatively impact ratings. Outstanding debt fails to significantly explain ratings. The composition of the local economy (e.g., the service sector) weights heavily in a rating, and proximity of the local government to areas with high municipal bond ratings increases ratings.
EDBL (Euskararen Datu-Base Lexikala) is a general-purpose lexical database used in Basque text-processing tasks. It is a large repository of lexical knowledge (currently around 80,000 entries) that acts as basis and support in a number of different NLP tasks, thus providing lexical information for several language tools: morphological analysis, spell checking and correction, lemmatization and tagging, syntactic analysis, and so on. It has been designed to be neutral in relation to the different linguistic formalisms, and flexible and open enough to accept new types of information. A browser-based user interface makes the job of consulting the database, correcting and updating entries, adding new ones, etc. easy to the lexicographer.
In the past, a divide could be seen between ’deep ’ parsers on the one hand, which construct a semantic representation out of their input, but usually have significant coverage problems, and more robust parsers on the other hand, which are usually based on a (statistical) model derived from a treebank and have larger coverage,
Unsupervised morphological analysis is the task of segmenting words into prefixes, suffixes and stems without prior knowledge of language-specific morphotactics and morpho-phonological rules. This paper introduces a simple, yet highly effective algorithm for unsupervised morphological learning for Bengali, an Indo–Aryan language that is highly inflectional in nature. When evaluated on a set of 4,110 human-segmented Bengali words, our algorithm achieves an F-score of 83%, substantially outperforming Linguistica, one of the most widely-used unsupervised morphological parsers, by about 23%.
In this thesis lateralization of olfactory functions was investigated by both behavioral and electrophysiological assessment, the latter with the olfactory event-related potential (OERP) technique. The olfactory sense is primarily ipsilateral in that a stimulus that is presented to one nostril is initially processed in the same hemisphere. This makes it possible to observe differences between stimulated nostrils as an indication of hemispheric difference. Study I explored differences in olfactory cognitive functions with respect to side of rhinal stimulation and demonstrated that familiarity ratings are higher at right- compared to left-nostril stimulation. No differences were found in episodic recognition memory or free identification, possibly reflecting inter-hemispheric interactions in higher cognitive functions. Effects of repetition priming were present in odor identification and tended to be more pronounced when tested via left nostril. Study II further investigated the effect of previous exposure in odor identification by a different experimental set-up, and demonstrated effects of repetition priming when tested via left- but not right-nostril stimulation. This finding indicates the importance of reconsidering possible sequential effects in olfactory research. Study III examined methodological aspects of an OERP protocol with respect to stimulus duration, which was used in Study IV. No differences in amplitudes or latencies where found between the stimulus durations of 150, 200 and 250 ms, suggesting the commonly used duration of 200 ms in a standard protocol. Study IV investigated laterality effects in OERPs with respect to side of stimulation and electrode site. The results showed consistent amplitudes and latencies regardless of rhinal side of stimulation. Larger amplitudes were demonstrated on left hemisphere and midline compared to right hemisphere, possibly explained by smaller N1/P2 amplitudes at the right-hemisphere sites at left-nostril stimulation. Apart from a proposed OERP protocol, the findings support the notions of a right-hemisphere predominance in processes related to olfactory perception and indicate, in accordance with other findings, a left-side advantage in conceptual repetition priming.
Abstract We present a new method for learning to parse a bilingual sentence using Inversion Transduction Grammar trained on a parallel corpus and a monolingual treebank. The method produces a parse tree for a bilingual sentence, showing the shared syntactic structures of individual sentence and the differences of word order within a syntactic structure. The method involves estimating lexical translation probability based on a word-aligning strategy and inferring probabilities for CFG rules. At runtime, a bottom-up CYK-styled parser is employed to construct the most probable bilingual parse tree for any given sentence pair. We also describe an implementation of the proposed method. The experimental results indicate the proposed model produces word alignments better than those produced by Giza++, a state-of-the-art word alignment system, in terms of alignment error rate and F-measure. The bilingual parse trees produced for the parallel corpus can be exploited to extract bilingual phrases and train a decoder for statistical machine translation.
This technical report introduces the Multi-lingual (Korean, Japanese, Chinese and English) information processing which has been carried out by the Institute of Korean Culture of Korea University since the year 2000. The aim of this Project is to develop a syntactic and morphological analyser to analyse multi-lingual sentences and construct a Multilingual Lexical Database containing 50,000 entries until August, 2006.
Body image has been shown to be influenced by weight loss. Little attention however has been devoted to personal evaluations of physical fitness (fitness evaluation) or the extent to which individuals psychologically invest in improving fitness (fitness orientation) during periods of weight loss. PURPOSE: The purpose of the present study was to examine the relation between fitness evaluation (FE) and fitness orientation (FO) subscales with other body image ratings and physiological measures of fitness during a twelve-week behavioral weight loss program. METHODS: Thirty overweight, sedentary women (age= 42.5 ± 8.5 years, BMI= 29.3 kg/m2 ± 3.2) participated in a twelve-week behavioral weight loss program which reduced energy intake to 1200–1500 kcal/day and dietary fat to <30% of total calories. Subjects were progressively increased to 40 min/day, 5 days/week of home-based walking exercise. Body image was assessed using Multidimensional Body-Self Relations Questionnaire (MBSRQ) subscales. Cardiorespiratory fitness was measured using time to reach 85% of age-predicted maximal heart rate during a submaximal test on a treadmill. RESULTS: Time to reach 85% of age-predicted maximal heart rate significantly increased over the 12 week intervention (10.3+3.0 min vs. 12.5+3.4 min, p< 0.00). Baseline FE and FO were positively associated with baseline cardiorespiratory fitness (r=0.55 and 0.53, respectively). Change in FO was positively associated with baseline FO, changes in appearance evaluation and orientation and change in cardiorespiratory fitness (r=0.40, 0.65, 0.37 and 0.39, respectively). Furthermore, positive change in FO was associated with 12-week scores of appearance orientation, health orientation and overweight preoccupation (r=0.38, 0.37, and 0.51, respectively). CONCLUSION: These findings demonstrate that individuals with higher scores of FO at baseline had the greatest change in FO at 12 weeks, which was positively associated with changes in other perceived personal fitness constructs and fitness improvements. Exercise interventions should target strategies for improving FO and FE as this may have implications for exercise adoption and maintenance during weight loss. Supported by a Research Incentive Grant from the University of Louisville
Abstract Attempting to automatically learn to identify verb complements from natural language corpora without the help of sophisticated linguistic resources like grammars, parsers or treebanks leads to a significant amount of noise in the data. In machine learning terms, where learning from examples is performed using class-labelled feature-value vectors, noise leads to an imbalanced set of vectors: assuming that the class label takes two values (in this work complement/non-complement), one class (complements) is heavily underrepresented in the data in comparison to the other. To overcome the drop in accuracy when predicting instances of the rare class due to this disproportion, we balance the learning data by applying one-sided sampling to the training corpus and thus by reducing the number of non-complement instances. This approach has been used in the past in several domains (image processing, medicine, etc) but not in natural language processing. For identifying the examples that are safe to remove, we use the value difference metric, which proves to be more suitable for nominal attributes like the ones this work deals with, unlike the Euclidean distance, which has been used traditionally in one-sided sampling. We experiment with different learning algorithms which have been widely used and their performance is well known to the machine learning community: Bayesian learners, instance-based learners and decision trees. Additionally we present and test a variation of Bayesian belief networks, the COr-BBN (Class-oriented Bayesian belief network). The performance improves up to 22% after balancing the dataset, reaching 73.7% f-measure for the complement class, having made use only a phrase chunker and basic morphological information for preprocessing.
During auditory perception, neural oscillations are known to entrain to acoustic dynamics but their role in the processing of auditory information remains unclear. As a complex temporal structure that can be parameterized acoustically, music is particularly suited to address this issue. In a combined behavioral and EEG experiment in human participants, we investigated the relative contribution of temporal (acoustic dynamics) and nontemporal (melodic spectral complexity) dimensions of stimulation on neural entrainment, a stimulus-brain coupling phenomenon operationally defined here as the temporal coherence between acoustical and neural dynamics. We first highlight that low-frequency neural oscillations robustly entrain to complex acoustic temporal modulations, which underscores the fine-grained nature of this coupling mechanism. We also reveal that enhancing melodic spectral complexity, in terms of pitch, harmony, and pitch variation, increases neural entrainment. Importantly, this manipulation enhances activity in the theta (5 Hz) range, a frequency-selective effect independent of the note rate of the melodies, which may reflect internal temporal constraints of the neural processes involved. Moreover, while both emotional arousal ratings and neural entrainment were positively modulated by spectral complexity, no direct relationship between arousal and neural entrainment was observed. Overall, these results indicate that neural entrainment to music is sensitive to the spectral content of auditory information and indexes an auditory level of processing that should be distinguished from higher-order emotional processing stages.<b>NEW & NOTEWORTHY</b> Low-frequency (<10 Hz) cortical neural oscillations are known to entrain to acoustic dynamics, the so-called neural entrainment phenomenon, but their functional implication in the processing of auditory information remains unclear. In a behavioral and EEG experiment capitalizing on parameterized musical textures, we disentangle the contribution of stimulus dynamics, melodic spectral complexity, and emotional judgments on neural entrainment and highlight their respective spatial and spectral neural signature.
A grammatical method of combining two kinds of speech repair cues is presented. One cue, prosodic disjuncture, is detected by a decision tree-based ensemble classifier that uses acoustic cues to identify where normal prosody seems to be interrupted (Lickley, 1996). The other cue, syntactic parallelism, codifies the expectation that repairs continue a syntactic category that was left unfinished in the reparandum (Levelt, 1983). The two cues are combined in a Treebank PCFG whose states are split using a few simple tree transformations. Parsing performance on the Switchboard and Fisher corpora suggests that these two cues help to locate speech repairs in a synergistic way.
In this paper, we present a generic analysis of etymological data intended to provide a uniform framework for modeling such information in lexical databases. Based on the explicit reification of etymons and of the links between them, our proposal provides means to state additional constraints for both, as well as a preliminary set of standard descriptors to be used to this end. The model has been extensively tested on a variety of concrete cases extracted from the TLFi (Trésor de la Langue Française informatisé), allowing us to identify further mechanisms (alternatives, multiple links, composition, etc.) needed for etymological data representation. Taking into account the current standardization efforts within ISO committee TC 37/SC 4 to define a specification platform for lexical data (aka LMF, Lexical Markup Framework), we try to show that our proposal could be a possible contribution to this project
The Leximancer system is a relatively new method for transforming lexical co-occurrence information from natural language into semantic patterns in an unsupervised manner. It employs two stages of co-occurrence information extraction—semantic andrelational—using a different algorithm for each stage. The algorithms used are statistical, but they employ nonlinear dynamics and machine learning. This article is an attempt to validate the output of Leximancer, using a set of evaluation criteria taken from content analysis that are appropriate for knowledge discovery tasks.
TransBooster is a wrapper technology designed to improve the performance of wide-coverage machine translation \nsystems. Using linguistically motivated syntactic information, it automatically decomposes source language sentences into shorter and syntactically simpler chunks, and recomposes their translation to form target language sentences. This generally improves both the word order \nand lexical selection of the translation. To date, TransBooster has been successfully applied to rule-based MT, statistical MT, and multi-engine MT. This paper presents \nthe application of TransBooster to Example-Based Machine Translation. In an experiment conducted on test sets \nextracted from Europarl and the Penn II Treebank we show that our method can raise the BLEU score up to 3.8% relative \nto the EBMT baseline. We also conduct a manual evaluation, showing that TransBooster-enhanced EBMT produces \na better output in terms of fluency than the baseline EBMT in 55% of the cases and in terms of accuracy in 53% of the \ncases.
In data-oriented English-Chinese machine translation, knowledge source is the very important basis for translation processing. This paper presents a kind of construction strategy for knowledge source which contains affluent grammatical and syntactical information. Firstly, taking lexical function grammar as the theoretical basis, treebank including parse trees converted from every sentence in the source language corpus is acquired. Secondly, based on the decomposition algorithm, the corresponding fragment-bank composed of all the legal fragments extracted from the treebank is constructed. Finally, based on the combination algorithm, the fragment-combination-bank including all the possible fragment-combination forms of every parse tree in the treebank is built. Based on the successful construction of the knowledge source, the whole machine translation process can be implemented efficiently and accurately.
Abstract The aim of this study was to examine associations between children's temperament, parent–child goodness-of-fit, and the emotional content of parent–child conversations about past events. Fifty one New Zealand 5- and 6-year-old children and their parents discussed 4 emotional past events. Parents rated children's temperament along 15 dimensions associated with effortful control, extroversion, and negative affect. Parents then rated their own expectations of children's temperament, from which parent–child goodness-of-fit was assessed. Children rated as higher in effortful control were involved in more emotional past-event conversations. Children whose negative affect ratings corresponded more closely with parental expectations were also involved in more emotional conversations. Our findings provide preliminary support for an association between temperament and narrative content.
We use integrations and combinations of taggers to improve the tagging accuracy of Icelandic text. The accuracy of the best performing integrated tagger, which consists of our linguistic rule-based tagger for initial disambiguation and a trigram tagger for full disambiguation, is 91.80%. Combining five different taggers, using simple voting, results in 93.34% accuracy. By adding two linguistically motivated rules to the combined tagger, we obtain an accuracy of 93.48%. This method reduces the error rate by 20.5%, with respect to the best performing tagger in the combination pool.
This paper explores techniques to take advantage of the fundamental difference in structure between hidden Markov models (HMM) and hierarchical hidden Markov models (HHMM). The HHMM structure allows repeated parts of the model to be merged together. A merged model takes advantage of the recurring patterns within the hierarchy, and the clusters that exist in some sequences of observations, in order to increase the extraction accuracy. This paper also presents a new technique for reconstructing grammar rules automatically. This work builds on the idea of combining a phrase extraction method with HHMM to expose patterns within English text. The reconstruction is then used to simplify the complex structure of an HHMM The models discussed here are evaluated by applying them to natural language tasks based on CoNLL-2004 1 and a sub-corpus of the Lancaster Treebank 2.
Preface Part I: Background Chapter 1: Factors Influencing the Acquisition and Refinement of Communication Skills. Chapter 2: Communication Access: Overview and Issues Chapter 3: Issues in Assessment and Intervention Part II Chapter 4: Pre-Language Communication Chapter 5: Pre-Language Assessment and Intervention Chapter 6: Later Language Development Chapter 7: Adolescent Language: Assessment and Intervention Appendix A: Resource List for Children with Auditory Processing Disorders (APD) Appendix B: Troubleshooting Technology Appendix C: Technology Resources Appendix D: Recommendations for Comprehensive Assessment: Pragmatics and Semantics. Appendix E: Familiarity and Transparency Ratings for 100 Idioms Appendix F: Familiarity Ratings for 107 Proverbs Appendix G: Pronunciation Skill Inventory Index
The paper focuses on the description of the process of “mining” lexical semantic information from published dictionaries of Brazilian Portuguese language. Specifically, it is described the manual approach of compiling and filtering hyperonymy/hyponymy and holonymy/meronym logic-conceptual relations of concrete nouns. These relations, filtering from the dictionaries, will be use to organize part of the nouns of Wordnet.Br lexical database.
This presentation reports the methodology followed and the results attained on an on-going project aiming at building a large lexical database of corpus-extracted multiword (MW) expressions for the Portuguese language. MW expressions were automatically extracted from a balanced 50 million word corpus compiled for this project, furthermore statistically interpreted using lexical association measures and are undergoing a manual validation process. The lexical database covers different types of MW expressions, from named entities to lexical associations with different degrees of cohesion, ranging from totally frozen idioms to favoured co-occurring forms, like collocations. We aim to achieve two main objectives with this resource: to build on the large set of data of different types of MW expressions to revise existing typologies of collocations and to integrate them in a larger theory of MW units; to use the extensive hand-checked data as training data to evaluate existing statistical lexical association measures.
In this paper, we present a novel approach to combine the outputs of multiple MT engines into a consensus translation. In contrast to previous Multi-Engine Machine \nTranslation (MEMT) techniques, we do not rely on word alignments of output hypotheses, but prepare the input sentence for multi-engine processing. We do this by using a recursive decomposition algorithm that produces simple chunks as input to the MT engines. A consensus translation \nis produced by combining the best chunk translations, selected through majority voting, a trigram language model \nscore and a confidence score assigned to each MT engine. We report statistically significant relative improvements \nof up to 9% BLEU score in experiments (English→Spanish) carried out on an 800-sentence test set extracted from the Penn-II Treebank.
In this work, we apply statistical learning algorithms to Lexicalized Tree Adjoining Grammar (LTAG) parsing, as an effort toward statistical analysis over deep structures. LTAG parsing is a well known hard problem. Statistical methods successfully applied to LTAG parsing could also be used in many other structure prediction problems in NLP. For the purpose of achieving accurate and efficient LTAG parsing, we will investigate two aspects of the problem, the data structure and the algorithm. 1. We introduce LTAG-spinal, a variant of LTAG with very desirable linguistic, computational and statistical properties. It can be shown that LTAG-spinal with adjunction constraints is weakly equivalent to the traditional LTAG. For the purpose of statistical processing, we extract an LTAG-spinal treebank from the Penn Treebank with Propbank annotation. 2. We not only explore various parsing strategies, but also investigate the reranking approach. (a) We first propose a left-to-right incremental parser for LTAG-spinal, as an attempt to dynamically incorporate supertagging and dependency analysis. A perceptron like discriminative learning algorithm is used for training. We further investigate a bidirectional dependency parser for LTAG-spinal, in order to overcome the limitation of left-to-right processing. We propose a novel algorithm for graph-based incremental construction, and apply this algorithm to LTAG style dependency parsing. (b) We also explore learning algorithms for parse reranking, as well as other NLP problems, e.g. Machine Translation. We propose a novel reranking strategy, Ordinal Regression with Uneven Margins (ORUM), which achieves state-of-the-art performance on parse reranking for CFG parsing and MT reranking. To sum up, we have accomplished the following achievements. (i) A new formalism, LTAG-spinal, which is weakly equivalent to LTAG. (ii) An LTAG-spinal Treebank extracted from the PTB with the Propbank annotation. (iii) A left-to-right incremental parser for LTAG-spinal. (iv) A bidirectional LTAG-spinal dependency parser. (v) A novel graph-based incremental construction algorithm, which could be applied to many structure prediction problem in NLP, e.g. semantic role labeling. (vi) A novel discriminative reranking algorithm, ORUM, which has been successfully applied to parse reranking as well as other tasks, e.g. MT reranking.
BACKGROUND: The Bethesda System (TBS) along with its companion atlas was updated in 2001 to improve standardization, clarity, and reproducibility of cervical cytology reporting. METHODS: The authors used a novel web-based format to compare assessments of 77 images demonstrating a range of classical and borderline cytologic changes by a self-selected group of United States cytotechnologists (n = 216) and pathologists (n = 185). RESULTS: Participants were highly experienced, with 71.2% of cytotechnologists and 53.0% of pathologists reporting >10 years of practice. The mean percentage of exact agreement with the panel was slightly though significantly higher for cytotechnologists (57.0%) compared with pathologists (53.4%), adjusted for experience (P =.004); cervical cytology percentage effort (P =.0005); or cervical accession volume (P =.0002). Compared with the TBS panel, exact agreement was achieved for 55.1% of image ratings compared with 82.3% agreement at the level of Negative vs non-Negative for images with a single-panel interpretation. Agreement with the panel was highest for images classified as Low-Grade Squamous Intraepithelial Lesion and lowest for Atypical Squamous Cells qualified as either of Undetermined Significance or Cannot Exclude a High-Grade Squamous Intraepithelial Lesion. Reviewers were less sensitive in identifying high-grade glandular lesions than they were in identifying high-grade squamous lesions at any threshold (P <.001). CONCLUSIONS: Morphologic appearances of images were more important determinants than participants' academic or professional degrees with regard to interobserver reproducibility in classifying cervical cytology images. Experienced cytotechnologists and pathologists performed similarly. Participants achieved higher sensitivity for identifying high-grade squamous lesions than they did for high-grade glandular lesions. These findings demonstrated that web-based studies may be useful in assessing interobserver agreement in classifying images.
The Dutch spelling system, like other European spelling systems, represents a certain balance between preserving the spelling of morphemes (the morphological principle) and obeying letter-to-sound regularities (the phonological principle). We present experimental results with artificial learners that show a competition effect between the two principles: adhering more to one principle leads to more violations of the other. The artificial learners, memory-based learning algorithms, are trained (1) to convert written words to their phonemic counterparts and (2) to analyze written words on their morphological composition, based on data extracted from the CELEX lexical database. As an exception to the competition effect we show that introducing the schwa as a letter in the spelling system causes both morphology and phonology to be learnt better by the artificial learners. In general we argue that artificial learning studies are a tool in obtaining objective measurements on a spelling system that may be of help in spelling reform processes. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Pronunciation dictionaries for speech synthesis. The Swedish Library of Talking Books and Braille (TPB) is producing talking books for university students by speech synthesis. To enhance the pronunciation, TPB has started compiling a large electronic pronunciation dictionary. In the initial phase, existing printed dictionaries are imported to the lexical database. As these dic-tionaries were made for human reading, they are transformed to meet the quite different require-ments from the speech synthesis systems.
On the norm of current Chinese characters,象 像 of non-noun mainly the verb lexical meaning have not got the ideal social effect yet in the division and usage.And the present state is ambiguous.By reviewing the differences in non-noun lexical meaning and evidence of their division,it is thought that only divided their usage clearly can they be used normally.
We propose a bootstrapping approach to creating a phrase-level alignment over a sentence-aligned parallel corpus, reporting concrete treebank annotation work performed on a sample of sentence tuples from the Europarl corpus, currently for English, French, German, and Spanish. The manually annotated seed data will be used as the basis for automatically labelling the rest of the corpus. Some preliminary experiments addressing the bootstrapping aspects are presented. The representation format for syntactic correspondence across parallel text that we propose as the starting point for a process of successive refinement emphasizes correspondences of major constituents that realize semantic arguments or modifiers; language-particular details of morphosyntactic realization are intentionally left largely unlabelled. We believe this format is a good basis for training NLP tools for multilingual application contexts in which consistency across languages is more central than fine-grained details in specific languages (in particular, syntax-based statistical Machine Translation). 1.
Queneau's language has been analysed many times, mostly from a linguistic point of view, with special attention being paid to such procedures as phonetic transcription, lexical and syntactic mistakes or vocabulary typical of colloquial speech. However, Queneau's aim is not simply to imitate spoken discourse. Underlining of the oral aspect of a literary text emphasises its ludic character, i.e., its being -- in a sense -- the author's intellectual game with the reader. Queneau's linguistic experiments are not just limited to the most frequently mentioned techniques, by means of which he introduces the spoken discourse into literature. Simultaneously, Queneau employs very sophisticated, precise or even technical vocabulary as well as varied stylistic figures, often very complex. The present article analyses this play of linguistic registers, which constitutes the originality of Queneau's style and demonstrates that it is the conscious strategy of the author, who, rejecting established linguistic norms and literary conventions, plays with the reader.
At first,the prediction and evaluation technology of image quality in the world are reviewed in the paper.The model of simulation and prediction of image quality is established.Then,The development process of image rating scale and the development method of image interpretability rating scale are introduced.Finally,the general image quality equation (GIQE) between objective image quality parameters and image interpretability is established in the paper.
We exploit the resources in the Arabic Treebank (ATB) and Arabic Gigaword (AG) to determine the best features for the novel task of automatically creating lexical semantic verb classes for Modern Standard Arabic (MSA). The verbs are classified into groups that share semantic elements of meaning as they exhibit similar syntactic behavior. The results of the clustering experiments are compared with a gold standard set of classes, which is approximated by using the noisy English translations provided in the ATB to create Levin-like classes for MSA. The quality of the clusters is found to be sensitive to the inclusion of syntactic frames, LSA vectors, morphological pattern, and subject animacy. The best set of parameters yields an Fβ=1 score of 0.456, compared to a random baseline of an Fβ=1 score of 0.205.
The correct attachment of prepositional phrases (PPs) is a central disambiguation problem in parsing natural languages. This paper compares the baseline situation in English, German and Swedish based on manual PP attachments in various treebanks for these languages. We argue that cross-language comparisons of the disambiguation results in previous research is impossible because of the different selection procedures when building the training and test sets. We perform uniform treebank queries and show that English has the highest noun attachment rate followed by Swedish and German. We also show that the high rate in English is dominated by the preposition of. From our study we derive a list of criteria for profiling data sets for PP attachment experiments.
In this paper, we construct a biomedical semantic role labeling (SRL) system that can be used to facilitate relation extraction. First, we construct a proposition bank on top of the popular biomedical GENIA treebank following the PropBank annotation scheme. We only annotate the predicate-argument structures (PAS's) of thirty frequently used biomedical predicates and their corresponding arguments. Second, we use our proposition bank to train a biomedical SRL system, which uses a maximum entropy (ME) model. Thirdly, we automatically generate argument-type templates which can be used to improve classification of biomedical argument types. Our experimental results show that a newswire SRL system that achieves an F-score of 86.29% in the newswire domain can maintain an F-score of 64.64% when ported to the biomedical domain. By using our annotated biomedical corpus, we can increase that F-score by 22.9%. Adding automatically generated template features further increases overall F-score by 0.47% and adjunct arguments (AM) F-score by 1.57%, respectively.
In recent years, researchers have shown an interest in face-to-face or online intercultural communication as a way of understanding cultural similarities and differences. The present research shows how intercultural communication is carried out on the Internet through an analysis of e-mail discourse data of Korean and Australian students. In this paper, we discuss communication and miscommunication between Anglophone Australians and Koreans using English as a second language, along with the concepts of culture, intercultural communication and pragmatic strategies. We compare choice of lexical words used in their e-mail letters and analyze their different writing styles, in terms of uncertainty and negotiation. Our research shows that in cultural contacts on the Internet, cultural differences are reflected as in face-to-face conversations, and that people negotiate to avoid conflicts and misunderstandings. In other words, native cultural norms and language conventions affect intercultural interactions on the Internet as well. Nevertheless, globalization may lead to a multi-cultural perspective on human communication with East and West combined, and thus understanding differences between cultures is necessary for effective intercultural communication.
Standard techniques used in multilingual terminology management fail to describe legal terminologies as they are bound to different legal systems and terms do not share a common meaning. In the LexALP project, we use a technique defined for general lexical databases to achieve cross language interoperability between lan-guages of the Alpine Convention. In this paper we present the methodology and tools developed for the collection, de-scription and harmonisation of the legal terminology of spatial planning and sus-tainable development in the four lan-guages of the countries of the Alpine Space. 1
This paper discusses a novel probabilistic synchronous TAG formalism, synchronous Tree Substitution Grammar with sister adjunction (TSG+SA). We use it to parse a language for which there is no training data, by leveraging off a second, related language for which there is abundant training data. The grammar for the resource-rich side is automatically extracted from a treebank; the grammar on the resource-poor side and the synchronization are created by handwritten rules. Our approach thus represents a combination of grammar-based and empirical natural language processing. We discuss the approach using the example of Levantine Arabic and Standard Arabic.
The present dissertation addresses a set of questions about processes involved in lexical access and literacy and psycholinguistic factors (such as word learning age, word frequency etc.) that affect them in monolingual and bilingual speakers. Four experiments examined these issues for nouns in native English speakers and bilingual Hindī- English speakers with a developmental perspective. Experiments 1 and 2 were conducted with English monolinguals in San Diego. In experiment 1, age of acquisition norms were collected from college-age adults. In experiment 2, online picture naming data was collected from four age groups of English monolinguals (5-7, 8-10, 11-13 and college-age adults). Experiments 3 and 4 were conducted on Hindī-English bilinguals in India. In experiment 3, age of acquisition and word frequency norms were collected from college-age adults. In experiment 4, online picture naming and word reading data were collected from three age groups of Hindī-English bilinguals (8-10, 11-13 and college-age adults). Comparisons of performance on two lexical access tasks (on-line picture naming and word reading) in monolingual English speakers and bilingual Hindī-English speakers, were conducted. Results and discussion are aimed at addressing issues of language processing, lexical access and development. Overall, results indicate that there is developmental improvement on the lexical access tasks. In addition, the predictor- outcome relationships are generally similar for both monolinguals and bilinguals. Age of acquisition is the most consistent predictor of both picture naming and word reading behavior, in both monolingual and bilingual speaker. There are differential effects of frequency in the languages of the bilingual in the word reading task, with orthographic differences interacting with frequency effects. However, there are interesting differences that arise between the monolinguals and within the bilinguals, because of language dominance and proficiency. Results and discussion focus on quantitative analyses, examining lexical access processes in monolinguals and bilinguals, and examining the relationship between the psycholinguistic variables (such as age of acquisition, frequency, and syllable length) and performance on the language production tasks. Future directions focus on highlighting some limitations of this research, in addition to discussing the need for more in depth qualitative analyses, and extending these paradigms to clinical populations