Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
International audience
The basic colour terms for black and white are studied in four archaic and two contemporary linguistic norms of the Chinese language. It is presented that studied Chinese linguistic norms use a common term for white and three different terms for black. It is suggested that the different basic colour terms for black might originate from different source languages. The study supports a panchronic language development instead of a diachronic one, and includes introductions to histories of the Chinese linguistic norms
We demonstrate a novel, robust vision-tolanguage generation system called Midge. Midge is a prototype system that connects computer vision to syntactic structures with semantic constraints, allowing for the automatic generation of detailed image descriptions. We explain how to connect vision detections to trees in Penn Treebank syntax, which provides the scaffolding necessary to further refine data-driven statistical generation approaches for a variety of end goals. 1
This dissertation answers the question what is and what is not ellipsis and specifies criteria for identification of elliptical sentences. It reports on an analysis of types of ellipsis from the point of view of semantic (semantico-syntactic) representation of sentences. It does not deal with conditions and causes of the constitution of elliptical positions in sentences (when and why is it possible to omit something in a sentence) but it focuses exclusively on the identification of elliptical positions (if there is something omitted and what) and on their semantic representation, specifically on their representation on the tectogrammatical level of the Prague Dependency Treebanks. In this dissertation, the dependency approach (used in the Prague Dependency Treebanks) is also compared with the generative approach (used in the Penn Treebank). It is possible to utilize this comparison in the (automatic) conversion from constituency trees to dependency trees.
We present a new way to get more morphologically and syntactically annotated data. We have developed an annotation editor tailored to school children to involve them in text annotation. Using this editor, they practice morphology and dependency-based syntax in the same way as they normally do at (Czech) schools, without any special training. Their annotation is then automatically transformed into the target annotation schema. The editor is designed to be language independent, however the subsequent transformation is driven by the annotation framework we are heading for. In our case, the object language is Czech and the target annotation scheme corresponds to the Prague Dependency Treebank annotation framework.
This paper presents a higher-order model for constituent parsing aimed at utilizing more local structural context to decide the score of a grammar rule instance in a parse tree. Experiments on English and Chinese treebanks confirm its advantage over its first-order version. It achieves its best F1 scores of 91.86% and 85.58 % on the two languages, respectively, and further pushes them to 92.80% and 85.60 % via combination with other highperformance parsers. 1
This paper analyze ultimate bearing capacity, transfer mechanism, failure process and influence of aperture, steel bar diameter, concrete rating and on elastic bearing capacity and ultimate bearing capacity of PBL shear connectors by the finite element analysis software ANSYS. This paper can provide a reference for further design of PBL shear connectors as well as push-out or pull-out test. The results show that the elasticity bearing capacity of PBL shear connectors is determined primarily by concrete while the ultimate bearing capacity is determined primarily by perforative steel bar.
BACKGROUND: Words can shape or reinforce a patient's coping strategies. We measured the emotional content of hand surgery words and some synonyms or alternatives in five categories (19 words total). METHODS: Healthy adult companions of 100 patients presenting to an orthopedic hand surgical practice were asked to score five hand surgery words and some synonyms and alternatives (19 total words) on three dimensions: affective/emotional (ranging from pleasant to unpleasant), arousal (ranging from calm to aroused), and dominance/control (ranging from dominated to feeling in control) using a validated methodology. Ratings were done using the self-assessment manikin-a validated graphic affective rating system. RESULTS: The emotional reaction to "discomfort" and "ache" was more positive than "pain." The words "tear" and "defect" were more positive than "rupture." The words "tight" and "stiff" were more positive than "locked" and "frozen." The word "faded" was more positive than "degenerated," "diminished," and "wasted". The words "overused" and "worn" were more positive than "cracked," "inflamed," and "broken." CONCLUSIONS: Some common hand surgery words have a relatively negative emotional content. Given that psychological distress is an important predictor of pain intensity and disability, additional research is merited to develop optimally positive language for describing musculoskeletal pathology.
In this study, we investigated the role of facial cues in cooperator and defector recognition. First, a face image database was constructed from pairs of full face portraits of target subjects taken at the moment of decision-making in a prisoner's dilemma game (PDG) and in a preceding neutral task. Image pairs with no deficiencies (n = 67) were standardized for orientation and luminance. Then, confidence in defector and cooperator recognition was tested with image rating in a different group of lay judges (n = 62). Results indicate that (1) defectors were better recognized (58% vs. 47%), (2) they looked different from cooperators (p <.01), (3) males but not females evaluated the images with a relative bias towards the cooperator category (p <.01), and (4) females were more confident in detecting defectors (p <.05). According to facial microexpression analysis, defection was strongly linked with depressed lower lips and less opened eyes. Significant correlation was found between the intensity of micromimics and the rating of images in the cooperator-defector dimension. In summary, facial expressions can be considered as reliable indicators of momentary social dispositions in the PDG. Females may exhibit an evolutionary-based overestimation bias to detecting social visual cues of the defector face.
We introduce the task of word and phrase-level polarity annotation for German as part of an attempt to develop a compositional theory of clause-level polarity determination. Thus, annotations should give access to the nested building blocks, the structural strata of polarity composition. Therefore and in contrast to existing polarity-tagged corpora, we annotate not exclusively on the basis of surface strings, but argue that proper polarity annotation of complex phrases requires access to their syntactic structures. We discuss the principles of our treebank design, and present the inter-annotator agreement of our kick-off annotations on a test suite of 270 sentences that was compiled specifically to contain interesting polarity combinations.
Cataplexy is pathognomonic of narcolepsy with cataplexy, and defined by a transient loss of muscle tone triggered by strong emotions. Recent researches suggest abnormal amygdala function in narcolepsy with cataplexy. Emotion treatment and emotional regulation strategies are complex functions involving cortical and limbic structures, like the amygdala. As the amygdala has been shown to play a role in facial emotion recognition, we tested the hypothesis that patients with narcolepsy with cataplexy would have impaired recognition of facial emotional expressions compared with patients affected with central hypersomnia without cataplexy and healthy controls. We also aimed to determine whether cataplexy modulates emotional regulation strategies. Emotional intensity, arousal and valence ratings on Ekman faces displaying happiness, surprise, fear, anger, disgust, sadness and neutral expressions of 21 drug-free patients with narcolepsy with cataplexy were compared with 23 drug-free sex-, age- and intellectual level-matched adult patients with hypersomnia without cataplexy and 21 healthy controls. All participants underwent polysomnography recording and multiple sleep latency tests, and completed depression, anxiety and emotional regulation questionnaires. Performance of patients with narcolepsy with cataplexy did not differ from patients with hypersomnia without cataplexy or healthy controls on both intensity rating of each emotion on its prototypical label and mean ratings for valence and arousal. Moreover, patients with narcolepsy with cataplexy did not use different emotional regulation strategies. The level of depressive and anxious symptoms in narcolepsy with cataplexy did not differ from the other groups. Our results demonstrate that narcolepsy with cataplexy accurately perceives and discriminates facial emotions, and regulates emotions normally. The absence of alteration of perceived affective valence remains a major clinical interest in narcolepsy with cataplexy, and it supports the argument for optimal behaviour and social functioning in narcolepsy with cataplexy.
We describe the architecture we set up during the SANCL shared task for parsing usergenerated texts, that deviate in various ways from linguistic conventions used in available training treebanks. This architecture focuses in coping with such a divergence. It relies on the PCFG-LA framework (Petrov and Klein, 2007), as implemented by Attia et al. (2010). We explore several techniques to augment robustness: (i) a lexical bridge technique (Candito et al., 2011) that uses unsupervised word clustering (Koo et al., 2008); (ii) a special instanciation of self-training aimed at coping with POS tags unknown to the training set; (iii) the wrapping of a POS tagger with rulebased processing for dealing with recurrent non-standard tokens; and (iv) the guiding of out-of-domain parsing with predicted part-ofspeech tags for unknown words and unknown (word, tag) pairs. Our systems ranked second and third out of eight in the constituency parsing track of the SANCL competition. 1
This paper examines both linguistic behavior and practical implications of empty argument insertion in the Hindi PropBank. The Hindi PropBank is annotated on the Hindi Dependency Treebank, which contains some empty categories but rarely the empty arguments of verbs. In this paper, we analyze four kinds of empty arguments, *PRO*, *REL*, *GAP*, *pro*, and suggest effective ways of annotating these arguments. Empty arguments such as *PRO * and *REL * can be inserted deterministically; we present linguistically motivated rules that automatically insert these arguments with high accuracy. On the other hand, it is difficult to find deterministic rules to insert *GAP * and *pro*; for these arguments, we introduce a new annotation scheme that concurrently handles both semantic role labeling and empty category insertion, producing fast and high quality annotation. In addition, we present algorithms for finding antecedents of *REL * and *PRO*, and discuss why finding antecedents for some types of *PRO * is difficult.
The article discusses problems related to translating titles of films, literary works and academic publications from German to Polish and vice-versa. The authors emphasize aspects common to translation of all titles, and analyse typical problems occurring while rendering film and literary titles. In case of titles of academic publications formal and content-related changes are considerably smaller than, for instance, in film titles. This is due to the fact that while working on the latter translators take into account not only linguistic norms and conventions, but also references to the cultural background. A title has to be equally catchy in both source and target languages, and that is why translators perform certain operations when they are convinced that literal translation might yield poor results. Producing a title not related to the original is a borderline occurrence. The translator frequently needs to seek the golden mean between the content and form of the original and the requirements of the target language (linguistic norms, conventions, traditions).
Lexicography in the Arab world has had important effects on the development of the Arabic language. The origin and subsequent development and refinement of traditional Arabic grammatical theory—as early as the eighth century- had intimate links with the practice of writing dictionaries. The Kitab al-ʿayn, by al-Khalil ibn Ahmad (d.c.786), which is the first full-scale dictionary in the Arab world, marked a significant milestone in the history of grammatical thought and set the tone for more works on Arabic grammar. For many centuries, the general mode of the theory has acknowledged a "closed" corpus of Qur'anic diction and pre-Islamic poetry and prose as the major source of Arabic lexicographic works. The main credo is that Arabic dictionaries should contain the "unattained" forms of the language and remain impervious to external persuasions; namely colloquialisms, borrowings, neologisms, and coinages. Arabic dictionaries continued to resist the slightest reform as to the codification of lexical innovations and the treatment of lexical gaps exhausting themselves to the point of stagnation. Today, English-Arabic dictionary editors have to deal with a huge number of lexical gaps that have cumulated over time. The lexical gap-filling process is carried out in a very unsystematic way that is far from creating an atmosphere of cooperation that ultimately contributes to creating unified English-Arabic lexical databases for lexicographic purposes. The paper explains how a modern English-Arabic dictionary can fall short in its modernizing role, and gives a snapshot of the most salient microstructural issues that characterize the Al-Mawrid Al-Hadeeth: A Modern English-Arabic Dictionary (2010).
When a video of someone speaking is paused, the stationary image of the speaker typically appears less flattering than the video, which contained motion. We call this the frozen face effect (FFE). Here we report six experiments intended to quantify this effect and determine its cause. In Experiment 1, video clips of people speaking in naturalistic settings as well as all of the static frames that composed each video were presented, and subjects rated how flattering each stimulus was. The videos were rated to be significantly more flattering than the static images, confirming the FFE. In Experiment 2, videos and static images were inverted, and the videos were again rated as more flattering than the static images. In Experiment 3, a discrimination task measured recognition of the static images that composed each video. Recognition did not correlate with flattery ratings, suggesting that the FFE is not due to better memory for particularly distinct images. In Experiment 4, flattery ratings for groups of static images were compared with those for videos and static images. Ratings for the video stimuli were higher than those for either the group or individual static stimuli, suggesting that the amount of information available is not what produces the FFE. In Experiment 5, videos were presented under four conditions: forward motion, inverted forward motion, reversed motion, and scrambled frame sequence. Flattery ratings for the scrambled videos were significantly lower than those for the other three conditions. In Experiment 6, as in Experiment 2, inverted videos and static images were compared with upright ones, and the response measure was changed to perceived attractiveness. Videos were rated as more attractive than the static images for both upright and inverted stimuli. Overall, the results suggest that the FFE requires continuous, natural motion of faces, is not sensitive to inversion, and is not due to a memory effect.
We consider a specific class of tree structures that can represent basic structures in linguistics and computer science such as XML documents, parse trees, and treebanks, namely, finite node-labeled sibling-ordered trees. We present axiomatizations of the monadic second-order logic (MSO), monadic transitive closure logic (FO(TC1)) and monadic least fixed-point logic (FO(LFP1)) theories of this class of structures. These logics can express important properties such as reachability. Using model-theoretic techniques, we show by a uniform argument that these axiomatizations are complete, i.e., each formula that is valid on all finite trees is provable using our axioms. As a backdrop to our positive results, on arbitrary structures, the logics that we study are known to be non-recursively axiomatizable.
We present the Prague Dependency Treebank 2.5, the newest version of PDT and the first to be released under a free license. We show the benefits of PDT 2.5 in comparison to other state-of-the-art treebanks. We present the new features of the 2.5 release, how they were obtained and how reliably they are annotated. We also show how they can be used in queries and how they are visualised with tools released alongside the treebank.
Annotated corpora such as treebanks are important for the development of parsers, language applications as well as understanding of the\nlanguage itself. Only very few languages possess these scarce resources. In this paper, we describe our efforts in syntactically annotating\na small corpora (600 sentences) of Tamil language. Our annotation is similar to Prague Dependency Treebank (PDT) and consists of\nannotation at 2 levels or layers: (i) morphological layer (m-layer) and (ii) analytical layer (a-layer). For both the layers, we introduce\nannotation schemes i.e. positional tagging for m-layer and dependency relations for a-layers. Finally, we discuss some of the issues in\ntreebank development for Tamil.
The annotation of large corpora is usually restricted to syntactic structure and word class. Pure lexical information and information on the structure of words are stored in specialized dictionaries (Baayen et al., 1995). Both data structures ‐ dictionary and text corpus ‐ can be matched to get e.g. a distribution of certain (restricted) lexical information from a text. This procedure works fine for synchronic corpora. What is missing, however, is either a special mark-up in texts linking each of the items to a certain time or a diachronic lexical database that allows for the matching of the items over time. In what follows, we take the latter approach and present a tool set (MoreXtractor, Morphilizer, MorQuery), a database (Morphilo-DB) and the architecture of a platform (Morphorm) for a sustainable use of diachronic linguistic data for Middle English, Early Modern English and Modern English.
In this paper we describe semi-automatical extending of the Czech WordNet lexical database (48,000 literals in 28,000 synsets) by translation of English literals from existing synsets in Princeton WordNet. We make use of a machine-readable bilingual dictionary to extract English-Czech translation pairs, search the English literals in Princeton WordNet and in case of a high-confidence match we transfer the literal into Czech WordNet. Along with literals, new synsets parallel to the English ones and identified by ILI are introduced into CzechWordNet, including information on their ILR (Internal Language Relations) such as hypernymy/hyponymy. The paper describes the parsing of the dictionary data, extraction of translation pairs and the criteria used for estimating the confidence level of a match. Results of the work are 36,228 added literals and 12,403 created synsets. An overview of previous similar attempts for other languages is also included.
Syntactic analysis is a valuable addition to corpus annotation, since it enables the retrieval of structural information which is otherwise difficult to access. The Norwegian Infrastructure for the Exploration of Syntax and Semantics is adding syntactic information to the Norwegian Newspaper Corpus and is effectively producing a treebank as a parsed corpus.
A morphological analyser only recognizes words that it already knows in the lexical database. It needs, however, a way of sensing significant changes in the language in the form of newly borrowed or coined words with high frequency. We develop a finite-state morphological guesser in a pipelined methodology for extracting unknown words, lemmatizing them, and giving them a priority weight for inclusion in a lexicon. The processing is performed on a large contemporary corpus of 1,089,111,204 words and passed through a machine-learning-based annotation tool. Our method is tested on a manually-annotated gold standard of 1,310 forms and yields good results despite the complexity of the task. Our work shows the usability of a highly non-deterministic finite state guesser in a practical and complex application. 1
Unknown words, or out of vocabulary words (OOV), cause a significant problem to morphological analysers, syntactic parses, MT systems and other NLP applications. Unknown words make up 29 % of the word types in in a large Arabic corpus used in this study. With today&apos;s corpus sizes exceeding 10 9 words, it becomes impossible to manually check corpora for new words to be included in a lexicon. We develop a finite-state morphological guesser and integrate it with a machine-learning-based pre-annotation tool in a pipeline architecture for extracting unknown words, lemmatizing them, and giving them a priority weight for inclusion in a lexical database. The processing is performed on a corpus of contemporary Arabic of 1,089,111,204 words. Our method is tested on a manually-annotated gold standard and yields encouraging results despite the complexity of the task. Our work shows the usability of a highly
Treebanking a large corpus of relatively structured speech transcribed from various Arabic Broadcast News (BN) sources has allowed us to begin to address the many challenges of annotating and parsing a speech corpus in Arabic. The now completed Arabic Treebank BN corpus consists of 432,976 source tokens (517,080 tree tokens) in 120 files of manually transcribed news broadcasts. Because news broadcasts are predominantly scripted, most of the transcribed speech is in Modern Standard Arabic (MSA). As such, the lexical and syntactic structures are very similar to the MSA in written newswire data. However, because this is spoken news, cross-linguistic speech effects such as restarts, fillers, hesitations, and repetitions are common. There is also a certain amount of dialect data present in the BN corpus, from on-the-street interviews and similar informal contexts. In this paper, we describe the finished corpus and focus on some of the necessary additions to our annotation guidelines, along with some of the technical challenges of a treebanked speech corpus and an initial parsing evaluation for this data. This corpus will be available to the community in 2012 as an LDC publication.
State-of-the-art dependency representations such as the Stanford Typed Dependencies may represent the grammatical relations in a sentence as directed, possibly cyclic graphs. Querying a syntactically annotated corpus for grammatical structures that are represented as graphs requires graph matching, which is a non-trivial task. In this paper, we present an algorithm for graph matching that is tailored to the properties of large, syntactically annotated corpora. The implementation of the algorithm is built on top of the popular IMS Open Corpus Workbench, allowing corpus linguists to re-use existing infrastructure. An evaluation of the resulting software, CWB-treebank, shows that its performance in real world applications, such as a web query interface, compares favourably to implementations that rely on a relational database or a dedicated graph database while at the same time offering a greater expres-sive power for queries. An intuitive graphical interface for building the query graphs is available via the Treebank.info project.
The goal of the presented parallel phrase extraction algorithm is to provide rich and robust set of translation syntactic patterns. To make this approach feasible, we consider the phrase-to-phrase alignments of a bilingual treebank annotated with syntactic constituents. For the intended purpose, the extracted phrasal nodes are encoded by the syntactical information of their components, highlighting some special constructs such as the functional words.
In this article, we first present the overall structure of the Pralex lexical database, the work with the data entry form and Praled’s functions. We then focus on the general principles of the elementary processing of database entries, which are subsequently specified according to the individual word classes and entry types. The article ends with a specific example of the processing of an entry in the Pralex lexical database.
This paper presents the work of the Hong Kong Polytechnic University (PolyUCOMP) team which has participated in the Semantic Textual Similarity task of SemEval-2012. The PolyUCOMP system combines semantic vectors with skip bigrams to determine sentence similarity. The semantic vector is used to compute similarities between sentence pairs using the lexical database WordNet and the Wikipedia corpus. The use of skip bigram is to introduce the order of words in measuring sentence similarity. 1
This paper describes the support for mouth activity annotation provided by the iLex annotation workbench on a holistic level connected to the lexical database, on a feature level, as well as in the context of semi-automatic annotation.
Existing logic-based querying tools for dependency treebanks use first order logic or monadic second order logic. We introduce a very fast model checker based on hybrid logic with operators ↓, @ and A and show that it is much faster than an existing querying tool for dependency treebanks based on first order logic, and much faster than an existing general purpose hybrid logic model checker. The querying tool is made publicly available.
Continuous self-reported emotion expressed by four pieces of music were collected on a two-dimensional (valence and arousal) emotion space in a repeated measures (test-retest conditions) design. Initial orientation time (IOT), test-retest reliability and afterglow were examined. Median IOT was 8 seconds. Valence ratings took up to 25 (median 4), and for arousal up to 35 (median 12) seconds. Slower tempi seemed to require longer IOT. Test-retest reliability examined correlation coefficients, and compared periods of sample-by-sample good agreement in response between Test and Retest condition. About 80% of responses were reliable in both the Test and Retest conditions regardless of response dimension. Pearson correlations demonstrated better test-retest reliability for arousal responses than for valence. Retest condition ratings were within 8% of Test condition rating within participant. Average standard deviations for ratings collapsed across dimension, stimulus and conditions was 12.2% of the ratings scale range. Afterglow effects – large outliers in spread of scores just after the end of a piece – were identified. The reliability of continuous emotional response is therefore considered to be quite good, but caution must be taken as to how to deal with the opening and ending of continuous emotional response data.
Query expansion is a crucial step in recall-oriented domains such as Patent Searching. Currently, automatic query expansion in patent search is mostly based on statistical measures. Additional query terms are extracted from the query documents based on entropy measures. To automate query expansion in patent searching, we acquire lexical knowledge from Query Logs of USPTO Patent Examiners. Results show good performance in query expansion and patent searching using the lexical database. This will help improving (semi-) automated query expansion in patent searching.
The lack of annotated corpora brings limitations in research of discourse classification for many languages. In this paper, we present the first effort towards recognizing ambiguities of discourse connectives, which is fundamental to discourse classification for resource-poor language such as Chinese. A language independent framework is proposed utilizing bilingual dictionaries, Penn Discourse Treebank and parallel data between English and Chinese. We start from translating the English connectives to Chinese using a bi-lingual dictionary. Then, the ambiguities in terms of senses a connective may signal are estimated based on the ambiguities of English connectives and word alignment information. Finally, the ambiguity between discourse usage and non-discourse usage were disambiguated using the co-training algorithm. Experimental results showed the proposed method not only built a high quality connective lexicon for Chinese but also achieved a high performance in recognizing the ambiguities. We also present a discourse corpus for Chinese which will soon become the first Chinese discourse corpus publicly available.
Après un bref rĂŠsumĂŠ de la thĂŠorie de Topic-Focus Articulation (TFA), la prĂŠsente ĂŠtude dĂŠmontre, à l'aide de plusieurs exemples illustrant l'annotation de principaux traits de TFA sur un large corpus (the Prague Dependency Treebank), que l'annotation du corpus apporte une valeur ajoutĂŠe au corpus, si deux conditions sont rĂŠunies: (i) le schĂŠma de l'annotation est basĂŠ sur une thĂŠorie linguistique solide, (ii) le procĂŠdĂŠ d'annotation est ĂŠtabli avec soin (c'est-à-dire de façon systĂŠmatique et cohĂŠrente). Une telle annotation est importante non seulement pour la structure de surface de la phrase mais encore davantage pour la structure phrastique sous-jacente, car elle est susceptible de mettre en ĂŠvidence les phĂŠnomènes cachĂŠs au niveau de la structure de surface, mais incontournables lors de la reprĂŠsentation du sens et du fonctionnement de la phrase.
Chinese word structure annotation is potentially useful for many NLP tasks, especially for Chinese word segmentation. Li and Zhou (2012) have presented an annotation for word structures in the Penn Chinese Treebank. But they only consider words that have productive affixes, which covers 35% of word types in that corpus. In this paper, we propose a linguistically inspired annotation that covers various morphological derivations of Chinese in a more general way, such that almost all multiple-character words can be structurally analyzed. As manual annotation is expensive, we propose a semi-supervised approach to automatic annotation, which combines the maximum entropy learning and the EM iteration for the Gaussian mixture model. The proposed method has achieved an accuracy of 90% on the testing set. © 2021 CLP 2012 - 2nd CIPS-SIGHAN Joint Conference on Chinese Language Processing. All Rights Reserved.
We present results of a study investigating evaluative learning in dementia patients with a classic evaluative conditioning paradigm. Picture pairs of three unfamiliar faces with liked, disliked, or neutral faces, that were rated prior to the presentation, were presented 10 times each to a group of dementia patients (N = 15) and healthy controls (N = 14) in random order. Valence ratings of all faces were assessed before and after presentation. In contrast to controls, dementia patients changed their valence ratings of unfamiliar faces according to their pairing with either a liked or disliked face, although they were not able to explicitly assign the picture pairs after the presentation. Our finding suggests preserved evaluative conditioning in dementia patients. However, the result has to be considered preliminary, as it is unclear which factors prevented the predicted rating changes in the expected direction in the control group.
Texts The Prague Czech-English Dependency Treebank 2.0 (PCEDT 2.0) is a major update of the Prague Czech-English Dependency Treebank 1.0 (LDC2004T25). It is a manually parsed Czech-English parallel corpus sized over 1.2 million running words in almost 50,000 sentences for each part. Data The English part contains the entire Penn Treebank - Wall Street Journal Section (LDC99T42). The Czech part consists of Czech translations of all of the Penn Treebank-WSJ texts. The corpus is 1:1 sentence-aligned. An additional automatic alignment on the node level (different for each annotation layer) is part of this release, too. The original Penn Treebank-like file structure (25 sections, each containing up to one hundred files) has been preserved. Only those PTB documents which have both POS and structural annotation (total of 2312 documents) have been translated to Czech and made part of this release. Each language part is enhanced with a comprehensive manual linguistic annotation in the PDT 2.0 style (LDC2006T01, Prague Dependency Treebank 2.0). The main features of this annotation style are: dependency structure of the content words and coordinating and similar structures (function words are attached as their attribute values) semantic labeling of content words and types of coordinating structures argument structure, including an argument structure ("valency") lexicon for both languages ellipsis and anaphora resolution. This annotation style is called tectogrammatical annotation and it constitutes the tectogrammatical layer in the corpus. For more details see below and documentation. Annotation of the Czech part Sentences of the Czech translation were automatically morphologically annotated and parsed into surface-syntax dependency trees in the PDT 2.0 annotation style. This annotation style is sometimes called analytical annotation; it constitutes the analytical layer of the corpus. The manual tectogrammatical (deep-syntax) annotation was built as a separate layer above the automatic analytical (surface-syntax) parse. A sample of 2,000 sentences was manually annotated on the analytical layer. Annotation of the English part The resulting manual tectogrammatical annotation was built above an automatic transformation of the original phrase-structure annotation of the Penn Treebank into surface dependency (analytical) representations, using the following additional linguistic information from other sources: PropBank (LDC2004T14) VerbNet NomBank (LDC2008T23) flat noun phrase structures (by courtesy of D. Vadas and J.R. Curran) For each sentence, the original Penn Treebank phrase structure trees are preserved in this corpus together with their links to the analytical and tectogrammatical annotation.
The strong association between music and speech has been supported by recent research focusing on musicians' superior abilities in second language learning and neural encoding of foreign speech sounds. However, evidence for a double association-the influence of linguistic background on music pitch processing and disorders-remains elusive. Because languages differ in their usage of elements (e.g., pitch) that are also essential for music, a unique opportunity for examining such language-to-music associations comes from a cross-cultural (linguistic) comparison of congenital amusia, a neurogenetic disorder affecting the music (pitch and rhythm) processing of about 5% of the Western population. In the present study, two populations (Hong Kong and Canada) were compared. One spoke a tone language in which differences in voice pitch correspond to differences in word meaning (in Hong Kong Cantonese, /si/ means 'teacher' and 'to try' when spoken in a high and mid pitch pattern, respectively). )
Major findings in attractiveness such as the role of averageness and symmetry have emerged primarily from neutral static visual stimuli. However it has increasingly been shown that ratings of attractiveness can be modulated within unisensory and multisensory modes by factors including emotional expression or by additional information about the person. For example, previous research has indicated that humorous individuals are rated as more desirable than their non-humorous equivalents (Bressler and Balshine, 2006). In two experiments we measured within and cross-sensory modulation of the attractiveness of unfamiliar faces. In Experiment 1 we examined if manipulating the number and type of expressions shown across a series of images of a person influences the attractiveness rating for that person. Results indicate that for happy expressions, ratings of attractiveness gradually increase as the proportional number of happy facial expressions increase, relative to the number of neutral expressions. In contrast, an increase in the proportion of angry expressions was not assocated with an increase in attractiveness ratings. In Experiment 2 we investigated if perceived attractiveness can be influenced by multisensory information provided during exposure to the face image. Ratings are compared across face images which were presented with or without voice information. In addition we provided either an auditory emotional cue (e.g., laughter) or neutral (e.g., coughing) cue to assess whether social information affects perceived attractiveness. Results shows that multisensory information about a person can increase attractiveness ratings, but that the emotional content of the cross-modal information can effect preference for some faces over others.
The case in Czech is the basic morphological means by which nouns express their function in a sentence. The objective of this thesis is to describe, from a frequency point of view, the relation between form and function of nouns, or, more precisely, how frequently cases (both simple and prepositional) are used to realise syntactic functions in sentences. The thesis is based on one of the largest corpora of written synchronic Czech: 100-million-token corpus SYN2005. In order to obtain data on frequencies of syntactic functions of nouns in relation to their cases, we annotated the corpus SYN2005 with a dependency syntactic annotation. For this annotation, we adopted the format of the analytical layer of the Prague Dependency Treebank. The syntactic annotation has been performed by a stochastic parser: the MST parser. Since the reliability of this annotation was not high enough, we have built an automatic correction module, which identifies errors of syntactic annotation in the output of the stochastic parser and corrects these errors by means of linguistic rules. We have implemented 26 different rules, but annotation errors have been reduced by merely 6-8%. However, this correction module can be further developed. It can be used to correct the output of any dependency parser trained on the data from...
It is no of doubt; the output of Example Based Machine Translation is always consequence of adaptation procedure over retrieved examples. There are always scopes for refining the resources structure of retriever examples. In this paper, we will describe an efficient approach to organize English-Hindi parallel examples taken from online available machine translation system instead of any linguistic database and its implementation towards English-Hindi Example-Based
Обсуждаются перспективы использования лингвистических онтологий WordNet 2.1 и WordNet 3.0, разработанных на базе национального корпуса британского варианта английского языка, для обучения лексике английского языка и проведения лингвистических исследований в классах лингвистического профиля старшей школыThe article focuses on future practical usage of WordNet 2.1 and WordNet 3.0, two large lexical databases of English based on the British National Corpus. They are both to be applied to teaching vocabulary and carrying out linguistic research amongst ELT high school students.