Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
It is difficult to deny the existence of Nigerian English, just as it is difficult to deny the existence of British English and American English. This claim arises from the fact that what characterizes a people’s linguistic norm takes its rise from the totality of the people’s sociocultural practices of which meaningful verbal sounds are of great essence. The meaningful verbal sounds serve as the chief instrument for communicating meanings, feelings, ideas, and
Penn Discourse Treebank style discourse parsing is a composite task of detecting explicit and non-explicit discourse relations, their connective and argument spans, and assigning a sense to these relations. Due to the composite nature of the task, the end-to-end performance is greatly affected by the error propagation. This paper describes the end-to-end discourse parser for English submitted to the CoNLL 2016 Shared Task on Shallow Discourse Parsing with the main focus of the parser being on argument spans and the reduction of global error through model selection. In the end-to-end closed-track evaluation the parser achieves F-measure of 0.2510 outperforming the best system of the previous year.
Discourse parsing is an important task in Language Understanding with applications to human-human and human-machine communication modeling. However, most of the research has focused on written text, and parsers heavily rely on syntactic parsers that themselves have low performance on dialog data. In our work, we address the problem of analyzing the semantic relations between discourse units in human-human spoken conversations. In particular, in this paper we focus on the detection of discourse connectives which are the predicate of such relations. The discourse relations are drawn from the Penn Discourse Treebank annotation model and adapted to a domain-specific Italian human-human spoken conversations. We study the relevance of lexical and acoustic context in predicting discourse connectives. We observe that both lexical and acoustic context have mixed effect on the prediction of specific connectives. While the oracle of using lexical and acoustic contextual feature combinations is F1 = 68.53, the lexical context alone significantly outperforms the baseline by more than 10 points with F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sub> = 64.93.
Abstract We present a work in progress aimed at extracting translation pairs of source and target dependency treelets to be used in a dependency-based machine translation system. We introduce a novel unsupervised method for parallel tree segmentation based on Gibbs sampling. Using the data from a Czech-English parallel treebank, we show that the procedure converges to a dictionary containing reasonably sized treelets; in some cases, the segmentation seems to have interesting linguistic interpretations.
Resumen El presente trabajo tiene como objetivo analizar con qué elementos textuales y mediante qué procedimientos y estrategias se explicita la norma especialmente en el diccionario bilingüe de Lucio Ambruzzi: Nuovo dizionario spagnolo-italiano e italiano-spagnolo (1948-49). Adquiere, para ello, un gran relieve el análisis de las dos secciones del diccionario, en las que se observa la presencia del legado de la tradición nacional en la que se colocan: el diccionario académico usual de 1925 y el manual e ilustrado de 1927 en el ámbito español; los diccionarios de Panzini, Migliorini, Monelli e Jàcono en el ámbito italiano. En el diccionario bilingüe de Lucio Ambruzzi (DBA) se analizarán las marcas de usos y los comentarios normativos en la microestructura de los neologismos, especialmente de los extranjerismos. El trabajo, en su desarrollo, tratará de indicar las características de los comentarios normativos explícitos, situándolos en el contexto histórico y cultural de la época a la que pertenecen, para lo cual se analizarán los diccionarios españoles e italianos consultados por Ambruzzi, que reflejan la mentalidad general y la ideología vigente respecto a los extranjerismos. Palabras clave: Ambruzzi, diccionario bilingüe, norma, neologismos, extranjerismos. Abstract This paper aims to analyse the textual elements and the procedures and strategies that Lucio Ambruzzi’s bilingual dictionary [Nuovo dizionario spagnolo-italiano e italiano-spagnolo (1948-49)] uses to specify the linguistic norm. For this reason, it is very important to study the two dictionary sections, where we can notice the national tradition legacy in which are placed (Spanish academic diccionary: “usual” 1925 and “manual e ilustrado” 1927. Italian dictionaries: Panzini, Migliorini, Monelli and Jàcono). In Lucio Ambruzzi’s bilingual dictionary (DBA) we will analyse the usage labels and the linguistic norm comments in the microstructure of neologisms, specially foreign terms. The paper points to indicate and catalogue the labels and explicit linguistic norm comments characteristics, establishing the historical and cultural context of the period they belong to. For this purpose, we will study the Spanish and Italian dictionaries Ambruzzi consulted, which reflect the general mind-set and the author’s ideology regarding foreign words. Keywords: Ambruzzi, bilingual dictionary, linguistic norm, neologism, foreignisms.
The present study compared the effects of: (a) PETTLEP imagery (e.g. imaging in the environment), (b) prior-observation (i.e. observing prior to imaging), and (c) traditional imagery (e.g. imaging sat in a quiet room) on the ease and vividness of external visual imagery (EVI), internal visual imagery (IVI), and kinaesthetic imagery (KI) of movements. Fifty-two participants (28 female, 24 male, Mage = 19.60 years, SD = 1.59) imaged the movements described in the Vividness of Movement Imagery Questionnaire-2 under the three conditions in a counterbalanced order. Vividness and ease of imaging ratings were recorded for each movement. A repeated measure MANOVA revealed that ease and vividness ratings for EVI, IVI, and KI were higher during the PETTLEP imagery condition compared to the traditional imagery condition, and vividness of EVI was higher during the observation imagery condition compared to traditional imagery. Findings indicate that incorporating PETTLEP elements into the imagery instructions leads to easier and more vivid movement EVI, IVI, and KI imagery.
Convert natural language text into tokens. Includes tokenizers for shingled n-grams, skip n-grams, words, word stems, sentences, paragraphs, characters, shingled characters, lines, Penn Treebank, regular expressions, as well as functions for counting characters, words, and sentences, and a function for splitting longer texts into separate documents, each with the same number of words. The tokenizers have a consistent interface, and the package is built on the 'stringi' and 'Rcpp' packages for fast yet correct tokenization in 'UTF-8'.
This paper illustrates the similarity between Thai and Laotian, and between Malay and Indonesian, based on an investigation on raw parallel data from Asian Language Treebank. The cross-lingual similarity is investigated and demonstrated on metrics of correspondence and order of tokens, based on several standard statistical machine translation techniques. The similarity shown in this study suggests a possibility on harmonious annotation and processing of the language pairs in future development.
Recently, there has been an explosion in the availability of large, good-quality cross-linguistic databases such as WALS (Dryer & Haspelmath, 2013), Glottolog (Hammarstrom et al., 2015) and Phoible (Moran & McCloy, 2014). Databases such as Phoible contain the actual segments used by various languages as they are given in the primary language descriptions. However, this segment-level representation cannot be used directly for analyses that require generalizations over classes of segments that share theoretically interesting features. Here we present a method and the associated R (R Core Team, 2014) code that allows the exible denition of such meaningful classes and that can identify the sets of segments falling into such a class for any language inventory. The method and its results are important for those interested in exploring cross-linguistic patterns of phonetic and phonological diversity and their relationship to extra-linguistic factors and processes such as climate, economics, history or human genetics.
This paper investigates the problem of cross-lingual transfer parsing, aiming at inducing dependency parsers for low-resource languages while using only training data from a resource-rich language (e.g., English). Existing model transfer approaches typically don't include lexical features, which are not transferable across languages. In this paper, we bridge the lexical feature gap by using distributed feature representations and their composition. We provide two algorithms for inducing cross-lingual distributed representations of words, which map vocabularies from two different languages into a common vector space. Consequently, both lexical features and non-lexical features can be used in our model for cross-lingual transfer. Furthermore, our framework is flexible enough to incorporate additional useful features such as cross-lingual word clusters. Our combined contributions achieve an average relative error reduction of 10.9% in labeled attachment score as compared with the delexicalized parser, trained on English universal treebank and transferred to three other languages. It also significantly outperforms state-of-the-art delexicalized models augmented with projected cluster features on identical data. Finally, we demonstrate that our models can be further boosted with minimal supervision (e.g., 100 annotated sentences) from target languages, which is of great significance for practical usage.
International audience
In this paper we provide a systematic and comprehensive set of modeling principles for representing etymological data in digital dictionaries using TEI. The purpose is to integrate in one coherent framework both digital representations of legacy dictionaries and born-digital lexical databases that are constructed manually or semi-automatically. We provide examples from many different types of etymological phenomena from traditional lexicographic practice, as well as analytical approaches from functional and cognitive linguistics such as metaphor, metonymy, and grammaticalization, which in many lexicographical and formal linguistic circles have not often been treated as truly etymological in nature, and have thus been largely left out of etymological dictionaries. In order to fully and accurately express the phenomena and their structures, we have made several proposals for expanding and amending some aspects of the existing TEI framework. Finally, with reference to both synchronic and diachronic data, we also demonstrate how encoders may integrate semantic web/linked open data information resources into TEI dictionaries as a basis for the sense, and/or the semantic domain, of an entry and/or an etymon.
This study describes a change in which relative clause extraposition is in the process of being lost in English, Icelandic, French, and Portuguese. This current change in progress has never been observed before, probably because it is so slow that it is undetectable without the aid of multiple diachronic parsed corpora (treebanks) with time depths of over 500 years each. Building on insights from Kiparsky (1995), the study shows that the change may date as far back as the innovation of Proto-Germanic and Proto-Romance relative clauses, as these varieties differentiated from Proto-Indo-European. It also shows that the unusually slow speed of the change is due to partial specialization of the construction along the dimension of prosodic weight, following the argument made at greater length in Fruehwald & Wallenberg 2016. Finally, the change is shown to have important consequences for the syntax of extraposition, supporting the adjunction analysis of Culi- cover and Rochemont (1990). The article also discusses the implications of Sauerland's (2003) analysis of English relative clauses, and while modern English data supports his analysis, the diachronic extraposition data is not yet fine-grained enough to bear on the ‘raising’ analysis of relatives in general. This is identified as an important question for further research on this change.
We present an investigation of the perception of authenticity in audiovisual laughter, in which we contrast spontaneous and volitional samples and examine the contributions of unimodal affective information to multimodal percepts. In a pilot study, we demonstrate that listeners perceive spontaneous laughs as more authentic than volitional ones, both in unimodal (audio-only, visual-only) and multimodal contexts (audiovisual). In the main experiment, we show that the discriminability of volitional and spontaneous laughter is enhanced for multimodal laughter. Analyses of relationships between affective ratings and the perception of authenticity show that, while both unimodal percepts significantly predict evaluations of audiovisual laughter, it is auditory affective cues that have the greater influence on multimodal percepts. We discuss differences and potential mismatches in emotion signalling through voices and faces, in the context of spontaneous and volitional behaviour, and highlight issues that should be addressed in future studies of dynamic multimodal emotion processing.
Speakers respond more slowly when naming pictures presented with taboo (i.e., offensive/embarrassing) than with neutral distractor words in the picture-word interference paradigm. Over four experiments, we attempted to localize the processing stage at which this effect occurs during word production and determine whether it reflects the socially offensive/embarrassing nature of the stimuli. Experiment 1 demonstrated taboo interference at early stimulus onset asynchronies of -150 ms and 0 ms although not at 150 ms. In Experiment 2, taboo distractors sharing initial phonemes with target picture names eliminated the interference effect. Using additive factors logic, Experiment 3 demonstrated that taboo interference and phonological facilitation effects do not interact, indicating that the two effects originate at different processing levels within the speech production system. In Experiment 4, interference was observed for masked taboo distractors, including those sharing initial phonemes with the target picture names, indicating that the effect cannot be attributed to a processing level involving responses in an output buffer. In two of the four experiments, the magnitude of the interference effect correlated significantly with arousal ratings of the taboo words. However, no significant correlations were found for either offensiveness or valence ratings. These findings are consistent with a locus for the taboo interference effect prior to the processing stage responsible for word form encoding. We propose a pre-lexical account in which taboo distractors capture attention at the expense of target picture processing due to their high arousal levels.
In this paper, we propose a new annotation approach to Chinese word segmentation, part-of-speech (POS) tagging and dependency labelling that aims to overcome the two major issues in traditional morphology-based annotation: Inconsistency and data sparsity. We re-annotate the Penn Chinese Treebank 5.0 (CTB5) and demonstrate the advantages of this approach compared to the original CTB5 annotation through word segmentation, POS tagging and machine translation experiments.
Accurate automatic processing of Web queries is important for high-quality information retrieval from the Web. While the syntactic structure of a large portion of these queries is trivial, the structure of queries with question intent is much richer. In this paper we therefore address the task of statistical syntactic parsing of such queries. We first show that the standard dependency grammar does not account for the full range of syntactic structures manifested by queries with question intent. To alleviate this issue we extend the dependency grammar to account for segments -independent syntactic units within a potentially larger syntactic structure. We then propose two distant supervision approaches for the task. Both algorithms do not require manually parsed queries for training. Instead, they are trained on millions of (query, page title) pairs from the Community Question Answering (CQA) domain, where the CQA page was clicked by the user who initiated the query in a search engine. Experiments on a new treebank 1 consisting of 5,000 Web queries from the CQA domain, manually parsed using the proposed grammar, show that our algorithms outperform alternative approaches trained on various sources: tens of thousands of manually parsed OntoNotes sentences, millions of unlabeled CQA queries and thousands of manually segmented CQA queries.
Syntactic information, obtainable through syntactical analysis, plays an important role in many areas of NLP. Researches of Indonesian constituent parser have been very limited with the currently available yields very poor performance of 38.89% and 47.22% using Earley and CYK algorithm respectively. With the availability of the newly introduced Indonesian treebank corpus, we evaluate the performance of Indonesian constituent parser using Trance parser, a language independent constituent parser which employs deep learning. The parser achieved a respectable f-score of 74.91%, a very significant improvement to previous researches.
Abstract Drawing on an analogy between discourse and syntactic trees, this paper chooses 359 Wall Street Journal articles with multiple paragraphs from the Rhetorical Structure Theory (RST) Discourse Treebank, and converts each discourse tree into three additional dependency ones, at discourse, paragraph and sentence levels, with exclusively elementary discourse units of clauses, sentences and paragraphs, respectively. It empirically tests and visually presents the genre-specific “summary+details” or “inverted pyramid” structuring of news discourse. It further extends the idea of inverted pyramid structuring to the paragraph and sentence levels. It proves that the body of the report also has a similar schematic top-down installment organization with macro-propositions on top. It also visually and statistically presents the rhetorical structures at sentence level, which differ to some extent from grammatical structures. Operated in line with the compositionality criterion and hierarchy principle of RST, the converted trees provide unique analytical advantages and constitute new research prospects.
BACKGROUND: Curious parallels between the processes of species and language evolution have been observed by many researchers. Retracing the evolution of Indo-European (IE) languages remains one of the most intriguing intellectual challenges in historical linguistics. Most of the IE language studies use the traditional phylogenetic tree model to represent the evolution of natural languages, thus not taking into account reticulate evolutionary events, such as language hybridization and word borrowing which can be associated with species hybridization and horizontal gene transfer, respectively. More recently, implicit evolutionary networks, such as split graphs and minimal lateral networks, have been used to account for reticulate evolution in linguistics. RESULTS: Striking parallels existing between the evolution of species and natural languages allowed us to apply three computational biology methods for reconstruction of phylogenetic networks to model the evolution of IE languages. We show how the transfer of methods between the two disciplines can be achieved, making necessary methodological adaptations. Considering basic vocabulary data from the well-known Dyen's lexical database, which contains word forms in 84 IE languages for the meanings of a 200-meaning Swadesh list, we adapt a recently developed computational biology algorithm for building explicit hybridization networks to study the evolution of IE languages and compare our findings to the results provided by the split graph and galled network methods. CONCLUSION: We conclude that explicit phylogenetic networks can be successfully used to identify donors and recipients of lexical material as well as the degree of influence of each donor language on the corresponding recipient languages. We show that our algorithm is well suited to detect reticulate relationships among languages, and present some historical and linguistic justification for the results obtained. Our findings could be further refined if relevant syntactic, phonological and morphological data could be analyzed along with the available lexical data.
Inferring implicit discourse relations in natural language text is the most difficult subtask in discourse parsing. Surface features achieve good performance, but they are not readily applicable to other languages without semantic lexicons. Previous neural models require parses, surface features, or a small label set to work well. Here, we propose neural network models that are based on feedforward and long-short term memory architecture without any surface features. To our surprise, our best configured feedforward architecture outperforms LSTM-based model in most cases despite thorough tuning. Under various fine-grained label sets and a cross-linguistic setting, our feedforward models perform consistently better or at least just as well as systems that require hand-crafted surface features. Our models present the first neural Chinese discourse parser in the style of Chinese Discourse Treebank, showing that our results hold cross-linguistically.
We present a novel annotation framework for representing predicate-argument structures, which uses dependency trees to encode the syntactic and semantic roles of a sentence simultaneously. The main contribution is a semantic role transmission model, which eliminates the structural gap between syntax and shallow semantics, making them compatible. A Chinese semantic treebank was built under the proposed framework, and the first release containing about 14K sentences is made freely available. The proposed framework enables semantic role labeling to be solved as a sequence labeling task, and experiments show that standard sequence labelers can give competitive performance on the new treebank compared with state-of-the-art graph structure models.
Morphological segmentation has traditionally been modeled with non-hierarchical models, which yield flat segmentations as output. In many cases, however, proper morphological analysis requires hierarchical structureespecially in the case of derivational morphology. In this work, we introduce a discriminative, joint model of morphological segmentation along with the orthographic changes that occur during word formation. To the best of our knowledge, this is the first attempt to approach discriminative segmentation with a context-free model. Additionally, we release an annotated treebank of 7454 English words with constituency parses, encouraging future research in this area. 1
Lexical information, including surface word form and part-of-speech (POS) information, plays a crucial role when predicting ambiguous dependency relationships in dependency parsing. However, for resolving dependency ambiguities, surface word information may be too sparse, while POS information may be too coarse. Supertags, which are lexical templates that represent rich syntactic information, have been shown to provide effective features at an intermediate level on the coarse-to-fine scale. In this work, we present a supertag design framework that allows us to instantiate various supertag sets based on the dependency structures. Using this framework, we instantiate various supertag sets and utilize them as features in transition-based dependency parsing systems. Performing experiments on the Penn Treebank and Universal Dependencies data sets, we show that our supertags are effective for transition-based parsers in multilingual parsing as well as English parsing. The comparison of the results of the different supertag sets shows that it is crucial to incorporate the head directionality, head labels, and dependent possession information in supertags to improve the parser performance.
Historical treebanks tend to be manually annotated, which is not surprising, since state-of-the-art parsers are not accurate enough to ensure high-quality annotation for historical texts. We test whether automatic parsing can be an efficient pre-annotation tool for Old East Slavic texts. We use the TOROT treebank from the PROIEL treebank family. We convert the PROIEL format to the CONLL format and use MaltParser to create syntactic pre-annotation. Using the most conservative evaluation method, which takes into account PROIEL-specific features, MaltParser by itself yields 0.845 unlabelled attachment score, 0.779 labelled attachment score and 0.741 secondary dependency accuracy (note, though, that the test set comes from a relatively simple genre and contains rather short sentences). Experiments with human annotators show that preparsing, if limited to sentences where no changes to word or sentence boundaries are required, increases their annotation rate. For experienced annotators, the speed gain varies from 5.80% to 16.57%, for inexperienced annotators from 14.61% to 32.17% (using conservative estimates). There are no strong reliable differences in the annotation accuracy, which means that there is no reason to suspect that using preparsing might lower the final annotation quality.
This paper questions the nature of the communicative event that takes place in online contexts between doctors and web-users, showing computer-mediated linguistic norms and discussing the nature of the participants’ roles. Based on an analysis of 1005 posts occurring between doctors and the users of health service websites, I analyse how doctor–patient communication is affected by the medium and how health professionals overcome issues concerning the virtual medical visit. Results suggest that (a) online medical answers offer a different service from that expected by users, as doctors cannot always fulfill patient requests, and (b) net consultations use aspects of traditional doctor–patient exchange and yet present a language and a style that are affected by the computer-mediated environment. Additionally, it seems that this new form leads to a different model of doctor–patient relationship. The findings are intended to provide new insights into web-based discourse in doctor–patient communication and to demonstrate the emergence of a new style in medical communication.
In werkwoordsgroepen met een vervangende infinitief (IPP) wordt in het Nederlands de keuze van het hulpwerkwoord van de voltooide tijd doorgaans bepaald door de IPP, zoals in heeft kunnen komen, maar de keuze kan ook door het hoofdwerkwoord bepaald worden, zoals in is kunnen komen. Gebruik makend van een aantal treebanks en corpora (CGN, Lassy, SoNaR) hebben we onderzocht welke IPP’s deze alternantie vertonen. Dat blijken er naast kunnen nog minstens 13 andere te zijn. Voor de twee meest frequente (moeten en kunnen) hebben we vervolgens nagegaan wat de verhouding is van de voorkomens met hebben en zijn in die gevallen waarin de alternantie mogelijk is. Daarbij is gebleken dat in gemiddeld 80% van de gevallen de keuze van het hulpwerkwoord door de IPP wordt bepaald. Een belangrijke factor bij de keuze is de aard van het hoofdwerkwoord.
We investigate mutual benefits between syntax and semantic roles using neural network models, by studying a parsingSRL pipeline, a SRLparsing pipeline, and a simple joint model by embedding sharing. The integration of syntactic and semantic features gives promising results in a Chinese Semantic Treebank, demonstrating large potentials of neural models for joint parsing and semantic role labeling.
The article presents methodological analyses of topical ideas of the famous modern linguist – E.Cosseriu. The authors argue that incorporation of theoretical ideas of E.Cosseriu could substentially extand the euristic potential of the conept of norm in the sphere of linguistics. The key to solvation of the problem lies in the necessity of changing of modern theoretical context of the question. The authors of the article consider the history of operationalization of the concept of norm in linguistics from the point of view of a specific hermeneutic approach in relation to other linguistic techniques. In the course of study of the role of linguistic norms in interaction of content and expression the article presents examples of extrapolation of the concept of norm from one discipline to another. In case of extrapolation of the concept of norm from the other disciplines into linguistic investigations the structural and functional dichotomy of language turns out that it is impossible to avoid considering spiritual as the world of objects, so that speech and language are considered as two different things. The rapid development of information technologies made possible to calculate many of the aspects of Humboldtian ideas. Computer statistics created conditions where the idea of "language picture of the world" and of the "inner form of the language" is gradually losing its original romantic charge and turns in a very trivial thing. Yet despite the fact that global standardization significantly enhances the processing and automatic addition of translation, which once gave beginning to hermeneutics the hermeneutic potential of linguistic norm still preseves many promissing prospects from the epistemological point of view. Key words: Linguistic norm and variability, E. Cosseriu, hermeneutics, euristic potential. В статье представлен методологический анализ актуальных идей известного современного лингвиста - E. Коссериу. Авторы утверждают, что введение теоретических идей E. Коссериу могло бы существенно расширить эвристический потенциал нормы в области лингвистики. Ключ к решению проблемы заключается в необходимости изменения современного теоретического контекста вопроса. Авторы статьи рассматривают историю ввода в действие понятия нормы в лингвистике с точки зрения конкретного герменевтического подхода по отношению к другим языковым методам. В ходе изучения роли языковых норм во взаимодействии содержания и выражения, в статье представлены примеры экстраполяции понятия нормы от одной дисциплины к другой. В случае экстраполяции понятия нормы с других дисциплин в лингвистические исследования структурно-функциональной дихотомии языка оказывается, что нельзя не рассматривать духовное как мир объектов, так как речь и язык рассматриваются как две разные вещи. Быстрое развитие информационных технологий сделало возможным для расчета многие аспекты гумбольдтовских идей. Статистика компьютера создала условия, в которых идея «языковой картины мира» и «внутренней формы языка» постепенно теряет свой первоначальный романтический заряд и превращается в очень тривиальную вещь. Тем не менее, несмотря на то, что глобальная стандартизация значительно улучшает обработку и автоматический перевод, который когда-то дал начало герменевтике, герменевтический потенциал языковой нормы хранит еще много перспектив с гносеологической точки зрения.Ключевые слова: лингвистическая норма и изменчивость, E.Коссериу, герменевтика, эвристический по- тенциал.
This paper presents the Universal Dependencies tagset (UD v1) as a new annotation scheme for Russian treebanks. The universal list of dependency relations was adopted and extended to comply with certain language-specific syntactic constructions. The tagset was validated, converting two Russian treebanks into the UD format, UD-Russian-SynTagRus and UD-Russian-Google.
Drawing on the work of Lawrence Abu Hamdan, a British-Lebanese artist and researcher currently based in Beirut, this essay examines the juridical and conceptual field of critical forensis which is situated at the juncture of security studies, art, and architecture. Abu Hamdan extends forensics to the area of “new audibilities,” with a focus on the politics of juridical hearing in situations of legal-identity profiling and voice authentication (the “shibboleth test”). Abu Hamdan's projects investigate how accent monitoring and audio surveillance, voice recognition, translation technologies, sovereign acts of listening, and court determinations of linguistic norms emerge as so many technical constraints on “freedom of speech,” itself a malleable term ascribed to discrepant claims and principles, yet taking on performative force in site-specific situations.
We present a new, sizeable dataset of nounnoun compounds with their syntactic analysis (bracketing) and semantic relations. Derived from several established linguistic resources, such as the Penn Treebank, our dataset enables experimenting with new approaches towards a holistic analysis of noun-noun compounds, such as jointlearning of noun-noun compounds bracketing and interpretation, as well as integrating compound analysis with other tasks such as syntactic parsing.
Objectives: This study investigated the role of response style biases in the assessment of positive and negative affect in aging research; it addressed whether response styles (a) are associated with age-related changes in cognitive abilities, (b) lead to distorted conclusions about age differences in affect, and (c) reduce the convergent and predictive validity of affect measures in relation to health outcomes. Method: A multidimensional item response theory model was used to extract response styles from affect ratings provided by respondents to the psychosocial questionnaire (n = 6,295; aged 50-100 years) in the Health and Retirement Study (HRS). Results: The likelihood of extreme response styles (disproportionate use of "not at all" and "very much" response categories) increased significantly with age, and this effect was mediated by age-related decreases in HRS cognitive test scores. Removing response styles from affect measures did not alter age patterns in positive and negative affect; however, it consistently enhanced the convergent validity (relationships with concurrent depression and mental health problems) and predictive validity (prospective relationships with hospital visits, physical illness onset) of the affect measures. Discussion: The results support the importance of detecting and controlling response styles when studying self-reported affect in aging research.
The paper evaluates the differences between two currently leading annotation schemes for dependency treebanks. By relying on four treebanks, we demonstrate that the treatment of conjunctions and adpositions represents the core difference between the two schemes and that this impacts the topological properties of the linguistic networks induced from the treebanks. We also show that such properties are reflected in the performances of four probabilistic dependency parsers trained on the treebanks. L’articolo valuta le differenze tra i due principali schemi di annotazione a dipenden-ze in uso. Sulla base di quattro treebank, l’articolo dimostra che il trattamento delle congiunzioni e delle pre/postposizioni rappresenta la differenza principale tra i due schemi e che ciò comporta delle conseguenze sulle proprietà topologiche dei net-work indotti dalle treebank. Inoltre, si dimostra come tali proprietà siano riflesse nell’accuratezza di quattro parser probabilistici a dipendenze addestrati sulle treebank.
Abstract We are investigating methods by which data from dependency syntax treebanks of ancient Greek can be applied to questions of authorship in ancient Greek historiography. From the Ancient Greek Dependency Treebank were constructed syntax words (sWords) by tracing the shortest path from each leaf node to the root for each sentence tree. This paper presents the results of a preliminary test of the usefulness of the sWord as a stylometric discriminator. The sWord data was subjected to clustering analysis. The resultant groupings were in accord with traditional classifications. The use of sWords also allows a more fine-grained heuristic exploration of difficult questions of text reuse. A comparison of relative frequencies of sWords in the directly transmitted Polybius book 1 and the excerpted books 9–10 indicate that the measurements of the two texts are generally very close, but when frequencies do vary, the differences are surprisingly large. These differences reveal that a certain syntactic simplification is a salient characteristic of Polybius’ excerptor, who leaves conspicuous syntactic indicators of his modifications.
PURPOSE: The focus of this study was to examine the influence of fundamental frequency (F0) and vocal tract length (VTL) modifications on speaker gender recognition in cochlear implant (CI) recipients for different stimulus types. METHOD: Single words and sentences were manipulated using isolated or combined F0 and VTL cues. Using an 11-point rating scale, CI recipients and listeners with normal hearing rated the maleness/femaleness of the corresponding voice. RESULTS: Speaker gender ratings for combined F0 and VTL modifications were similar across all stimulus types in both CI recipients and listeners with normal hearing, although the CI recipients showed a somewhat larger ambiguity. In contrast to listeners with normal hearing, F0-VTL and F0-only modifications revealed similar ratings in the CI recipients when using words as stimuli. However, when sentences were used, a difference was found between F0-VTL-based and F0-based ratings. Modifying VTL cues alone did not affect ratings in the CI group. CONCLUSIONS: Whereas speaker gender ratings by listeners with normal hearing relied on combined VTL and F0 cues, CI recipients made only limited use of VTL cues, which might be one reason behind problems with identifying the speaker on the basis of voice. However, use of the voice cues depended on stimulus type, with the greater information in sentences allowing a more detailed analysis than single words in both listener groups.
Abstract Three studies examined gender differences in the effect of storytelling ability on perceptions of a person's attractiveness as a short‐term and long‐term romantic partner. In Study 1, information about a potential partner's storytelling ability was provided. Study 2 participants read a good or poor story supposedly written by a potential partner. Results suggested that only women's attractiveness assessments of men as a long‐term date increased for good storytellers. Storytelling ability did not affect men's ratings of women nor did it affect ratings of short‐term partners. Study 3 suggested that the effect of storytelling ability on long‐term attractiveness for male targets may be mediated by perceived status. Storytelling ability appears to increase perceived status and thus helps men attract long‐term partners.
We propose a framework to model human comprehension of discourse connectives. Following the Bayesian pragmatic paradigm, we advocate that discourse connectives are interpreted based on a simulation of the production process by the speaker, who, in turn, considers the ease of interpretation for the listener when choosing connectives. Evaluation against the sense annotation of the Penn Discourse Treebank confirms the superiority of the model over literal comprehension. A further experiment demonstrates that the proposed model also improves automatic discourse parsing.
The Universal Dependencies (UD) Project seeks to build a cross-lingual studies of treebanks, linguistic structures and parsing. Its goal is to create a set of multilingual harmonized treebanks that are designed according to a universal annotation scheme. In this paper, we report on the conversion of the Uyghur dependency treebank to a UD version of the treebank which we term the Uyghur Universal Dependency Treebank (UyDT). We present the mapping of the Uyghur dependency treebank’s labelling scheme to the UD scheme, along with a clear description of the structural changes required in this conversion.
In accordance with the compositionality criterion and hierarchy principle of Rhetorical Structure Theory (RST), this study reframes each tree in the RST Discourse Treebank into three new dependency trees with ultimate nodes being clauses, sentences, and paragraphs, respectively, which also draw on an analogy between syntactic and discourse trees. Detailed percentages of various RST relations at the three granularity levels are examined, illuminating the discourse processes of organizing units of one granularity level into those of the next upper level and suggesting certain homogeneity and interaction across levels in the Treebank, particularly at the two upper levels. The study demonstrates the applicability of RST analysis between same-level terminal units. With unique analytical advantages, the newly constructed discourse dependency trees provide new research prospects.
We propose a classification framework for semantic type identification of compounds in Sanskrit. We broadly classify the compounds into four different classes namely, Avyayībhāva, Tatpuruṣa, Bahuvrīhi and Dvandva. Our classification is based on the traditional classification system followed by the ancient grammar treatise Adṣṭādhyāyī, proposed by Pāṇini 25 centuries back. We construct an elaborate features space for our system by combining conditional rules from the grammar Adṣṭādhyāyī, semantic relations between the compound components from a lexical database Amarakoṣa and linguistic structures from the data using Adaptor Grammars. Our in-depth analysis of the feature space highlight inadequacy of Adṣṭādhyāyī, a generative grammar, in classifying the data samples. Our experimental results validate the effectiveness of using lexical databases as suggested by Amba Kulkarni and Anil Kumar, and put forward a new research direction by introducing linguistic patterns obtained from Adaptor grammars for effective identification of compound type. We utilise an ensemble based approach, specifically designed for handling skewed datasets and we %and Experimenting with various classification methods, we achieve an overall accuracy of 0.77 using random forest classifiers.
The Internet-scale open source software (OSS) production in various communities are generating abundant reusable resources for software developers. However, how to retrieve and reuse the desired and mature software from huge amounts of candidates is a great challenge: there are usually big gaps between the user application contexts (that often used as queries) and the OSS key words (that often used to match the queries). In this paper, we define the scenario-based query problem for OSS retrieval, and then we propose a novel approach to reformulate the raw query by leveraging the crowd wisdom from millions of developers to improve the retrieval results. We build a software-specific domain lexical database based on the knowledge in open source communities, by which we can expand and optimize the input queries. The experiment results show that, our approach can reformulate the initial query effectively and outperforms other existing search engines significantly at finding mature software.