Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Background: Disclosing medical errors is considered necessary by patients, ethicists, and health care professionals. Literature insists on the framing of this disclosure and describes the apology as appropriate and necessary. However, this policy seems difficult to put into practice. Few works have explored the function and meaning of the apology. Objective: The aim of this study was to explore the role ascribed to apology in communication between healthcare professionals and patients when disclosing a medical error, and to discuss these findings using a linguistic and philosophical perspective. Methods: Qualitative exploratory study, based on face-to-face semi-structured interviews, with seven physicians in a neonatal unit in France. Discourse analysis. Results: Four themes emerged. Difference between apology in everyday life and in the medical encounter; place of the apology in the process of disclosure together with explanations, regrets, empathy and ways to avoid repeating the)
Social skills training, performed by human trainers, is a well-established method for obtaining appropriate skills in social interaction. Previous work automated the process of social skills training by developing a dialogue system that teaches social communication skills through interaction with a computer avatar. Even though previous work that simulated social skills training only considered acoustic and linguistic information, human social skills trainers take into account visual and other non-verbal features. In this paper, we create and evaluate a social skills training system that closes this gap by considering the audiovisual features of the smiling ratio and the head pose (yaw and pitch). In addition, the previous system was only tested with graduate students; in this paper, we applied our system to children or young adults with autism spectrum disorders. For our experimental evaluation, we recruited 18 members from the general population and 10 people with autism spectrum dis)
Complex networks are often organized in groups or communities of agents that share the same features and/or functions, and this structural organization is built naturally with the formation of the system. In social networks, we argue that the dynamic of linguistic interactions of agreement among people can be a crucial factor in generating this community structure, given that sharing opinions with another person bounds them together, and disagreeing constantly would probably weaken the relationship. We present here a computational model of opinion exchange that uncovers the community structure of a network. Our aim is not to present a new community detection method proper, but to show how a model of social communication dynamics can reveal the (simple and overlapping) community structure in an emergent way. Our model is based on a standard Naming Game, but takes into consideration three social features: trust, uncertainty and opinion preference, that are built over time as agents comm)
The judgement of skill experience and its levels is ambiguous though it is crucial for decision-making in sport sciences studies. We developed a fuzzy decision support system to classify experience of non-elite distance runners. Two Mamdani subsystems were developed based on expert running coaches’ knowledge. In the first subsystem, the linguistic variables of training frequency and volume were combined and the output defined the quality of running practice. The second subsystem yielded the level of running experience from the combination of the first subsystem output with the number of competitions and practice time. The model results were highly consistent with the judgment of three expert running coaches (r>0.88, p<0.001) and also with five other expert running coaches (r>0.86, p<0.001). From the expert’s knowledge and the fuzzy model, running experience is beyond the so-called "10-year rule" and depends not only on practice time, but on the quality of practice (training volume and)
We report the results of a bilingual continuous recognition memory task during which single- and multi-neuron activity was recorded in human subjects with intracranial microwire implants. Subjects (n = 5) were right-handed Spanish-English bilinguals who were undergoing evaluation prior to surgery for severe epilepsy. Subjects were presented with Spanish and English words and the task was to determine whether any given word had been seen earlier in the testing session, irrespective of the language in which it had appeared. Recordings in the left and right hippocampus revealed notable laterality, whereby both Spanish and English items that had been seen previously in the other language (switch trials) triggered increased neural firing in the left hippocampus. Items that had been seen previously in the same language (repeat trials) triggered increased neural firings in the right hippocampus. These results are consistent with theories that propose roles of both the left- and right-hemisph)
Objective: Knowing which specific verbal techniques “good” therapists use in their daily work is important for training and evaluation purposes. In order to systematize what is being practiced in the field, our aim was to empirically identify verbal techniques applied in psychodynamic sessions and to differentiate them according to their basic semantic features using a bottom-up, qualitative approach. Method: Mixed-Method-Design: In a comprehensive qualitative study, types of techniques were identified at the level of utterances based on transcribed psychodynamic therapy sessions using Qualitative Content Analysis (4211 utterances). The definitions of the identified categories were successively refined and modified until saturation was achieved. In a subsequent quantitative study, inter-rater reliability was assessed both at the level of utterances (n = 8717) and at the session level (n = 38). The convergent validity of the categories was investigated by analyzing associations with )
Today, a considerable proportion of the public political discourse on nationwide elections proceeds in Online Social Networks. Through analyzing this content, we can discover the major themes that prevailed during the discussion, investigate the temporal variation of positive and negative sentiment and examine the semantic proximity of these themes. According to existing studies, the results of similar tasks are heavily dependent on the quality and completeness of dictionaries for linguistic preprocessing, entity discovery and sentiment analysis. Additionally, noise reduction is achieved with methods for sarcasm detection and correction. Here we report on the application of these methods on the complete corpus of tweets regarding two local electoral events of worldwide impact: the Greek referendum of 2015 and the subsequent legislative elections. To this end, we compiled novel dictionaries for sentiment and entity detection for the Greek language tailored to these events. We subsequen)
This work is the first to take advantage of recurrent neural networks to predict influenza-like illness (ILI) dynamics from various linguistic signals extracted from social media data. Unlike other approaches that rely on timeseries analysis of historical ILI data and the state-of-the-art machine learning models, we build and evaluate the predictive power of neural network architectures based on Long Short Term Memory (LSTMs) units capable of nowcasting (predicting in “real-time”) and forecasting (predicting the future) ILI dynamics in the 2011 – 2014 influenza seasons. To build our models we integrate information people post in social media e.g., topics, embeddings, word ngrams, stylistic patterns, and communication behavior using hashtags and mentions. We then quantitatively evaluate the predictive power of different social media signals and contrast the performance of the-state-of-the-art regression models with neural networks using a diverse set of evaluation metrics. Finally, we )
Because a biomass gasification station includes various hazard factors, hazard assessment is needed and significant. In this article, the cloud model (CM) is employed to improve set pair analysis (SPA), and a novel hazard assessment method for a biomass gasification station is proposed based on the cloud model-set pair analysis (CM-SPA). In this method, cloud weight is proposed to be the weight of index. In contrast to the index weight of other methods, cloud weight is shown by cloud descriptors; hence, the randomness and fuzziness of cloud weight will make it effective to reflect the linguistic variables of experts. Then, the cloud connection degree (CCD) is proposed to replace the connection degree (CD); the calculation algorithm of CCD is also worked out. By utilizing the CCD, the hazard assessment results are shown by some normal clouds, and the normal clouds are reflected by cloud descriptors; meanwhile, the hazard grade is confirmed by analyzing the cloud descriptors. After that)
We examined if external cues such as other agents' actions can influence the choice of language during voluntary and cued object naming in bilinguals in three experiments. Hindi–English bilinguals first saw a cartoon waving at a color patch. They were then asked to either name a picture in the language of their choice (voluntary block) or to name in the instructed language (cued block). The colors waved at by the cartoon were also the colors used as language cues (Hindi or English). We compared the influence of the cartoon’s choice of color on naming when speakers had to indicate their choice explicitly before naming (Experiment 1) as opposed to when they named directly on seeing the pictures (Experiment 2 and 3). Results showed that participants chose the language indicated by the cartoon greater number of times (Experiment 1 and 3). Speakers also switched significantly to the language primed by the cartoon greater number of times (Experiment 1 and 2). These results suggest that choi)
Physical capacity and coordination cannot alone predict success in team sports such as soccer. Instead, more focus has been directed towards the importance of cognitive abilities, and it has been suggested that executive functions (EF) are fundamentally important for success in soccer. However, executive functions are going through a steep development from adolescence to adulthood. Moreover, more complex EF involving manipulation of information (higher level EF) develop later than simple executive functions such as those linked to simple working memory capacity (Core EF). The link between EF and success in young soccer players is therefore not obvious. In the present study we investigated whether EF are associated with success in soccer in young elite soccer players. We performed tests measuring core EF (a demanding working memory task involving a variable n-back task; dWM) and higher level EF (Design Fluency test; DF). Color-Word Interference Test and Trail Making Test were performed)
The increasing growth of literature in biodiversity presents challenges to users who need to discover pertinent information in an efficient and timely manner. In response, text mining techniques offer solutions by facilitating the automated discovery of knowledge from large textual data. An important step in text mining is the recognition of concepts via their linguistic realisation, i.e., terms. However, a given concept may be referred to in text using various synonyms or term variants, making search systems likely to overlook documents mentioning less known variants, which are albeit relevant to a query term. Domain-specific terminological resources, which include term variants, synonyms and related terms, are thus important in supporting semantic search over large textual archives. This article describes the use of text mining methods for the automatic construction of a large-scale biodiversity term inventory. The inventory consists of names of species, amongst which naming variati)
Humans are highly adept at categorizing visual stimuli, but studies of human categorization are typically validated by verbal reports. This makes it difficult to perform comparative studies of categorization using non-human animals. Interpretation of comparative studies is further complicated by the possibility that animal performance may merely reflect reinforcement learning, whereby discrete features act as discriminative cues for categorization. To assess and compare how humans and monkeys classified visual stimuli, we trained 7 rhesus macaques and 41 human volunteers to respond, in a specific order, to four simultaneously presented stimuli at a time, each belonging to a different perceptual category. These exemplars were drawn at random from large banks of images, such that the stimuli presented changed on every trial. Subjects nevertheless identified and ordered these changing stimuli correctly. Three monkeys learned to order naturalistic photographs; four others, close-up sectio)
For the past 50 years, acknowledgments have been studied as important paratextual traces of research practices, collaboration, and infrastructure in science. Since 2008, funding acknowledgments have been indexed by Web of Science, supporting large-scale analyses of research funding. Applying advanced linguistic methods as well as Correspondence Analysis to more than one million acknowledgments from research articles and reviews published in 2015, this paper aims to go beyond funding disclosure and study the main types of contributions found in acknowledgments on a large scale and through disciplinary comparisons. Our analysis shows that technical support is more frequently acknowledged by scholars in Chemistry, Physics and Engineering. Earth and Space, Professional Fields, and Social Sciences are more likely to acknowledge contributions from colleagues, editors, and reviewers, while Biology acknowledgments put more emphasis on logistics and fieldwork-related tasks. Conflicts of intere)
Abstract: In this study, we compare statistical properties of ancient and modern Chinese within the framework of weighted complex networks. We examine two language networks based on different Chinese versions of the Records of the Grand Historian. The comparative results show that Zipf’s law holds and that both networks are scale-free and disassortative. The interactivity and connectivity of the two networks lead us to expect that the modern Chinese text would have more phrases than the ancient Chinese one. Furthermore, by considering some of the topological and weighted quantities, we find that expressions in ancient Chinese are briefer than in modern Chinese. These observations indicate that the two languages might have different linguistic mechanisms and combinatorial natures, which we attribute to the stylistic differences and evolution of written Chinese. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied )
Despite popular media portraying hoarding to be a problem of extremely poor housekeeping, most hoarded homes are relatively clean – large amounts of stuff just prevent the home from being functional. Some hoarded homes, however, develop poor living conditions like filth or disrepair. To date, little is known about how homes end up this way. The current study identified unique predictors and generated ideas about complex processes involved in the development of poor living conditions in hoarding. Three community agencies shared in-home assessment data for mainly involuntary clients with problematic living conditions, such as hoarded or filthy homes. These community agencies were the Metropolitan Boston Housing Partnership (n=115) in Boston, MA, the Hoarding Action Response Team (n=137) in Vancouver, BC, and the Hamilton Gatekeepers Program (n=209) in Hamilton, ON. Each site completed in-home assessments from 2010-2014 to evaluate client characteristics (lack of insight, social isolation, state of mind) and conditions of the home (number of pets, clutter accumulation, unusable bathrooms or kitchens) using the HOMES: Multidisciplinary Hoarding Risk Assessment, the Clutter Image Rating Scale, or a similar measure. Site-specific regression analyses identified unique predictors of poor living conditions. Clients with high clutter accumulation were at increased risk for squalor at all three sites, while kitchen or bathroom problems uniquely predicted squalor at two sites. Within two agencies, number of pets was also a consistent predictor of one indicator of squalor, the presence of urine or feces. Few clients had household disrepair (9-12% within sites), but findings hint that disrepair is associated with high clutter accumulation. Findings related to poor insight being a predictor of squalor were mixed. This is the first study to directly examine poor living conditions in hoarding. Replicated study findings across sites suggest that common features of hoarding, such as clutter accumulation and unusable rooms, are unique predictors for squalor. Results from this study can help community agencies that deal with problematic living situations prioritize intervention goals, especially if staff believe clients are at risk for poor living conditions.
Vowel reduction is a prominent feature of American English, as well as other stress-timed languages. As a phonological process, vowel reduction neutralizes multiple vowel quality contrasts in unstressed syllables. For bilinguals whose native language is not characterized by large spectral and durational differences between tonic and atonic vowels, systematically reducing unstressed vowels to the central vowel space can be problematic. Failure to maintain this pattern of stressed-unstressed syllables in American English is one key element that contributes to a “foreign accent” in second language speakers. Reduced vowels, or “schwas,” have also been identified as particularly vulnerable to the co-articulatory effects of adjacent consonants. The current study examined the effects of adjacent sounds on the spectral and temporal qualities of schwa in word-final position. Three groups of English-speaking adults were tested: Miami-based monolingual English speakers, early Spanish-English bil)
A numeral classifier is required between a numeral and a noun in Chinese, which comes in two varieties, sortal classifer (C) and measural classifier (M), also known as ‘classifier’ and ‘measure word’, respectively. Cs categorize objects based on semantic attributes and Cs and Ms both denote quantity in terms of mathematical values. The aim of this study was to conduct a psycholinguistic experiment to examine whether participants process C/Ms based on their mathematical values with a semantic distance comparison task, where participants judged which of the two C/M phrases was semantically closer to the target C/M. Results showed that participants performed more accurately and faster for C/Ms with fixed values than the ones with variable values. These results demonstrated that mathematical values do play an important role in the processing of C/Ms. This study may thus shed light on the influence of the linguistic system of C/Ms on magnitude cognition. [ABSTRACT FROM AUTHOR], Copyright o)
Patients with Parkinson’s disease (PD) display a variety of impairments in motor and non-motor language processes; speech is decreased on motor aspects such as amplitude, prosody and speed and on linguistic aspects including grammar and fluency. Here we investigated whether verbal monitoring is impaired and what the relative contributions of the internal and external monitoring route are on verbal monitoring in patients with PD relative to controls. Furthermore, the data were used to investigate whether internal monitoring performance could be predicted by internal speech perception tasks, as perception based monitoring theories assume. Performance of 18 patients with Parkinson’s disease was measured on two cognitive performance tasks and a battery of 11 linguistic tasks, including tasks that measured performance on internal and external monitoring. Results were compared with those of 16 age-matched healthy controls. PD patients and controls generally performed similarly on the lingui)
Health organizations are increasingly using social media, such as Twitter, to disseminate health messages to target audiences. Determining the extent to which the target audience (e.g., age groups) was reached is critical to evaluating the impact of social media education campaigns. The main objective of this study was to examine the separate and joint predictive validity of linguistic and metadata features in predicting the age of Twitter users. We created a labeled dataset of Twitter users across different age groups (youth, young adults, adults) by collecting publicly available birthday announcement tweets using the Twitter Search application programming interface. We manually reviewed results and, for each age-labeled handle, collected the 200 most recent publicly available tweets and user handles’ metadata. The labeled data were split into training and test datasets. We created separate models to examine the predictive validity of language features only, metadata features only, l)
Human language is composed of sequences of reusable elements. The origins of the sequential structure of language is a hotly debated topic in evolutionary linguistics. In this paper, we show that sets of sequences with language-like statistical properties can emerge from a process of cultural evolution under pressure from chunk-based memory constraints. We employ a novel experimental task that is non-linguistic and non-communicative in nature, in which participants are trained on and later asked to recall a set of sequences one-by-one. Recalled sequences from one participant become training data for the next participant. In this way, we simulate cultural evolution in the laboratory. Our results show a cumulative increase in structure, and by comparing this structure to data from existing linguistic corpora, we demonstrate a close parallel between the sets of sequences that emerge in our experiment and those seen in natural language. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the)
The amount of data from languages spoken all over the world is rapidly increasing. Traditional manual methods in historical linguistics need to face the challenges brought by this influx of data. Automatic approaches to word comparison could provide invaluable help to pre-analyze data which can be later enhanced by experts. In this way, computational approaches can take care of the repetitive and schematic tasks leaving experts to concentrate on answering interesting questions. Here we test the potential of automatic methods to detect etymologically related words (cognates) in cross-linguistic data. Using a newly compiled database of expert cognate judgments across five different language families, we compare how well different automatic approaches distinguish related from unrelated words. Our results show that automatic methods can identify cognates with a very high degree of accuracy, reaching 89% for the best-performing method Infomap. We identify the specific strengths and weaknes)
We investigated categorical perception of rising and falling pitch contours by tonal and non-tonal listeners. Specifically, we determined minimum durations needed to perceive both contours and compared to those of production, how stimuli duration affects their perception, whether there is an intrinsic F0 effect, and how first language background, duration, directions of pitch and vowel quality interact with each other. Continua of fundamental frequency on different vowels with 9 duration values were created for identification and discrimination tasks. Less time is generally needed to effectively perceive a pitch direction than to produce it. Overall, tonal listeners’ perception is more categorical than non-tonal listeners. Stimuli duration plays a critical role for both groups, but tonal listeners showed a stronger duration effect, and may benefit more from the extra time in longer stimuli for context-coding, consistent with the multistore model of categorical perception. Within a cer)
The extraction of abstract structures from speech (or from gestures in the case of sign languages) has been claimed to be a fundamental mechanism for language acquisition. In the present study we registered the neural responses that are triggered when a violation of an abstract, token-independent rule is detected. We registered ERPs while presenting participants with trisyllabic CVCVCV nonsense words in an oddball paradigm. Standard stimuli followed an ABB rule (where A and B are different syllables). Importantly, to distinguish neural responses triggered by changes in surface information from responses triggered by changes in the underlying abstract structure, we used two types of deviant stimuli. Phoneme deviants differed from standards only in their phonemes. Rule deviants differed from standards in both their phonemes and their composing rule. We observed a significant positivity as early as 300 ms after the presentation of deviant stimuli that violated the abstract rule (Rule dev)
This paper explores how information flow properties of a network affect the formation of categories shared between individuals, who are communicating through that network. Our work is based on the established multi-agent model of the emergence of linguistic categories grounded in external environment. We study how network information propagation efficiency and the direction of information flow affect categorization by performing simulations with idealized network topologies optimizing certain network centrality measures. We measure dynamic social adaptation when either network topology or environment is subject to change during the experiment, and the system has to adapt to new conditions. We find that both decentralized network topology efficient in information propagation and the presence of central authority (information flow from the center to peripheries) are beneficial for the formation of global agreement between agents. Systems with central authority cope well with network top)
We learn language from our social environment. In general, the more sources we have, the less informative each of them is, and the less weight we should assign it. If this is the case, people who interact with fewer others should be more susceptible to the influence of each of their interlocutors. This paper tests whether indeed people who interact with fewer other people have more malleable phonological representations. Using a perceptual learning paradigm, this paper shows that individuals who regularly interact with fewer others are more likely to change their boundary between /d/ and /t/ following exposure to an atypical speaker. It further shows that the effect of number of interlocutors is not due to differences in ability to learn the speaker’s speech patterns, but specific to likelihood of generalizing the learned pattern. These results have implications for both language learning and language change, as they suggest that individuals with smaller social networks might play an )
Collective behaviour is a fascinating and easily observable phenomenon, attractive to a wide range of researchers. In biology, computational models have been extensively used to investigate various properties of collective behaviour, such as: transfer of information across the group, benefits of grouping (defence against predation, foraging), group decision-making process, and group behaviour types. The question ‘why,’ however remains largely unanswered. Here the interest goes into which pressures led to the evolution of such behaviour, and evolutionary computational models have already been used to test various biological hypotheses. Most of these models use genetic algorithms to tune the parameters of previously presented non-evolutionary models, but very few attempt to evolve collective behaviour from scratch. Of these last, the successful attempts display clumping or swarming behaviour. Empirical evidence suggests that in fish schools there exist three classes of behaviour; swarmi)
Recent theories propose that language comprehension can influence perception at the low level of perceptual system. Here, we used an adaptation paradigm to test whether processing language caused color adaptation in the visual system. After prolonged exposure to a color linguistic context, which depicted red, green, or non-specific color scenes, participants immediately performed a color detection task, indicating whether they saw a green color square in the middle of a white screen or not. We found that participants were more likely to perceive the green color square after listening to discourses denoting red compared to discourses denoting green or conveying non-specific color information, revealing that language comprehension caused an adaptation aftereffect at the perceptual level. Therefore, semantic representation of color may have a common neural substrate with color perception. These results are in line with the simulation view of embodied language comprehension theory, which )
Paintings have high cultural and commercial value, so that needs to be preserved. Many techniques have been attempted to analyze properties of paintings, including X-ray analysis and optical coherence tomography (OCT) methods, and enable conservation of paintings from forgeries. In this paper, we suggest a simple and accurate optical analysis system to protect them from counterfeit which is comprised of fiber optics reflectance spectroscopy (FORS) and line laser-based topographic analysis. The system is designed to fully cover the whole area of paintings regardless of its size for the accurate analysis. For additional assessments, a line laser-based high resolved OCT was utilized. Some forgeries were created by the experts from the three different styles of genuine paintings for the experiments. After measuring surface properties of paintings, we could observe the results from the genuine works and the forgeries have the distinctive characteristics. The forgeries could be distinguishe)
This is the first study to examine the effect of phonetic contexts on children’s lexical tone production. Mandarin tones in disyllabic words produced by forty-four 2- to 6-year-old children and twelve mothers were low-pass filtered to eliminate lexical information. Native Mandarin-speaking adults categorized the tones based on the pitch information in the filtered stimuli. All mothers’ tones were categorized with ceiling accuracy. Counter to the findings in most previous studies on children’s tone acquisition and the prevailing assumption in models of speech development that children acquire suprasegmental features much earlier than segmental features, this study found that children as old as six years of age have not mastered the production of Mandarin tones. Children’s tones were judged with significantly lower accuracy than mothers’ productions. Tone accuracy improved, while cross subject variability in tone accuracy decreased, with age. Children’s tone accuracy was affected by the)
Word recognition includes the activation of a range of syntactic and semantic knowledge that is relevant to language interpretation and reference. Here we explored whether or not the number of arguments a verb takes impinges negatively on verb processing time. In this study, three experiments compared the dynamics of spoken word recognition for verbs with different preferred argument structure. Listeners’ eye movements were recorded as they searched an array of pictures in response to hearing a verb. Results were similar in all the experiments. The time to identify the referent increased as a function of the number of arguments, above and beyond any effects of label appropriateness (and other controlled variables, such as letter, phoneme and syllable length, phonological neighborhood, oral and written lexical frequencies, imageability and rated age of acquisition). The findings indicate that the number of arguments a verb takes, influences referent identification during spoken word re)
Sentence reading involves multiple linguistic operations including processing of lexical and compositional semantics, and determining structural and grammatical relationships among words. Previous studies on Indo-European languages have associated left anterior temporal lobe (aTL) and left interior frontal gyrus (IFG) with reading sentences compared to reading unstructured word lists. To examine whether these brain regions are also involved in reading a typologically distinct language with limited morphosyntax and lack of agreement between sentential arguments, an FMRI study was conducted to compare passive reading of Chinese sentences, unstructured word lists and disconnected character lists that are created by only changing the order of an identical set of characters. Similar to previous findings from other languages, stronger activation was found in mainly left-lateralized anterior temporal regions (including aTL) for reading sentences compared to unstructured word and character li)
Though metaphoric language comprehension has previously been investigated with event-related potentials, little attention has been devoted to extending this research from the monolingual to the bilingual context. In the current study, late proficient unbalanced Polish (L1)–English (L2) bilinguals performed a semantic decision task to novel metaphoric, conventional metaphoric, literal, and anomalous word pairs presented in L1 and L2. The results showed more pronounced P200 amplitudes to L2 than L1, which can be accounted for by differences in the subjective frequency of the native and non-native lexical items. Within the early N400 time window (300–400 ms), L2 word dyads evoked delayed and attenuated amplitudes relative to L1 word pairs, possibly indicating extended lexical search during foreign language processing, and weaker semantic interconnectivity for L2 compared to L1 words within the memory system. The effect of utterance type was observed within the late N400 time window (400–)
Objectives: The present study explored tone perception ability in school age Mandarin-speaking children with otitis media with effusion (OME) in noisy listening environments. The study investigated the interaction effects of noise, tone type, age, and hearing status on monaural tone perception, and assessed the application of a hierarchical clustering algorithm for profiling hearing impairment in children with OME. Methods: Forty-one children with normal hearing and normal middle ear status and 84 children with OME with or without hearing loss participated in this study. The children with OME were further divided into two subgroups based on their severity and pattern of hearing loss using a hierarchical clustering algorithm. Monaural tone recognition was measured using a picture-identification test format incorporating six sets of monosyllabic words conveying four lexical tones under speech spectrum noise, with the signal-to-noise ratio (SNR) conditions ranging from -9 to -21 dB. Re)
We present a new open source software tool called BEASTling, designed to simplify the preparation of Bayesian phylogenetic analyses of linguistic data using the BEAST 2 platform. BEASTling transforms comparatively short and human-readable configuration files into the XML files used by BEAST to specify analyses. By taking advantage of Creative Commons-licensed data from the Glottolog language catalog, BEASTling allows the user to conveniently filter datasets using names for recognised language families, to impose monophyly constraints so that inferred language trees are backward compatible with Glottolog classifications, or to assign geographic location data to languages for phylogeographic analyses. Support for the emerging cross-linguistic linked data format (CLDF) permits easy incorporation of data published in cross-linguistic linked databases into analyses. BEASTling is intended to make the power of Bayesian analysis more accessible to historical linguists without strong programmi)
Despite the ongoing growth in the number of published randomized controlled trials (RCTs) and increased quality assessment of RCTs, the association between the quality and characteristics in the text has not been sufficiently studied. We are interested in a specific question: what kind of sentences is a good indicator of high quality RCTs? To help researchers to efficiently screen articles worth reading, this study aims 1) to quantify the linguistic features of articles and 2) to build a document assessment model to evaluate quality of RCTs using only the abstract. All RCTs that were conducted in Japan in 2010 as original articles were included in the analysis. Data were independently assessed by two reviewers using a risk-of-bias tool. Three aspects of linguistic style were quantitatively measured, and a document model was constructed to evaluate the RCTs. A total of 302 RCTs were selected for quality assessment. Of these, 255 articles were assessed as high quality and 47 as low qual)
PDTSC 1.0 is a multi-purpose corpus of spoken language. 768,888 tokens, 73,374 sentences and 7,324 minutes of spontaneous dialog speech have been recorded, transcribed and edited in several interlinked layers: audio recordings, automatic and manual transcription and manually reconstructed text. PDTSC 1.0 is a delayed release of data annotated in 2012. It is an update of Prague Dependency Treebank of Spoken Language (PDTSL) 0.5 (published in 2009). In 2017, Prague Dependency Treebank of Spoken Czech (PDTSC) 2.0 was published as an update of PDTSC 1.0.
This study introduces the Sentiment Analysis and Cognition Engine (SEANCE), a freely available text analysis tool that is easy to use, works on most operating systems (Windows, Mac, Linux), is housed on a user’s hard drive (as compared to being accessed via an Internet interface), allows for batch processing of text files, includes negation and part-of-speech (POS) features, and reports on thousands of lexical categories and 20 component scores related to sentiment, social cognition, and social order. In the study, we validated SEANCE by investigating whether its indices and related component scores can be used to classify positive and negative reviews in two well-known sentiment analysis test corpora. We contrasted the results of SEANCE with those from Linguistic Inquiry and Word Count (LIWC), a similar tool that is popular in sentiment analysis, but is pay-to-use and does not include negation or POS features. The results demonstrated that both the SEANCE indices and component scores outperformed LIWC on the categorization tasks.
This work explores the feasibility of a crowd-based pair-wise comparison evaluation to get feedback on machine translation progress for under-resourced languages. Specifically, we propose a task based on simple work units to compare the outputs of five English-to-Basque systems, which we implement in a web application. In our design, we put forward two key aspects that we believe community collaboration initiatives should consider in order to attract and maintain participants, that is, providing both a community challenge and a personal challenge. We describe how these aspects can comply with a strict methodology to ensure research validity. In particular, we consider the evaluation set size and the characteristics of the test sentences, the number of evaluators per comparison pair, and a mechanism to identify dishonest participation (or participants with insufficient linguistic knowledge). We also describe our dissemination effort, which targeted both general users and interest groups. Over 500 people participated actively in the Ebaluatoia campaign and we were able to collect over 35,000 evaluations in a short period of 10 days. From the results, we complete the ranking of the systems under evaluation and establish whether the difference in quality between the systems is significant.
This study investigates L2 acquisition of English present perfect by Greek Cypriot Greek speakers. One hundred Greek Cypriot university students took part in the study, the first part of which examined the sensitivity to grammatical norms (a passage correction task, based on Odlin et al. 2006), and the other part was focused on the production of English present perfect (elicitation of natural discourse, essays about personal experience). The results showed that L2 learners used more non-target tense forms (present simple and past simple) than the target present perfect in typical contexts, which is due to transfer from L1 Cypriot Greek (CG). The data only partially supports the Inherent Lexical Aspect Hypothesis (Andersen and Shirai 1996; Bardovi-Harlig 1999), as L2 learners used perfective and past tense morphology with both punctual-telic predicates (achievements or accomplishments) and atelic or durative predicates (state or activity), though their production of target present perfect improves with more years of exposure to L2 English and there is a decrease in the use of stative and activity verbs with perfective and past tense marking.
The objective of this paper is to investigate the possibility of using key word analysis of corpus linguistics methodology in the translation quality assessment in legal genre. The conventional approach in translation quality assessment focuses on the achievement of equivalence between the source text and the target text. However, this kind of strong emphasis on preserving the letter of law often disrupts the understanding of the target reader, thus causing default in securing the same legal effect intended in the translated legal text. Against this backdrop, this study adopts more reader-oriented approach in legal translation quality assessment based on the concept of textual fit proposed mainly by Biel (2014). Focusing more on the expectancy norm of the target audience, this mode of assessment evaluates the extent to which the translation fits into the non-translated convention of the relevant sub-genre of legal language, thus lessening the cognitive processing efforts of the target reader. This paper suggests a model for assessing textual fit by examining the overused lexical, grammatical and semantic patterns of English-translated Korean statutes compared to non-translated English statutes, based on hierarchical key word analysis results provided by Wmatrix.
We present a widely applicable methodology to bring machine translation (MT) to under-resourced languages in a cost-effective and rapid manner. Our proposal relies on web crawling to automatically acquire parallel data to train statistical MT systems if any such data can be found for the language pair and domain of interest. If that is not the case, we resort to (1) crowdsourcing to translate small amounts of text (hundreds of sentences), which are then used to tune statistical MT models, and (2) web crawling of vast amounts of monolingual data (millions of sentences), which are then used to build language models for MT. We apply these to two respective use-cases for Croatian, an under-resourced language that has gained relevance since it recently attained official status in the European Union. The first use-case regards tourism, given the importance of this sector to Croatia’s economy, while the second has to do with tweets, due to the growing importance of social media. For tourism, we crawl parallel data from 20 web domains using two state-of-the-art crawlers and explore how to combine the crawled data with bigger amounts of general-domain data. Our domain-adapted system is evaluated on a set of three additional tourism web domains and it outperforms the baseline in terms of automatic metrics and/or vocabulary coverage. In the social media use-case, we deal with tweets from the 2014 edition of the soccer World Cup. We build domain-adapted systems by (1) translating small amounts of tweets to be used for tuning by means of crowdsourcing and (2) crawling vast amounts of monolingual tweets. These systems outperform the baseline (Microsoft Bing) by 7.94 BLEU points (5.11 TER) for Croatian-to-English and by 2.17 points (1.94 TER) for English-to-Croatian on a test set translated by means of crowdsourcing. A complementary manual analysis sheds further light on these results.
Removal of boilerplate is one of the essential tasks in web corpus construction and web indexing. Boilerplate (redundant and automatically inserted material like menus, copyright notices, navigational elements, etc.) is usually considered to be linguistically unattractive for inclusion in a web corpus. Also, search engines should not index such material because it can lead to spurious results for search terms if these terms appear in boilerplate regions of the web page. In this paper, I present and evaluate a supervised machine-learning approach to general-purpose boilerplate detection for languages based on Latin alphabets using Multi-Layer Perceptrons (MLPs). It is both very efficient and very accurate (between 95 % and \(99\,\%\) correct classifications, depending on the input language). I show that language-specific classifiers greatly improve the accuracy of boilerplate detectors. The single features used for the classification are evaluated with regard to the merit they contribute to the classification. Furthermore, I show that the accuracy of the MLP is on a par with that of a wide range of other classifiers. My approach has been implemented in the open-source texrex web page cleaning software, and large corpora constructed using it are available from the COW initiative, including the CommonCOW corpora created from CommonCrawl datasets.
When linguistically annotated data is scarce, as is the case for many under-resourced languages, one has to resort to less complete forms of annotations obtained from crawled dictionaries and/or through cross-lingual transfer. Several recent works have shown that learning from such partially supervised data can be effective in many practical situations. In this work, we review two existing proposals for learning with ambiguous labels which extend conventional learners to the weakly supervised setting: a history-based model using a variant of the perceptron, on the one hand; an extension of the Conditional Random Fields model on the other hand. Focusing on the part-of-speech tagging task, but considering a large set of ten languages, we show (a) that good performance can be achieved even in the presence of ambiguity, provided however that both monolingual and bilingual resources are available; (b) that our two learners exploit different characteristics of the training set, and are successful in different situations; (c) that in addition to the choice of an adequate learning algorithm, many other factors are critical for achieving good performance in a cross-lingual transfer setting.
In line with the dimensional theory of emotional space, we developed affective norms for words rated in terms of valence, arousal and dominance in a group of older adults to complete the adaptation of the Affective Norms for English Words (ANEW) for Italian and to aid research on aging. Here, as in the original Italian ANEW database, participants evaluated valence, arousal, and dominance by means of the Self-Assessment Manikin (SAM) in a paper-and-pencil procedure. We observed high split-half reliabilities within the older sample and high correlations with the affective ratings of previous research, especially for valence, suggesting that there is large agreement among older adults within and across-languages. More importantly, we found high correlations between younger and older adults, showing that our data are generalizable across different ages. However, despite this across-ages accord, we obtained age-related differences on three affective dimensions for a great number of words. In particular, older adults rated as more arousing and more unpleasant a number of words that younger adults rated as moderately unpleasant and arousing in our previous affective norms. Moreover, older participants rated negative stimuli as more arousing and positive stimuli as less arousing than younger participants, thus leading to a less-curved distribution of ratings in the valence by arousal space. We also found more extreme ratings for older adults for the relationship between dominance and arousal: older adults gave lower dominance and higher arousal ratings for words rated by younger adults with middle dominance and arousal values. Together, these results suggest that our affective norms are reliable and can be confidently used to select words matched for the affective dimensions of valence, arousal and dominance across younger and older participants for future research in aging.
The present event-related potential (ERP) study investigated for the first time whether children with early-onset social anxiety disorder (SAD) process affective facial expressions of varying intensities differently than non-anxious controls. Participants were 15 SAD patients and 15 non-anxious controls (mean age of 9 years). They were presented with schematic faces displaying anger and happiness at four intensity levels (25%, 50%, 75%, and 100%), as well as with neutral faces. ERPs in early and later time windows (P100, N170, late positivity [LP]), as well as affective ratings (valence and arousal) for the faces, were recorded. SAD patients rated the faces as generally more arousing, regardless of the type of emotion and intensity. Moreover, they displayed enhanced right-parietal LP (350-650 ms). Both arousal ratings and LP reflect stimulus intensity. Therefore, this study provides first evidence of an intensity amplification bias in pediatric SAD during facial affect processing.
Mondzish (Mangish) lexical database, including transcriptions of my audio recordings collected in China in from 2012-2015.
One particular problem in large vocabulary continuous speech recognition for low-resourced languages is finding relevant training data for the statistical language models. Large amount of data is required, because models should estimate the probability for all possible word sequences. For Finnish, Estonian and the other fenno-ugric languages a special problem with the data is the huge amount of different word forms that are common in normal speech. The same problem exists also in other language technology applications such as machine translation, information retrieval, and in some extent also in other morphologically rich languages. In this paper we present methods and evaluations in four recent language modeling topics: selecting conversational data from the Internet, adapting models for foreign words, multi-domain and adapted neural network language modeling, and decoding with subword units. Our evaluations show that the same methods work in more than one language and that they scale down to smaller data resources.
This paper presents an overview of studies on automated hand gesture analysis, which is mainly concerned with recognition and segmentation issues related to functional types and gesture phases. The issues selected for discussion have been arranged in a way that takes account of problems within the Theory of Gestures that each study seeks to address. Their principal computational factors that were involved in conducting the analysis of automated hand gesture have been examined, and an analysis of open research issues has been carried out for each application dealt with in the studies.