Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Evenks and Evens, Tungusic-speaking reindeer herders and hunter-gatherers, are spread over a wide area of northern Asia, whereas their linguistic relatives the Udegey, sedentary fishermen and hunter-gatherers, are settled to the south of the lower Amur River. The prehistory and relationships of these Tungusic peoples are as yet poorly investigated, especially with respect to their interactions with neighbouring populations. In this study, we analyse over 500 complete mtDNA genome sequences from nine different Evenk and even subgroups as well as their geographic neighbours from Siberia and their linguistic relatives the Udegey from the Amur-Ussuri region in order to investigate the prehistory of the Tungusic populations. These data are supplemented with analyses of Y-chromosomal haplogroups and STR haplotypes in the Evenks, Evens, and neighbouring Siberian populations. We demonstrate that whereas the North Tungusic Evenks and Evens show evidence of shared ancestry both in the maternal )
The body-specificity hypothesis (BSH) predicts that right-handers and left-handers allocate positive and negative concepts differently on the horizontal plane, i.e., while left-handers allocate negative concepts on the right-hand side of their bodily space, right-handers allocate such concepts to the left-hand side. Similar research shows that people, in general, tend to allocate positive and negative concepts in upper and lower areas, respectively, in relation to the vertical plane. Further research shows a higher salience of the vertical plane over the horizontal plane in the performance of sensorimotor tasks. The aim of the paper is to examine whether there should be a dominance of the vertical plane over the horizontal plane, not only at a sensorimotor level but also at a conceptual level. In Experiment 1, various participants from diverse linguistic backgrounds were asked to rate the words “up”, “down”, “left”, and “right”. In Experiment 2, right-handed participants from two ling)
Evidence that the motor and the linguistic systems share common syntactic representations would open new perspectives on language evolution. Here, crossing disciplinary boundaries, we explore potential parallels between the structure of simple actions and that of sentences. First, examining Typically Developing (TD) children displacing a bottle with or without knowledge of its weight prior to movement onset, we provide kinematic evidence that the sub-phases of this displacing action (reaching + moving the bottle) manifest a structure akin to linguistic embedded dependencies. Then, using the same motor task, we reveal that children suffering from specific language impairment (SLI), whose core deficit affects syntactic embedding and dependencies, manifest specific structural motor anomalies parallel to their linguistic deficits. In contrast to TD children, SLI children performed the displacing-action as if its sub-phases were juxtaposed rather than embedded. The specificity of SLI’s str)
In second language acquisition research, the critical period hypothesis (cph) holds that the function between learners' age and their susceptibility to second language input is non-linear. This paper revisits the indistinctness found in the literature with regard to this hypothesis's scope and predictions. Even when its scope is clearly delineated and its predictions are spelt out, however, empirical studies–with few exceptions–use analytical (statistical) tools that are irrelevant with respect to the predictions made. This paper discusses statistical fallacies common in cph research and illustrates an alternative analytical method (piecewise regression) by means of a reanalysis of two datasets from a 2010 paper purporting to have found cross-linguistic evidence in favour of the cph. This reanalysis reveals that the specific age patterns predicted by the cph are not cross-linguistically robust. Applying the principle of parsimony, it is concluded that age patterns in second language a)
The combined knowledge of word meanings and grammatical rules does not allow a listener to grasp the intended meaning of a speaker’s utterance. Pragmatic inferences on the part of the listener are also required. The present work focuses on the processing of ironic utterances (imagine a slow day being described as “really productive”) because these clearly require the listener to go beyond the linguistic code. Such utterances are advantageous experimentally because they can serve as their own controls in the form of literal sentences (now imagine an active day being described as “really productive”) as we employ techniques from electrophysiology (EEG). Importantly, the results confirm previous ERP findings showing that irony processing elicits an enhancement of the P600 component (Regel et al., 2011). More original are the findings drawn from Time Frequency Analysis (TFA) and especially the increase of power in the gamma band in the 280–400 time-window, which points to an integration a)
The cortical regions involved in the different stages of speech production are relatively well-established, but their spatio-temporal dynamics remain poorly understood. In particular, the available studies have characterized neural events with respect to the onset of the stimulus triggering a verbal response. The core aspect of language production, however, is not perception but action. In this context, the most relevant question may not be how long after a stimulus brain events happen, but rather how long before the production act do they occur. We investigated speech production-related brain activity time-locked to vocal onset, in addition to the common stimulus-locked approach. We report the detailed temporal interplay between medial and left frontal activities occurring shortly before vocal onset. We interpret those as reflections of, respectively, word selection and word production processes. This medial-lateral organization is in line with that described in non-linguistic action)
Recently, sentence comprehension in languages other than European languages has been investigated from a cross-linguistic perspective. In this paper, we examine whether and how animacy-related semantic information is used for real-time sentence comprehension in a SOV word order language (i.e., Japanese). Twenty-three Japanese native speakers participated in this study. They read semantically reversible and non-reversible sentences with canonical word order, and those with scrambled word order. In our results, the second argument position in reversible sentences took longer to read than that in non-reversible sentences, indicating that animacy information is used in second argument processing. In contrast, for the predicate position, there was no difference in reading times, suggesting that animacy information is NOT used in the predicate position. These results are discussed using the sentence comprehension models of an SOV word order language. [ABSTRACT FROM AUTHOR], Copyright of PLo)
Understanding the patterns and causes of differential structural stability is an area of major interest for the study of language change and evolution. It is still debated whether structural features have intrinsic stabilities across language families and geographic areas, or if the processes governing their rate of change are completely dependent upon the specific context of a given language or language family. We conducted an extensive literature review and selected seven different approaches to conceptualising and estimating the stability of structural linguistic features, aiming at comparing them using the same dataset, the World Atlas of Language Structures. We found that, despite profound conceptual and empirical differences between these methods, they tend to agree in classifying some structural linguistic features as being more stable than others. This suggests that there are intrinsic properties of such structural features influencing their stability across methods, language )
Universal linguistic constraints seem to govern the organization of sound sequences in words. However, our understanding of the origin and development of these constraints is incomplete. One possibility is that the development of neuromuscular control of articulators acts as a constraint for the emergence of sequences in words. Repetitions of the same consonant observed in early infancy and an increase in variation of consonantal sequences over months of age have been interpreted as a consequence of the development of neuromuscular control. Yet, it is not clear how sequential coordination of articulators such as lips, tongue apex and tongue dorsum constrains sequences of labial, coronal and dorsal consonants in words over the course of development. We examined longitudinal development of consonant-vowel-consonant(-vowel) sequences produced by Japanese children between 7 and 60 months of age. The sequences were classified according to places of articulation for corresponding consonants)
Learning the functional properties of objects is a core mechanism in the development of conceptual, cognitive and linguistic knowledge in children. The cerebral processes underlying these learning mechanisms remain unclear in adults and unexplored in children. Here, we investigated the neurophysiological patterns underpinning the learning of functions for novel objects in 10-year-old healthy children. Event-related fields (ERFs) were recorded using magnetoencephalography (MEG) during a picture-definition task. Two MEG sessions were administered, separated by a behavioral verbal learning session during which children learned short definitions about the “magical” function of 50 unknown non-objects. Additionally, 50 familiar real objects and 50 other unknown non-objects for which no functions were taught were presented at both MEG sessions. Children learned at least 75% of the 50 proposed definitions in less than one hour, illustrating children's powerful ability to rapidly map new funct)
People can implicitly learn a connection between linguistic forms and meanings, for example between specific determiners (e.g. this, that…) and the type of nouns to which they apply. Li et al (2013) recently found that transfer of form-meaning connections from a concrete domain (height) to an abstract domain (power) was achieved in a metaphor-consistent way without awareness, showing that unconscious knowledge can be abstract and flexibly deployed. The current study aims to determine whether people transfer knowledge of form-meaning connections not only from a concrete domain to an abstract one, but also vice versa, consistent with metaphor representation being bi-directional. With a similar paradigm as used by Li et al, participants learnt form- meaning connections of different domains (concrete vs. abstract) and then were tested on two kinds of generalizations (same and different domain generalization). As predicted, transfer of form-meaning connections occurred bidirectionally when)
Identifying metaphorical language-use (e.g., sweet child) is one of the challenges facing natural language processing. This paper describes three novel algorithms for automatic metaphor identification. The algorithms are variations of the same core algorithm. We evaluate the algorithms on two corpora of Reuters and the New York Times articles. The paper presents the most comprehensive study of metaphor identification in terms of scope of metaphorical phrases and annotated corpora size. Algorithms’ performance in identifying linguistic phrases as metaphorical or literal has been compared to human judgment. Overall, the algorithms outperform the state-of-the-art algorithm with 71% precision and 27% averaged improvement in prediction over the base-rate of metaphors in the corpus. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's ex)
It is well-known that word frequencies arrange themselves according to Zipf's law. However, little is known about the dependency of the parameters of the law and the complexity of a communication system. Many models of the evolution of language assume that the exponent of the law remains constant as the complexity of a communication systems increases. Using longitudinal studies of child language, we analysed the word rank distribution for the speech of children and adults participating in conversations. The adults typically included family members (e.g., parents) or the investigators conducting the research. Our analysis of the evolution of Zipf's law yields two main unexpected results. First, in children the exponent of the law tends to decrease over time while this tendency is weaker in adults, thus suggesting this is not a mere mirror effect of adult speech. Second, although the exponent of the law is more stable in adults, their exponents fall below 1 which is the typical value of)
Objectives: Increasingly, medical research involves patients who complete outcomes in different languages. This occurs in countries with more than one common language, such as Canada (French/English) or the United States (Spanish/English), as well as in international multi-centre collaborations, which are utilized frequently in rare diseases such as systemic sclerosis (SSc). In order to pool or compare outcomes, instruments should be measurement equivalent (invariant) across cultural or linguistic groups. This study provides an example of how to assess cross-language measurement equivalence by comparing the Center for Epidemiologic Studies Depression (CES-D) scale between English-speaking Canadian and Dutch SSc patients. Methods: The CES-D was completed by 922 English-speaking Canadian and 213 Dutch SSc patients. Confirmatory factor analysis (CFA) was used to assess the factor structure in both samples. The Multiple-Indicator Multiple-Cause (MIMIC) model was utilized to assess the amo)
We investigated whether and how comprehending sentences that describe a social context influences our motor behaviour. Our stimuli were sentences that referred to objects having different connotations (e.g., attractive/ugly vs. smooth/prickly) and that could be directed towards the self or towards “another person” target (e.g., “The object is ugly/smooth. Bring it to you/Give it to another person”). Participants judged whether each sentence was sensible or non-sensible by moving the mouse towards or away from their body. Mouse movements were analysed according to behavioral and kinematics parameters. In order to enhance the social meaning of the linguistic stimuli, participants performed the task either individually (Individual condition) or in a social setting, in co-presence with the experimenter. The experimenter could either act as a mere observer (Social condition) or as a confederate, interacting with participants in an off-line modality at the end of task execution (Joint condi)
During sentence production, linguistic information (semantics, syntax, phonology) of words is retrieved and assembled into a meaningful utterance. There is still debate on how we assemble single words into more complex syntactic structures such as noun phrases or sentences. In the present study, event-related potentials (ERPs) were used to investigate the time course of syntactic planning. Thirty-three volunteers described visually animated scenes using naming formats varying in syntactic complexity: from simple words (‘W’, e.g., “triangle”, “red”, “square”, “green”, “to fly towards”), to noun phrases (‘NP’, e.g., “the red triangle”, “the green square”, “to fly towards”), to a sentence (‘S’, e.g., “The red triangle flies towards the green square.”). Behaviourally, we observed an increase in errors and corrections with increasing syntactic complexity, indicating a successful experimental manipulation. In the ERPs following scene onset, syntactic complexity variations were found in a P3)
Due to its pivotal geographical location and proximity to transcontinental migratory routes, Iran has played a key role in subsequent migrations, both prehistoric and historic, between Africa, Asia and Europe. To shed light on the genetic structure of the Iranian population as well as on the expansion patterns and population movements which affected this region, the complete mitochondrial genomes of 352 Iranians were obtained. All Iranian populations studied here exhibit similarly high diversity values comparable to the other groups from the Caucasus, Anatolia and Europe. The results of AMOVA and MDS analyses did not associate any regional and/or linguistic group of populations in the Anatolia/Caucasus and Iran region pointing to close genetic positions of Persians and Qashqais to each other and to Armenians, and Azeris from Iran to Georgians. By reconstructing the complete mtDNA phylogeny of haplogroups R2, N3, U1, U3, U5a1g, U7, H13, HV2, HV12, M5a and C5c we have found a previously)
The processing of notes and chords which are harmonically incongruous with their context has been shown to elicit two distinct late ERP effects. These effects strongly resemble two effects associated with the processing of linguistic incongruities: a P600, resembling a typical response to syntactic incongruities in language, and an N500, evocative of the N400, which is typically elicited in response to semantic incongruities in language. Despite the robustness of these two patterns in the musical incongruity literature, no consensus has yet been reached as to the reasons for the existence of two distinct responses to harmonic incongruities. This study was the first to use behavioural and ERP data to test two possible explanations for the existence of these two patterns: the musicianship of listeners, and the resolved or unresolved nature of the harmonic incongruities. Results showed that harmonically incongruous notes and chords elicited a late positivity similar to the P600 when they)
Background: Acknowledgment of all serious limitations to research evidence is important for patient care and scientific progress. Formal research on how biomedical authors acknowledge limitations is scarce. Objectives: To assess the extent to which limitations are acknowledged in biomedical publications explicitly, and implicitly by investigating the use of phrases that express uncertainty, so-called hedges; to assess the association between industry support and the extent of hedging. Design: We analyzed reporting of limitations and use of hedges in 300 biomedical publications published in 30 high and medium -ranked journals in 2007. Hedges were assessed using linguistic software that assigned weights between 1 and 5 to each expression of uncertainty. Results: Twenty-seven percent of publications (81/300) did not mention any limitations, while 73% acknowledged a median of 3 (range 1–8) limitations. Five percent mentioned a limitation in the abstract. After controlling for confounders,)
This study examined whether the degree of complexity of a grammatical component in a language would impact on its representation in the brain through identifying the neural correlates of grammatical morpheme processing associated with nouns and verbs in Chinese. In particular, the processing of Chinese nominal classifiers and verbal aspect markers were investigated in a sentence completion task and a grammaticality judgment task to look for converging evidence. The Chinese language constitutes a special case because it has no inflectional morphology per se and a larger classifier than aspect marker inventory, contrary to the pattern of greater verbal than nominal paradigmatic complexity in most European languages. The functional imaging results showed BA47 and left supplementary motor area and superior medial frontal gyrus more strongly activated for classifier processing, and the left posterior middle temporal gyrus more responsive to aspect marker processing. We attributed the activ)
Ethnic Belarusians make up more than 80% of the nine and half million people inhabiting the Republic of Belarus. Belarusians together with Ukrainians and Russians represent the East Slavic linguistic group, largest both in numbers and territory, inhabiting East Europe alongside Baltic-, Finno-Permic- and Turkic-speaking people. Till date, only a limited number of low resolution genetic studies have been performed on this population. Therefore, with the phylogeographic analysis of 565 Y-chromosomes and 267 mitochondrial DNAs from six well covered geographic sub-regions of Belarus we strove to complement the existing genetic profile of eastern Europeans. Our results reveal that around 80% of the paternal Belarusian gene pool is composed of R1a, I2a and N1c Y-chromosome haplogroups – a profile which is very similar to the two other eastern European populations – Ukrainians and Russians. The maternal Belarusian gene pool encompasses a full range of West Eurasian haplogroups and agrees wel)
The aims of this study were (1) to document the recognition performance of environmental sounds (ESs) in Mandarin-speaking children with cochlear implants (CIs) and to analyze the possible associated factors with the ESs recognition; (2) to examine the relationship between perception of ESs and receptive vocabulary level; and (3) to explore the acoustic factors relevant to perceptual outcomes of daily ESs in pediatric CI users. Forty-seven prelingually deafened children between ages 4 to 10 years participated in this study. They were divided into pre-school (group A: age 4–6) and school-age (group B: age 7 to 10) groups. Sound Effects Recognition Test (SERT) and the Chinese version of the revised Peabody Picture Vocabulary Test (PPVT-R) were used to assess the auditory perception ability. The average correct percentage of SERT was 61.2% in the preschool group and 72.3% in the older group. There was no significant difference between the two groups. The ESs recognition performance of ch)
For DNA sequences of various species we construct the Google matrix of Markov transitions between nearby words composed of several letters. The statistical distribution of matrix elements of this matrix is shown to be described by a power law with the exponent being close to those of outgoing links in such scale-free networks as the World Wide Web (WWW). At the same time the sum of ingoing matrix elements is characterized by the exponent being significantly larger than those typical for WWW networks. This results in a slow algebraic decay of the PageRank probability determined by the distribution of ingoing elements. The spectrum of is characterized by a large gap leading to a rapid relaxation process on the DNA sequence networks. We introduce the PageRank proximity correlator between different species which determines their statistical similarity from the view point of Markov chains. The properties of other eigenstates of the Google matrix are also discussed. Our results establish sc)
When parents select similar sounding names for their children, do they set themselves up for more speech errors in the future? Questionnaire data from 334 respondents suggest that they do. Respondents whose names shared initial or final sounds with a sibling’s reported that their parents accidentally called them by the sibling’s name more often than those without such name overlap. Having a sibling of the same gender, similar appearance, or similar age was also associated with more frequent name substitutions. Almost all other name substitutions by parents involved other family members and over 5% of respondents reported a parent substituting the name of a pet, which suggests a strong role for social and situational cues in retrieving personal names for direct address. To the extent that retrieval cues are shared with other people or animals, other names become available and may substitute for the intended name, particularly when names sound similar. [ABSTRACT FROM AUTHOR], Copyright )
While there has been a fair amount of research investigating children's syntactic processing during spoken language comprehension, and a wealth of research examining adults' syntactic processing during reading, as yet very little research has focused on syntactic processing during text reading in children. In two experiments, children and adults read sentences containing a temporary syntactic ambiguity while their eye movements were monitored. In Experiment 1, participants read sentences such as, 'The boy poked the elephant with the long stick/trunk from outside the cage' in which the attachment of a prepositional phrase was manipulated. In Experiment 2, participants read sentences such as, 'I think I'll wear the new skirt I bought tomorrow/yesterday. It's really nice' in which the attachment of an adverbial phrase was manipulated. Results showed that adults and children exhibited similar processing preferences, but that children were delayed relative to adults in their detection of i)
Background: Within the structural and grammatical bounds of a common language, all authors develop their own distinctive writing styles. Whether the relative occurrence of common words can be measured to produce accurate models of authorship is of particular interest. This work introduces a new score that helps to highlight such variations in word occurrence, and is applied to produce models of authorship of a large group of plays from the Shakespearean era. Methodology: A text corpus containing 55,055 unique words was generated from 168 plays from the Shakespearean era (16th and 17th centuries) of undisputed authorship. A new score, CM1, is introduced to measure variation patterns based on the frequency of occurrence of each word for the authors John Fletcher, Ben Jonson, Thomas Middleton and William Shakespeare, compared to the rest of the authors in the study (which provides a reference of relative word usage at that time). A total of 50 WEKA methods were applied for Fletcher, Jons)
Sexual selection has resulted in sex-based size dimorphism in many mammals, including humans. In Western societies, average to taller stature men and comparatively shorter, slimmer women have higher reproductive success and are typically considered more attractive. This size dimorphism also extends to vocalisations in many species, again including humans, with larger individuals exhibiting lower formant frequencies than smaller individuals. Further, across many languages there are associations between phonemes and the expression of size (e.g. large /a, o/, small /i, e/), consistent with the frequency-size relationship in vocalisations. We suggest that naming preferences are a product of this frequency-size relationship, driving male names to sound larger and female names smaller, through sound symbolism. In a 10-year dataset of the most popular British, Australian and American names we show that male names are significantly more likely to contain larger sounding phonemes (e.g. “Thomas)
This study examines the intergenerational transfer of human communication systems. It tests if human communication systems evolve to be easy to learn or easy to use (or both), and how population size affects learnability and usability. Using an experimental-semiotic task, we find that human communication systems evolve to be easier to use (production efficiency and reproduction fidelity), but harder to learn (identification accuracy) for a second generation of naïve participants. Thus, usability trumps learnability. In addition, the communication systems that evolve in larger populations exhibit distinct advantages over those that evolve in smaller populations: the learnability loss (from the Initial signs) is more muted and the usability benefits are more pronounced. The usability benefits for human communication systems that evolve in a small and large population is explained through guided variation reducing sign complexity. The enhanced performance of the communication systems tha)
Text tokenization is a fundamental pre-processing step for almost all the information processing applications. This task is nontrivial for the scarce resourced languages such as Urdu, as there is inconsistent use of space between words. In this paper a morpheme matching based approach has been proposed for Urdu text tokenization, along with some other algorithms to solve the additional issues of boundary detection of compound words, affixation, reduplication, names and abbreviations. This study resulted into 97.28% precision, 93.71% recall, and 95.46% F1-measure; while tokenizing a corpus of 57000 words by using a morpheme list with 6400 entries. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles for individual use. This abstract may be abridged. No)
In the philosophical theory of communicative action, rationality refers to interpersonal communication rather than to a knowing subject. Thus, a social view of rationality is suggested. The theory differentiates between two kinds of rationality, the emancipative communicative and the strategic or instrumental reasoning. Using experimental designs in an fMRI setting, recent studies explored similar questions of reasoning in the social world and linked them with a neural network including prefrontal and parietal brain regions. Here, we employed an fMRI approach to highlight brain areas associated with strategic and communicative reasoning according to the theory of communicative action. Participants were asked to assess different social scenarios with respect to communicative or strategic rationality. We found a network of brain areas including temporal pole, precuneus, and STS more activated when participants performed communicative reasoning compared with strategic thinking and a cont)
Activity in the ventral striatum has frequently been associated with retrieval success, i.e., it is higher for hits than correct rejections. Based on the prominent role of the ventral striatum in the reward circuit, its activity has been interpreted to reflect the higher subjective value of hits compared to correct rejections in standard recognition tests. This hypothesis was supported by a recent study showing that ventral striatal activity is higher for correct rejections than hits when the value of rejections is increased by external incentives. These findings imply that the striatal response during recognition is context-sensitive and modulated by the adaptive significance of “oldness” or “newness” to the current goals. The present study is based on the idea that not only external incentives, but also other deviations from standard recognition tests which affect the subjective value of specific response types should modulate striatal activity. Therefore, we explored ventral striat)
This paper describes the annotation process and linguistic properties of the Persian syntactic dependency treebank. The treebank consists of approximately 30,000 sentences annotated with syntactic roles in addition to morpho-syntactic features. One of the unique features of this treebank is that there are almost 4800 distinct verb lemmas in its sentences making it a valuable resource for educational goals. The treebank is constructed with a bootstrapping approach by means of available tagging and parsing tools and manually correcting the annotations. The data is splitted into standard train, development and test set in the CoNLL dependency format and is freely available to researchers. 1
We report data from an internet questionnaire of sixty number trivia. Participants were asked for the number of cups in their house, the number of cities they know and 58 other quantities. We compare the answers of familial sinistrals – individuals who are left-handed themselves or have a left-handed close blood-relative – with those of pure familial dextrals – right-handed individuals who reported only having right-handed close blood-relatives. We show that familial sinistrals use rounder numbers than pure familial dextrals in the survey responses. Round numbers in the decimal system are those that are multiples of powers of 10 or of half or a quarter of a power of 10. Roundness is a gradient concept, e.g. 100 is rounder than 50 or 200. We show that very round number like 100 and 1000 are used with 25% greater likelihood by familial sinistrals than by pure familial dextrals, while pure familial dextrals are more likely to use less round numbers such as 25, 60, and 200. We then use Si)
We evaluated the influence of speed–accuracy trade-offs on performance in the sustained attention to response task (SART), a task often used to evaluate the effectiveness of techniques designed to improve sustained attention. In the present study, we experimentally manipulated response delay in a variation of the SART and found that commission errors, which are commonly used as an index of lapses in sustained attention, were a systematic function of manipulated differences in response delay. Delaying responses to roughly 800 ms after stimulus onset reduced commission errors substantially. We suggest the possibility that any technique that affects response speed will indirectly alter error rates independently of improvements in sustained attention. Investigators therefore need to carefully explore, report, and correct for changes in response speed that accompany improvements in performance or, alternatively, to employ tasks that control for response speed.
‘Lexical bundles’ as a category of word combinations are words which follow each other more frequently than expected by chance. This corpus-based study attempts to compare the frequencies of three- and four-word lexical bundles in research articles of three disciplines: physics, computer engineering, and applied linguistics. Moreover, it aims to scrutinize them between native and nonnative research articles of applied linguistics to see whether Iranian authors who publish articles in English, use lexical bundles in the same way as native authors. To this end, three native corpora and a non-native corpus of research articles were collected, each including approximately one million words. All the analyses were conducted through Wordsmith Tools (Scott, 2010) and Hyland’s (2008) taxonomy of most frequent academic lexical bundles. The results show that there are relatively significant differences between the frequencies of the lexical bundles employed across the disciplines. In addition, they differ significantly between the native and nonnative articles of applied linguistics. It is also revealed that lexical bundles are realized differently across different disciplines and that non-natives do not follow the norms of natives appropriately. Findings can be used to improve writing in different disciplines and create more cohesive and coherent texts.
Wordnets are built of synsets, not of words. A synset consists of words. Synonymy is a relation between words. Words go into a synset because they are synonyms. Later, a wordnet treats words as synonymous because they belong in the same synset\(\ldots\) Such circularity, a well-known problem, poses a practical difficulty in wordnet construction, notably when it comes to maintaining consistency. We propose to make a wordnet a net of words or, to be more precise, lexical units. We discuss our assumptions and present their implementation in a steadily growing Polish wordnet. A small set of constitutive relations allows us to construct synsets automatically out of groups of lexical units with the same connectivity. Our analysis includes a thorough comparative overview of systems of relations in several influential wordnets. The additional synset-forming mechanisms include stylistic registers and verb aspect.
Previous studies examining binocular coordination during reading have reported conflicting results in terms of the nature of disparity (e.g. Kliegl, Nuthmann, & Engbert (Journal of Experimental Psychology General 135:12-35, 2006); Liversedge, White, Findlay, & Rayner (Vision Research 46:2363-2374, 2006). One potential cause of this inconsistency is differences in acquisition devices and associated analysis technologies. We tested this by directly comparing binocular eye movement recordings made using SR Research EyeLink 1000 and the Fourward Technologies Inc. DPI binocular eye-tracking systems. Participants read sentences or scanned horizontal rows of dot strings; for each participant, half the data were recorded with the EyeLink, and the other half with the DPIs. The viewing conditions in both testing laboratories were set to be very similar. Monocular calibrations were used. The majority of fixations recorded using either system were aligned, although data from the EyeLink system showed greater disparity magnitudes. Critically, for unaligned fixations, the data from both systems showed a majority of uncrossed fixations. These results suggest that variability in previous reports of binocular fixation alignment is attributable to the specific viewing conditions associated with a particular experiment (variables such as luminance and viewing distance), rather than acquisition and analysis software and hardware.
Anticipating Computer Language — On Some Conventions in the Burmese Inscriptions Rudolf A. Yanson (bio) The problem of space was always a pressing one for Old Burmese scribes. For inscriptions, they used stone slabs or palm leaves. Stone slabs were difficult to carve while palm leaves were an expensive recording technology because turning palm leaves into materials to write upon required toilsome preparation. So it was natural that the scribes when writing tried to save space and therefore introduced different conventions into orthography. In the course of time, after orthography became more or less stabilized, and the practice of writing inscriptions became more widespread, scribes started to experiment with orthography, introducing into inscriptions their own judgment regarding spellings of some common words/markers. Some conventions these scribes implemented became the norm of present orthography, but some remain the peculiarity of Old Burmese. Those used in Modern Burmese (MB) unfortunately are never found in dictionaries. Although the origin of some of them is explained in different publications, in several cases I find the provided explanations unconvincing.1 Besides, the relevant publications contain just lists of abbreviations met in palm leaf manuscripts and parabaiks covering the period beginning in the 18th and subsequent centuries. [End Page 391] My analysis and examples are based on the analysis of inscriptions published in five volumes in Myanmar.2 In each case, after presenting an example, I will put in the number of the volume, page, and the line from above where the example appears. I shall start by describing the use of numerals in the context of some grammaticals and lexicals, as well as some peculiar cases. The first numeral to be described in an unusual function is 2. Besides its direct function, it was widely used for a variety of purposes, such as reduplication of verbs to intensify their meaning or to express the action as multiple one, e.g.,3 "who (whenever) comes to the monastery" (2,17,2). The verb la, "to come," is followed by the numeral 2. The phrase holds good without the numeral, but the meaning would change and sound like "whoever may come to the monastery." In the following example, the numeral is put after the word "very" to intensify its meaning: "very-very" (2,35,7). The numeral 2 was also used after nouns for emphatic plurality, e.g., "places where there is pure water" (1,289,12). The numeral is put after "place," and the context meaning of the phrase is "wherever one gets to, there will be pure water." Without the numeral, the meaning would be "a place where there is pure water." One more example: "worlds." Formal analysis leads to the meaning "the two worlds," but the contextual meaning is "whatever future worlds may be" (2,21,12). Usually this meaning is expressed by repeating the word "world," but quite often it is expressed with the numeral postponed to the word. [End Page 392] The described functions of the numeral 2 can be traced already in the earliest inscriptions, suggesting that this pattern dates to the early stages of the evolution of writing practices. Quite interesting is the following example: "cows and water buffaloes sacrificed in 3 groups" (3,35,10). Here the numeral is used instead of the word "to sacrifice." Words for "two" and "sacrifice" are spelled the same, so why not use the simple numeral in the place of though not so long but yet complicated for performing on stone word? Some more examples of the unusual use of the numeral 2 are as follows: "may (he) submerge in Hell" (4,113,9–10). The words "two" and "to submerge" are spelled the same. In the example, the word "to submerge" is spelled with numeral 2 and the last symbol of the words "two" and "to submerge", i.e., ca with virama. It appears that the numeral is used instead of symbol hna, which actually does not save space, but inscribing this last symbol is much more complicated than that of numeral. The next two examples also present the mixture in spelling of the word for "two" and the corresponding numeral: "two cows," "two pahsos (dress)" (4,113,4). The word "two...
Modern communication environments have changed the cognitive patterns of individuals, who are now used to the interaction of information encoded in different semiotic modalities, especially visual and linguistic. Despite this, the main premise of Corpus Linguistics is still ruling: our perception of and experience with the world is conveyed in texts, which nowadays need to be studied from a multimodal perspective. Therefore, multimodal corpora are becoming extremely useful to extract specialized knowledge and explore the insights of specialized language and its relation to non-language-specific representations of knowledge. It is our assertion that the analysis of the image-text interface can help us understand the way visual and linguistic information converge in subject-field texts. In this article, we use Frame-based terminology to sketch a novel proposal to study images in a corpus rich in pictorial representations for their inclusion in a terminological resource on the environment. Our corpus-based approach provides the methodological underpinnings to create meaning within terminographic entries, thus facilitating specialized knowledge transfer and acquisition through images.
Metaphor makes our thoughts more vivid and fills our communication with richer imagery. Furthermore, according to the conceptual metaphor theory (CMT) of Lakoff and Johnson (Metaphors we live by. University of Chicago Press, Chicago, 1980), metaphor also plays an important structural role in the organization and processing of conceptual knowledge. According to this account, the phenomenon of metaphor is not restricted to similarity-based extensions of meanings of individual words, but instead involves activating fixed mappings that reconceptualize one whole area of experience in terms of another. CMT produced a significant resonance in the fields of philosophy, linguistics, cognitive science and artificial intelligence and still underlies a large proportion of modern research on metaphor. However, there has to date been no comprehensive corpus-based study of conceptual metaphor, which would provide an empirical basis for evaluating the CMT using real-world linguistic data. The annotation scheme and the empirical study we present in this paper is a step towards filling this gap. We test our annotation procedure in an experimental setting involving multiple annotators and estimate their agreement on the task. The goal of the study is to investigate (1) how intuitive the conceptual metaphor explanation of linguistic metaphors is for human annotators and whether it is possible to consistently annotate interconceptual mappings; (2) what are the main difficulties that the annotators experience during the annotation process; (3) whether one conceptual metaphor is sufficient to explain a linguistic metaphor or whether a chain of conceptual metaphors is needed. The resulting corpus annotated for conceptual mappings provides a new, valuable dataset for linguistic, computational and cognitive experiments on metaphor.
Gaze-contingent displays provide a valuable method in visual research for controlling visual input and investigating its visual and cognitive processing. Although the body of research using gaze-contingent retinal stabilization techniques has grown considerably during the last decade, only few studies have been concerned with the reliability of the specific real-time simulations applied. Using a Landolt ring discrimination task, we present a behavioral validation of gaze-contingent central scotoma simulation in healthy observers. Importantly, behavioral testing is necessary to show whether the simulation impairs foveal processing of visual information. This test becomes even more crucial when researchers are faced with null results in a task performed with the scotoma, as compared with a control condition. It must be ruled out that the lack of behavioral effects results from a type II error caused by improper implementation before conclusions about foveal contributions to the given task may be drawn. In our experiment, the scotoma effectively prevented foveal processing of the visual stimuli, leading to significantly reduced response accuracies, as compared with unimpaired vision. Moreover, the final fixation at the time of the participants’ responses was placed close to the target position in the unimpaired condition, whereas the distance to the target was enhanced with the scotoma, indicating that the observers were not able to discriminate visual target stimuli from distractors, due to the scotoma. The present work presents a validated behavioral testing method for the efficiency of gaze-contingent scotoma simulations, including code for implementation. In addition, solutions for common methodological problems are discussed.
In recent years, sentiment analysis (SA) has emerged as a rapidly expanding field of application and research in the area of information retrieval. In order to facilitate the task of selecting lexical resources for automated SA systems, this paper sets out a detailed analysis of four widely used sentiment lexica. The analysis provides an overview of the coverage of each lexicon individually, the overlap and consistency of the four resources and a corpus analysis of the distribution of the resources’ lexical contents in general and specialised language. This work aims to explore the characteristics of affective language as represented by these lexica and the implications of the findings for developers of SA systems.
The experience of a user of major search engines or other web information retrieval services looking for information in the Basque language is far from satisfactory: they only return pages with exact matches but no inflections (necessary for an agglutinative language like Basque), many results in other languages (no search engine gives the option to restrict its results to Basque), etc. This paper proposes using morphological query expansion and language-filtering words in combination with the APIs of search engines as a very cost-effective solution to build appropriate web search services for Basque. The implementation details of the methodology (choosing the most appropriate language-filtering words, the number of them, the most frequent inflections for the morphological query expansion, etc.) have been specified by corpora-based studies. The improvements produced have been measured in terms of precision and recall both over corpora and real web searches. Morphological query expansion can improve recall up to 47 % and language-filtering words can raise precision from 15 % to around 90 %, although with a loss in recall of about 30–35 %. The proposed methodology has already been successfully used in the Basque search service Elebila (http://www.elebila.eu) and the web-as-corpus tool CorpEus (http://www.corpeus.org), and the approach could be applied to other morphologically rich or under-resourced languages as well.
The CLARIN Metadata Infrastructure (CMDI) that is being developed in Common Language Resources and Technology Infrastructure (CLARIN) is a computer-supported framework that combines a flexible component approach with the explicit declaration of semantics. The goal of the Dutch CLARIN project “Creating & Testing CLARIN Metadata Components” was to create metadata components and profiles for a wide variety of existing resources housed at two data centres according to the CMDI specifications. In doing so the principles of the framework were tested. The results of the project are of benefit to other CLARIN-projects that are expected to adhere to the CMDI framework and its accompanying tools.
This paper explores the feasibility of modelling concept concreteness perceived by humans and representing it in computational semantic lexicons, addressing an issue at the crossroads of computational linguistics, lexicography, and psycholinguistics. The inherent distinction between concrete words and abstract words in psychology has relied mostly on subjective human ratings. This practice is hardly scalable and does not consider the effect of polysemy. In view of this, we attempt to obtain a measure of concreteness from dictionary definitions comparable to human judgement, capitalising on conventional lexicographic assumptions and the regularities exhibited in the surface structures of sense definitions. The structural pattern of a definition is analysed and scored on a 7-point scale of concreteness ratings. The definition scores turned out to be quite effective for a dichotomous distinction between concrete and abstract concepts and more consistent with human ratings for the former. Beyond the two-way distinction, however, the results were more variable. The study has thus revealed the potentials and limitations of our approach, suggesting that different defining styles probably reflect the describability of concepts, and describability alone may not be sufficient for differentiating the degree of concreteness. The range of definition patterns has to be reconsidered, in combination with other inseparable factors constituting our perception of concreteness, for better modelling on a finer scale of concreteness distinction to enrich semantic lexicons for natural language processing.
In the current event-related potential (ERP) study, we investigated how speech rhythm impacts speech segmentation and facilitates the resolution of syntactic ambiguities in auditory sentence processing. Participants listened to syntactically ambiguous German subject- and object-first sentences that were spoken with either regular or irregular speech rhythm. Rhythmicity was established by a constant metric pattern of three unstressed syllables between two stressed ones that created rhythmic groups of constant size. Accuracy rates in a comprehension task revealed that participants understood rhythmically regular sentences better than rhythmically irregular ones. Furthermore, the mean amplitude of the P600 component was reduced in response to object-first sentences only when embedded in rhythmically regular but not rhythmically irregular context. This P600 reduction indicates facilitated processing of sentence structure possibly due to a decrease in processing costs for the less-preferre)
Workload capacity, an important concept in many areas of psychology, describes processing efficiency across changes in workload. The capacity coefficient is a function across time that provides a useful measure of this construct. Until now, most analyses of the capacity coefficient have focused on the magnitude of this function, and often only in terms of a qualitative comparison (greater than or less than one). This work explains how a functional extension of principal components analysis can capture the time-extended information of these functional data, using a small number of scalar values chosen to emphasize the variance between participants and conditions. This approach provides many possibilities for a more fine-grained study of differences in workload capacity across tasks and individuals.
In this paper we provide an account of the cross-lingual lexical substitution task run as part of SemEval-2010. In this task both annotators (native Spanish speakers, proficient in English) and participating systems had to find Spanish translations for target words in the context of an English sentence. Because only translations of a single lexical unit were required, this task does not necessitate a full blown translation system. This we hope encouraged those working specifically on lexical semantics to participate without a requirement for them to use machine translation software, though they were free to use whatever resources they chose. In this paper we pay particular attention to the resources used by the various participating systems and present analyses to demonstrate the relative strengths of the systems as well as the requirements they have in terms of resources. In addition to the analyses of individual systems we also present the results of a combined system based on voting from the individual systems. We demonstrate that the system produces better results at finding the most frequent translation from the annotators compared to the highest ranked translation provided by individual systems. This supports our other analyses that the systems are heterogeneous, with different strengths and weaknesses.