Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Latencies of buttonpresses are a staple of cognitive science paradigms. Often keyboards are employed to collect buttonpresses, but their imprecision and variability decreases test power and increases the risk of false positives. Response boxes and data acquisition cards are precise, but expensive and inflexible, alternatives. We propose using open-source Arduino microcontroller boards as an inexpensive and flexible alternative. These boards connect to standard experimental software using a USB connection and a virtual serial port, or by emulating a keyboard. In our solution, an Arduino measures response latencies after being signaled the start of a trial, and communicates the latency and response back to the PC over a USB connection. We demonstrated the reliability, robustness, and precision of this communication in six studies. Test measures confirmed that the error added to the measurement had an SD of less than 1 ms. Alternatively, emulation of a keyboard results in similarly precise measurement. The Arduino performs as well as a serial response box, and better than a keyboard. In addition, our setup allows for the flexible integration of other sensors, and even actuators, to extend the cognitive science toolbox.
Researchers have long sought to distinguish between single-process and dual-process cognitive phenomena, using responses such as reaction times and, more recently, hand movements. Analysis of a response distribution’s modality has been crucial in detecting the presence of dual processes, because they tend to introduce bimodal features. Rarely, however, have bimodality measures been systematically evaluated. We carried out tests of readily available bimodality measures that any researcher may easily employ: the bimodality coefficient (BC), Hartigan’s dip statistic (HDS), and the difference in Akaike’s information criterion between one-component and two-component distribution models (AICdiff). We simulated distributions containing two response populations and examined the influences of (1) the distances between populations, (2) proportions of responses, (3) the amount of positive skew present, and (4) sample size. Distance always had a stronger effect than did proportion, and the effects of proportion greatly differed across the measures. Skew biased the measures by increasing bimodality detection, in some cases leading to anomalous interactive effects. BC and HDS were generally convergent, but a number of important discrepancies were found. AICdiff was extremely sensitive to bimodality and identified nearly all distributions as bimodal. However, all measures served to detect the presence of bimodality in comparison to unimodal simulations. We provide a validation with experimental data, discuss methodological and theoretical implications, and make recommendations regarding the choice of analysis.
This chapter examines the theory of norms and exploitations in relation to word meaning, anthropology, and the philosophy of language, looking in particular at the work of Aristotle, Ludwig Wittgenstein, Hilary Putnam, and H. P. Grice, as well as that of Bronisław Malinowski, Eleanor Rosch, and Michael Tomasello. It also discusses the lexicon and theories of language, lexical semantics, the attempt by thinkers such as John Wilkins and Gottfried Wilhelm Leibniz to make language precise during the Age of Enlightenment in Europe, and semantic primitives in preference semantics.
In this article, we validate an experimental paradigm, SPaM, that we first described elsewhere (Luke & Christianson, Memory & Cognition 40:628–641, 2012). SPaM is a synthesis of self-paced reading and masked priming. The primary purpose of SPaM is to permit the study of sentence context effects on early word recognition. In the experiment reported here, we show that SPaM successfully reproduces results from both the self-paced reading and masked-priming literatures. We also outline the advantages and potential uses of this paradigm. For users of E-Prime, the experimental program can be downloaded from our lab website, http://epl.beckman.illinois.edu/.
Serial cognitive assessment is conducted to monitor changes in the cognitive abilities of patients over time. At present, mainly the regression-based change and the ANCOVA approaches are used to establish normative data for serial cognitive assessment. These methods are straightforward, but they have some severe drawbacks. For example, they can only consider the data of two measurement occasions. In this article, we propose three alternative normative methods that are not hampered by these problems—that is, multivariate regression, the standard linear mixed model (LMM), and the linear mixed model combined with multiple imputation (LMM with MI) approaches. The multivariate regression method is primarily useful when a small number of repeated measurements are taken at fixed time points. When the data are more unbalanced, the standard LMM and the LMM with MI methods are more appropriate because they allow for a more adequate modeling of the covariance structure. The standard LMM has the advantage that it is easier to conduct and that it does not require a Monte Carlo component. The LMM with MI, on the other hand, has the advantage that it can flexibly deal with missing responses and missing covariate values at the same time. The different normative methods are illustrated on the basis of the data of a large longitudinal study in which a cognitive test (the Stroop Color Word Test) was administered at four measurement occasions (i.e., at baseline and 3, 6, and 12 years later). The results are discussed and suggestions for future research are provided.
In the field of constituency parsing, there exist multiple human-labeled treebanks which are built on non-overlapping text samples and follow different annotation standards. Due to the extreme cost of annotating parse trees by human, it is desirable to automatically convert one treebank (called source treebank) to the standard of another treebank (called target treebank) which we are interested in. Conversion results can be manually corrected to obtain higher-quality annotations or can be directly used as additional training data for building syntactic parsers. To perform automatic treebank conversion, we divide constituency parses into two separate levels: the part-of-speech (POS) and syntactic structure (bracketing structures and constituent labels), and conduct conversion on these two levels respectively with a feature-based approach. The basic idea of the approach is to encode original annotations in a source treebank as guide features during the conversion process. Experiments on two Chinese treebanks show that our approach can convert POS tags and syntactic structures with the accuracy of 96.6 and 84.8 %, respectively, which are the best reported results on this task.
In this article, we examine the effectiveness of bootstrapping supervised machine-learning polarity classifiers with the help of a domain-independent rule-based classifier that relies on a lexical resource, i.e., a polarity lexicon and a set of linguistic rules. The benefit of this method is that though no labeled training data are required, it allows a classifier to capture in-domain knowledge by training a supervised classifier with in-domain features, such as bag of words, on instances labeled by a rule-based classifier. Thus, this approach can be considered as a simple and effective method for domain adaptation. Among the list of components of this approach, we investigate how important the quality of the rule-based classifier is and what features are useful for the supervised classifier. In particular, the former addresses the issue in how far linguistic modeling is relevant for this task. We not only examine how this method performs under more difficult settings in which classes are not balanced and mixed reviews are included in the data set but also compare how this linguistically-driven method relates to state-of-the-art statistical domain adaptation.
Short texts are typically composed of small number of words, most of which are abbreviations, typos and other kinds of noise. This makes the noise to signal ratio relatively high for this specific category of text. A high proportion of noise in the data is undesirable for analysis procedures as well as machine learning applications. Text normalization techniques are used to reduce the noise and improve the quality of text for processing and analysis purposes. In this work, we propose a combination of statistical and rule-based techniques to normalize short texts. More specifically, we focus our attention on SMS messages. We base our normalization approach on a statistical machine translation system which translates from noisy data to clean data. This system is trained on a small manually annotated set. Then, we study several automatic methods to extract more general rules from the normalizations generated with the statistical machine translation system. We illustrate the proposed methodology by conducting some experiments with a SMS Haitian-Créole data collection. In order to evaluate the performance of our methodology we use several Haitian-Créole dictionaries, the well-known perplexity criteria and the achieved reduction of vocabulary.
Between 1988 and 2010, the renowned British physicist Stephen Hawking wrote five popular science books aimed at bringing physics closer to a wider audience than the mere academia.The operation proved very successful -with his best-seller alone (A Brief History of Time, 1988) reported to have sold over 10 million copies 1 (Paris 2007) -and made him into an acclaimed popular author.This study considers the books Hawking wrote especially for popularizing purposes, presenting reflections on the relationship between specialized and popular discourse.It focuses in particular on Hawking's first such work, A Brief History of Time, which was made into an even more popular adaptation titled A Briefer History of Time (2005).The chapter details how the subject has been adapted and transferred from a high into a popular (writing) and an even more popular (re-writing) level.This is done by comparing the works against the general features of specialized/scientific discourse, to single out their variation from -or conformity to -the established norms thereof, providing samples of textual analysis and highlighting relevant lexical and syntactic phenomena.An interpretation of such phenomena is proposed according to Critical Discourse Analysis methodology, i.e. considering language in light of the many social, cultural and economic variables informing this type of communication.
OBJECTIVE: Lexical fluency tests are frequently used to assess language and executive function in clinical practice. We investigated the influences of age, gender, and education on lexical verbal fluency in an educationally-diverse, elderly Korean population and provided its' normative information. METHODS: We administered the lexical verbal fluency test (LVFT) to 1676 community-dwelling, cognitively normal subjects aged 60 years or over. RESULTS: In a stepwise linear regression analysis, education (B=0.40, SE=0.02, standardized B=0.506) and age (B=-0.10, SE=0.01, standardized B=-0.15) had significant effects on LVFT scores (p<0.001), but gender did not (B=0.40, SE=0.02, standardized B=0.506, p>0.05). Education explained 28.5% of the total variance in LVFT scores, which was much larger than the variance explained by age (5.42%). Accordingly, we presented normative data of the LVFT stratified by age (60-69, 70-74, 75-79, and ≥80 years) and education (0-3, 4-6, 7-9, 10-12, and ≥13 years). CONCLUSION: The LVFT norms should provide clinically useful data for evaluating elderly people and help improve the interpretation of verbal fluency tasks and allow for greater diagnostic accuracy.
This paper describes the generation of temporally anchored infobox attribute data from the Wikipedia history of revisions. By mining (attribute, value) pairs from the revision history of the English Wikipedia we are able to collect a comprehensive knowledge base that contains data on how attributes change over time. When dealing with the Wikipedia edit history, vandalic and erroneous edits are a concern for data quality. We present a study of vandalism identification in Wikipedia edits that uses only features from the infoboxes, and show that we can obtain, on this dataset, an accuracy comparable to a state-of-the-art vandalism identification method that is based on the whole article. Finally, we discuss different characteristics of the extracted dataset, which we make available for further study.
Multilingual posts can potentially affect the outcomes of content analysis on microblog platforms. To this end, language identification can provide a monolingual set of content for analysis. We find the unedited and idiomatic language of microblogs to be challenging for state-of-the-art language identification methods. To account for this, we identify five microblog characteristics that can help in language identification: the language profile of the blogger (blogger), the content of an attached hyperlink (link), the language profile of other users mentioned (mention) in the post, the language profile of a tag (tag), and the language of the original post (conversation), if the post we examine is a reply. Further, we present methods that combine these priors in a post-dependent and post-independent way. We present test results on 1,000 posts from five languages (Dutch, English, French, German, and Spanish), which show that our priors improve accuracy by 5 % over a domain specific baseline, and show that post-dependent combination of the priors achieves the best performance. When suitable training data does not exist, our methods still outperform a domain unspecific baseline. We conclude with an examination of the language distribution of a million tweets, along with temporal analysis, the usage of twitter features across languages, and a correlation study between classifications made and geo-location and language metadata fields.
This paper presents an empirical evaluation of coreference resolution that covers several interrelated dimensions. The main goal is to complete the comparative analysis from the SemEval-2010 task on Coreference Resolution in Multiple Languages. To do so, the study restricts the number of languages and systems involved, but extends and deepens the analysis of the system outputs, including a more qualitative discussion. The paper compares three automatic coreference resolution systems for three languages (English, Catalan and Spanish) in four evaluation settings, and using four evaluation measures. Given that our main goal is not to provide a comparison between resolution algorithms, these are merely used as tools to shed light on the different conditions under which coreference resolution is evaluated. Although the dimensions are strongly interdependent, making it very difficult to extract general principles, the study reveals a series of interesting issues in relation to coreference resolution: the portability of systems across languages, the influence of the type and quality of input annotations, and the behavior of the scoring measures.
This paper presents GATE Teamware—an open-source, web-based, collaborative text annotation framework. It enables users to carry out complex corpus annotation projects, involving distributed annotator teams. Different user roles are provided (annotator, manager, administrator) with customisable user interface functionalities, in order to support the complex workflows and user interactions that occur in corpus annotation projects. Documents may be pre-processed automatically, so that human annotators can begin with text that has already been pre-annotated and thus making them more efficient. The user interface is simple to learn, aimed at non-experts, and runs in an ordinary web browser, without need of additional software installation. GATE Teamware has been evaluated through the creation of several gold standard corpora and internal projects, as well as through external evaluation in commercial and EU text annotation projects. It is available as on-demand service on GateCloud.net, as well as open-source for self-installation.
ABSTRACT The divide long believed to exist between outer and expanding circle Englishes has recently been called into question. Gilquin and Granger (2011) point out that the exposure to and use of English varies substantially within EFL countries, and Hilbert and Krug (2012) and Edwards (forthcoming) demonstrate that, within varieties, characteristics of EFL and ESL can coexist. As Hundt and Vogel (2011: 161) write, ‘increasing globalization might eventually blur the distinction between ENL, ESL and EFL varieties’. However, although studies comparing inner, outer and expanding circle countries are indeed emerging (e.g. Hundt & Vogel 2011, Wulff & Römer 2009), due to the nature of the corpus data available from the expanding circle (e.g. ICLE), such studies tend to focus on student writing only. Against this backdrop, we expand this scope of genres and take a first step in answering Davydova’s (2012) call for indigenised and learner varieties to be investigated on the same grounds, with the potential for variation of a similar nature depending on their variable extra-linguistic backgrounds. We seek to shed further light on the nature of the continuum across EFL, ESL and ENL. Our data come from the written components of the International Corpus of English (ICE) for Great Britain and the USA (ENL) and Hong Kong, India and Singapore (ESL), as well as from a comparable corpus of Dutch English (EFL). The latter, to our knowledge, is the first expanding circle corpus encompassing all ICE text categories, thus allowing for comparisons across a range of genres. Inspired by Gilquin and Granger’s (2011) work with ICLE, we take the preposition into as a case study, conducting a quantitative and qualitative analysis of its syntactic patterns, semantic distribution, lexical variation, phraseological uses and non-standard uses. We aim to test recent claims that the cline to be found in terms of norm orientation is ENL > EFL > ESL (Hundt & Vogel 2011, Van Rooy 2006) and, within ESL, that the more advanced varieties in Schneider’s (2003) dynamic model will be the most dissimilar to ENL (Mukherjee & Gries 2009). In light of these claims, we hypothesise that Dutch English will be closest to the native norm, and Singapore English most distant. Hierarchical cluster analyses in fact reveal Singapore English to be the most norm oriented, thus supporting Hundt and Vogel’s (2011) assertion that such ESL varieties can show lingering exonormative trends. Moreover, Dutch English is not markedly distinct from the New Englishes, but clusters with them in different ways depending on the focus of the analysis, e.g. like Indian English, it shows a relatively higher proportion of intransitive patterning with into than the other corpora. These results support Davydova’s (2012) claim that learner and New Englishes should be approached in an integrated fashion. In our view, they should be seen as existing on a continuum along which individuals and groups can move depending on their norm orientation as well as their levels of proficiency in and exposure to English.
Previous studies examining binocular coordination during reading have reported conflicting results in terms of the nature of disparity (e.g. Kliegl, Nuthmann, & Engbert (Journal of Experimental Psychology General 135:12-35, 2006); Liversedge, White, Findlay, & Rayner (Vision Research 46:2363-2374, 2006). One potential cause of this inconsistency is differences in acquisition devices and associated analysis technologies. We tested this by directly comparing binocular eye movement recordings made using SR Research EyeLink 1000 and the Fourward Technologies Inc. DPI binocular eye-tracking systems. Participants read sentences or scanned horizontal rows of dot strings; for each participant, half the data were recorded with the EyeLink, and the other half with the DPIs. The viewing conditions in both testing laboratories were set to be very similar. Monocular calibrations were used. The majority of fixations recorded using either system were aligned, although data from the EyeLink system showed greater disparity magnitudes. Critically, for unaligned fixations, the data from both systems showed a majority of uncrossed fixations. These results suggest that variability in previous reports of binocular fixation alignment is attributable to the specific viewing conditions associated with a particular experiment (variables such as luminance and viewing distance), rather than acquisition and analysis software and hardware.
With the involvement of Banks, Letter of Credit overcomes the limitation of time and space distance, provides financial undertaking for both exporter and importer, and balance the potential risk of both parties in terms of international payment.From the view of genre, L/C text belongs to specific category of English for Specific Purpose(hereinafter, ESP) research. Various typical features can be figured out corresponding to the levels of lexical, syntax and text. Based on the rich practice of international trade, this paper explores those features in light of abundant L/C material. From the perspective of lexical, terminologies are frequently employed and many a regular vocabulary is to be attached very professional meaning rather than general meaning as common ground usage. Legal language plays another important role in L/C text since it has much connection with legal action in sense of norm stipulation. Archaism also appears very often, making the solemn effect. From the perspective of syntax, parallel structure and passive voice are used at high frequency in L/C text. Ellipsis is common with respect to those sentence structure under L/C terms. In text level, L/C has typical features, such as staid configuration and sequence.In accordance with those features in the lexical, syntax and text levels, it's hope, in turn, to give suggestions concerning the translation of L/C against those different levels.
Semantic lexical resources are a mainstay of various Natural Language Processing applications. However, comprehensive and reliable resources are rare and not often freely available. Handcrafted resources are too costly for being a general solution while automatically-built resources need to be validated by experts or at least thoroughly evaluated. We propose in this paper a picture of the current situation with regard to lexical resources, their building and their evaluation. We give an in-depth description of Wiktionary, a freely available and collaboratively built multilingual dictionary. Wiktionary is presented here as a promising raw resource for NLP. We propose a semi-automatic approach based on random walks for enriching Wiktionary synonymy network that uses both endogenous and exogenous data. We take advantage of the wiki infrastructure to propose a validation “by crowds”. Finally, we present an implementation called WISIGOTH, which supports our approach.
We evaluated the influence of speed–accuracy trade-offs on performance in the sustained attention to response task (SART), a task often used to evaluate the effectiveness of techniques designed to improve sustained attention. In the present study, we experimentally manipulated response delay in a variation of the SART and found that commission errors, which are commonly used as an index of lapses in sustained attention, were a systematic function of manipulated differences in response delay. Delaying responses to roughly 800 ms after stimulus onset reduced commission errors substantially. We suggest the possibility that any technique that affects response speed will indirectly alter error rates independently of improvements in sustained attention. Investigators therefore need to carefully explore, report, and correct for changes in response speed that accompany improvements in performance or, alternatively, to employ tasks that control for response speed.
Studies of lexical–semantic relations aim to understand the mechanism of semantic memory and the organization of the mental lexicon. However, standard paradigmatic relations such as “hypernym” and “hyponym” cannot capture connections among concepts from different parts of speech. WordNet, which organizes synsets (i.e., synonym sets) using these lexical–semantic relations, is rather sparse in its connectivity. According to WordNet statistics, the average number of outgoing/incoming arcs for the hypernym/hyponym relation per synset is 1.33. Evocation, defined as how much a concept (expressed by one or more words) brings to mind another, is proposed as a new directed and weighted measure for the semantic relatedness among concepts. Commonly applied semantic relations and relatedness measures do not seem to be fully compatible with data that reflect evocations among concepts. They are compatible but evocation captures MORE. This work aims to provide a reliable and extendable dataset of concepts evoked by, and evoking, other concepts to enrich WordNet, the existing semantic network. We propose the use of disambiguated free word association data (first responses to verbal stimuli) to infer and collect evocation ratings. WordNet aims to represent the organization of mental lexicon, and free word association which has been used by psycholinguists to explore semantic organization can contribute to the understanding. This work was carried out in two phases. In the first phase, it was confirmed that existing free word association norms can be converted into evocation data computationally. In the second phase, a two-stage association-annotation procedure of collecting evocation data from human judgment was compared to the state-of-the-art method, showing that introducing free association can greatly improve the quality of the evocation data generated. Evocation can be incorporated into WordNet as directed links with scales, and benefits various natural language processing applications.
La traduction juridique est un processus complexe au cours duquel le traducteur doit prendre une série de décisions. En effet, il doit traduire les mots et le texte tout en laissant les normes inchangées. Sʼil sʼagit de traduire des normes juridiques, la traduction consiste à faire en sorte que le texte produise dans la langue cible les mêmes effets que dans la langue source. Ainsi, dans le présent article, nous tenterons de préciser en quoi la notion de norme intéresse la traduction juridique. Par ailleurs, la notion de norme implique inévitablement une théorie de lʼéquivalence. En effet, ce type de traduction exige des connaissances particulières dans le domaine juridique dans la langue source comme dans la langue cible. Aussi, partant dʼun corpus, nous nous demanderons si le traducteur linguiste observe les mêmes règles dʼéquivalence que le traducteur juriste. Rendent-ils à lʼidentique lʼidentité de sens quelles que soient les divergences de structures (grammaticales, stylistiques ou lexicales) qui sʼétablissent entre les textes sources et cibles? Partant du modèle de Toury concernant lʼactivité traduisante, nous lʼappliquerons aux traductions des textes juridiques, mettant ainsi en évidence les similitudes ou divergences entre traducteurs linguistes et traducteurs juristes.
Many authors adhere to the rule that test reliabilities should be at least .70 or .80 in group research. This article introduces a new standard according to which reliabilities can be evaluated. This standard is based on the costs or time of the experiment and of administering the test. For example, if test administration costs are 7 % of the total experimental costs, the efficient value of the reliability is .93. If the actual reliability of a test is equal to this efficient reliability, the test size maximizes the statistical power of the experiment, given the costs. As a standard in experimental research, it is proposed that the reliability of the dependent variable be close to the efficient reliability. Adhering to this standard will enhance the statistical power and reduce the costs of experiments.
Web 2.0 provides user-friendly tools that allow persons to create and publish content online. User generated content often takes the form of short texts (e.g., blog posts, news feeds, snippets, etc). This has motivated an increasing interest on the analysis of short texts and, specifically, on their categorisation. Text categorisation is the task of classifying documents into a certain number of predefined categories. Traditional text classification techniques are mainly based on word frequency statistical analysis and have been proved inadequate for the classification of short texts where word occurrence is too small. On the other hand, the classic approach to text categorization is based on a learning process that requires a large number of labeled training texts to achieve an accurate performance. However labeled documents might not be available, when unlabeled documents can be easily collected. This paper presents an approach to text categorisation which does not need a pre-classified set of training documents. The proposed method only requires the category names as user input. Each one of these categories is defined by means of an ontology of terms modelled by a set of what we call proximity equations. Hence, our method is not category occurrence frequency based, but highly depends on the definition of that category and how the text fits that definition. Therefore, the proposed approach is an appropriate method for short text classification where the frequency of occurrence of a category is very small or even zero. Another feature of our method is that the classification process is based on the ability of an extension of the standard Prolog language, named Bousi~Prolog , for flexible matching and knowledge representation. This declarative approach provides a text classifier which is quick and easy to build, and a classification process which is easy for the user to understand. The results of experiments showed that the proposed method achieved a reasonably useful performance.
This paper presents a detailed analysis of the use of crowdsourcing services for the Text Summarization task in the context of the tourist domain. In particular, our aim is to retrieve relevant information about a place or an object pictured in an image in order to provide a short summary which will be of great help for a tourist. For tackling this task, we proposed a broad set of experiments using crowdsourcing services that could be useful as a reference for others who want to rely also on crowdsourcing. From the analysis carried out through our experimental setup and the results obtained, we can conclude that although crowdsourcing services were not good to simply gather gold-standard summaries (i.e., from the results obtained for experiments 1, 2 and 4), the encouraging results obtained in the third and sixth experiments motivate us to strongly believe that they can be successfully employed for finding some patterns of behaviour humans have when generating summaries, and for validating and checking other tasks. Furthermore, this analysis serves as a guideline for the types of experiments that might or might not work when using crowdsourcing in the context of text summarization.
Wordnets are built of synsets, not of words. A synset consists of words. Synonymy is a relation between words. Words go into a synset because they are synonyms. Later, a wordnet treats words as synonymous because they belong in the same synset\(\ldots\) Such circularity, a well-known problem, poses a practical difficulty in wordnet construction, notably when it comes to maintaining consistency. We propose to make a wordnet a net of words or, to be more precise, lexical units. We discuss our assumptions and present their implementation in a steadily growing Polish wordnet. A small set of constitutive relations allows us to construct synsets automatically out of groups of lexical units with the same connectivity. Our analysis includes a thorough comparative overview of systems of relations in several influential wordnets. The additional synset-forming mechanisms include stylistic registers and verb aspect.
‘Lexical bundles’ as a category of word combinations are words which follow each other more frequently than expected by chance. This corpus-based study attempts to compare the frequencies of three- and four-word lexical bundles in research articles of three disciplines: physics, computer engineering, and applied linguistics. Moreover, it aims to scrutinize them between native and nonnative research articles of applied linguistics to see whether Iranian authors who publish articles in English, use lexical bundles in the same way as native authors. To this end, three native corpora and a non-native corpus of research articles were collected, each including approximately one million words. All the analyses were conducted through Wordsmith Tools (Scott, 2010) and Hyland’s (2008) taxonomy of most frequent academic lexical bundles. The results show that there are relatively significant differences between the frequencies of the lexical bundles employed across the disciplines. In addition, they differ significantly between the native and nonnative articles of applied linguistics. It is also revealed that lexical bundles are realized differently across different disciplines and that non-natives do not follow the norms of natives appropriately. Findings can be used to improve writing in different disciplines and create more cohesive and coherent texts.
Opinion mining on conversational telephone speech tackles two challenges: the robustness of speech transcriptions and the relevance of opinion models. The two challenges are critical in an industrial context such as marketing. The paper addresses jointly these two issues by analyzing the influence of speech transcription errors on the detection of opinions and business concepts. We present both modules: the speech transcription system, which consists in a successful adaptation of a conversational speech transcription system to call-centre data and the information extraction module, which is based on a semantic modeling of business concepts, opinions and sentiments with complex linguistic rules. Three models of opinions are implemented based on the discourse theory, the appraisal theory and the marketers’ expertise, respectively. The influence of speech recognition errors on the information extraction module is evaluated by comparing its outputs on manual versus automatic transcripts. The F-scores obtained are 0.79 for business concepts detection, 0.74 for opinion detection and 0.67 for the extraction of relations between opinions and their target. This result and the in-depth analysis of the errors show the feasibility of opinion detection based on complex rules on call-centre transcripts.
Appropriate evaluation of referring expressions is critical for the design of systems that can effectively collaborate with humans. A widely used method is to simply evaluate the degree to which an algorithm can reproduce the same expressions as those in previously collected corpora. Several researchers, however, have noted the need of a task-performance evaluation measuring the effectiveness of a referring expression in the achievement of a given task goal. This is particularly important in collaborative situated dialogues. Using referring expressions used by six pairs of Japanese speakers collaboratively solving Tangram puzzles, we conducted a task-performance evaluation of referring expressions with 36 human evaluators. Particularly we focused on the evaluation of demonstrative pronouns generated by a machine learning-based algorithm. Comparing the results of this task-performance evaluation with the results of a previously conducted corpus-matching evaluation (Spanger et al. in Lang Resour Eval, 2010b), we confirmed the limitation of a corpus-matching evaluation and discuss the need for a task-performance evaluation.
In describing motion events verbs of manner provide information about the speed of agents or objects in those events. We used eye tracking to investigate how inferences about this verb-associated speed of motion would influence the time course of attention to a visual scene that matched an event described in language. Eye movements were recorded as participants heard spoken sentences with verbs that implied a fast (“dash”) or slow (“dawdle”) movement of an agent towards a goal. These sentences were heard whilst participants concurrently looked at scenes depicting the agent and a path which led to the goal object. Our results indicate a mapping of events onto the visual scene consistent with participants mentally simulating the movement of the agent along the path towards the goal: when the verb implies a slow manner of motion, participants look more often and longer along the path to the goal; when the verb implies a fast manner of motion, participants tend to look earlier at the goal)
In a critical review of the heuristics used to deal with zero word frequencies, we show that four are suboptimal, one is good, and one may be acceptable. The four suboptimal strategies are discarding words with zero frequencies, giving words with zero frequencies a very low frequency, adding 1 to the frequency per million, and making use of the Good–Turing algorithm. The good algorithm is the Laplace transformation, which consists of adding 1 to each frequency count and increasing the total corpus size by the number of word types observed. A strategy that may be acceptable is to guess the frequency of absent words on the basis of other corpora and then increasing the total corpus size by the estimated summed frequency of the missing words. A comparison with the lexical decision times of the English Lexicon Project and the British Lexicon Project suggests that the Laplace transformation gives the most useful estimates (in addition to being easy to calculate). Therefore, we recommend it to researchers.
Researchers studying infants’ spontaneous allocation of attention have traditionally relied on hand-coding infants’ direction of gaze from videos; these techniques have low temporal and spatial resolution and are labor intensive. Eye-tracking technology potentially allows for much more precise measurement of how attention is allocated at the subsecond scale, but a number of technical and methodological issues have given rise to caution about the quality and reliability of high temporal resolution data obtained from infants. We present analyses suggesting that when standard dispersal-based fixation detection algorithms are used to parse eye-tracking data obtained from infants, the results appear to be heavily influenced by interindividual variations in data quality. We discuss the causes of these artifacts, including fragmentary fixations arising from flickery or unreliable contact with the eyetracker and variable degrees of imprecision in reported position of gaze. We also present new algorithms designed to cope with these problems by including a number of new post hoc verification checks to identify and eliminate fixations that may be artifactual. We assess the results of our algorithms by testing their reliability using a variety of methods and on several data sets. We contend that, with appropriate data analysis methods, fixation duration can be a reliable and stable measure in infants. We conclude by discussing ways in which studying fixation durations during unconstrained orienting may offer insights into the relationship between attention and learning in naturalistic settings.
The efficient processing of sentences in native speakers is the result of great automaticity and speed both in lexical retrieval and in structural computations. In lexical retrieval, all interpretations of a word are accessed, and non-convergent interpretations are quickly pruned. Garden paths also suggest the autonomy of syntactic computations from contextual knowledge, although there is semantic feedback on proposed syntactic attachments at every stage of processing. Early effects of both lexical and contextual semantic knowledge also point to immediate discourse-semantics processing. Sentence processing, therefore, involves computations in various sub-modules, in the limits of their interfaces (Crocker, 1996; inter alia). Sentence processing includes structural computations that are blind to other sources of knowledge. This blindness is a presumed source of efficiency. However, this efficiency comes at the price of a certain dumbness as the processor seems to be unable to learn from its mistakes, taking the same routes over and over even if they are dead-ends leading to garden paths (Fodor, 1983, 2000). Grammatical research argues that a generative computational system specialized for human language (CHL) plays a significant role in giving language its expressive power. CHL crucially mediates between lexical information and the conceptual intentional system (CI-system) that interfaces with CHL at the level of Logical Form (LF). Thus, constraints on movements and on binding appear to be specific to natural-language grammars. Indeed, formal logical systems do not have such constraints. Grammatical research in the generative paradigm has attempted to understand the role of CHL, with all its idiosyncrasies, in language design in terms of mental constitution. Hence, research on the grammar of anaphora (Reuland, 2001; Reinhart & Reuland, 1993) and research on the processing of movement dependencies (Gibson, 1998, 2000; Gibson & Warren, 2004) both conclude that the computation of referential dependencies in syntax plays a central role in the management of the global processing load. Binding reduces the number of assignments of values to variables (Reuland, 2001; Reinhart & Reuland, 1993) and movement traces refresh the activation of referents in discourse-semantics (Gibson, 1998, 2000; Gibson & Warren, 2004). Quirky grammatical dependencies are pervasive in human languages and their target-like acquisition is not trivial for the second language (L2) learner. Formal grammatical rules constitute a non-negligible portion of what needs to be acquired, in addition to vocabulary items. Beyond either the perceived or real needs of L2 learners to approximate the target-language norms or their personal desire to do so, one may wonder whether there are any benefits to formal grammatical rules in L2 acquisition. CHL computations of grammatical rules clearly involve costs, but the dependencies that these computations establish might also eke out efficiencies in the CI-system in return. Benefits to discourse-semantics processing, if they can be found, could offer insights into the role of UGconstrained grammatical states in L2 cognition. Given that a range of cognitive abilities are available
We set forth to show that lexical connectivity plays a role in understanding early word learning. By considering words that are learned in temporal proximity to one another to be related, we are able to better predict the words next learned by toddlers. We build conditional probability models based on data from the growing vocabularies of 77 toddlers, followed longitudinally for a year. This type of conditional probability model outperforms the current norms based on baseline probabilities of learning given age alone. This is a first step to capturing the interaction between a child’s productive vocabulary and their learning environment in order to understand what words a child might learn next. We also test different types of variants of this conditional probability and find that not only is there information in words that are learned in proximity to one another but that it matters how models integrate this information. The application of this work may provide better cognitive models of acquisition and perhaps allow us to detect children at risk for enduring language difficulties earlier and more accurately.
Abstract This chapter looks at more sophisticated versions of the ambiguity theory. We might say that “ought” is context sensitive rather than lexically ambiguous. And we can try wide-scoping. There are many ways of making (T) and (J) both come out true. But what we really want is to accept the norms. We don’t really just want the truth of the two sentences on some interpretation or another. But even in its more sophisticated versions, all the ambiguity theory can provide is the truth of the two sentences. The norms themselves remain inconsistent.
Gamedesire (GD), as a meeting place for multilingual communities, is an inexhaustible resource for linguistic research on Computer-Mediated Discourse (CMD). GD remains linguistically under-researched though it gives rise to a plethora of linguistic issues relating in particular to written English. This paper is a seminal work that takes as its keynote the linguistic analysis of one key issue — namely 'euphemism of nicknaming'. The present study seeks to examine the various language tools users employ in their creation of their own nicknames on the URL http://www.gamedesire.com. A corpus of 200 nicknames has randomly been collected in 2008 and tested against an existing model of euphemism by Warren (1992) [7]. The study shows that a large number of connotatively dysphemistic nicknames are denotatively euphemized by deviating from language norms and employing a wide range of linguistic and paralinguistic devices, including word-formation, orthographic modification, borrowing and semantic innovation. Some of these nicknames were not subsumable under the original model and, therefore, necessitated developing a new rendition. The new rendition of the model created other mechanisms not developed by the original model. The study concludes that language users employ different styles of nicknaming marked by grammatical (change of word grammar), lexical (creation of new words), phonological (irregularity of pronunciation), orthographic (irregularity of spelling), and semantic deviations (transference of meaning).
espanolEste trabajo describe el sistema de normalizacion de tuits en espanol desarrollado por el Grupo de Lengua Y Sociedad de la Informacion (LYS) de la Universidade da Coruna para el Tweet-Norm 2013. Se trata de un sistema conceptualmente sencillo y flexible que emplea pocos recursos y que aborda el problema desde un punto de vista lexico. EnglishThis work describes the system for the normalization of tweets in Spanish designed by the Language in the Information Society (LYS) Group of the University of A Coruna for Tweet-Norm 2013. It is a conceptually simple and flexible system, which uses few resources and that faces the problem from a lexical point of view.
The language mechanisms of the substantivizing and lexicalization of the Russian pronoun -nashi‖ (-our‖) are analysed with respect to semantic derivation. The potentiality of using the substantivized and lexicalized pronoun-nashi‖ (-our‖) in language conceptualization in the sphere of norms, ideals and values in the national picture of the world is considered.
Translation is a kind of a trial for the target language, a test of its expressive possibilities, but also an exam of the abilities and skills of a translator. Even the best translators refrained from translating “holy books”, due to the challenges of uniqueness of the form and as a precaution of potential sin and (or) blasphemy which translation can cause. However, translating the Word of God, in this case, the Qur’an, is a necessity, and for Bosniaks, it is a national mission, it testifies to their religious tradition written in Bosnian language at a given time. Therefore, translations of the Qur'an deserve a serious scientific analysis and a responsible, multifaceted research approach, about which we have not had a chance to read a lot in linguistics, in particular Bosnistics. The exceptions are the books of Dž. Latić, PhD, and his scientific and professional papers published in the Proceedings of FIS. This study comprises a corpus of four well-known Bosnian translations of the Qur’an, as follows: Besim Korkut, Mustafa Mlivo (whose authenticity is disputed, i.e. its direct translation from the Arabic original), Enes Karić and Esad Duraković. Furthermore, due to the volume of the material, objects of interest are focused on the first and thirtieth juz (first twenty and last twenty pages of translations of the Qur'an) from which all examples of specific linguistic phenomena and regularities have been taken. The main objectives of this paper are: initiation and actualisation of lexicological and general semantic research of translations of the sacred text, which are grammatically interesting and stylogenic, the research of specific lexical-semantic level of linguistic structure of the translations of the Qur'an and a scientific contribution to the study of this kind of discourse. The task of the paper is to describe, or reinterpret the theoretical principles of lexical semantics of Bosnian language by using examples from the corpus, then to affirm the Bosnian language standard by highlighting examples which contribute to the strengthening of linguistic norm in all segments. For the purpose of achieving the objectives and tasks, different methods have been used: monographic, descriptive, comparative, contrastive and lexical-stylistic method. The Qur'anic text is a real repository for stylistic interpretation as well (and not only stylistic, of course) and its literary perfection is a proof of its divine origin. In this paper, a repertoire of semantic figures – tropes, typical contexts in which they operate, their meaning and use have been noted. Stylistics, no matter how successful it is, cannot penetrate into the secret of Qur’anic ijaza (supernatural origin), “as anatomy cannot penetrate into the mystery of creation.” Keywords: tropes, lexical-semantic figures, stylem, metaphor, metonymy, synecdoche, periphrasis, epithet, personification, simile
Artiklis tulevad vaatluse alla 17. sajandi kiriklike teoste tõlkija ja keelenormi kujundaja Heinrich Stahli tekstides kasutatud kaheksa haruldast tüvisõna ja seitse tuletist, mis (1) esinevad tema teostes vaid ühe korra (nn hapax legomenon’id), (2) esinevad põhjaeesti kirjakeeles esimest korda just Stahli teostes, (3) ei ole tänapäeva kirjakeeles sellises vormis ja/või tähenduses kasutusel, (4) ei ole otselaenud (alam)- saksa keelest. Käsitluse eesmärgiks on välja tuua Stahli haruldast sõnavara, mis pole tänapäeva eesti keeles enam kas tüve, moodustusviisi või tähenduse poolest läbipaistev. Niisugused lekseemid võimaldavad täpsemat pilguheitu Stahli teoste eripärasele sõnavarale ja selle kasutamise motiividele. Artiklis käsitletud tüvisõnad peegeldavad arhailist sõnavarakihti, mis on enamjaolt rahvakeelne ja võib osaliselt pärineda varasematest, Stahli-eelsetest allikatest. Tuletised avavad lisaks ka mehhanisme, kuidas Stahl on kasutanud teksti vajadustest lähtudes produktiivseid tuletusvõimalusi või toetunud analoogiamallidele. Selline täpseid esinemissagedusi arvestav uurimus on võimalikuks saanud pärast Stahli tekstide korpuse lemmatiseeritud kuju valmimist 2013. aastal.Rare words from the works of Heinrich Stahl. The article takes a look at 8 stem words and 7 derivations used in the works of Heinrich Stahl, the 17th century translator of religious texts and shaper of language norms. The stem words and derivations discussed in the article (1) only occur once in his texts, (2) first appear in Literary North-Estonian in Stahl’s texts, (3) are not used in Modern Estonian in the same form and/or meaning, (4) are not direct loans from (Low) German. These lexemes reveal the peculiarity of the lexicon of Stahl’s works. Because of their rarity, the lemmatising of such units while coding the corpus of Stahl’s texts has been problematic. These archaic words are not transparent in stem, form or meaning in Modern Estonian. The lexical stems discussed in the article reflect an archaic layer of the lexicon which is mostly vernacular and may partly originate in earlier, pre-Stahl sources. Derivations, in addition to revealing archaic lexicon, also reveal the mechanisms of how Stahl used productive derivation or analogy patterns depending on the demands of the text.
The article appraises gender representation in the 1999 Nigerian Constitution using insights from critical discourse analysis, feminism and systemic functional linguistics, with particular emphasis on grammatical cohesion. Specifically, it examines lexical and grammatical expressions that encode gender in the Constitution, the ideological positions evident in these expressions, and their impact on gender parity and socio-political equity. The focus is on the reference-antecedent cohesion of gender-marked pronouns and nouns used to refer to individuals and social/political positions. Our findings show a preponderance of generic masculine noun and pronoun references, tracking antecedents that refer to social and political positions open to eligible individuals in Nigeria, while the single feminine referent was a marked case. These findings buttress the ‘male-as-norm’ ideology and the relegation to anonymity of the female gender in this important national document. For equity and fairness, the article recommends revising the Constitution with epicene expressions to expunge gender biases.
on, Tweets Abstract: The lexical richness and its ease of access to large volumes of information converts the Web 2.0 into an important resource for Natural Language Processing. Nevertheless, the frequent presence of non-normative linguistic phenomena that can make any automatic processing challenging. In this paper is described the partici- pation in the Text Normalisation Workshop at the SEPLN conference (Tweet-norm 2013). The Workshop includes one unique task focused on the normalisation of Spa- nish tweets. For this task we have used TENOR, a multilingual lexical normalisation tool for Web 2.0 texts.
The present study investigates whether a minimal manipulation in task demands can induce core linguistic combinatorial mechanisms to extend beyond the bounds of normal grammatical phrases. Using magnetoencephalography, we measured neural activity evoked by the processing of adjective-noun phrases in canonical (red cup) and reversed order (cup red). During a task not requiring composition (verification against a color blob and shape outline), we observed significant combinatorial activity during canonical phrases only - as indexed by minimum norm source activity localized to the left anterior temporal lobe at 200-250 ms(cf. [1], [2]). When combinatorial task demands were introduced (by simply combining the blob and outline into a single colored shape) we observed significant combinatorial activity during reversed sequences as well. These results demonstrate the first direct evidence that basic linguistic combinatorial mechanisms can be deployed outside of normal grammatical expressions in response to task demands, independent of changes in lexical or attentional factors.
Unlike some varieties of English in Southeast Asia, the notion that there is a ‘Thai English’ is debatable. This paper examines distinctive non-native features of a lexicon found in contemporary Thai writing in English to ascertain if English in this Expanding Circle country is developing its own linguistic norms. An analysis of features of lexical creativity in five short stories and novels is carried out to determine whether the characteristics found indicate that a Thai English vocabulary exists. An ‘integrated framework’ which combines concepts in World Englishes by Braj B, Kachru, Peter Strevens, and Edgar W. Schneider is adopted in this study. It appears that certain categories of lexical creativity in the fiction examined represent five indicators of Thai English - contextualization, innovation, nativization, transcultural creativity, and localization - and reveals a developing non-native variety of English.
This paper addresses the problem of semantic entity resolution (SER), which aims to determine whether some or none of the entities in a knowledge base is mentioned in a given web document. The lexical features, e.g., words and phrases, which are critical to the resolution of the semantic entities are typically of a small amount compared to all lexical features in the web document, and therefore can be modeled as sparse signals. Two techniques leveraging the principles of sparse signal recovery are proposed to identify the sparse, salient lexical features: one technique, based on the Lasso algorithm with the l2-norm distance metric, attempts to recover all the salient lexical features at once; the other technique, namely Posterior Probability Pursuit (PPP), sequentially identifies salient features one after one using the negative log posterior probability as the distance metric. Using a knowledge base consisting of about 100 million entities, we show that the proposed techniques exploiting the sparsity nature underlying SER deliver substantial performance improvement over baseline methods without sparsity consideration, demonstrating the potentials of sparse signal techniques in entity-centric web information processing.
This paper examines the sociolinguistic import of an emerging hybrid street language by a group known as the Ágábá Boys in Calabar South, Cross River State, South-eastern Nigeria. The paper explores the lexically and contextually driven ingroup code of the Ágábá Boys which is manifested in slang, metaphors and a variety of taboo expressions embodied in expletives, profanities, insults, curses and swear words. The group uses its peculiar language in addition to other socially constructed dialects to reinforce anti-establishment behaviour, conceptualize identity, enhance solidarity and foster group integration. Youth language in Calabar South is full of improvization and allows enormous creative possibilities. It is adjudged to be generally deviant and exotic and mainly perceived as a mark of poor parentage, unemployment, limited education and low social orientation. It enables youth (Ágábá Boys) to identify their individuality and express their deviant tendencies against established norms and conventions. The paper highlights how youth in Calabar South produce and reinforce their marginal/deviant status through iterative and creative language use which is essential for the creation of urban subculture and their group dynamics.
=7) control groups. Following nine sessions combining computerized rapid accelerated-reading program (RAP), which individually tailors rate of written text presentation to comprehension criterion (80%), and self-regulated strategies for attending and engaging, the treated group significantly outperformed the wait-listed group before treatment on (a) a grade-normed, silent sentence reading rate task requiring lexical- and syntactic level processing to decide which of three sentences makes sense; and (b) RAP presentation rates yoked to comprehension accuracy level. Each group improved significantly on these same outcomes from before to after instruction. Attention ratings and working memory for written words predicted post-treatment accuracy, which correlated significantly with the silent sentence reading rate score. Implications are discussed for (a) preventing silent reading disabilities during the transition to increasing emphasis on silent reading, (b) evidence-based approaches for making accommodation of extra time on timed tests requiring silent reading, and
Question Time is a distinctive daily parliamentary routine. Its aim is to hold Ministers of the State accountable for the actions and decisions of the Government. However, in many Parliaments, including the New Zealand and Australian Federal Houses of Representatives, it is more of a theatrical performance where parties try their best to score political points. As any performance, Question Time is governed by certain rules and regulations outlined in an official document Standing Orders. As there is not much action, Standing Orders mainly describe language norms and specify „unparliamentary language‟. This research looks at and analyses the use of formulaic vocabulary used by MPs in the year preceding general elections in New Zealand and Australia. The formulaic language includes phrasal lexical items and formulae for asking / answering questions, for raising points of order and the Speakers‟ idiolectal phrasal vocabulary for quelling disorder in the Chambers and regulating the work of the House. The framework developed for this research consisted of the following steps: an ethnographic study of Question Time as a communicative performance which included the development of a database containing all the empirical material; a xii linguistic study of Question Time including genrelect study, parliamentary formulae study and disorder analysis before the elections. As a result this research has shown that Question Time is a communicative performance event in New Zealand and Australia with significant cultural, historic and linguistic differences in spite of the common origins of the two Parliaments. It has identified 60 Question Time genre-specific phrasal lexical items that MPs use in the two Parliaments, studied their structure and meaning (where necessary). It has also looked at the strategies the MPs employ for creating disorder in the House, and the ways of quelling disorder by the Speakers of the two Parliaments.
There is considerable ethno-linguistic and genetic variation among human populations in Asia, although tracing the origins of this diversity is complicated by migration events. Thailand is at the center of Mainland Southeast Asia (MSEA), a region within Asia that has not been extensively studied. Genetic substructure may exist in the Thai population, since waves of migration from southern China throughout its recent history may have contributed to substantial gene flow. Autosomal SNP data were collated for 438,503 markers from 992 Thai individuals. Using the available self-reported regional origin, four Thai subpopulations genetically distinct from each other and from other Asian populations were resolved by Neighbor-Joining analysis using a 41,569 marker subset. Using an independent Principal Components-based unsupervised clustering approach, four major MSEA subpopulations were resolved in which regional bias was apparent. A major ancestry component was common to these MSEA subpopula)
Abstract Semantic substitution errors (slips of the tongue) naturally occurring in Russian normal speech were analyzed for word frequency, word length, target-error cooccurrence strength, and word association norms. Target word frequencies were found to be significantly lower than error word frequencies; besides, there is a very significant positive correlation between target and error frequency values. Contrary to the view that the frequency effect is located at the stage of phonological encoding, the results suggest that frequency is coded at an earlier stage of lexical selection. Word length is a significant variable that determines the outcome of the error for non-cohyponym target-error pairs but not for cohyponym pairs. At the same time, cohyponym target-error pairs are characterized by much higher cooccurrence measures and stronger associative links compared to non-cohyponym pairs. Theoretical implications of these findings are discussed
Pragmatic markers are an important part of the grammar of conversation and not simply markers of disfluency. They have a number of functions that help the speaker to organize the conversation and to express feelings and attitudes. Advanced EFL learners use frequent pragmatic markers such as well. However their use of well diverges from the native speaker norm. The present study uses data from the Swedish component of the LINDSEI corpus and its native speaker counterpart (LOCNEC) to examine similarities and differences between native and non-native speakers. The overall picture is that Swedish learners overuse well, although there are considerable individual differences. Thus learners use well above all as a fluency device to cope with speech management problems but underuse it for attitudinal purposes. Pragmatic markers cannot be taught in the same way as other lexical items but it is important to discuss how and where they are used.