Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
These days creating the corpus of texts for Uzbek language, creating and developing linguistic databases, search-engine systems – are one of the crucial tasks of computational linguistics. Particularly, electronic dictionary-thesauruses, semantic dictionaries are one of them. Dictionary-thesaurus formation structure for Uzbek language, transferring the terminological dictionary into the e-version and implementing rules for establishing semantic relations between words where it gives a chance to establish automation linguistic processes of dictionary-thesauruses, which is the foundation of linguistic databases. Analyzing logical structure of paper-based dictionary thesauruses has given a chance to formalize its structure and creating rules for converting to e-version of dictionary-thesaurus syllables by using predicates language. Descriptors system is suggested in PROLOG language rules set for constructing e-version of dictionary – syllables.
Dependency parsing can provide the connection of linguistic unit (words) by a directed links. This paper presents annotating a general domain corpus by using unsupervised approach by applying Universal part-of-speech (U-POS) to build Treebank for unsupervised dependency parsing of Myanmar Language. Up to now it is still hard task to obtain complete syntactic structures for Myanmar Language. Dependency structures of words in Myanmar sentences are also presented of general words and phrases orders and the relations of basic sentence structures. To annotate by using U-POS, UDPipe is used. Moreover, the preliminary results of annotated trees and parsing experiment are presented. Parsing experiments are evaluated by UDPipe in terms of unlabeled and labeled attachment scores: (UAS) and (LAS), which are 93.20%, and 91.21% in test experiment respectively.
Most language users agree that some words sound harsh (e.g. grotesque) whereas others sound soft and pleasing (e.g. lagoon). While this prominent feature of human language has always been creatively deployed in art and poetry, it is still largely unknown whether the sound of a word in itself makes any contribution to the word’s meaning as perceived and interpreted by the listener. In a large-scale lexicon analysis, we focused on the affective substrates of words’ meaning (i.e. affective meaning) and words’ sound (i.e. affective sound); both being measured on a two-dimensional space of valence (ranging from pleasant to unpleasant) and arousal (ranging from calm to excited). We tested the hypothesis that the sound of a word possesses affective iconic characteristics that can implicitly influence listeners when evaluating the affective meaning of that word. The results show that a significant portion of the variance in affective meaning ratings of printed words depends on a number of spe)
Comprehending natural language quantifiers (like many, all, or some) involves linguistic and numerical abilities. However, the extent to which both factors play a role is controversial. In order to determine the specific contributions of linguistic and number skills in quantifier comprehension, we examined two groups of participants that differ in their language abilities while their number skills appear to be similar: Participants with Down syndrome (DS) and participants with Williams syndrome (WS). Compared to rather poor linguistic skills of individuals with DS, individuals with WS display relatively advanced language abilities. Participants with WS also outperformed participants with DS in a quantifier comprehension task while number knowledge did not differ between the two groups. When compared to typically developing (TD) children of the same mental age, participants with WS displayed similar levels regarding quantifier abilities, but participants with DS performed worse than th)
Word sense disambiguation (WSD) is the process of identifying an appropriate sense for an ambiguous word. With the complexity of human languages in which a single word could yield different meanings, WSD has been utilized by several domains of interests such as search engines and machine translations. The literature shows a vast number of techniques used for the process of WSD. Recently, researchers have focused on the use of meta-heuristic approaches to identify the best solutions that reflect the best sense. However, the application of meta-heuristic approaches remains limited and thus requires the efficient exploration and exploitation of the problem space. Hence, the current study aims to propose a hybrid meta-heuristic method that consists of particle swarm optimization (PSO) and simulated annealing to find the global best meaning of a given text. Different semantic measures have been utilized in this model as objective functions for the proposed hybrid PSO. These measures consis)
Background: In practical research, it was found that most people made health-related decisions not based on numerical data but on perceptions. Examples include the perceptions and their corresponding linguistic values of health risks such as, smoking, syringe sharing, eating energy-dense food, drinking sugar-sweetened beverages etc. For the sake of understanding the mechanisms that affect the implementations of health-related interventions, we employ fuzzy variables to quantify linguistic variable in healthcare modeling where we employ an integrated system dynamics and agent-based model. Methodology: In a nonlinear causal-driven simulation environment driven by feedback loops, we mathematically demonstrate how interventions at an aggregate level affect the dynamics of linguistic variables that are captured by fuzzy agents and how interactions among fuzzy agents, at the same time, affect the formation of different clusters(groups) that are targeted by specific interventions. Results:)
Introduction: Oral Anticoagulation therapy (OAC) is highly effective in the management of thromboembolic disorders. An adequate level of knowledge is important for self-management and optimizing clinical outcomes. The Anticoagulation Knowledge Tool (AKT) was developed to assess OAC knowledge and caters for both patients prescribed direct oral anticoagulants or vitamin K antagonist (VKA). However, evidence regarding its psychometric proprieties, validity and reliability are unavailable in non-English speaking settings. For this reason, the aim of this study is to provide further evidence of validity for AKT and also developing an Italian AKT version (I-AKT) supported by evidence of validity and reliability. Methods: A multiphase study was conducted which included the following: cultural and linguistic validity; i.e. content validity; construct validity; reliability assessment. The Construct validity was performed using the contrasted group approach using three groups comprised of hea)
The Newest Vital Sign (NVS) is a simple, quick and accurate screening test for health literacy (HL). It has been validated for different languages but, to date, not for the Croatian language. The aim of this study was to develop a linguistically validated Croatian version of the NVS and to use it at a later stage in a pilot study of health literacy assessment of hospital patients in Croatia. A full linguistic validation procedure was applied, including forward and backward translation, expert panel review, cognitive interview with 10 respondents from general population, and full involvement in the procedure of one of the screening test developers, the lead author of the NVS-UK version. HL testing on 100 hospital patients (55% women, median age 63.5 years) revealed 58% of patients had less than adequate HL level (scores less than 4), and mean NVS total score was 3.34. A positive significant association was observed between HL and educational level (p = 0.002). A high percentage of pati)
Scholarly studies and common accounts of national politics enjoy pointing out the resilience of ideological divides among populations. Building on the image of political cleavages and geographic polarization, the regionalization of politics has become a truism across Northern democracies. Left unquestioned, this geography plays a central role in shaping electoral and referendum campaigns. In Europe and North America, observers identify recurring patterns dividing local populations during national votes. While much research describes those patterns in relation to ethnicity, religious affiliation, historic legacy and party affiliation, current approaches in political research lack the capacity to measure their evolution over time or other vote subsets. This article introduces “Dyadic Agreement Modeling” (DyAM), a transdisciplinary method to assess the evolution of geographic cleavages in vote outcomes by implementing a metric of agreement/disagreement through Network Analysis. Unlike ex)
Over the past century, personality theory and research has successfully identified core sets of characteristics that consistently describe and explain fundamental differences in the way people think, feel and behave. Such characteristics were derived through theory, dictionary analyses, and survey research using explicit self-reports. The availability of social media data spanning millions of users now makes it possible to automatically derive characteristics from behavioral data—language use—at large scale. Taking advantage of linguistic information available through Facebook, we study the process of inferring a new set of potential human traits based on unprompted language use. We subject these new traits to a comprehensive set of evaluations and compare them with a popular five factor model of personality. We find that our language-based trait construct is often more generalizable in that it often predicts non-questionnaire-based outcomes better than questionnaire-based traits (e.g)
Introduction: Natural resource management uses expert judgement to estimate facts that inform important decisions. Unfortunately, expert judgement is often derived by informal and largely untested protocols, despite evidence that the quality of judgements can be improved with structured approaches. We attribute the lack of uptake of structured protocols to the dearth of illustrative examples that demonstrate how they can be applied within pressing time and resource constraints, while also improving judgements. Aims and methods: In this paper, we demonstrate how the IDEA protocol for structured expert elicitation may be deployed to overcome operational challenges while improving the quality of judgements. The protocol was applied to the estimation of 14 future abiotic and biotic events on the Great Barrier Reef, Australia. Seventy-six participants with varying levels of expertise related to the Great Barrier Reef were recruited and allocated randomly to eight groups. Each participant)
This paper focuses on a novel methodology of subjective speech quality measurement and repeatability of its results between laboratory conditions and simulated environmental conditions. A single set of speech samples was distorted by various background noises and low bit-rate coding techniques. This study aimed to compare results of subjective speech quality tests with and without a parallel task deploying the ITU-T P.835 methodology. Afterward, tests results performed with and without a parallel task were compared using Pearson correlation, CI95, and numbers of opposite pair-wise comparisons. The tests show differences in results in the case of a parallel task. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles for individual use. This abstract may)
Crowdsourcing services, such as MTurk, have opened a large pool of participants to researchers. Unfortunately, it can be difficult to confidently acquire a sample that matches a given demographic, psychographic, or behavioral dimension. This problem exists because little information is known about individual participants and because some participants are motivated to misrepresent their identity with the goal of financial reward. Despite the fact that online workers do not typically display a greater than average level of dishonesty, when researchers overtly request that only a certain population take part in an online study, a nontrivial portion misrepresent their identity. In this study, a proposed system is tested that researchers can use to quickly, fairly, and easily screen participants on any dimension. In contrast to an overt request, the reported system results in significantly fewer (near zero) instances of participant misrepresentation. Tests for misrepresentations were conducted by using a large database of past participant records (~45,000 unique workers). This research presents and tests an important tool for the increasingly prevalent practice of online data collection.
Up to now, the potential of eye tracking in science as well as in everyday life has not been fully realized because of the high acquisition cost of trackers. Recently, manufacturers have introduced low-cost devices, preparing the way for wider use of this underutilized technology. As soon as scientists show independently of the manufacturers that low-cost devices are accurate enough for application and research, the real advent of eye trackers will have arrived. To facilitate this development, we propose a simple approach for comparing two eye trackers by adopting a method that psychologists have been practicing in diagnostics for decades: correlating constructs to show reliability and validity. In a laboratory study, we ran the newer, low-cost EyeTribe eye tracker and an established SensoMotoric Instruments eye tracker at the same time, positioning one above the other. This design allowed us to directly correlate the eye-tracking metrics of the two devices over time. The experiment was embedded in a research project on memory where 26 participants viewed pictures or words and had to make cognitive judgments afterwards. The outputs of both trackers, that is, the pupil size and point of regard, were highly correlated, as estimated in a mixed effects model. Furthermore, calibration quality explained a substantial amount of individual differences for gaze, but not pupil size. Since data quality is not compromised, we conclude that low-cost eye trackers, in many cases, may be reliable alternatives to established devices.
Modern Ukrainian language is characterized by interrelated tendencies of synthetical character and analyticity, which are motivated by: 1) the folk-colloquial element of the literary norm; 2) book tradition; 3) the law of language economy, etc. One of the brightest analysts expresses the dynamic development of the prepositional system, which has recently been actively replenished by semantically specialized two-component, three-component, and other entities. The ultimate manifestation of such an analyticism is the presence of dissected prepositions of the sample від... до, з... до.The textual function of the prepositions within the texts appears in the inter-phrased / intra-phrased, intersentenced manifestations; at the same time, the expansion of the paradigmatic plane of prepositions, their participation in the creation of images, are apparent. Within the context the preposition appears as a relatively independent element with its own inventory of distributions, phonetic variants, and others. The solution of the question of the status of the text of the prepositions will enable understanding of the mechanisms of interaction between the components and components of the language system and the establishment of ways for the creation and functioning of syntaxemes, the extension of the formation of their own semantic potential of the latter.Lexical and grammatical meanings of prepositions coincide, which does not mean their identity. The grammatical meaning of the preposition is the realization of the form of syntactic relation between the words, and the lexical one should consider the designation of a certain relation between the objects, the action and the object, etc., which makes it possible to enumerate the prepositions to the words-relatives. The ability of prepositions to determine its lexical meaning in the noun (more broadly – in the name) is related not to the absence of this value, but to its corresponding specifics. Interpretation of prepositions should be based on the lexical environment, since they indicate the relation between the objects. The ratio implies the presence of not less than two quantities, therefore, prepositions can not exist without these values, that is, they can not be used independently. The meaning of the prepositions implies their functioning within the phrase. And by its nature the primary prepositions are many-valued, they are characterized by homonymy. In this case, the context is a diagnostic indicator of a certain value as a virtual-system.In analyzing the semantics of relations, it is necessary to take into account, as much as possible, the particularities of the lexical meaning of prepositions and the semantic features of words that form the left and right-side distribution of the prepositional-case design. Due to this, the following classification of intratextual semantic relations, expressed by the Ukrainian primary prepositions in artistic-fiction and journalistic language-linguistic discourse practices, is real: 1) spatial relationships that are the most researched in modern linguistics. Among them differentiate the local (place) and additive (direction). Locative relations are differentiated into: suppressive (finding above the surface) and invasive (finding inside).
Abstract The first part of this paper outlines the relevant aspects of functional structuralism serving lexicographers as a departure point for building a model of lexical meaning useable in the Dictionary of Contemporary Slovak Language. This section also points to some aspects of Klára Buzássyová’s research on lexis and wordformation that have enriched the functionalstructuralist paradigm. The second section shows other theoretical and methodological frameworks, such as linguistic pragmatics, cognitive linguistics and corpus linguistics (all of them departing in some respect from the structuralism and, in other aspects, being complementary with it) that can enhance the structuralist basis of the model. The third section outlines an extended model of lexical meaning that represents a synthesis of all those theoretical frameworks and, at the same time, represents a reflection of three language constituents: 1. The social constituent is present in consideration of communicative functions of utterances, naming functions of lexical units, functional styles and registers, language norms, and situational contexts; 2. The psychological component takes the form of consideration of the prototype effect, the abolition of boundaries between linguistic meaning and other parts of cognition; 3. Thanks to the structural/systematic component, a description of paradigmatic and syntagmatic behaviour of words can be performed, and an inventory of formalcontent units and categories (lexemes, lexies, wordforming and grammatical structures) can be provided. In our dictionary practice, the abovementioned model is reflected in the methodological procedures as follows: 1. Systemization of repetitive (regular, standardized) phenomena; 2. Prototypicalization of meaning description; 3. Contextualization/encyclopedization of meaning description; 4. Pragmatization of meaning description; 5. Continualized presentation of language phenomena, i.e., introduction of numerous phenomena of transient and indeterminate nature and indicating the existence of a semanticpragmatic and lexicalgrammatical continuum; 6. “Discretization” of combinatorial continuum, i.e., identification and description of entrenched word combinations with naming functions.
This paper describes Stanford's system at the CoNLL 2018 UD Shared Task. We introduce a complete neural pipeline system that takes raw text as input, and performs all tasks required by the shared task, ranging from tokenization and sentence segmentation, to POS tagging and dependency parsing. Our single system submission achieved very competitive performance on big treebanks. Moreover, after fixing an unfortunate bug, our corrected system would have placed the 2 nd, 1 st, and 3 rd on the official evaluation metrics LAS, MLAS, and BLEX, and would have outperformed all submission systems on lowresource treebank categories on all metrics by a large margin. We further show the effectiveness of different model components through extensive ablation studies. * These authors contributed roughly equally.
We propose Efficient Neural Architecture Search (ENAS), a fast and inexpensive approach for automatic model design. In ENAS, a controller learns to discover neural network architectures by searching for an optimal subgraph within a large computational graph. The controller is trained with policy gradient to select a subgraph that maximizes the expected reward on the validation set. Meanwhile the model corresponding to the selected subgraph is trained to minimize a canonical cross entropy loss. Thanks to parameter sharing between child models, ENAS is fast: it delivers strong empirical performances using much fewer GPU-hours than all existing automatic model design approaches, and notably, 1000x less expensive than standard Neural Architecture Search. On the Penn Treebank dataset, ENAS discovers a novel architecture that achieves a test perplexity of 55.8, establishing a new state-of-the-art among all methods without post-training processing. On the CIFAR-10 dataset, ENAS designs novel architectures that achieve a test error of 2.89%, which is on par with NASNet (Zoph et al., 2018), whose test error is 2.65%.
Recurrent neural nets (RNN) and convolutional neural nets (CNN) are widely used on NLP tasks to capture the long-term and local dependencies, respectively. Attention mechanisms have recently attracted enormous interest due to their highly parallelizable computation, significantly less training time, and flexibility in modeling dependencies. We propose a novel attention mechanism in which the attention between elements from input sequence(s) is directional and multi-dimensional (i.e., feature-wise). A light-weight neural net, "Directional Self-Attention Network (DiSAN)," is then proposed to learn sentence embedding, based solely on the proposed attention without any RNN/CNN structure. DiSAN is only composed of a directional self-attention with temporal order encoded, followed by a multi-dimensional attention that compresses the sequence into a vector representation. Despite its simple form, DiSAN outperforms complicated RNN models on both prediction quality and time efficiency. It achieves the best test accuracy among all sentence encoding methods and improves the most recent best result by 1.02% on the Stanford Natural Language Inference (SNLI) dataset, and shows state-of-the-art test accuracy on the Stanford Sentiment Treebank (SST), Multi-Genre natural language inference (MultiNLI), Sentences Involving Compositional Knowledge (SICK), Customer Review, MPQA, TREC question-type classification and Subjectivity (SUBJ) datasets.
Detection and correction of errors and inconsistencies in "gold treebanks" are becoming more and more central topics of corpus annotation. The paper illustrates a new incremental method for enhancing treebanks, with particular emphasis on the extension of error patterns across different textual genres and registers. Impact and role of corrections have been assessed in a dependency parsing experiment carried out with four different parsers, whose results are promising. For both evaluation datasets, the performance of parsers increases, in terms of the standard LAS and UAS measures and of a more focused measure taking into account only relations involved in error patterns, and at the level of individual dependencies.
This paper presents a methodology for rule based bottom up parsing technique forModern Standard Arabic (MSA) inContext Free Grammar (CFG) formalism in Phrase Structure Grammar (PSG) representation, where the grammar isautomatically extracted from a syntactically annotated corpus.The extracted grammar is used to build an automatic lexicon andgrammar rules module. Furthermore, the extracted CFG is further transformed into Probabilistic Context Free Grammar (PCFG)that could be used in a hybrid approach, which is also calculated automatically. The used corpus is the Penn ArabicTreebank(PATB)and algorithm implementation is performed with Natural Language Processing Toolkit (NLTK).The parsershowed that automatic extraction of grammar improved the grammar building phase in both coverage of structures and timeneeded, but still needs further manual constrains addition. Automatic extraction of grammar is able to enhance rule basedgrammar parsers and it will enable a new paradigm of statistically directed symbolic parsing.
To fasten treebank construction, it is necessary to design an integrated annotation tool that includes word segmenter, sentence parser for initial tree suggestion, tree visualizer, tree-structure editor, and collaborative functions. In the past, existing tools did not consider an integrated platform that provides preprocessing, automated or semi-automated mechanism for parse tree suggestion, as well as tagged corpus data management. This paper presents a so-called CF Planter, a toolset for semi-automatic Thai treebank construction that consist of word segmenter, part-of-speech tagger, statistical parser, a web-based GUI for syntactic tree refinement and management. Given an input sentence, its most likely syntactic tree is automatically suggested and visualized to an annotator for manual correction before adding into the treebank repository. Whenever a new syntactic tree is appended into the treebank, the treebank repository is iteratively refined by computing a set of newly revised grammar rules based on revised probabilities. Toolset is performed to severally illustrate with grammar frequencies. The toolset facilitates annotators to easily tag tree structure for an input sentence. Finally, the process of automatic suggestion of syntactic tree is evaluated.
We demonstrate that replacing an LSTM encoder with a self-attentive architecture can lead to improvements to a state-ofthe-art discriminative constituency parser. The use of attention makes explicit the manner in which information is propagated between different locations in the sentence, which we use to both analyze our model and propose potential improvements. For example, we find that separating positional and content information in the encoder can lead to improved parsing accuracy. Additionally, we evaluate different approaches for lexical representation. Our parser achieves new state-ofthe-art results for single models trained on the Penn Treebank: 93.55 F1 without the use of any external data, and 95.13 F1 when using pre-trained word representations. Our parser also outperforms the previous best-published accuracy figures on 8 of the 9 languages in the SPMRL dataset.
Chronic pain may alter both affect- and value-related behaviors, which represents a potentially treatable aspect of chronic pain experience. Current understanding of how chronic pain influences the function of brain reward systems, however, is limited. Using a monetary incentive delay task and functional magnetic resonance imaging (fMRI), we measured neural correlates of reward anticipation and outcomes in female participants with the chronic pain condition of fibromyalgia (N = 17) and age-matched, pain-free, female controls (N = 15). We hypothesized that patients would demonstrate lower positive arousal, as well as altered reward anticipation and outcome activity within corticostriatal circuits implicated in reward processing. Patients demonstrated lower arousal ratings as compared with controls, but no group differences were observed for valence, positive arousal, or negative arousal ratings. Group fMRI analyses were conducted to determine predetermined region of interest, nucleus accumbens (NAcc) and medial prefrontal cortex (mPFC), responses to potential gains, potential losses, reward outcomes, and punishment outcomes. Compared with controls, patients demonstrated similar, although slightly reduced, NAcc activity during gain anticipation. Conversely, patients demonstrated dramatically reduced mPFC activity during gain anticipation-possibly related to lower estimated reward probabilities. Further, patients demonstrated normal mPFC activity to reward outcomes, but dramatically heightened mPFC activity to no-loss (nonpunishment) outcomes. In parallel to NAcc and mPFC responses, patients demonstrated slightly reduced activity during reward anticipation in other brain regions, which included the ventral tegmental area, anterior cingulate cortex, and anterior insular cortex. Together, these results implicate altered corticostriatal processing of monetary rewards in chronic pain.
We explore dynamic evaluation, where sequence models are adapted to the recent sequence history using gradient descent, assigning higher probabilities to re-occurring sequential patterns. We develop a dynamic evaluation approach that outperforms existing adaptation approaches in our comparisons. We apply dynamic evaluation to outperform all previous word-level perplexities on the Penn Treebank and WikiText-2 datasets (achieving 51.1 and 44.3 respectively) and all previous character-level cross-entropies on the text8 and Hutter Prize datasets (achieving 1.19 bits/char and 1.08 bits/char respectively).
We present a novel abstractive summarization framework that draws on the recent development of a treebank for the Abstract Meaning Representation (AMR). In this framework, the source text is parsed to a set of AMR graphs, the graphs are transformed into a summary graph, and then text is generated from the summary graph. We focus on the graph-to-graph transformation that reduces the source semantic graph into a summary graph, making use of an existing AMR parser and assuming the eventual availability of an AMR-to-text generator. The framework is data-driven, trainable, and not specifically designed for a particular domain. Experiments on gold-standard AMR annotations and system parses show promising results. Code is available at: https://github.com/summarization
In this paper we address extractive summarization of long threads in online discussion fora. We present an elaborate user evaluation study to determine human preferences in forum summarization and to create a reference data set. We showed long threads to ten different raters and asked them to create a summary by selecting the posts that they considered to be the most important for the thread. We study the agreement between human raters on the summarization task, and we show how multiple reference summaries can be combined to develop a successful model for automatic summarization. We found that although the inter-rater agreement for the summarization task was slight to fair, the automatic summarizer obtained reasonable results in terms of precision, recall, and ROUGE. Moreover, when human raters were asked to choose between the summary created by another human and the summary created by our model in a blind side-by-side comparison, they judged the model’s summary equal to or better than the human summary in over half of the cases. This shows that even for a summarization task with low inter-rater agreement, a model can be trained that generates sensible summaries. In addition, we investigated the potential for personalized summarization. However, the results for the three raters involved in this experiment were inconclusive. We release the reference summaries as a publicly available dataset.
We evaluate corpus-based measures of linguistic complexity obtained using Universal Dependencies (UD) treebanks. We propose a method of estimating robustness of the complexity values obtained using a given measure and a given treebank. The results indicate that measures of syntactic complexity might be on average less robust than those of morphological complexity. We also estimate the validity of complexity measures by comparing the results for very similar languages and checking for unexpected differences. We show that some of those differences that arise can be diminished by using parallel treebanks and, more importantly from the practical point of view, by harmonizing the languagespecific solutions in the UD annotation.
Warriner, Shore, Schmidt, Imbault, and Kuperman, Canadian Journal of Experimental Psychology, 71; 71–88 (2017) have recently proposed a slider task in which participants move a manikin on a computer screen toward or further away from a word, and the distance (in pixels) is a measure of the word’s valence. Warriner, Shore, Schmidt, Imbault, and Kuperman, Canadian Journal of Experimental Psychology, 71; 71–88 (2017) showed this task to be more valid than the widely used rating task, but they did not examine the reliability of the new methodology. In this study we investigated multiple aspects of this task’s reliability. In Experiment 1 (Exps. 1.1–1.6), we showed that the sliding scale has high split-half reliability (r = .868 to .931). In Experiment 2, we also showed that the slider task elicits consistent repeated responses both within a single session (Exp. 2: r = .804) and across two sessions separated by one week (Exp. 3: r = .754). Overall, the slider task, in addition to having high validity, is highly reliable.
Collective behaviors are observed throughout nature, from bacterial colonies to human societies. Important theoretical breakthroughs have recently been made in understanding why animals produce group behaviors and how they coordinate their activities, build collective structures, and make decisions. However, standardized experimental methods to test these findings have been lacking. Notably, easily and unambiguously determining the membership of a group and the responses of an individual within that group is still a challenge. The radial arm maze is presented here as a new standardized method to investigate collective exploration and decision-making in animal groups. This paradigm gives individuals within animal groups the opportunity to make choices among a set of discrete alternatives, and these choices can easily be tracked over long periods of time. We demonstrate the usefulness of this paradigm by performing a set of refuge-site selection experiments with groups of fish. Using an open-source, robust custom image-processing algorithm, we automatically counted the number of animals in each arm of the maze to identify the majority choice. We also propose a new index to quantify the degree of group cohesion in this context. The radial arm maze paradigm provides an easy way to categorize and quantify the choices made by animals. It makes it possible to readily apply the traditional uses of the radial arm maze with single animals to the study of animal groups. Moreover, it opens up the possibility of studying questions specifically related to collective behaviors.
Theory-driven text analysis has made extensive use of psychological concept dictionaries, leading to a wide range of important results. These dictionaries have generally been applied through word count methods which have proven to be both simple and effective. In this paper, we introduce Distributed Dictionary Representations (DDR), a method that applies psychological dictionaries using semantic similarity rather than word counts. This allows for the measurement of the similarity between dictionaries and spans of text ranging from complete documents to individual words. We show how DDR enables dictionary authors to place greater emphasis on construct validity without sacrificing linguistic coverage. We further demonstrate the benefits of DDR on two real-world tasks and finally conduct an extensive study of the interaction between dictionary size and task performance. These studies allow us to examine how DDR and word count methods complement one another as tools for applying concept dictionaries and where each is best applied. Finally, we provide references to tools and resources to make this method both available and accessible to a broad psychological audience.
In this paper, we describe our project of building a phrase structure treebank for Persian. The treebank consists of approximately 30000 sentences. With the help of this treebank, the researcher can investigate syntactic phenomena, extract grammars, train and test parsers, etc. In addition to these motivations, as another advantage of it we can refer to the fact that the sentences of this treebank are selected from an available dependency treebank. So the final treebank has two syntactic representations: phrase structure and dependency structure. The treebank is built using a bootstrapping approach, which converts a dependency structure tree to a phrase structure tree and the annotations are corrected manually. Using the new phrase structure treebank, we train models for constituency parsers. The treebank is freely available for educational purposes <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>.
In this study we developed and evaluated a crowdsourcing-based latent semantic analysis (LSA) approach to computerized summary scoring (CSS). LSA is a frequently used mathematical component in CSS, where LSA similarity represents the extent to which the to-be-graded target summary is similar to a model summary or a set of exemplar summaries. Researchers have proposed different formulations of the model summary in previous studies, such as pregraded summaries, expert-generated summaries, or source texts. The former two methods, however, require substantial human time, effort, and costs in order to either grade or generate summaries. Using source texts does not require human effort, but it also does not predict human summary scores well. With human summary scores as the gold standard, in this study we evaluated the crowdsourcing LSA method by comparing it with seven other LSA methods that used sets of summaries from different sources (either experts or crowdsourced) of differing quality, along with source texts. Results showed that crowdsourcing LSA predicted human summary scores as well as expert-good and crowdsourcing-good summaries, and better than the other methods. A series of analyses with different numbers of crowdsourcing summaries demonstrated that the number (from 10 to 100) did not significantly affect performance. These findings imply that crowdsourcing LSA is a promising approach to CSS, because it saves human effort in generating the model summary while still yielding comparable performance. This approach to small-scale CSS provides a practical solution for instructors in courses, and also advances research on automated assessments in which student responses are expected to semantically converge on subject matter content.
The broad use of computer-supported collaborative-learning (CSCL) environments (e.g., instant messenger–chats, forums, blogs in online communities, and massive open online courses) calls for automated tools to support tutors in the time-consuming process of analyzing collaborative conversations. In this article, the authors propose and validate the cohesion network analysis (CNA) model, housed within the ReaderBench platform. CNA, grounded in theories of cohesion, dialogism, and polyphony, is similar to social network analysis (SNA), but it also considers text content and discourse structure and, uniquely, uses automated cohesion indices to generate the underlying discourse representation. Thus, CNA enhances the power of SNA by explicitly considering semantic cohesion while modeling interactions between participants. The primary purpose of this article is to describe CNA analysis and to provide a proof of concept, by using ten chat conversations in which multiple participants debated the advantages of CSCL technologies. Each participant’s contributions were human-scored on the basis of their relevance in terms of covering the central concepts of the conversation. SNA metrics, applied to the CNA sociogram, were then used to assess the quality of each member’s degree of participation. The results revealed that the CNA indices were strongly correlated to the human evaluations of the conversations. Furthermore, a stepwise regression analysis indicated that the CNA indices collectively predicted 54% of the variance in the human ratings of participation. The results provide promising support for the use of automated computational assessments of collaborative participation and of individuals’ degrees of active involvement in CSCL environments.
RESUMEN EN CASTELLANO Esta tesis contribuye al estudio de analisis linguistico del ingles antiguo con bases de datos lexicas basadas en corpus. Aunque la lematizacion es considerada una de las tareas necesarias para la creacion de diccionarios, no se dispone de corpus lematizados en ingles antiguo. Ademas, en el caso de este periodo historico del ingles, que presenta numerosas variantes morfologicas y carece de estandar ortografico, es imprescindible disponer de un corpus lematizado. Por ello, el objetivo de esta tesis es lematizar una parte del lexico verbal derivado del ingles antiguo, lo que combina aspectos de morfologia, lexicografia y analisis de corpus. El alcance se restringe a las clases verbales mas complejas morfologicamente del ingles antiguo, verbos irregulares y verbos reduplicativos, que incluyen los preterito-presentes, los anomalos, los contractos y los fuertes de la clase VII. Esto requiere, en primer lugar, la seleccion y el manejo de las fuentes de datos y de verificacion de resultados, y en segundo lugar, la formulacion y secuenciado de los pasos de las tareas de lematizacion. Este trabajo tambien plantea la cuestion de la automatizacion en el proceso de la lematizacion, sobre la que escasa bibliografia se ha encontrado. La metodologia combina busquedas automaticas en el lematizador Norna y la revision manual de los resultados con las fuentes lexicograficas disponibles. El lematizador esta basado en la version 2004 del corpus de The Dictionary of Old English (DOE), que contiene aproximadamente tres mil textos y tres millones de palabras. Las fuentes lexicograficas consultadas son, por un lado, la base de datos The Grid (Nerthus Project), y por otro lado, los diccionarios de ingles antiguo, icluyendo el DOE, Bosworth and Toller, Hall-Meritt, and Sweet. Se han tenido en cuenta dos enfoques diferentes para la lematizacion en esta investigacion. Los verbos fuertes de la clase VII se han lematizado aplicando un algoritmo de busqueda basado en las formas principales del verbo (Metola Rodriguez 2015). Este algoritmo se ha creado a partir de los radicales, las flexiones y los elementos preverbales de los verbos fuertes del ingles antiguo. Por otra parte, los verbos derivados de los preterito-presentes, contractos y anomalos se han buscado a partir de sus formas simples. En conclusion, esta tesis ofrece un inventario de lemas y formas flexivas de los verbos analizados. Desde el punto de vista de la aplicabilidad, este trabajo presenta diferentes procedimientos de lematizacion automatica y manual que pueden ser aplicados a los campos de la lexicografia y la linguistica de corpus. RESUMEN EN INGLES This thesis contributes to the research in the linguistic analysis of Old English with corpus-based lexical databases. Although lemmatisation is generally accepted as one of the necessary tasks of dictionary making, no lemmatised corpus is available in Old English. In the specific area of Old English, which presents numerous morphological variations and lacks a written standard, a lemmatised corpus is necessary. Thus, the aim of this thesis is to lemmatise a part of the derived verbal lexicon of Old English, combining aspects of Morphology, Lexicography and Corpus Analysis. The scope is restricted to the most morphologically complex verbal classes of Old English, including irregular verbs and reduplicative verbs, which comprise preterite-present, anomalous, contracted and strong VII verbs. This aim requires, firstly, the selection and management of the sources of data and verification of results; and secondly, the design and sequencing of the steps of the lemmatisation tasks. This research also raises the issue of the automatisation of the process of lemmatisation of Old English verbs, on which little previous literature has been found. The methodology comprises automatic searches on the lemmatiser Norna and the manual revision of the hits with the available lexicographical sources. The lemmatiser is based on the 2004 version of The Dictionary of Old English Corpus (DOE), which contains approximately three thousand texts and three million words. The lexicographical sources checked are, in the first place, the database The Grid (Nerthus Project), and secondly, the Old English dictionaries, including the DOE, Bosworth and Toller, Hall-Meritt, and Sweet. Two different approaches to lemmatisation have been taken in this research. On the one hand, the class VII strong verbs are lemmatised by means of a search algorithm that is based on the main forms of the verbs (Metola Rodriguez 2015). The search algorithm is created on the basis on the roots, the set of inflections and the preverbal items of the strong verbs of Old English. On the other hand, the derived preterite-present, anomalous and contracted verbs are searched by means of their simplexes. In conclusion, this thesis offers an inventory of inflectional forms and lemmas of the verbs under analysis. On the applied side, this work presents different procedures of automatic and manual lemmatisation that can be applied to the fields of Lexicography and Corpus Linguistics.
В статье представлен обзор исследований, посвященных проблеме территориальной диф- ференциации языка в ортологическом аспекте. Установлено, что в связи с языковой нормой лин- гвисты выделяют три типа регионального варьирования: 1) сосуществование отдельных локаль- но маркированных единиц в одной нормативной системе; 2) дивергенцию диатопических орто- логических комплексов; 3) взаимодействие нормативных реализаций с их ненормативными диалектными аналогами в рамках определенного национального языка. Особое внимание уделя- ется ортологическому подходу к изучению регионального варьирования в отечественной лин- гвистике, а также особенностям использования локализмов в современной российской массовой коммуникации.
Past studies examining how people judge faces for trustworthiness and dominance have suggested that they use particular facial features (e.g. mouth features for trustworthiness, eyebrow and cheek features for dominance ratings) to complete the task. Here, we examine whether eye movements during the task reflect the importance of these features. We here compared eye movements for trustworthiness and dominance ratings of face images under three stimulus configurations: Small images (mimicking large viewing distances), large images (mimicking face to face viewing), and a moving window condition (removing extrafoveal information). Whereas first area fixated, dwell times, and number of fixations depended on the size of the stimuli and the availability of extrafoveal vision, and varied substantially across participants, no clear task differences were found. These results indicate that gaze patterns for face stimuli are highly individual, do not vary between trustworthiness and dominance ratings, but are influenced by the size of the stimuli and the availability of extrafoveal vision.
In this paper, we propose Dynamic Self-Attention (DSA), a new self-attention mechanism for sentence embedding. We design DSA by modifying dynamic routing in capsule network (Sabouretal.,2017) for natural language processing. DSA attends to informative words with a dynamic weight vector. We achieve new state-of-the-art results among sentence encoding methods in Stanford Natural Language Inference (SNLI) dataset with the least number of parameters, while showing comparative results in Stanford Sentiment Treebank (SST) dataset.
Deep neural networks (DNNs) have achieved impressive predictive performance due to their ability to learn complex, non-linear relationships between variables. However, the inability to effectively visualize these relationships has led to DNNs being characterized as black boxes and consequently limited their applications. To ameliorate this problem, we introduce the use of hierarchical interpretations to explain DNN predictions through our proposed method, agglomerative contextual decomposition (ACD). Given a prediction from a trained DNN, ACD produces a hierarchical clustering of the input features, along with the contribution of each cluster to the final prediction. This hierarchy is optimized to identify clusters of features that the DNN learned are predictive. Using examples from Stanford Sentiment Treebank and ImageNet, we show that ACD is effective at diagnosing incorrect predictions and identifying dataset bias. Through human experiments, we demonstrate that ACD enables users both to identify the more accurate of two DNNs and to better trust a DNN's outputs. We also find that ACD's hierarchy is largely robust to adversarial perturbations, implying that it captures fundamental aspects of the input and ignores spurious noise.
Failing to recognize one's mirror image can signal an abnormality in one's sense of self. In dissociative identity disorder (DID), individuals often report that their mirror image can feel unfamiliar or distorted. They also experience some of their own thoughts, emotions, and bodily sensations as if they are nonautobiographical and sometimes as if instead, they belong to someone else. To assess these experiences, we designed a novel backwards masking paradigm in which participants were covertly shown their own face, masked by a stranger's face. Participants rated feelings of familiarity associated with the strangers' faces. 21 control participants without trauma-generated dissociation rated masks, which were covertly preceded by their own face, as more familiar compared to masks preceded by a stranger's face. In contrast, across two samples, 28 individuals with DID and similar clinical presentations (DSM-IV Dissociative Disorder Not Otherwise Specified type 1) did not show increased familiarity ratings to their own masked face. However, their familiarity ratings interacted with self-reported identity state integration. Individuals with higher levels of identity state integration had response patterns similar to control participants. These data provide empirical evidence of aberrant self-referential processing in DID/DDNOS and suggest this is restored with identity state integration.
We present PAWS, a multi-lingual parallel treebank with coreference annotation. It consists of English texts from the Wall Street Journal translated into Czech, Russian and Polish. In addition, the texts are syntactically parsed and word-aligned. PAWS is based on PCEDT 2.0 and continues the tradition of multilingual treebanks with coreference annotation. The paper focuses on the coreference annotation in PAWS and its language-specific differences. PAWS offers linguistic material that can be further leveraged in cross-lingual studies, especially on coreference.
We propose a novel neural network model for joint part-of-speech (POS) tagging and dependency parsing. Our model extends the well-known BIST graph-based dependency parser (Kiperwasser and Goldberg, 2016) by incorporating a BiLSTM-based tagging component to produce automatically predicted POS tags for the parser. On the benchmark English Penn treebank, our model obtains strong UAS and LAS scores at 94.51% and 92.87%, respectively, producing 1.5+% absolute improvements to the BIST graph-based parser, and also obtaining a state-of-the-art POS tagging accuracy at 97.97%. Furthermore, experimental results on parsing 61 "big" Universal Dependencies treebanks from raw texts show that our model outperforms the baseline UDPipe (Straka and Straková, 2017) with 0.8% higher average POS tagging score and 3.6% higher average LAS score. In addition, with our model, we also obtain state-of-the-art downstream task scores for biomedical event extraction and opinion analysis applications. Our code is available together with all pre-trained models at: https://github.com/datquocnguyen/jPTDP
We unify recent neural approaches to one-shot learning with older ideas of associative memory in a model for metalearning. Our model learns jointly to represent data and to bind class labels to representations in a single shot. It builds representations via slow weights, learned across tasks through SGD, while fast weights constructed by a Hebbian learning rule implement one-shot binding for each new task. On the Omniglot, Mini-ImageNet, and Penn Treebank one-shot learning benchmarks, our model achieves state-of-the-art results.
بنك المشجّرات محلّل حاسوبيّ للظّواهر التّركيبيّة في اللّغة العربيّة، استثمر مبادئ نظريّة التّحكّم والرّبط التّوليديّة، وحوسباتها وتصوّراتها للنّحو الكلّيّ، غايته في ذلك بناء نظام حوسبيّ آليّ، يحاكي في اشتغاله النّظام الحوسبيّ اللّغويّ الطّبيعيّ. وقد حقّق بنك المشجّرات نتائج مهمّة في هذا الشّأن، تتمثّل في بلوغه الانتظام والتّناسق في معالجة الأبنية الإعرابيّة، لكنّ العمل لم يخل من هنات، أهمّها عدم اتّسام السّيرورة الاشتقاقيّة بالخاصّيّة التّكراريّة المميّزة للّغة البشريّة، وخرق حوسبة النّقل للقيود الجزبريّة التي أقرّتها النّظريّة اللّسانيّة، وهو ما يجعلنا نشكّك في كفايته الوصفيّة لسانيّا.
The contrast between the contextual and general meaning of a word serves as an important clue for detecting its metaphoricity. In this paper, we present a deep neural architecture for metaphor detection which exploits this contrast. Additionally, we also use cost-sensitive learning by re-weighting examples, and baseline features like concreteness ratings, POS and WordNet-based features. The best performing system of ours achieves an overall F1 score of 0.570 on All POS category and 0.605 on the Verbs category at the Metaphor Shared Task 2018.
This paper is an analysis of the Hill‟s Strategy Development Framework and application of the framework to a Fast Food Restaurant business which is operating in Brunei Darussalam.The study of this article concentrates on corporate objectives, marketing strategy, order qualifiers, order winners, and the operations strategy within the company.A review of the relevant literature conducted on corporate goals, competitive priorities, Order Qualifier and Order Winner. The methodology based on a desk review of secondary data and non-participant observations research approach.The findings demonstrated that Fast Food Restaurant Business in Brunei Darussalam required more considerable attention to focus on improving the Customer Service Relationship (CSR) value.It is crucial to concentrate on CSR that would enhance the business brand value, a better image rating and thereby contributing to company‟s sales and gaining a new customer.The study recommended that Fast Food Restaurant Business create a new market such as selling the frozen product.Fast Food Restaurant Business potentially can sell this in their store or export their patent right on the product internationally.The contribution of this paper is to provide provides a review of the application of the strategic framework to uncover the potential customer benefits package of the strategy.
Anger is considered a unique high-arousal and approach-related negative emotion. The influence of individual differences in trait anger on the processing of visual stimuli is relevant to questions about emotional processing and remains to be explored. Using functional magnetic resonance imaging (fMRI), we explored the neural responses to standardized images, selected based on valence and arousal ratings in a group of men with high trait anger compared to those with normative to low anger scores (controls). Results show increased activation in the left-lateralized ventral fronto-parietal attention network to unpleasant images by individuals with high trait anger. There was also a group by arousal interaction in the left thalamus/pulvinar such that individuals with high trait anger had increased pulvinar activation to the high-arousal (versus low arousal) unpleasant images as compared to controls. Thus, individual differences in trait anger in men are associated with brain regions subserving executive attentional and sensory integration during the processing of unpleasant emotional stimuli, particularly to high arousal images.
English in Singapore has always presented a balancing act for its founders. The colonial era saw a distinct role for English, i.e. to produce English-speaking officers for the British administration, while modern Singapore sees English being used as both a national and international lingua franca and as a major language that connects the island city-state to the world. ‘English-knowing bilingualism’ has gained ascendancy in Singapore and may become a core competency for the 21st-century world with the rise in status of English as a global language. However, the path to English-knowing bilingualism in the pluri-lingual and heterogeneous country was often marked by paradoxical debates surrounding the issues of language maintenance and shift, identity and the transmission of values, equity and meritocracy, as well as balancing between local versus global linguistic norms and standards. This paper focuses on the continuing debates, from the past to the present, as new challenges arise and argues how a new balance has to be achieved in the language strategy, policy and management for future-readiness in Singapore.