Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The use of Internet panels to collect survey data is increasing because it is cost-effective, enables access to large and diverse samples quickly, takes less time than traditional methods to obtain data for analysis, and the standardization of the data collection process makes studies easy to replicate. A variety of probability-based panels have been created, including Telepanel/CentERpanel, Knowledge Networks (now GFK KnowledgePanel), the American Life Panel, the Longitudinal Internet Studies for the Social Sciences panel, and the Understanding America Study panel. Despite the advantage of having a known denominator (sampling frame), the probability-based Internet panels often have low recruitment participation rates, and some have argued that there is little practical difference between opting out of a probability sample and opting into a nonprobability (convenience) Internet panel. This article provides an overview of both probability-based and convenience panels, discussing potential benefits and cautions for each method, and summarizing the approaches used to weight panel respondents in order to better represent the underlying population. Challenges of using Internet panel data are discussed, including false answers, careless responses, giving the same answer repeatedly, getting multiple surveys from the same respondent, and panelists being members of multiple panels. More is to be learned about Internet panels generally and about Web-based data collection, as well as how to evaluate data collected using mobile devices and social-media platforms.
Imagination inflation is where imaginative elaboration of possible childhood experiences inflates (increases) participants’ estimation that these events actually occurred, as indicated by pre- to post-manipulation ratings changes. This research primarily uses the Life Events Inventory (LEI), listing possible experiences that could have happened during childhood (Garry, Manning, Loftus, & Sherman, Psychonomic Bulletin & Review, 3, 208–214, 1996). Although imagination inflation research has spawned more than 50 investigations, no normative ratings exist on individual items contained in the LEI. To address this, we present descriptive statistics (mean, median, standard deviation, confidence interval) for 124 LEI items on occurrence (how likely is it that this experience happened to you), plausibility (how plausible is it that this event could have happened to someone), and desirability (how desirable is this experience). Occurrence and plausibility showed similar patterns of mean item ratings and were highly correlated, whereas desirability was moderately correlated with plausibility and unrelated to occurrence. These data should facilitate a more informed selection of specific LEI items to use in further research and can assist in clarifying the contributions of normative occurrence, plausibility, and desirability to imagination inflation effects.
Past research finds that people prefer to sit next to others who are similar to them in a variety of dimensions such as race, sex, and physical appearance. This preference for similarity in seating arrangements is called aggregation and is most commonly measured with the aggregation index (Campbell, Kruskal, & Wallace, Sociometry 29, 1–15, 1966). The aggregation index compares the observed dissimilarity in seating with the amount of dissimilarity that would be expected if seats were chosen randomly. However, the current closed-form equations for this method limit the ease, flexibility, and inferences that researchers have. This paper presents a new approach for studying aggregation that uses bootstrapped resampling of the seating environment to estimate the aggregation index parameters. This method, compiled as an executable program, SocialAggregation, reads a seating chart matrix provided by the researcher and automatically computes the observed number of dissimilar adjacencies, and simulates random seating preferences. The current method’s estimates not only converge with those of the original method, but it also handles a wider variety of situations and also allows for more precise hypothesis testing by directly modeling the distribution of the seating arrangements. Developing a better measure of aggregation opens new possibilities for understanding intergroup biases, and allows researchers to examine aggregation more efficiently.
Mutual eye contact is a key aspect that accompanies any social interaction. Through mutual gaze we establish a communicative link with another person and inform him/her of our goals and motivations. While much attention has been directed to studying the mechanisms of perception / classification of gaze direction, we know very little on the temporal aspects of mutual gaze. We have all likely experienced instances of uncomfortable eye contact, while for example speaking with a stranger or standing in front of someone in an elevator. Here we studied in a very large subject pool (>400 participants) what constitutes a preferred time of mutual eye contact and how these estimates of preferred mutual gaze relate to participant eye behaviour. Participants viewed movies of actors (4 female, 4 male) establishing eye contact with them for variable amounts of time. At the end of each movie participants classified the period of mutual gaze as being “uncomfortably short” or “uncomfortably long”, thus yielding an estimate of “preferred” time of mutual gaze. We also collected ratings on a set face traits of the actor viewed in the clips. We found that the preferred period of mutual eye contact varied as a function of subjective ratings of actor threat, trustworthiness & attractiveness. Threatening faces were associated with lower periods of preferred eye contact, while conversely trustworthy faces were associated with longer periods of preferred mutual gaze. Analysis of patterns of eye fixations showed that fixations tend to be more concentrated in the actor’s eye region in participants exhibiting longer preferred periods of mutual gaze, suggesting that these participants are more likely to reciprocate the eye behaviour of the actor. Finally we also observed that the concentration of fixations in the actor’s eye region was also associated with higher dominance ratings. Meeting abstract presented at VSS 2015
Tswana, a Bantu language in the Sotho group, is characterised by an agglutinative morphology and a disjunctive orthography, which mainly affects the verb category. In particular, verbal prefixes are usually written disjunctively, while suffixes follow a conjunctive writing style. Therefore, Tswana tokenisation cannot be based solely on whitespace, as is the case in many alphabetic, segmented languages, including the conjunctively written Nguni group of South African Bantu languages. This paper shows how a combination of two finite state tokeniser transducers and a finite state morphological analyser are combined to solve the Tswana (verb) tokenisation problem. The approach has the important advantage of bringing the processing of Tswana, beyond the morphological analysis level, in line with what is appropriate for the Nguni languages. This means that the challenge of the disjunctive orthography is met at the tokenisation/morphological analysis level and does not in principle propagate to subsequent levels of analysis such as POS tagging and shallow parsing, etc. The tokenisation approach is novel and, when implemented and evaluated, yields an F1-score of 95 % with respect to a hand tokenised gold standard.
The main goal of this research is to build a sentiment analysis system which automatically determines user opinions of the Stanford Sentiment Treebank in terms of three sentiments such as positive, negative, and neutral. Firstly, sentiment sentences are POS tagged and parsed to dependency structures. All nodes of the Treebank and their polarities are automatically extracted from the Treebank. We train two Support Vector Machines models. One is for a node level classification and the other is for a sentence level. We have tried various type of features such as word lexicons, POS tags, Sentiment lexicons, head-modifier relations, and sibling relations. Though we acquired 74.2% in accuracy on the test set for 3 class node level classification and 67.0% for 3 class sentence level classification, our experimental results for 2 class classification are comparable to those of the state of art system using the same corpus.
Developing an app version of a printed dictionary is a new challenge faced by lexicographers. Lexicographers involved in the app development process must consider fundamental lexicographic aspects as well as learn to understand technological and usage issues inherent to the new media. An inevitable question is how closely the content and layout can be made to match the printed dictionary while still offering ‘digital’ functionality such as linking, collapsed sections, audio, etc. Only a few reports discussing these issues have so far been published. The aim of our paper is to further advance the exchange of knowledge and experience by sharing our observations made during the development of a new app corresponding to the comprehensive printed dictionary, Svensk ordbok utgiven av Svenska Akademien (the Contemporary Dictionary of the Swedish Academy, 2009). The app is the result of close cooperation between the financer The Swedish Academy, lexicographers and system developers at the Department of Swedish, University of Gothenburg and Isolve AB, a Stockholm-based app development agency specializing in dictionary apps.
Theory of Mind (ToM) has repeatedly been defined as the ability to understand that others believe their own things based on their own subjective interpretations and experiences, and that their thoughts are determined independently from your own. In this study, we wanted to see if individual differences in ToM are capable of causing different perceptions of an individual's interactions with human like robotics and highlight whether or not individual differences in ToM account for different levels of how individuals experience what is called the "Uncanny Valley phenomenon" and to see whether or not having a fully developed theory of mind is essential to the perception of the interaction. This was assessed by inquiring whether or not individuals with Autism Spectrum Disorder (ASD) perceive robotics and artificially intelligent technology in the same ways that typically developed individuals do; we focused on the growing use of social robotics in ASD therapies. Studies have indicated that differences of ToM exist between individuals with ASD and those who are typically developed. Comparably, we were also curious to see if differences in empathy levels also accounted for differences in ToM and thus a difference in the perceptions of human like robotics. A robotic image rating survey was administered to a group of University of central Florida students, as well as 2 surveys - the Autism Spectrum Quotient (ASQ) and the Basic Empathy Scale (BES), which helped optimize a measurement for theory of mind. Although the results of this study did not support the claim that individuals with ASD do not experience the uncanny valley differently than typically developed individuals, there were significant enough results to conclude that different levels of empathy may account for individual differences in the uncanny valley. People with low empathy seemed to have experienced less of an uncanny valley feeling, while people with higher recorded empathy showed to experience more of an uncanny valley sensitivity.
One key criterion when creating a representation of the lexicon of any language within a dictionary or lexical database is that it must be decided which groups of idiosyncratic and systematically modified variants together form a lexeme. Few researchers have, however, attempted to outline such principles as they might apply to sign languages. As a consequence, some sign language dictionaries and lexical databases appear to be mixed collections of phonetic, phonological, morphological, and lexical variants of lexical signs (e.g. Brien 1992) which have not addressed what may be termed as the lemma dilemma. In this paper, we outline the lemmatisation practices used in the creation of BSL SignBank (Fenlon, Cormier et al. 2014), a lexical database and dictionary of British Sign Language based on signs identified within the British Sign Language Corpus ( http://www.bslcorpusproject.org ). We argue that the principles outlined here should be considered in the creation of any sign language lexical database and ultimately any sign language dictionary and reference grammar.
This paper is interested in the assisted interrogation of lexical databases designed according to the LMF standard (Lexical Markup Framework) ISO-24613. The proposed solution is based on a requirement-based lexical web service generation approach that makes easier the task of engineers when developing NLP (Natural Language Processing) systems. Using this approach, the developer will not deal with the database content or its structure. Also, he will not use any language query. For each lexical requirement, we generate through a lexical web service for interrogating LMF databases. Also we generate with it its enhanced WSDL description file that we have enriched by semantic data describing data categories used in the requirement specification.
In a multi-lingual country like India where every language has its own phonological system, there is a need for language specific articulation test. Although in recent times there has been increasing awareness among parents for early intervention in children with articulation problems within the regional areas of the nation, the availability of articulation tests in the regional languages is very limited. The present study makes a preliminary attempt at developing an assessment tool to assess the articulatory skills of Tulu speaking children, who form a significant population in South India. Word list was developed based on familiarity rating and was administered on 50 children, aged 3-8 years. The target speech sounds were embedded in words which were presented in picture form to elicit responses from the participants. The responses were analysed qualitatively and in terms of production accuracy across age groups.
The paper refers to the specifics of the French-Polish automated translation. The author summarizes briefly the rules of creating electronic lexical databases in the context of the Object Oriented Approach. This particular approach, created for the purpose of translation, especially for automated translation, by Wiesław Banyś, offers a specific approach towards the language description with the help of object classes seen through their semantic relations. The purpose of this description is to determine correctly the meaning of the words and polysemic expressions of the source language in the translation into the target language. Therefore, the main question of the paper concerns the role of frames and/or scripts criterion in the process of the natural language words’ disambiguation for automated translation. By analysing the example of the French word collier the Author shows how it is possible to solve the problem of near synonyms in automated translation with the use of the above-mentioned criterion.
Parsing models for all Universal Depenencies 1.2 Treebanks, created solely using UD 1.2 data (http://hdl.handle.net/11234/1-1548). To use these models, you need Parsito binary, which you can download from http://hdl.handle.net/11234/1-1584.
The article introduces two internet sources designated to the study of Older Czech language (13th to 18th centuries); both have been designed and run by The Department of Language Development at The Institute of the Czech Language at the Academy of Sciences of the Czech Republic. The first source, Vokabulář webový [Web Vocabulary] (http://vokabular.ujc.cas.cz), makes texts, images and audio materials available to the study of Older Czech language. The accessible materials are, primarily, both modern and historical dictionaries, amongst which the most salient is the, gradually growing, Elektronický slovník staré češtiny [Electronic Old-Czech Vocabulary] that treats Old-Czech lexicon from the dawn of Czech language to the end of the 15th century. Furthermore, Vokabulář includes electronic editions of the works originating in the period from the 13th century to the beginning of the 19th century, presented both as continuous texts and in the corpus version; digitalized copies of Older-Czech grammar books; basic scientific literature; audiobooks of Older-Czech texts; and software tools utilized for the work with historical texts. The second source is Lexikální databáze hu-manistické a barokní češtiny [Lexical Database of Humanistic and Baroque Czech] (http://madla.ujc.cas.cz). It records the Czech vocabulary of the 16th to 18th centuries based on the excerption of the authentic contemporary texts (both old prints and manuscripts): Lexical database illustrates the Czech vocabulary with direct quotations, including stating the source. Thus, Lexical Database partly substitutes the missing Czech vocabulary of the mentioned period.
Children entering institutional education can activate the linguistic and non-linguistic norms they bring from home. These norms, however, are very diverse in their nature and children from the same age-group remain at different levels in terms of language-acquisition. In our paper we seek to identify where and what sort of problems may arise in acquiring the mother-tongue, what factors may hinder the acquisition of the first language, what are the symptoms of backwardness in the field of linguistic–communication and which areas measuring tests tend to focus upon. Finally by presenting an indication system we would like to show what opportunities observation may have in purposeful development.Keywords: native language acquisition, linguistic competence, measurementof communication skills, observation of children’s language performance.
This work describes a system that performs morphological analysis and generation of Pali words. The system works with regular inflectional paradigms and a lexical database. The generator is used to build a collection of inflected and derived words, which in turn is used by the analyzer. Generating and storing morphological forms along with the corresponding morphological information allows for efficient and simple look up by the analyzer. Indeed, by looking up a word and extracting the attached morphological information, the analyzer does not have to compute this information. As we must, however, assume the lexical database to be incomplete, the system can also work without the dictionary component, using a rule-based approach.
The article discusses a number of relatively little-known aspects of Marcus Tullius Cicero’s linguistic views. Special attention is paid to his ideas about language learning, issues of purity of language, as well as linguistic norms and anomalies. These conclusions are of interest to researchers in the field of ancient linguistic thought.
This paper analyses several points of interlingual dependency mismatch on the material of a parallel Czech-English dependency treebank. Particularly, the points of alignment mismatch between the valency frame arguments of the corresponding verbs are observed and described. The attention is drawn to the question whether such mismatches stem from the inherent semantic properties of the individual languages, or from the character of the used linguistic theory. Comments are made on the possible shifts in meaning. The authors use the findings to make predictions about possible machine translation implementation of the data.
This paper examines L2 learners’ familiarity and correct use of formulaic sequences in English. In particular, the extent to which L2 proficiency level and collocational frequency affected L2 learners’ knowledge and use was investigated. Thirty L2 learners of English were tested on their correct use of 32 formulaic sequences in English, made of up ‘Verb + out’ and asked to rate their familiarity with these formulaic sequences. The results showed that the familiarity ratings given by the low L2 proficiency participants showed a significant positive correlation with their accuracy scores, but no such correlation was found for the high proficiency group. Familiarity ratings correlated positively with both word frequency and formulaic sequence frequency, but neither word frequency nor formulaic sequence frequency showed a significant correlation with accuracy of use, although formulaic sequence frequency showed a trend toward a positive correlation. Effects of frequency and decomposability on various types of formulaic sequences are discussed, along with the disparity in L2 learners’ exposure to formulaic sequences and their ability to use them accurately.
We address some theoretical and practical issues relating to generation, processing, and management of Translation Corpus (TC) in Indian languages, which is developed in a consortiummode project (ILCIII) 1 under the DeitY, Govt. of India. Issues are discussed here for the first time keeping in mind the ready application of TC in various domains of computational and applied linguistics. We first define what is a TC; describe the process of its construction; identify its features; exemplify the processes of text alignment in TC; discuss methods of text analysis; propose for restructuring of translational units; define the process of extraction of translational equivalents; propose for generating bilingual lexical database and TermBank from a structured TC; and finally identify areas where a TC and information extracted from it may be utilized. Since construction of TC in Indian languages is full of hurdles, we try to construct a roadmap with a focus on techniques and methodologies that may be applied for achieving the task. The issues are brought under focus to justify the work that generated TC for some Indian languages for future reference and application. 1. What is a Translation Corpus? Theoretically, a Translation Corpus (TC) suggests that it contains texts and their translation. It is entitled to include bilingual (and multilingual) texts as well as texts that may fit under translation. A TC, by virtue of its character and composition, is made of two parts: a text from a source language (SL) and its translation from a target language (TL) (15) (24), (39). Although, a TC is normally bilingual and bidirectional (28), it can be multilingual and multi� directional as well (37), as it actually happens in case of the ILCII and ILCIII projects for the Indian language s. In these two projects a new strategy is adopted where Hindi is treated as the only SL and several other Indian languages are treated as the TL (Fig. 1). The issue of multidirectionality can be understood if all the target languages can establish linguistic links with each other as they are linked up with SL. Since the ILCII TC has not tried to venture into this direction, it makes sense to keep the present discussion confined within a scheme of bilingualism and bidirectionality, with, for examp le, Hindi
Given the concentration of economic growth and power in science fields and the current levels of racial stratification in schooling, this study examined (1) the effects of race on students’ connectedness to science and career aspirations, (2) the extent to which these effects were moderated by school racial composition and racialized tracking, and (3) the differences in modeling effects using separate variables for race and gender (i.e., White, Black, Hispanic, female) versus race/gender (e.g., White female, Black male, etc.). Using the lens of racial formation theory, this study situated access to science knowledge as a racial project, conferring and denying access to resources along racial lines. Reviews of the literature on science self-efficacy, identity, engagement, and career aspirations revealed an under-emphasis on school institutional factors, such as racial composition and racialized tracking (which are important in sociological literature), as shaping student outcomes. The study analyzed data from the nationally representative High School Longitudinal Study that surveyed students in 2009 during their freshman year in high school and again in 2012 during most students’ junior year (n = 6,998). Affective ratings (in self-efficacy, identity, engagement) and career aspirations for students measured in 2012 were examined as dependent variables and a variable for racialized tracking was estimated given schools’ placement of students in advanced science coursework in 2012. Although school racial composition was not found to moderate race on outcome effects, primary analyses demonstrated that the presence of racialized tracking in the students’ schools did moderate these effects. Overall these results suggested that the student subgroups most often at a disadvantage compared to White students for the science outcomes studied were Hispanic males and females; Black students’ ratings and aspirations were largely on par or exceeded those of their White counterparts. In addition, results indicated that racialized tracking served to exacerbate gaps for Hispanic students and may also diminish career aspirations for Black students. Finally, while examining effects by race/gender did provide some additional insight and nuance in the interpretation of these results, there were clear instances where these more detailed analyses were not needed or may have obscured results that were clearer when aggregated by race. Given these results, implications for policy, practice, and future research are discussed.
This paper explores the interaction between eventive information and morpho-syntax based on Chinese VV compounds. Chinese VV Compounds’ identical morpho-syntactic structure represents different event relations between the two component words and the correct interpretation of the meaning of these compounds relies on the prediction on their event relations. Without overt syntactic clues, we propose that ontology-based conceptual classification can be used to predict the event relation between the two component words. Compounding is the most productive way to research multi-word expressions in Mandarin Chinese. A Mandarin VV compound can be classified according to the eventive relation between two simplex verbs, which specifies how the eventive meanings of the two simplex verbs combine to form the meaning of the compound. The way in which two events combine with each other depends upon their event types, and the three types of eventive relations that we deal with in this paper are coordinate, modificational, and resultative. Using an ontology-based prediction approach, we hypothesized that the eventive relations could be predicted by the conceptual classification of the two simplex verbs’ event types. First, we utilized SUMO and Sinica BOW to classify each simplex verb. Next, the correlation between the ontology-based classification of each verb position and each eventive type was scored using a manually tagged lexical database and a training set was established. Finally, we encoded the ontological information of each VV compound in a 3-tuple based on these correlation scores. This 3-tuple was represented as a three-dimensional vector and was used to predict the eventive type of the new VV compounds. The results of our findings show that the classification experiments on event relation of unknown VV compounds can be reliably predicted based on the ontological classification of their component words.
The vocative is a residuary case in most Indo-European languages, mirroring a particular Proto-Indo-European status. Its syntactical function is preserved in the descendant languages, but the morphological aspects are strongly simplified. In Latin, not unlike the cognate languages, the general tendency is toward a formal overlapping with the nominative case. The Romanian vocative is, in the Romance frame, surprisingly multifarious. It displays four distinct variants: desinence and intonation; desinence, intonation and prolongation of the final vowel; intonation and vowel prolongation; solely intonation. Old Romanian texts attest the tendency of gradually replacing the vocative form with the nominative form, perceived as more expressive. On the other hand, there is an observable development of the formal marks specific to this syntactical function; these marks are only partially inherited from Latin. In nowadays Romanian language the formal specificity of the vocative case is not diminishing – on the contrary, some colloquial vocative forms (not yet acceptable in the frame of the linguistic norm) emphasize an unambiguous linguistic will to maintain this case, while the general tendency is to reduce as much as possible the differences between the actual two cases of the Romanian language, nominative-accusative and genitive-dative.
Online reviews have a profound impact on the customer or “newbie” who wishes to purchase or consume a product via Web 2.0 e-commerce. Online reviews contain features that form half of the analysis in opinion mining. Most of today’s systems work on the basis of summarization, looking at the average obtained features and their sentiments, leading to structured review information being generated. Often, the context surrounding a feature, which helps the sentiment of the review to be classified clearly, is overlooked. The Web 3.0-based machine interpretable Resource Description Framework (RDF) can be used to structure these unstructured reviews into features and sentiments, which are obtained via traditional preprocessing and extraction techniques. Here, data about the context is also provided for future ontology-based analysis, with support from the WordNet lexical database for word sense disambiguation and SentiWordNet scores for sentiment word extraction. Many popular RDF vocabularies are helpful for obtaining such machine-processable data. This work forms the basis for creating/upgrading the (available) OWL Ontology that can be used as a structured data model with rich semantics for supervised machine learning. With this method, the classified sentiment categories are validated in relation to precise sentiments and are sent back to the interface in corresponding “feature/sentiment” pairs so that reviews are filtered clearly, which helps to satisfy the feature set of the customer.
We present our work on semi-supervised parsing of natural language sentences, focusing on multi-source crosslingual transfer of delexicalized dependency parsers. We first evaluate the influence of treebank annotation styles on parsing performance, focusing on adposition attachment style. Then, we present KLcpos3, an empirical language similarity measure, designed and tuned for source parser weighting in multi-source delexicalized parser transfer. And finally, we introduce a novel resource combination method, based on interpolation of trained parser models.
In this paper we present an enhanced algorithm with modified approach to extricate various Triplets i.e. subject-predicate-object from Natural language sentences. The Treebank Structure and the Typed Dependencies obtained from Stanford Parser are used to elicit multiple triplets from English Sentences. Typed Dependencies represents grammatical connections among the words of any sentence and represents how triplets are associated. The intended interpretation behind the extraction of Triplets is that the subject is acting on the object in a way described by the predicate. In graphical form it can be considered that subject and object will be acting as nodes i.e. entities and predicate as edges i.e. relationship. The resulting triplets and relations can be useful for building and analysis of a social network graph and for generating communication pattern and Information retrieval.
The importance of the problem under investigation is determined by the increasing importance of learning English, which sometimes raises a number of linguistic, cultural and pedagogical issues that can be linked with students’ understanding of the English language itself. The purpose of the article is to reveal some historicocultural aspects of language teaching, which include acquisition of knowledge, shaping skills needed for cross-cultural communication; as well as to highlight the most important features to ensure competence-based approach. The leading approach to the study of the problem is systematic, involving basic content of teaching the history of the English language, which represents a merger of two distinct-subdisciplines of linguistics: sociolinguistics and culturology and focusing on cognitive and communicative components of linguoculturological competence. The paper presents an overview that any foreign language should be viewed not only as a system of linguistic norms, but also as a system of social norms and behavior, spiritual values; language is central to historical and social interaction in every society. The materials of this paper can be recommended for use in modern practice of educational institutions, as well as in the system of teacher training.
The low success rate when retrieving information through web searches could be verified virtually in all areas of knowledge, due to the large amount of information available which raises the selection complexity for relevant articles. A query consists in chosen terms to drive the search for related documents. However, if new terms could be added in order to expand the relevance of the search, then there is what is called query semantic enrichment. This paper presents a semantic enrichment model to improve the quality of results for medical articles queries. This model knows the search context by using a repository of articles which is previously subjected to Latent Semantic Analysis and is supported by the National Cancer Institute ontology and the WordNet lexical database. In this way, new terms which are semantically related to the conducted search context, could be proposed to help raising precision when retrieving relevant articles.
Similarity calculation between Business Process Models (BPM) has an important role in the process of managing BPM repository. One of its uses is to facilitate the searching process of a model in the repository. Similarity calculation between business processes is closely related with semantic string similarity. Semantic string similarity is usually performed by utilizing a lexical database, such as WordNet, to find the semantic meaning of words. The problem in WordNet is that this lexical database contains terms wich have more than one meaning or polysemous. Selecting the wrong meaning will decrease the accuracy of similarity calculation process. In this study, we will try to improve the accuracy of similarity calculation of business processes using Word Sense Disambiguation (WSD). The main purpose is to eliminate the ambiguity of polysemous words before calculating the similarity value. WSD is performed by unsupervised methods based on the value of graph connectivity. Then, we used a lexical database that is focused in the business and industry field. The results from this study is able to achieve higher accuracy of the sense selection process for terms especially terms that are related to business and industrial domains. It will also increase the accuracy of similarity value calculation between the business process models.
This paper describes a parsing model that combines the exact dynamic programming of CRF parsing with the rich nonlinear featurization of neural net approaches. Our model is structurally a CRF that factors over anchored rule productions, but instead of linear potential functions based on sparse features, we use nonlinear potentials computed via a feedforward neural network. Because potentials are still local to anchored rules, structured inference (CKY) is unchanged from the sparse case. Computing gradients during learning involves backpropagating an error signal formed from standard CRF sufficient statistics (expected rule counts). Using only dense features, our neural CRF already exceeds a strong baseline CRF model (Hall et al., 2014). In combination with sparse features, our system achieves 91.1 F1 on section 23 of the Penn Treebank, and more generally outperforms the best prior single parser results on a range of languages.
Accurate identification of phrasal translation equivalents is critical to both phrase-based and syntax-based machine translation systems. We show that the extraction of many phrasal translation equivalents is made impossible by word alignments done without taking syntactic structures into consideration. To address the problem, we propose a new annotation scheme where word alignment and the alignment of non-terminal nodes (i.e., phrases) are done simultaneously to avoid conflicts between word alignments and syntactic structures. Relying on this new alignment approach, we construct a Hierarchically Aligned Chinese-English Parallel Treebank (HACEPT), and show that all phrasal translation equivalents can be automatically extracted based on the phrase alignments in HACEPT.
We develop novel first- and second-order features for dependency parsing based on the Google Syntactic Ngrams corpus, a collection of subtree counts of parsed sentences from scanned books. We also extend previous work on surface $n$-gram features from Web1T to the Google Books corpus and from first-order to second-order, comparing and analysing performance over newswire and web treebanks. Surface and syntactic $n$-grams both produce substantial and complementary gains in parsing accuracy across domains. Our best system combines the two feature sets, achieving up to 0.8% absolute UAS improvements on newswire and 1.4% on web text.
Margaret Thatcher was the first woman to become Prime Minister of the UK. It has been claimed, however, that she did little for the cause of women. Part of the problem is Thatcher made clear that while she was a woman she thought of herself as a politician first. In this chapter we consider the linguistic consequences of adopting such a position, and we argue that Thatcher used specific discourse structures conducive to the adversarial style of the British parliament. As this style has been equated with male discourse patterns some argue that Thatcher adopted male linguistic norms. However, adversarial styles are not inherently “male” and we consider whether Thatcher was speaking like a man or merely as a politician.
In this paper, vowel distributions of 3260 English monomorphemic verbs are examined to verify the claim that vowel height plays a considerable role in stress assignment in English. The claim was made by research on English nouns from a lexical database, CELEX(Baayen, Piepenbrock and Gulikers, 1995). It has found that vowel quality, especially vowel height, has a significant influence on stress location. Results of the current study provide further evidence for the claim that vowel height is related to stress location: First, lowness attracts stress; Second, a high front lax vowel, [?], is the second frequent vowel in unstressed syllables following a schwa.; Third, in stressed light syllables, a mid vowel, [?], appears most frequently while a low vowel [ae] is the most common vowel winning stress over a heavy syllable. These findings offer insights into the relations between phonetic properties of vowel phonemes and the nature of stress.
This article explains why XML format has become established as the standard format for multilevel hierarchical structuring of linguistic databases and how an XML Schema can be used to manage the formal structure and content of elements in a dictionary database. Various aspects that must be taken into account when structuring complex dictionary databases in XML format are presented: the lexicographic or content aspect, the practical aspect, and the technical aspect. Decision-making is illustrated with the example of designing an XML Schema for the Dictionary of Slovenian Synonyms.
Redundancy is an important psycholinguistic concept which is often used for explanations of language change, but is notoriously difficult to operationalize and measure. Assuming that the reconstruction of a syntactic structure by a parser can be used as a rough model of the understanding of a sentence by a human hearer, I propose a method for estimating redundancy. The key idea is to compare performances of a parser on a given treebank before and after artificially removing all information about a certain grammeme from the morphological annotation. The change in performance can be used as an estimate for the redundancy of the grammeme. I perform an experiment, applying MaltParser to an Old Church Slavonic treebank to estimate grammeme redundancy in Proto-Slavic. The results show that those Old Church Slavonic grammemes within the case, number and tense categories that were estimated as most redundant are those that disappeared in modern Russian. Moreover, redundancy estimates serve as a good predictor of case grammeme frequencies in modern Russian. The small sizes of the samples do not allow to make definitive conclusions for number and tense.
Paraphrase identification is a semantic text similarity task which is an important part of many natural language processing applications. Existing methods use vector space models, word co-occurrence information, lexical databases, parsers and machine translation (MT) evaluation metrics to find text similarity. However, other aspects such as negations, inverse relations and semantic roles of the sentences are also very much important in identifying paraphrases. Furthermore, the semantics of the sentences are hidden when the sentences are complex. We propose an approach to find similarity between pair of texts by considering all these factors. We have used an approach to determine set of clauses present in the texts by resolving conjunctions in complex sentences that identify hidden triples from the text. The approach extracts clause-based similarity features namely concept score, relation score, proposition score and word score from the texts. We have combined these similarity features along with MT metrics features to identify whether the texts are paraphrases or not using Support Vector Machine model. We have evaluated our methodology to measure the paraphrase similarity for Microsoft Research corpus. The statistical tests namely |$k$|-fold paired |$t$|-test and McNemar's test show that including clause-based features significantly improved the performance. Also, our approach outperforms state-of-the-art methods in terms of accuracy, |$F$|1-measure and |$f$|1-measure.
Previous studies indicate that emotion regulation may occur unconsciously, without the cost of cognitive efforts; and that conscious acceptance effectively reduces the emotional consequences of negative events. However, it has yet to be determined how conscious and unconscious acceptance strategies differ in behavioral and physiological consequences of emotion regulation. As unconscious regulation occurs with little cost of cognitive resources, the current study hypothesizes that unconscious acceptance regulates the emotional consequence of negative events more effectively compared to conscious acceptance. Subjects were randomly assigned to conscious acceptance, unconscious acceptance and control conditions. A frustrating arithmetic task was used to induce negative emotion. Emotional experiences were assessed by the positive affect and negative affect scale (PANAS) while emotion-related physiological activation was assessed by the heart-rate reactivity. The results showed that unconscious acceptance produced less reductions of positive affect ratings compared to conscious acceptance during frustration. In addition, both conscious and unconscious acceptance strategies significantly decreased emotion-related heart-rate activity (to a similar extent) in comparison with the control condition. Moreover, heart-rate reactivity showed a trend of positive correlation with negative affect rating and a trend of negative correlation with positive affect rating during frustration compared to baseline phases. Thus, unconscious acceptance is not only able to decrease emotion-related physiological activity, but also able to produce better emotional experiences compared to conscious acceptance. This suggests that it is practically important to consider unconscious acceptance for emotion regulation in real-life settings.
The spinal tree adjoining grammar (TAG) parsing model of [Carreras 08] achieves the current state-of-the-art constituent parsing accuracy on the commonly used English Penn Treebank evaluation setting. Unfortunately, the model has the serious drawback of low parsing efficiency since its Eisner-CKY style parsing algorithm needs O(n4) computation time for input length n. This paper investigates a more practical solution and presents a beam search shift-reduce algorithm for spinal TAG parsing. Since the algorithm works in O(bn) (b is beam width), it can be expected to provide a significant improvement in parsing speed. However, to achieve faster parsing, it needs to prune a large number of candidates in an exponentially large search space and often suffers from severe search errors. In fact, our experiments show that the basic beam search shift-reduce parser does not work well for spinal TAGs. To alleviate this problem, we extend the proposed shift-reduce algorithm with two techniques: Dynamic Programming of [Huang 10a] and Supertagging. The proposed extended parsing algorithm is about 8 times faster than the Berkeley parser, which is well-known to be fast constituent parsing software, while offering state-of-the-art performance. Moreover, we conduct experiments on the Keyaki Treebank for Japanese to show that the good performance of our proposed parser is language-independent.
Word order differences between source and target languages pose a serious challenge to statistical machine translation (SMT). Pre-ordering, an approach that reorders source words into a target-word-like order as a preprocessing step, has been shown effective in handling word order between different languages and improving translation performance of SMT. In this paper, we propose a novel word reordering method based on the pre-ordering framework. Instead of using a supervised parser trained on a monolingual treebank, our method extracts bilingual structural information for reordering from automatically wordaligned sentence pairs into dependency-tree-like structures, then learns a reordering model by training a dependency parser on this extracted pseudo-treebank. Experiment results show that our pre-ordering method is effective in permuting source words to resemble word order of the target language, and improving translation quality.
In communication a great deal of meaning is exchanged through body language, including gaze, posture, hand gestures and body movements. Body language is largely culture-specific, and rests, for its comprehension, on people's sharing socio-cultural and linguistic norms. In cross-cultural communication, L2 speakers' use of body language may convey meaning that is not understood or misinterpreted by the interlocutors, affecting the pragmatics of communication. In spite of its importance for cross-cultural communication, body language is neglected in ESL/EFL teaching. This paper argues that the study of body language should be integrated in the syllabus of ESL/EFL teaching and learning. This is done by: 1) reviewing literature showing the tight connection between language, speech and gestures and the problems that might arise in cross-cultural communication when speakers use and interpret body language according to different conventions; 2) reporting the data from two pilot studies showing that L2 learners transfer L1 gestures to the L2 and that these are not understood by native L2 speakers; 3) reporting an experience teaching body language in an ESL/EFL classroom. The paper suggests that in multicultural ESL/EFL classes teaching body language should be aimed primarily at raising the students' awareness of the differences existing across cultures.
In this paper we explore different statistical dependency parsers for parsing Telugu. We consider five popular dependency parsers namely, MaltParser, MSTParser, TurboParser, ZPar and Easy-First Parser. We experiment with different parser and feature settings and show the impact of different settings. We also provide a detailed analysis of the performance of all the parsers on major dependency labels. We report our results on test data of Telugu dependency treebank provided in the ICON 2010 tools contest on Indian languages dependency parsing. We obtain state-of-the art performance of 91.8% in unlabeled attachment score and 70.0% in labeled attachment score. To the best of our knowledge ours is the only work which explored all the five popular dependency parsers and compared the performance under different feature settings for Telugu.
Recently, dependency parsing has been used for development of dependency parsers. There are many parsers built in the area of NLP for grammatical information extraction. These parsers can be used to build treebanks which can serve as resources for research purposes. This paper describes the various parsers based on different parsing methodology for different languages. One of the advantages of the dependency parsing is that it resolves ambiguity. In this paper a comparative table of different parsers is proposed for better analysis.