Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Part of Speech (POS) Tagging can be applied by several tools and several programming languages. This work focuses on the Natural Language Toolkit (NLTK) library in the Python environment and the gold standard corpora installable. The corpora and tagging methods are analyzed and com- pared by using the Python language. Different taggers are analyzed according to their tagging ac- curacies with data from three different corpora. In this study, we have analyzed Brown, Penn Treebank and NPS Chat corpuses. The taggers we have used for the analysis are; default tagger, regex tagger, n-gram taggers. We have applied all taggers to these three corpuses, resultantly we have shown that whereas Unigram tagger does the best tagging in all corpora, the combination of taggers does better if it is correctly ordered. Additionally, we have seen that NPS Chat Corpus gives different accuracy results than the other two corpuses.
Abstract This article characterizes aspect‐perception as a distinct form of judgment in Kant's sense: a distinct way in which the mind contacts world and applies concepts. First, aspect‐perception involves a mode of thinking about things apart from any established routine of conceptualizing them. It is thus a form of concept application that is essentially reflection about language. Second, this mode of reflection has an experiential, sometimes perceptual, element: in aspect‐perception, that is, we experience meanings—bodies of norms. Third, aspect‐perception can be “preparatory”: it may help us to decide what linguistic norms to develop and how to conceptualize—make the world thinkable. Fourth, the article discusses the forms of justification for which aspect‐perception allows—the necessity and normativity involved in employing this form of judgment.
We apply the well-known parsing technique of self-training to a new type of text: language-learner text. This type of text often contains grammatical and other errors which can cause problems for traditional treebank-based parsers. Evaluation on a small test set of student data shows improvement over the baseline, both by training on native or non-native text. The main contribution of this paper adds additional support for the claim that the new self-trained parser has improved over the baseline by carrying out a qualitative linguistic analysis of the kinds of differences between two parsers on non-native text. We show that for a number of linguistically interesting cases, the self-trained parser is able to provide better analyses, despite the sometimes ungrammatical nature of the text. 1
International audience
INTRODUCTION: To examine the long-term impact of graphic health-warning labels (GHWL) on adolescents' cognitive processing of warning labels and cigarette pack perceptions. METHODS: Cross-sectional school-based surveys of students aged 13-17 years residing in urban centers, conducted prior to GHWL introduction (2005) and 6 months (2006), 2 years (2008), and 5 years (2011) post-GHWL introduction. Students who had seen a cigarette pack in the previous 6 months or in 2006, who had seen GHWL were included in analyses (2005 n = 2,560; 2006 n = 1,306; 2008 n = 2,303; 2011 n = 2,716). Smoking stage, reported exposure to cigarette packs, cognitive processing of GHWL, and positive and negative perceptions of pack image were assessed. RESULTS: While cognitive processing of GHWL in 2006 and 2008 was greater than 2005 (p <.01), by 2011 scores had returned to 2005 levels. This pattern of change was consistent across smoking status groups. Pack image perceptions became more negative over time among all students, irrespective of smoking experience. While positive pack image ratings were lower in all subsequent years than 2005, the 2008 rating was higher than 2006 (p <.01). A significant interaction between survey time and smoking status (p <.01) showed that significant increases in positive pack ratings after 2006 only occurred among current and experimental smokers. CONCLUSIONS: When novel, GHWL on cigarette packs increase cognitive processing among adolescents. However, this effect diminishes after 5 years, suggesting more regular message refreshment is needed. Australia's adoption of plain packaging is intended to undermine positive pack appeal and increase warning salience.
This study investigates on building a better Chinese word segmentation model for statistical machine translation. It aims at leveraging word boundary information, automatically learned by bilingual character-based alignments, to induce a preferable segmentation model. We propose dealing with the induced word boundaries as soft constraints to bias the continuous learning of a supervised CRFs model, trained by the treebank data (labeled), on the bilingual data (unlabeled). The induced word boundary information is encoded as a graph propagation constraint. The constrained model induction is accomplished by using posterior regularization algorithm. The experiments on a Chinese-to-English machine translation task reveal that the proposed model can bring positive segmentation effects to translation quality.
AbstractIn this paper, we present some statistical data on the distribution of parts of speech and dependency relations in a large manually annotated Hungarian Treebank, the Szeged Dependency Treebank. We hypothesize that the domain of the text influences the distribution of the above elements, thus we pay special attention to differences between domains. We present the characteristic rank-frequency distributions of parts of speech and dependency relations in Hungarian and analyse the domain similarities and differences among sub-corpora as regards the above distributions. Our results reveal that the computer and newspaper texts are most similar to each other while the domains literature and compositions also exhibit some similarities. On the other hand, the business news and the law sub-corpora are unique, both having their own characteristics.
In this paper we explore the relationship between the genre of a text and the types of situations introduced by the clauses of the text, working from the perspective of the theory of discourse modes (Smith, 2003). The typology of situation types distinguishes between, for example, events, states, generic statements, and speech acts. We analyze texts of different genres from two English text corpora, the Penn Discourse TreeBank (PDTB) and the Manually Annotated SubCorpus (MASC) of the Open American National Corpus. Texts of different types – genres in the PDTB and subcorpora in MASC – are segmented into clauses, and each clause is labeled with the type of situation it introduces to the discourse. We then compare the distribution of situation types across different text types, finding systematic differences across genres. Our findings support predictions of the discourse modes theory and offer new insights into the relationship between text types and situation type distributions.
The world is completely working on digital data. The largest and prime or main collection of this digital data is web. The size of this web is increasing round-the-clock. The principal problem is to search this huge database for specific information. To state whether a web page is relevant to a search topic is a dilemma[l]. There are many techniques to state the relevancy but if focus on the users' perspective as key issue to guide search then semantic based web crawler are unsurpassed. Semantic based web crawlers maps relevancy with the help of lexical database. The crawler uses the senses provided by lexical database to discover relatedness among the search query and the web page being searched. Focused web crawler helps to find the similarity of web page to the search query without downloading that page. Thus focused web crawler is saving the bandwidth required to download a web page. This paper proposed and discuss one such approach to implement semantic based focused web crawler.
Motivated by methods used in language modeling and grammar induction, we propose the use of pragmatic constraints and perplexity as criteria to filter the unlabeled data used to generate the semantic similarity model. We investigate unsupervised adaptation algorithms of the semantic-affective models proposed in [1, 2]. Affective ratings at the utterance level are generated based on an emotional lexicon, which in turn is created using a semantic (similarity) model estimated over raw, unlabeled text. The proposed adaptation method creates task-dependent semantic similarity models and task-dependent word/term affective ratings. The proposed adaptation algorithms are tested on anger/distress detection of transcribed speech data and sentiment analysis in tweets showing significant relative classification error reduction of up to 10%.
In recent years, high increase in the amount of published web elements and the need to store, classify, restore, and process them have intensified the importance of natural language processing and its related tools such as automatic summarizers and machine translators. In this paper, a novel approach for evaluating automatic abstractive summarization system is proposed which can also be used in the other Natural Language Processing and Information Retrieval Applications. By comparing auto-abstracts (abstracts created by machine) with human abstracts (ideal abstracts created by human), the metrics introduced in the proposed tool can automatically measure the quality of auto-abstracts. Evidently, we can't semantically compare texts of abstractive summaries by comparison of just their words' appearance. So it is necessary to use a lexical database such as WordNet. We use FerdowsNet with a proper idea for Farsi language and it notably improves the evaluation results. This tool has been assessed by linguistic experts. This tool contains metric for determining the quality of summaries automatically by comparing them with summaries generated by humans (Ideal summaries). Evidently, we can't semantically compare texts of abstractive summaries by comparison of just their words' appearance and it is necessary to use a lexical database. We use this database with a proper idea together with Farsi parser in order to identify groups forming sentences and the results of evaluation improve significantly.
Concession is one of the trickiest semantic discourse relations appearing in natural language. Many have tried to sub-categorize Concession and to define formal criteria to both distinguish its subtypes as well as for distinguishing Concession from the (similar) semantic relation of Contrast. But there is still a lack of consensus among the different proposals. In this paper, we focus on those approaches, e.g. (Lagerwerf 1998), (Winter & Rimon 1994), and (Korbayova & Webber 2007), assuming that Concession features two primary interpretations, "direct" and "indirect". We argue that this two way classification falls short of accounting for the full range of variants identified in naturally occurring data. Our investigation of one thousand Concession tokens in the Penn Discourse Treebank (PDTB) reveals that the interpretation of concessive relations varies according to the source of expectation. Four sources of expectation are identified. Each is characterized by a different relation holding between the eventuality that raises the expectation and the eventuality describing the expectation. We report a) a reliable inter-annotator agreement on the four types of sources identified in the PDTB data, b) a significant improvement on the annotation of previous disagreements on Concession-Contrast in the PDTB and c) a novel logical account of Concession using basic constructs from Hobbs' (1998) logic. Our proposal offers a uniform framework for the interpretation of Concession while accounting for the different sources of expectation by modifying a single predicate in the proposed formulae.
Many algorithms for natural language processing rely on manual feature engineering. In this paper, we show that we can achieve state-of-the-art performance for part-of-speech tagging of Twitter microposts by solely relying on automatically inferred word embeddings as features and a neural network. By pre-training the neural network with large amounts of automatically labeled Twitter microposts to initialize the weights, we achieve a state-of-the-art accuracy of 88.9% when tagging Twitter microposts with Penn Treebank tags.
Discourse relation parsing is an impor-tant task with the goal of understanding text beyond the sentence boundaries. With the availability of annotated corpora (Penn Discourse Treebank) statistical discourse parsers were developed. In the litera-ture it was shown that the discourse pars-ing subtasks of discourse connective de-tection and relation sense classification do not generalize well across domains. The biomedical domain is of particular interest due to the availability of Biomedical Dis-course Relation Bank (BioDRB). In this paper we present cross-domain evaluation of PDTB trained discourse relation parser and evaluate feature-level domain adapta-tion techniques on the argument span ex-traction subtask. We demonstrate that the subtask generalizes well across domains. 1
We describe a representation scheme and an analysis engine using that scheme, both of which have been used to develop infrastructure for HLT. The Shakti Standard Format is a readable and robust representation scheme for analysis frameworks and other purposes. The representation is highly extensible. This representation scheme, based on the blackboard architectural model, allows a very wide variety of linguistic and non-linguistic information to be stored in one place and operated upon by any number of processing modules. We show how it has been successfully used for building machine translation systems for several language pairs using the same architecture. It has also been used for creation of language resources such as treebanks and for different kinds of annotation interfaces. There is even a query language designed for this representation. Easily wrappable into XML, it can be used equally well for distributed computing.
Syntactic parsing is a fundamental problem of natural language processing, and statistical syntactic parsing based on treebank gradually becomes the mainstream techniques of modern syntactic parsing following the building of large scale annotated treebanks. Firstly, the main treebanks and the main methods to measure syntactic parsing system performances were described. Secondly, the main statistical syntactic parsing models and the current researches on the Chinese syntactic parsing were presented and analyzed. Finally, the problems and the future study trends of statistical syntactic parsing were discussed and summarized. Is is pointed that Chinese syntax analysis methods are not suitable for the features of Chinese, and they do not effectively characterize the essential features of Chinese, thus the performances of syntactic parsing of Chinese are far below the performances of English. To integrate semantic information in syntactic parsing and to establish a joint syntactic and semantic statistical parsing model based on. semantic analysis will be a important study direction of syntactic parsing.
This journal article draws a distinction between the split and unified functions obtaining in the formation of Old English nouns and adjectives. The starting point of the discussion is an enlarged inventory of lexical functions that draw on Meaning-Text Theory and structural-functional grammars and explain the change of meaning caused by prefixation and suffixation in Old English. The extended inventory of lexical functions consists of 33 functions and has been applied to ca. 7,500 affixed nouns and adjectives extracted from the lexical database Nerthus (www.nerthusproject.com). The distinction between split and unified functions, in such a way that the former can be realized by both prefixes and suffixes and the latter by either prefixation or suffixation, allows for some generalizations. Firstly, the analysis proves that there are more functions involved in prefixation than in suffixation. Secondly, prefixation is meaning oriented while suffixation is class oriented.
Parsing the Arabic language is a difficult task given the specificities of this language and given the scarcity of digital resources (grammars and annotated corpora). In this paper, we suggest a method for Arabic parsing based on supervised machine learning. We used the SVMs algorithm to select the syntactic labels of the sentence. Furthermore, we evaluated our parser following the cross validation method by using the Penn Arabic Treebank. The obtained results are very encouraging.
This paper reviews the understanding divergence of the termsentence patternof Chinese grammar scholars;from the perspective of Chinese information processing, it analyses the lack of sentence pattern structure in current syntactic parsing and treebank construction in this field; and gives a recent review of formalization research of Li Jinxi's grammar system, indicating its strengths and still shortcomings on sentence pattern structure; it uses Li Jinxi's diagrammatic parsing method as a prototype design of a new type of diagrammatic parsing method of Chinese syntactic structure, specifically including a diagrammatic representation of the syntactic structure and structured XML storage format.
In this paper we present several approaches towards constructing joint ensemble models for mor-phosyntactic tagging and dependency parsing for a morphologically rich language – Bulgarian. In our experiments we use state-of-the-art taggers and dependency parsers to obtain an extended version of the treebank for Bulgarian, BulTreeBank, which, in addition to the standard CoNLL fields, contains predicted morphosyntactic tags and dependency arcs for each word. In order to select the most suitable tag and arc from the proposed ones, we use several ensemble techniques, the result of which is a valid dependency tree. Most of these approaches show improvement over the results achieved individually by the tools for tagging and parsing. 1
The architecture of writing systems metaphor has special relevance for understanding the structural nature of the Japanese writing system, and, more specifically, for appreciating how the 2,136 kanji of the 常用漢字表 /jō-yō-kan-ji-hyō/* ‘List of characters for general use’ function as the core building blocks in the orthographic representation of a considerable proportion of the Japanese lexicon. In seeking to illuminate the multiple layers of internal structure within Japanese kanji, the Japanese lexicon, and the Japanese writing system, the paper draws on insights and observations gained from an ongoing project to construct a large-scale Japanese lexical database system. Reflecting structural distinctions within the database, the paper consists of three main sections addressing the different structural levels of kanji components, jōyō kanji, and the lexicon. Keywords: Japanese writing system; building blocks; jōyō kanji; components; orthographic structure; database
We investigate the usefulness of syntactic knowledge in estimating the quality of English-French translations. We find that dependency and constituency tree kernels perform well but the error rate can be further reduced when these are combined with hand-crafted syntactic features. Both types of syntactic features provide information which is complementary to tried-and-tested nonsyntactic features. We then compare source and target syntax and find that the use of parse trees of machine translated sentences does not affect the performance of quality estimation nor does the intrinsic accuracy of the parser itself. However, the relatively flat structure of the French Treebank does appear to have an adverse effect, and this is significantly improved by simple transformations of the French trees. Finally, we provide further evidence of the usefulness of these transformations by applying them in a separate task ‐ parser accuracy prediction.
The aim of this article is to measure the indexes of productivity of the prefix ful - and the suffix - ful in Old English adjective formation. This analysis is based on Baayen’s framework, which comprises different measures on productivity. The major sources of the analysis are The Dictionary of Old English Corpus and the lexical database of Old English Nerthus. This study of productivity allows for a diachronic perspective on the evolution of these affixes from the Old English period to the present. The main conclusion drawn from this analysis is that the suffix -ful is more productive than its prefixal counterpart, which implies that more productive patterns are still maintained in Present-day English in contradistinction to the less productive ones.
We investigate how the granularity of POS tags influences POS tagging, and furthermore, how POS tagging performance relates to parsing results.For this, we use the standard "pipeline" approach, in which a parser builds its output on previously tagged input.The experiments are performed on two German treebanks, using three POS tagsets of different granularity, and six different POS taggers, together with the Berkeley parser.Our findings show that less granularity of the POS tagset leads to better tagging results.However, both too coarse-grained and too fine-grained distinctions on POS level decrease parsing performance.
In this article, we propose the first work that investigates the feasibility of Arabic discourse segmentation into elementary discourse units within the segmented discourse representation theory framework. We first describe our annotation scheme that defines a set of principles to guide the segmentation process. Two corpora have been annotated according to this scheme: elementary school textbooks and newspaper documents extracted from the syntactically annotated Arabic Treebank. Then, we propose a multiclass supervised learning approach that predicts nested units. Our approach uses a combination of punctuation, morphological, lexical, and shallow syntactic features. We investigate how each feature contributes to the learning process. We show that an extensive morphological analysis is crucial to achieve good results in both corpora. In addition, we show that adding chunks does not boost the performance of our system.
Chunking or shallow syntactic parsing is proving to be a task of interest to many natural language processing applications. The problem gets worse for the Arabic language because of its specific features that make it quite different and even more ambiguous than other natural languages when processed. In this paper, we present a method for chunking Arabic texts based on supervised learning. We use the Conditional Random Fields algorithm and the Penn Arabic Treebank to train the model. For the experimentation, we use over than 10,100 sentences as training data and 2,524 sentences for the test. The evaluation of the method consists of the calculation of the generated model accuracy and the results are very encouraging.
This paper investigates the recognition of unknown words in Chinese parsing. Two methods are proposed to handle this problem. One is the modification of a character-based model. We model the emission probability of an unknown word using the first and last characters in the word. It aims to reduce the POS tag ambiguities of unknown words to improve the parsing performance. In addition, a novel method, using graph-based semisupervised learning (SSL), is proposed to improve the syntax parsing of unknown words. Its goal is to discover additional lexical knowledge from a large amount of unlabeled data to help the syntax parsing. The method is mainly to propagate lexical emission probabilities to unknown words by building the similarity graphs over the words of labeled and unlabeled data. The derived distributions are incorporated into the parsing process. The proposed methods are effective in dealing with the unknown words to improve the parsing. Empirical results for Penn Chinese Treebank and TCT Treebank revealed its effectiveness.
Natural language is a fundamental thing of human-society to communicate and interact with one another. In this globalization era, we interact with different regional people as per our interest in social, cultural, economical, educational and professional domain. There are thousands of natural languages exist in our earth. It is quite tough, rather impossible to know all the languages. So we need a computerized approach to convert one natural language to another as per our necessity. This computerized conversion among multiple languages is known as multilingual machine translation. But in this paper we work with a bilingual model, where we concern with two languages: English and Bengali. We use soft computational approach where fuzzy If-Then rule is applied to choose a lemma from prior knowledge; Penn TreeBank PoS tags and HMM tagger are used as lexical class marker to each word in corpora.
This paper introduces a new technique for phrase-structure parser analysis, catego-rizing possible treebank structures by inte-grating regular expressions into derivation trees. We analyze the performance of the Berkeley parser on OntoNotes WSJ and the English Web Treebank. This provides some insight into the evalb scores, and the problem of domain adaptation with the web data. We also analyze a “test-on-train ” dataset, showing a wide variance in how the parser is generalizing from differ-ent structures in the training material. 1
With growing interest in the creation and search of linguistic annotations that form general graphs (in contrast to formally simpler, rooted trees), there also is an increased need for infrastructures that support the exploration of such representations, for example logical-form meaning representations or semantic dependency graphs. In this work, we heavily lean on semantic technologies and in particular the data model of the Resource Description Framework (RDF) to represent, store, and efficiently query very large collections of text annotated with graph-structured representations of sentence meaning. Keywords:Semantic Dependency Graphs, Treebank Search, Resource Description Framework 1.
The web today is huge and enormous collection of data today and it goes on increasing day by day. Thus, searching for some particular data in this collection has a significant impact. Researches taking place give prominence to the relevancy and relatedness of the data that is found. Inspite of their relevance pages for any search topic, the results are still huge to be explored. Another important issue to be kept in mind is the users’ standpoint differs from time to time from topic to topic. Effective relevance prediction can help avoid downloading and visiting many irrelevant pages. The performance of a crawler depends mostly on the opulence of links in the specific topic being searched. This paper reviews the researches on web crawling algorithms used for searching. Keywords— Web Crawling Algorithms, Crawling Algorithm Survey, Search Algorithms, Lexical Database, Metadata, Semantic. __________________________________________________*****_________________________________________________
The sentiment mining approaches can typically be divided into lexicon and machine learning approaches. Recently there are an increasing number of approaches which combine both to improve the performance when used separately. However, this still lacks contextual understanding which led to the introduction of deep learning approaches which allows for semantic compositionality over a sentiment treebank. This paper enhances the deep learning approach with semantic lexicon so that scores can be computed in-stead merely nominal classification. Besides, neutral classification is also improved. Results suggest that the approach outperforms its original.
BACKGROUND: Careful observation of the longitudinal course of bipolar disorders is pivotal to finding optimal treatments and improving outcome. A useful tool is the daily prospective Life-Chart Method, developed by the National Institute of Mental Health. However, it remains unclear whether the patient version is as valid as the clinician version. METHODS: We compared the patient-rated version of the Lifechart (LC-self) with the Young-Mania-Rating Scale (YMRS), Inventory of Depressive Symptoms-Clinician version (IDS-C), and Clinical Global Impression-Bipolar version (CGI-BP) in 108 bipolar I and II patients who participated in the Naturalistic Follow-up Study (NFS) of the German centres of the Bipolar Collaborative Network (BCN; formerly Stanley Foundation Bipolar Network). For statistical evaluation, levels of severity of mood states on the Lifechart were transformed numerically and comparison with affective scales was performed using chi-square and t tests. For testing correlations Pearson´s coefficient was calculated. RESULTS: Ratings for depression of LC-self and total scores of IDS-C were found to be highly correlated (Pearson coefficient r = -.718; p <.001), whilst the correlation of ratings for mania with YMRS compared to LC-self were slightly less robust (Pearson coefficient r =.491; p =.001). These results were confirmed by good correlations between the CGI-BP IA (mania), IB (depression) and IC (overall mood state) and the LC-self ratings (Pearson coefficient r =.488, r =.721 and r =.65, respectively; all p <.001). CONCLUSIONS: The LC-self shows a significant correlation and good concordance with standard cross sectional affective rating scales, suggesting that the LC-self is a valid and time and money saving alternative to the clinician-rated version which should be incorporated in future clinical research in bipolar disorder. Generalizability of the results is limited by the selection of highly motivated patients in specialized bipolar centres and by the open design of the study.
This paper describes experiments for statistical dependency parsing using two different parsers trained on a recently extended dependency treebank for Greek, a language with a moderately rich morphology. We show how scores obtained by the two parsers are influenced by morphology and dependency types as well as sentence and arc length. The best LAS obtained in these experiments was 80.16 on a test set with manually validated POS tags and lemmas. 1
OBJECTIVE: Premenstrual dysphoric disorder (PMDD) is associated with increased pain, but there has been a lack of well-controlled research assessing pain responsivity, sex hormones, and their relationships in this group. This study was designed to address this gap in the literature. MATERIALS AND METHODS: Healthy, regularly cycling participants (14 PMDD, 14 non-PMDD) attended pain testing sessions during the mid-follicular, ovulatory, and late-luteal phases of the menstrual cycle (order counterbalanced) and salivary estradiol, progesterone, and testosterone were assessed at each testing session. Pain sensitivity was measured from electrocutaneous threshold/tolerance, ischemic threshold/tolerance, sensory and affective ratings of electrocutaneous and ischemic stimuli, and the nociceptive flexion reflex threshold (NFR, a measure of spinal nociception). RESULTS: Women with PMDD had higher sensory pain ratings of electrocutaneous stimuli and trends for lower ischemic thresholds and higher affective pain ratings of electrocutaneous stimuli. However, there were no group differences observed in NFR threshold. Testosterone levels were also lower during the mid-follicular and ovulatory phases in PMDD. Correlations between pain outcomes and estradiol and testosterone indicated that these hormones are hypoalgesic, with estradiol having a greater hypoalgesic effect within the PMDD group. DISCUSSION: Overall, women with PMDD may have a phase-independent hyperalgesia, with pain amplification likely occurring at the supraspinal level rather than the spinal level, given the lack of group differences in NFR threshold. Because testosterone was hypoalgesic and lower in women with PMDD, and there were strong associations between pain and estradiol in PMDD, sex hormones may play a role in PMDD-related hyperalgesia.
Linguistic norms emerge in human communities because people imitate each other. A shared linguistic system provides people with the benefits of shared knowledge and coordinated planning. Once norms are in place, why would they ever change? This question, echoing broad questions in the theory of social dynamics, has particular force in relation to language. By definition, an innovator is in the minority when the innovation first occurs. In some areas of social dynamics, important minorities can strongly influence the majority through their power, fame, or use of broadcast media. But most linguistic changes are grassroots developments that originate with ordinary people. Here, we develop a novel model of communicative behavior in communities, and identify a mechanism for arbitrary innovations by ordinary people to have a good chance of being widely adopted. To imitate each other, people must form a mental representation of what other people do. Each time they speak, they must also decide which form to produce themselves. We introduce a new decision function that enables us to smoothly explore the space between two types of behavior: probability matching (matching the probabilities of incoming experience) and regularization (producing some forms disproportionately often). Using Monte Carlo methods, we explore the interactions amongst the degree of regularization, the distribution of biases in a network, and the network position of the innovator. We identify two regimes for the widespread adoption of arbritrary innovations, viewed as informational cascades in the network. With moderate regularization of experienced input, average people (not well-connected people) are the most likely source of successful innovations. Our results shed light on a major outstanding puzzle in the theory of language change. The framework also holds promise for understanding the dynamics of other social norms.
The conceptions of a linguistic norm for non-linguists form the focus of this article. These conceptions manifest themselves as mental values in the speakers’ minds and can be viewed as implicit reference values of linguistic judgements. There will hereby be an attempt to reconstruct the conceptual contours of people without any academic background in linguistics and to illustrate which criteria play a role in the assessments given and which structural domains can even be assessed. With respect to the method employed here, the study builds on the qualitative content analysis used for extracting and interpreting data. The result of which forms a categorical system that has been constructed deductively and rechecked inductively by the material. The system of categories constructed, which is empirically based on 56 qualitative interviews, also forms the data’s interpretative framework. With the aid of the acquired results, it is shown that non-linguists certainly have a clear picture of what constitutes a good language or how a good language should be; it is closely oriented on written language, invariant, understandable, and has a high communicative scope. In this way, a complex and multilayered conception of linguistic norms can be established. The contours of which will be sketched in this article.
Languages that have no explicit word de-limiters often have to be segmented for sta-tistical machine translation (SMT). This is commonly performed by automated seg-menters trained on manually annotated corpora. However, the word segmentation (WS) schemes of these annotated corpora are handcrafted for general usage, and may not be suitable for SMT. An analysis was performed to test this hypothesis us-ing a manually annotated word alignment (WA) corpus for Chinese-English SMT. An analysis revealed that 74.60 % of the sentences in the WA corpus if segmented using an automated segmenter trained on the Penn Chinese Treebank (CTB) will contain conflicts with the gold WA an-notations. We formulated an approach based on word splitting with reference to the annotated WA to alleviate these con-flicts. Experimental results show that the refined WS reduced word alignment error rate by 6.82 % and achieved the highest BLEU improvement (0.63 on average) on the Chinese-English open machine trans-lation (OpenMT) corpora compared to re-lated work. 1
We present a new dependency parsing method for Korean applying cross-lingual transfer learning and domain adaptation techniques. Unlike existing transfer learning methods relying on aligned corpora or bilingual lexicons, we propose a feature transfer learning method with minimal supervision, which adapts an existing parser to the target language by transferring the features for the source language to the target language. Specifically, we utilize the Triplet/Quadruplet Model, a hybrid parsing algorithm for Japanese, and apply a delexicalized feature transfer for Korean. Experiments with Penn Korean Treebank show that even using only the transferred features from Japanese achieves a high accuracy (81.6%) for Korean dependency parsing. Further improvements were obtained when a small annotated Korean corpus was combined with the Japanese training corpus, confirming that efficient crosslingual transfer learning can be achieved without expensive linguistic resources.
The global spread of English and the advent of a need for English as an International Language has become one of the hotly-debated issues in recent years. This owes much to the fact that English speakers today are more likely to be non-native speakers of English than native speakers, and most likely to use English in communication with other non-native speakers of English than native speakers. A significant number of scholars (e.g., Honna, 2003; Widdowson, 2003) even believe that English is no longer the sole property of its native speakers. Nevertheless, majority of English language teaching coursebooks are still being published by major Anglo-American publishers and are based on the linguistic norms and cultures of native English speaking countries, mainly the USA and the UK. Inevitably, criticism regarding an accurate presentation of cultural information and images about a variety of norms and cultures beyond the Anglo-Saxon and European world has risen. In fact, the English presented in these coursebooks has been seen as mainly representing the linguistic norms and culture of its native speakers, thereby offering ‘English of Specific Cultures’. The current discussions on the English language teaching and culture axis, however, make possible an understanding of an English language that has become first international and then global, thereby creating possibilities of portrayal of linguistic norms and cultures of Outer and Expanding circle countries especially through ELT coursebooks. Commissioned as such, then, English can be regarded as a language through which access to Englishes and cultures of the world accompanies its pedagogy, hence ‘ English for Specific Cultures’ (Yano, 2009). Discussing at length the role of English as an International Language and its cultural implications, this article investigates the varieties of Englishes in a series of EIL-based coursebooks, inquiring whether they are based on English of Specific Cultures or English for Specific Cultures.
What factors contribute to subjective experiences of familiarity, and are these subject to unconscious selection? We investigated the circumstances under which judgments of familiarity are sensitive to task-irrelevant sources using the artificial grammar learning paradigm, a task known to be heavily reliant on familiarity-based responding. In 2 experiments, we manipulated ‘free-floating feelings of familiarity’ by subliminally priming participants with either a subjectively familiar stimulus (their surname) or unfamiliar stimulus (a random letter string). In Experiment 1, after training on an artificial grammar, participants were required to rate the familiarity of a new set of grammar strings where the subliminal priming manipulation preceded each rating. Under these instructions the manipulation significantly altered ratings of familiarity. In Experiment 2, the training, the request for familiarity ratings, and the subliminal manipulation were all unchanged. In addition, however, participants were informed about the presence of rules dictating the structure of the training strings and were required to judge both whether each test-string conformed to those rules and to report the basis for their judgment. This broader decision context eliminated the effect of subliminal primes on ratings of familiarity even when participants’ reported basis for their judgments revealed no conscious knowledge of the rule structure. These results demonstrate that unconscious sources of familiarity can be selected or excluded according to conscious task contexts. The findings are incompatible with theories that equate familiarity with automaticity and those that state people must always be aware of the structural antecedents of metacognition.
Part-of-speech (POS) taggers can be quite accurate, but for practical use, accuracy often has to be sacrificed for speed. For example, the maintainers of the Stanford tagger (Toutanova et al., 2003; Manning, 2011) recommend tagging with a model whose per tag error rate is 17% higher, relatively, than their most accurate model, to gain a factor of 10 or more in speed. In this paper, we treat POS tagging as a single-token independent multiclass classification task. We show that by using a rich feature set we can obtain high tagging accuracy within this framework, and by employing some novel feature-weight-combination and hypothesis-pruning techniques we can also get very fast tagging with this model. A prototype tagger implemented in Perl is tested and found to be at least 8 times faster than any publicly available tagger reported to have comparable accuracy on the standard Penn Treebank Wall Street Journal test set.
For languages such as English, several constituent-to-dependency conversion schemes are pro-posed to construct corpora for dependency parsing. It is hard to determine which scheme is better because they reflect different views of dependency analysis. We usually obtain dependen-cy parsers of different schemes by training with the specific corpus separately. It neglects the correlations between these schemes, which can potentially benefit the parsers. In this paper, we study how these correlations influence final dependency parsing performances, by proposing a joint model which can make full use of the correlations between heterogeneous dependencies, and finally we can answer the following question: parsing heterogeneous dependencies jointly or separately, which is better? We conduct experiments with two different schemes on the Penn Treebank and the Chinese Penn Treebank respectively, arriving at the same conclusion that joint-ly parsing heterogeneous dependencies can give improved performances for both schemes over the individual models.