Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Ontologies are recognised as important tools, not only for effective and efficient information sharing, but also for information extraction and text mining. In the biomedical domain, the need for a common ontology for information sharing has long been recognised, and several ontologies are now widely used. However, there is confusion among researchers concerning the type of ontology that is needed for text mining , and how it can be used for effective knowledge management, sharing, and integration in biomedicine. We argue that there are several different ways to define an ontology and that, while the logical view is popular for some applications, it may be neither possible nor necessary for text mining. We propose a text-centered approach for knowledge sharing, as an alternative to formal ontologies. We argue that a thesaurus (i.e. an organised collection of terms enriched with relations) is more useful for text mining applications than formal ontologies.
There is a strong relationship between evaluation and methods for automatically training language processing systems, where generally the same resource and metrics are used both to train system components and to evaluate them. To date, in dialogue systems research, this general methodology is not typically applied to the dialogue manager and spoken language generator. However, any metric for evaluating system performance can be used as a feedback function for automatically training the system. This approach is motivated with examples of the application of reinforcement learning to dialogue manager optimization, and the use of boosting to train the spoken language generator.
This study compared four common methods for scoring a popular working memory span task, Daneman and Carpenter’s (1980) reading span test. More continuous measures, such as the total number of words recalled or the proportion of words per set averaged across all sets, were more normally distributed, had higher reliability, and had higher correlations with criterion measures (reading comprehension and Verbal SAT) than did traditional span scores that quantified the highest set size completed or the number of words in correct sets. Furthermore, creation of arbitrary groups (e.g., high-span and low-span groups) led to poor reliability and greatly reduced predictive power. It is recommended that researchers score span tasks with continuous measures and avoid post hoc dichotomization of working memory span groups.
The paper presents details and comparison of two valuable language resources for Czech, two independent verb valency frames electronic dictionaries. The FIMU verb valency frames dictionary was designed during the EuroWordNet project and contains semantic roles and links to the Czech wordnet semantic network. The VALLEX 1.0 format is based on the formalism of the Functional Generative Description (FGD) and was developed during the Prague Dependency Treebank (PDT) project. We present the tools and approaches that were used within the process of adopting the FIMU Vallex format for the wordnet enriched valency frames. 1.
Toempower thegeneral massthrough access toinformation andknowledge, organized efforts arebeing madetodevelop relevant content inlocallanguages andprovide local language capabilities toutility software. Wehavedeveloped a Question Answering (QA)System forHindidocuments that wouldberelevant formassesusingHindiasprimary language ofeducation. Theusershould beabletoaccess information fromE-learning documents ina userfriendly way,that isbyquestioning thesystem intheir native language Hindi andthesystem will return theintended answer (also in Hindi) bysearching incontext fromtherepository ofHindi documents. Thelanguage constructs, querystructure, commonwords, etc.arecompletely different inHindias compared toEnglish. A novelstrategy, inaddition to conventional search andNLP techniques, wasusedto construct theHindi QAsystem. Thefocus isoncontext based retrieval ofinformation. Forthis purpose weimplemented a Hindi search engine that works onlocality-based similarity heuristics toretrieve relevant passages fromthecollection. It alsoincorporates language analysis modules like stemmer andmorphological analyzer aswellasself constructed lexical database ofsynonyms. Theexperimental results over corpus oftwoimportant domains ofagriculture andscience showeffectiveness ofourapproach.
Broad coverage, high quality parsers are available for only a handful of languages. A prerequisite for developing broad coverage parsers for more languages is the annotation of text with the desired linguistic representations (also known as “treebanking”). However, syntactic annotation is a labor intensive and time-consuming process, and it is difficult to find linguistically annotated text in sufficient quantities. In this article, we explore using parallel text to help solving the problem of creating syntactic annotation in more languages. The central idea is to annotate the English side of a parallel corpus, project the analysis to the second language, and then train a stochastic analyzer on the resulting noisy annotations. We discuss our background assumptions, describe an initial study on the “projectability” of syntactic relations, and then present two experiments in which stochastic parsers are developed with minimal human intervention via projection from English.
"They... Speak Better English Than the English Do":Colonialism and the Origins of National Linguistic Standardization in America Paul K. Longmore (bio) Recent scholarship has traced efforts to fashion an American national language through standardization of forms and usage. Christopher Looby, in Voicing America: Language, Literary Form, and the Origins of the United States, David Simpson, in The Politics of American English, 1776–1850, and Kenneth Cmiel, in Democratic Eloquence: The Fight Over Popular Speech in Nineteenth-Century America, all examine public debates about these matters and, in particular, the labors of linguistic reformers to shape the national tongue. But these important studies focus mainly on the revolutionary, early national, and antebellum periods, and although Cmiel recounts the impact of late eighteenth-century British prescriptivists on postrevolutionary American thinking he does not extensively consider colonial efforts to regulate the language.1 In fact, attempts to shape written and spoken American English according to ideas of correctness, propriety, and, most important, a national standard began before American Independence. But those ideas and that standard were British rather than American. The effort reflected colonial desire to copy metropolitan English linguistic norms in order to attain cultural legitimacy within the British Empire. Postrevolutionary exertions perpetuated attitudes and activities that began in the late colonial period. This essay examines the colonial origins of the movement to standardize and nationalize American English. The central fact of colonials' experience is that they act as agents of an expansionist imperial society. As one result, dominant colonial groups are acutely aware of the metropolitan standard of the language they share with the homeland. In developing an extraterritorial variety of that language, they often labor to match the metropolitan standard. Transplanted speakers of various dialects of a common tongue encounter one another in [End Page 279] new geographical and social environments. Contact often produces dialect mixing and leveling and a compromise dialect called a koine. Koineization largely involves unconscious modification of speech forms. But the attentiveness of many colonials to a metropolitan standard indicates that colonial koines arise from not just spontaneous changes but conscious shaping. As users of the koine "nativize" their common tongue, they continuously render normative judgments about alternative usages. Prescribing what is correct, they seek to standardize the extraterritorial version of the language (Siegel 8; Haas; Stein). North American British colonials, especially those in the elite and middling ranks, took as their model the written and spoken English of the imperial center. Like elite and middling Britons, higher-status colonials used this "proper" and "true" English to distinguish themselves from people below them in the social hierarchy. Nonetheless and again like socially ambitious Britons, many colonials wielded linguistic correctness as a tool of social mobility. Colonials of all ranks emulated metropolitan Standard English in order to elevate their standing within the Empire. In the long run in a pattern typical of colonies of settlement, their efforts unintentionally helped to create a common language that provided one basis for American nationhood. Colonials' adoption of the metropolitan standard of English and their manner of applying it appear in three kinds of evidence: contemporary observers' evaluations of colonial speech; higher-status colonials' descriptions of British immigrants' non-standard English speech; and colonials' formal efforts to educate themselves in metropolitan Standard English. Eighteenth-century observers praised Anglophone colonials for matching metropolitan linguistic norms. They focused on pronunciation and accent, vocabulary and phraseology. William Eddis, secretary to Maryland's royal governor (1769–1777), avowed, "[T]he pronunciation of the generality of the people has an accuracy and elegance that cannot fail of gratifying the most judicious ear" (33). Jonathan Boucher, a tutor and Anglican priest in the Chesapeake (1759–1775), asserted that colonials displayed "the purest Pronunciation of the English Tongue that is anywhere to be met with" (30). "Accuracy," "elegance," and "purity" referred to both colonials' emulation of metropolitan standard pronunciation and the absence from their speech of British regional accents. Lord Adam Gordon, [End Page 280] a Scot, made the same point about word usage and grammar. Describing mid-1760s Philadelphia, he admitted that "the propriety of Language here surprized me much, the English tongue being spoken by all ranks, in a degree of purity and perfection, surpassing...
Abstract In order to demonstrate elevated disgust sensitivity and facilitated disgust learning in patients suffering from blood injection injury phobia, 23 phobics and 20 controls underwent an evaluative conditioning experiment. They were presented with picture pairs consisting of affectively neutral pictures (CS), which were followed by either disgust-inducing, fear-inducing, pleasant, or neutral scenes (US). During the presentation we recorded the electromyogram (EMG) of the musculus levator labii as a specific disgust indicator. Affective ratings for the pictures were determined before and after conditioning. Also, CS-US contingency verbalisation (CV) was assessed. Phobics reported a greater overall disgust sensitivity, experienced stronger feelings of disgust, and showed greater EMG responses while viewing disgust-eliciting scenes than control subjects. Evaluative conditioning occurred equally in both groups and depended on CV.
We describe a method of constructing Thai WordNet, a lexical database in which Thai words are organized by their meanings. Our methodology takes WordNet and LEXiTRON machine-readable dictionaries into account. The semantic relations between English words in WordNet and the translation relations between English and Thai words in LEXiTRON are considered. Our methodology is operated via WordNet Builder system. This paper provides an overview of the WordNet Builder architecture and reports on some of our experience with the prototype implementation.
Abstract. The present paper focuses on representation of morphological meanings on the underlying syntactic level. The concept of semantic counterparts of morphological meanings, the so-called grammatemes, was introduced in Functional Generative Description in the 1960’s. We suggest an elaborated system of these grammatemes, which have become a part of the tectogrammatical level of the Prague Dependency Treebank.
Natural languages encode gender distinctions in various ways. We investigate the differences between English and Hebrew in this respect, our departure point being the relations that are defined between the feminine and the masculine realizations of nouns in the English WordNet. We define a number of distinct classes of English nouns which differ in the way they realize gender distinctions. We then define similar classes of Hebrew nouns and show how to map the Hebrew nouns (and relations defined over them) to the English structure. This establishes a systematic assignment of Hebrew nouns to WordNet synsets, which is consistent with the ideas underlying multilingual extensions of WordNet. The main result is a consistent Hebrew WordNet which is aligned with the English one, but an additional contribution is a set of desiderata for the correct encoding of (systematic) semantic differences among languages. 1
Translation tests are widely used for high school term tests and entrance examinations as well as university entrance examinations. Although a considerable number of papers point out the possibility of low reliability for scoring, little is known about the factors raters play in the reliability of scoring (Watanabe, 1994). This study examines how the professional backgrounds of raters affect rating criteria. The results indicated that novice raters tended to over-estimate examinees' comprehension whereas experienced raters were more likely to focus on the correctness of the Japanese sentence. In addition, it turned out that the difficulty of sentences affected the scoring of both experienced and novice raters. After administering a sorting task, the difficulty of the sentences showed that the perception of sentence difficulty did not correspond to the difficulty of examinees' translation. The paper closes by suggesting several pedagogical implications for administering translation tests. Of particular importance is that test developers should consider not only the complexity of sentence structures and vocabulary, but also the examinees' topic familiarity of the sentences to be translated.
Abstract This study compared the efficacy of measures of naming speed, verbal fluency and self-ratings for establishing language dominance in 25 bilingual English–Spanish adults with college degrees. Naming speed was measured by total naming times (in seconds) for five Alzheimer's Quick Test tasks (Wiig, Nielsen, Minthon & Warkentin, 2002) and verbal fluency with the Word Listing by Domain (Lambert, Havelka, & Crosby, 1958; Fishman & Cooper, 1969). Self-ratings of English–Spanish competence (listening, speaking, reading, and writing) and frequency of use of each spoken language served as standards for comparisons. For the aggregate sample, color–form, color–animal, and color–object naming times were significantly shorter for English than Spanish (p <.01). There was 100% agreement in language-dominance judgments between self-ratings of language competence and frequency of use, and color–form, color–animal, and color–object naming-time differences in the two languages. Word Listing by Domain quotients for language dominance showed a lower degree of agreement (52%) with self-ratings and naming-time differences. The findings suggest that cross-linguistic comparisons of naming times for color–form, color–animal, and color–object naming may be helpful in screening adults for language dominance for psychoeducational assessment purposes.
We formalize weighted dependency parsing as searching for maximum spanning trees (MSTs) in directed graphs. Using this representation, the parsing algorithm of Eisner (1996) is sufficient for searching over all projective trees in O(n3) time. More surprisingly, the representation is extended naturally to non-projective parsing using Chu-Liu-Edmonds (Chu and Liu, 1965; Edmonds, 1967) MST algorithm, yielding an O(n2) parsing algorithm. We evaluate these methods on the Prague Dependency Treebank using online large-margin learning techniques (Crammer et al., 2003; McDonald et al., 2005) and show that MST parsing increases efficiency and accuracy for languages with non-projective dependencies.
In this article, I explore the ways in which ethnic identity is expressed by following the formulaic socio-linguistic norm, the very method of which defies the authenticity of identity itself, thereby asserting the identity's multi-facetedness as sustained in performative linguistic practice. I look at multi-sited socio-linguistic interactions among Koreans in Japan, who claim their primary identity to be that of North Korea's overseas citizens even though none of them have North Korean passport or nationality. Their identity, in other words, is based on ideological commitment, which is in reality supported by their ongoing linguistic practice. A close look at their socio-linguistic life reveals their ethnicity's dual or multiple ontology, which challenges among other things the currently dominant assertion of Japanese self in the western academic discourse.
Given a sequence of samples from an unknown probability distribution, a statistical estimator aims at providing an approximate guess of the distribution by utilizing statistics from the samples. One crucial property of a `good' estimator is that its guess approaches the unknown distribution as the sample sequence grows large. This property is called consistency. This paper concerns estimators for natural language parsing under the Data- Oriented Parsing (DOP) model. The DOP model specifies how a probabilistic grammar is acquired from statistics over a given training treebank, a corpus of sentence-parse pairs. Recently, Johnson [15] showed that the BOP estimator (called DOPl) is biased and inconsistent. A second relevant problem with DOP1 is that it suffers from an overwhelming computational inefficiency. This paper presents the first (nontrivial) consistent estimator for the DOP model. The new estimator is based on a combination of held-out estimation and a bias toward parsing with shorter derivations. To justify the need for a biased estimator in the case of DOP we prove that every non-overfitting DOP estimator is statistically biased. Our choice for the bias toward shorter derivations is justified by empirical experience, mathematical convenience and efficiency considerations. In support of our theoretical results of consistency and computational efficiency, we also report experimental results with the new estimator.
This article is devoted to the problem of quantifying noun groups in German. After a thorough description of the phenom ena, the results of corpus-based investigations are described. Moreover, some examples are given that underline the necessity of integrating some kind of information other than grammar sensu stricto into the treebank. We argue that a more sophisticated and fine-grained annotation in the treebank would have very positve effects on stochastic parsers trained on the treebank and on grammars induced from the treebank, and it would make the treebank more valuable as a source of data for theoretical linguistic investigations. The information gained from corpus research and the analyses that are proposed are realized in the framework of SILVA, a parsing and extraction tool for German text corpora.
Traditional Chinese text chunking approach is to identify phrases using only one model and same features. It is shown that one model couldn't comprise each phrase's characteristics, and same features are not suitable to all phrases, data sparseness also appears. Multi-agent strategy uses several model and sensitive features of each phrase to identify different phrases. This paper describes the multi-agent strategy applied in the identification of Chinese phrases whose main features are: 1) easy and quick communication between phrases; 2) avoidance of data sparseness. Through testing on Chinese Penn Treebank, F score of Chinese text chunking using multi-agent strategy achieves to 95.82%, which is higher than the best result that has been reported.
This paper presents a novel approach to constructing multilingual lexical databases using semantic frames. Starting with the conceptual information contained in the English FrameNet database, we propose a corpus-based procedure for producing parallel lexicon fragments for Spanish, German, and Japanese, which mirror the English entries in breadth and depth. The resulting lexicon fragments are linked to each other via semantic frames, which function as interlingual representations. The resulting parallel FrameNets differ from other multilingual databases in three significant points: (1) they provide for each entry an exhaustive account of the semantic and syntactic combinatorial possibilities of each lexical unit; (2) they offer for each entry semantically annotated example sentences from large electronic corpora; (3) by employing semantic frames as interlingual representations, the parallel FrameNets make use of independently existing linguistic concepts that can be empirically verified.1 1 I am grateful to Charles Fillmore, Collin Baker, Carlos Subirats, Kyoko Hirose Ohara, Hans U. Boas, Jonathan Slocum, Inge De Bleecker, Jana Thompson, and three anonymous referees for very helpful comments on the material discussed in the article.
Three studies demonstrate the warm glow heuristic (Monin, 2003) without relying on aggregated ratings, and illustrate the important distinction between correlating average ratings versus averaging individual correlations. In Study 1, we re-analyze previous data correlating individual ratings with aggregates from another small sample of raters. In Study 2, we correlate individual familiarity ratings with normed attractiveness from a large sample of raters (n > 2,500). Study 3 bypasses the issue of aggregates altogether by having participants provide both attractiveness and familiarity ratings and computing correlations within participants. Despite this more conservative approach, the results of all three studies support the existence of the beautiful–is–familiar phenomenon.
We present a system that automatically identifies Attribution, an intra-sentential relation in the RST Treebank. The system uses uses syntactic information from Penn Treebank parse trees. It identifies Attributions as structures in which a verb takes an SBAR complement, and achieves a f-score of.92. This supports our claim that the Attribution relation should be eliminated from a discourse treebank, since it represents information that is already present in the Penn Treebank, in a different form. More generally, we suggest that intra-sentential relations in the RST Treebank might all be eliminable in this way. 1
In order to realize the full potential of dependency-based syntactic parsing, it is desirable to allow non-projective dependency structures. We show how a data-driven deterministic dependency parser, in itself restricted to projective structures, can be combined with graph transformation techniques to produce non-projective structures. Experiments using data from the Prague Dependency Treebank show that the combined system can handle non-projective constructions with a precision sufficient to yield a significant improvement in overall parsing accuracy. This leads to the best reported performance for robust non-projective parsing of Czech.
@conference{ai-giguet-2005-1, author = {Giguet, Emmanuel and Luquet, Pierre-Sylvain}, title = {Multilingual Lexical Database Generation from parallel texts with endogenous resources}, booktitle = {PAPILLON-2005 Workshop on Multilingual Lexical Databases}, year = {2005}, month = {December 12-14}, address = {Chiang Rai, Thaïland} }
In this paper we discuss the application of semi-supervised machine learning method-co-training on Chinese Text Chunking. Firstly, we give the definition of Chinese chunk,then the formalized definition of co-training algorithm.We proposed a example selection method based on the consistence, using two classifiers: Transductive HMM and fnTBL to combine a classification system to perform the Chinese text chunking task with the small-scale labled Chinese treebank and large-scale unlabled Chinese corpus. The result were compared with the self-training result and the result of the non co-training experiment in which we only used the small-scale Chinese treebank as training data and use one classifier(Transductive HMM or fnTBL) to recognize the Chinese chunk. The improvement is significant, the F value of the two classifiers reached 83.41%,85.34%, get a improvement of 2.13 points and 7.21 points respectively.
An automatic method for annotating the Penn-II Treebank (Marcus et al., 1994) with high-level Lexical Functional Grammar (Kaplan and Bresnan, 1982; Bresnan, 2001; Dalrymple, 2001) f-structure representations is presented by Burke et al. (2004b). The annotation algorithm is the basis for the automatic acquisition of wide-coverage and robust probabilistic approximations of LFG grammars (Cahill et al., 2004) and for the induction of subcategorisation frames (O’Donovan et al., 2004; O’Donovan et al., 2005). Annotation quality is, therefore, extremely important and to date has been measured against the DCU 105 and the PARC 700 Dependency Bank (King et al., 2003). The annotation algorithm achieves f-scores of 96.73% for complete f-structures and 94.28% for preds-only f-structures against the DCU 105 and 87.07% against the PARC 700 using the feature set of Kaplan et al. (2004). Burke et al. (2004a) provides detailed analysis of these results. \nThis paper presents an evaluation of the annotation algorithm against PropBank (Kingsbury and Palmer, \n2002). PropBank identifies the semantic arguments of each predicate in the Penn-II treebank and annotates their semantic roles. As PropBank was developed independently of any grammar formalism it provides a platform for making more meaningful comparisons between parsing technologies than was previously possible. PropBank also allows a much larger scale evaluation than the smaller DCU 105 and PARC 700 gold standards. In order to perform the evaluation, first, we automatically converted the PropBank annotations \ninto a dependency format. Second, we developed conversion software to produce PropBank-style semantic annotations in dependency format from the f-structures automatically acquired by the annotation algorithm from Penn-II. The evaluation was performed using the evaluation software of Crouch et al. (2002) and Riezler et al. (2002). Using the Penn-II Wall Street Journal Section 24 as the development set, currently we achieve an f-score of 76.58% against PropBank for the Section 23 test set.
We present a method for automatic RMRS semantics construction from dependency structures, following the semantic algebra of Copestake et al. (2001). We have applied this method to a subset of the TIGER Dependency Bank for German (Forst et al., 2004) to obtain a semantic treebank for (HPSG) parser evaluation. We describe the semantics construction mechanism and give evaluation figures from manual validation of the treebank. These indicate high precision of the automatic RMRS construction process.
Research on child language acquisition would benefit from the availability of a large body of syntactically parsed utterances between parents and children. We consider the problem of generating such a ``treebank'' from the CHILDES corpus, which currently contains primarily orthographically transcribed speech tagged for lexical category.
The Proposition Bank project takes a practical approach to semantic representation, adding a layer of predicate-argument information, or semantic role labels, to the syntactic structures of the Penn Treebank. The resulting resource can be thought of as shallow, in that it does not represent coreference, quantification, and many other higher-order phenomena, but also broad, in that it covers every instance of every verb in the corpus and allows representative statistics to be calculated. We discuss the criteria used to define the sets of semantic roles used in the annotation process and to analyze the frequency of syntactic/semantic alternations in the corpus. We describe an automatic system for semantic role tagging trained on the corpus and discuss the effect on its performance of various types of information, including a comparison of full syntactic parsing with a flat representation and the contribution of the empty “trace” categories of the treebank.
Modern day lexical databases are not constructive, but differential (Miller et al 1990).We outline the logical structure of the task of a conceptual analyst who wishes to construct a constructive lexicon based on a Universal Theory Model of Concepts.Elementary notions of descriptive and explanatory adequacy are developed within this model.The diverse streams of evidence available for the conceptual analyst to engage in empirical inquiry are reviewed in the domains of light and perception. Theories for Constructive LexiconsMajor projects have been conducted for centuries to record, in one form or another, representations of word meanings.This has been in the form of the lexicographer's dictionary or the more modern computational linguist's electronic databases (WordNet, FrameNet, Verb-Net) e.g. last year's symposium (Miller, Fillmore, Palmer, Lenat and Hayes 2004).The degree to which computational linguists and cognitive scientists draw upon these resources as accurate descriptions of lexical knowledge is astonishing.Researchers speak of "putting meaning in your trees", providing "deep semantics", automatically labeling "semantic roles", or using Word-Net as the authority for word sense disambiguation, to name some key projects.(Palmer et al 2002, Fillmore et al 2001, Gildea and Jurafsky 2002) Given the increasing reliance to these databases, it may appear that the terms meaning and semantic are not being used glibly and that the theoretical foundations of these databases were sound.Even the most cursory analysis shows that this is not so.Consider the distinction made by the originators of WordNet between a constructive vs. differential lexicon: In a differential theory of the lexicon, meanings can be represented by any symbols that enable a theorist to distinguish among them; In a constructive theory of the lexicon, the representation should "contain sufficient information to support an accurate construction of the concept (by either a person or a machine)" (Miller et al 1990).Today's dictionaries and all of today's lexical databases are differential: the intension of synsets of WordNet are just sufficient so that someone who already knows English can distinguish among synsets, while the thematic roles used in VerbNet and FrameNet are notoriously difficult to define: primitive terms such as Agent,
We explore the application of memorybased learning to morphological analysis and part-of-speech tagging of written Arabic, based on data from the Arabic Treebank. Morphological analysis -- the construction of all possible analyses of isolated unvoweled wordforms -- is performed as a letter-by-letter operation prediction task, where the operation encodes segmentation, part-of-speech, character changes, and vocalization. Part-of-speech tagging is carried out by a bi-modular tagger that has a subtagger for known words and one for unknown words. We report on the performance of the morphological analyzer and part-of-speech tagger. We observe that the tagger, which has an accuracy of 91.9% on new data, can be used to select the appropriate morphological analysis of words in context at a precision of 64.0 and a recall of 89.7.
Vese of the third generation,with its vehement passion,strives to break through the confines of entire discourse and to transcend the summit of poetic creation newly instituted by hazy poetry.In a word,it aims to grapple with language by disdaining and revolting against all fixed linguistic norms so as to shake off the grave sense of social and historical responsibility shouldered by poets of former generations,the ideological restraints imposed by tragic social reality,and above all to point to the very essence of life and inquire into the minute sense of existence as acutely felt by individuals.The crux lies in language,however and what verse of the third generation has endeavored to attain remains utopian in linguistic aspects in that the goal thus targeted would,in consequence,make verse readily understood by few readers and merely by the poet him/herself,hence the somewhat blocked channel of communication between poets and readers.Such a phenomenon has sounded the alarm for literature as follows: adequate attention must be paid to language in literary creation of any kind.
It is widely believed that the difference between regular and irregular verbs is restricted to form. This study questions that belief. We report a series of lexical statistics showing that irregular verbs cluster in denser regions in semantic space. Compared to regular verbs, irregular verbs tend to have more semantic neighbors that in turn have relatively many other semantic neighbors that are morphologically irregular. We show that this greater semantic density for irregulars is reflected in association norms, familiarity ratings, visual lexical-decision latencies, and word-naming latencies. Meta-analyses of the materials of two neuroimaging studies show that in these studies, regularity is confounded with differences in semantic density. Our results challenge the hypothesis of the supposed formal encapsulation of rules of inflection and support lines of research in which sensitivity to probability is recognized as intrinsic to human language.