Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Abstract Recently, there has been a growing interest in regional variation within African American English. This study reviews a work done on local speech in Pittsburgh, Pennsylvania, discussing trends for both African American and White ethnic groups. Just as scholars have found in other geographic regions, in Pittsburgh, African Americans and Whites share a number of feature characteristics of the local dialect, but remain distinct in a number of other ways. Research in Pittsburgh, as elsewhere, highlights the complexity, rather than the homogeneity, of African American speech across the country, as speakers exhibit alignment to both regional and supraregional ethnic linguistic norms.
In Anti-Individualism,11 S.C. Goldberg, Anti-Individualism: Mind and Language, Knowledge and Justification (Cambridge University Press, 2007, xiii + 265 pp. £42.75). Sanford Goldberg presents a theory of testimony and testimonial knowledge, and considers the implications of this theory for philosophy of language, philosophy of mind, and epistemology. According to Goldberg, plausible assumptions about testimonial knowledge provide important motivation for the view that linguistic meaning is not determined exclusively by purely internal states and processes, and that the same is true of the representational content of propositional attitudes. Instead of being determined by purely internal factors, such as the states and processes that are considered in cognitive neuroscience, or the states and processes that are studied by computational psychology, meaning and content depend, in part, on relations between individual agents and the linguistic communities that they inhabit. In short, according to Goldberg, considerations pertaining to testimonial knowledge provide motivation for semantic externalism. He also maintains that this motivation is largely independent of the arguments for externalism that are familiar from the writings of Kripke, Putnam, and Burge.22 S.A. Kripke, Naming and Necessity (Harvard University Press, 1980); H. Putnam, “The Meaning of ‘Meaning’,” reprinted in H. Putnam (ed.), Mind, Language, and Knowledge: Philosophical Papers, Vol. II, (Cambridge University Press, 1975), pp. 215–71; Burge, T., “ Individualism and the Mental”, Midwest Studies in Philosophy, 4 (1979), pp. 73– 121. The line of thought leading to these conclusions occupies the first half of the book. In the second half, Goldberg turns his attention to epistemological issues. What is it, he asks, for a testimonial belief to count as epistemically justified, and what is it for a testimonial belief to count as epistemically rational? He argues that these epistemic properties involve reliability in at least two ways. First, in order for testimonial beliefs to exemplify the properties, the relevant believers must be equipped with, and make use of, a reliable ability to determine whether testimony is called into question by defeaters—that is, by considerations that challenge testimony, suggesting that it may not be trustworthy. And second, the beliefs must derive from testimonial sources that are in fact reliable. Now, as is customary, Goldberg takes reliability to be intimately related to truth, and therefore to depend on factors that are external to the believer. It follows that the epistemic status of a testimonial belief is doubly dependent on external factors. According to Goldberg, then, when we reflect on the justification and rationality of testimonial beliefs, we gain a new appreciation of the plausibility of epistemological externalism—new in the sense that it is independent of the arguments for externalism that appear elsewhere in the literature. In sum, it is the author's view that plausible assumptions about testimony provide a foundation for fresh and largely autonomous arguments for semantic externalism and epistemological externalism. The book is an extended defense of this idea. The first chapter sets the stage for later developments by arguing for an initial characterization of testimonial knowledge. The characterization is based on four conditions. First, to acquire testimonial knowledge, an agent must form a belief that is testimonially grounded. In other words, the agent must form a belief in response to a piece of testimony, and in forming the belief, the agent must assume or presuppose that the testimony is reliable. Second, the agent's presupposition concerning the reliability of the testimony must be correct. Third, it must be the case that the agent arrives at a semantic interpretation of the testimony by a reliable process. That is to say, in Goldberg's words, the agent must recover “the proposition attested to by a reliable process of comprehension” (p. 31). And fourth, it must be true that the agent's acceptance of the testimony is “the upshot of a reliable capacity for distinguishing reliable from unreliable testimony” (p. 31). In all, then, reliability figures in Goldberg's initial characterization of testimonial knowledge in four different ways. Should this be a cause of concern for readers with strong internalist intuitions? Goldberg thinks not. One reason is that he is claiming only that the four reliability-involving requirements are necessary conditions of testimonial knowledge. He is allowing room for additional requirements, including requirements with a strong internalist flavor. Another reason is that, in chapter 1, at any rate, he is concerned only with knowledge. He thinks that even fairly radical internalists will be willing to allow that knowledge has a significant externalist dimension, though they would no doubt reject companion claims about epistemic justification or epistemic rationality. At all events, Goldberg offers justifications for his four conditions. The conditions are needed, he says, in order to honor the intuition that if a testimonial belief is to count as knowledge, then it cannot be a mere accident that the belief is true. Each of the second, third, and fourth conditions rules out a form of accidentality. Goldberg also uses examples to motivate the conditions. It is important to him that the conditions be accepted as correct, especially the third and fourth ones, for the book is in effect a lengthy meditation on their meaning and implications. Goldberg's argument for semantic externalism spans several chapters and is quite complex. I cannot do justice to all of its components here. In place of a complete account, I offer the following summary, which seems to capture a sizable portion of the core ideas: First premise: We have a great deal of testimonial knowledge. Second premise: Testimonial knowledge is possible only if the users of a language are equipped with procedures that enable them to recover propositions reliably from asserted sentences. (pp. 2–3) Third premise: In a wide range of cases, including a great many that arise in everyday life, we are confronted with the task of recovering a proposition from an asserted sentence in a context in which we have no real grasp of the beliefs and intentions of the speaker, and in which, therefore, we would be unable to determine which psychological states led the speaker to assert the sentence. (p. 41) Fourth premise: If, in many cases, we are unable to recover propositions from sentences by theorizing about the psychological states of speakers, then, assuming that we have the ability to recover propositions reliably, it must be true that we apply public linguistic norms in interpreting testimony—that is, fixed rules or principles that assign meanings to expressions, and that can be applied independently of assumptions about the purely psychological states and dispositions of specific speakers. (p. 41) Fifth premise: Since public linguistic norms can be applied independently of assumptions about the purely psychological states and dispositions that animate individual speakers, norms must be metaphysically independent of such states and dispositions, and it is therefore possible to imagine situations in which the operative norms are different than the ones that obtain in the actual case, but in which all purely psychological states and dispositions are the same. (pp. 104–5) Sixth premise: If the norms that assign meanings to sentences can vary while all purely psychological states and dispositions are held fixed, then semantic externalism is true. (p. 105) Conclusion: Semantic externalism is true. Although this line of thought touches most of the key bases, it also misses a few. Thus, while it is true that the fourth and fifth premises do justice to one of Goldberg's main reasons for maintaining that meaning is determined by public linguistic norms, he also has a second reason that is developed at some length. (My summary emphasizes chapter 2 at the expense of chapter 3.) In addition, I should note that Goldberg extends the present argument, which is concerned only with linguistic meanings, in such a way as to obtain the additional conclusion that the representational contents of propositional attitudes are determined externalistically. And there is also a third important omission. I will get to it in a moment. As I have represented him, Goldberg is concerned to derive semantic externalism from considerations having to do with testimonial knowledge, together with the observation that we are often unable to recover propositions reliably from asserted sentences by theorizing about the psychological states of individual speakers. If this is correct, then his line of thought is quite different than the arguments for semantic externalism that we find elsewhere in the literature. The other arguments make no mention of testimony, or of the risks that would be involved in interpreting assertions in terms of hypotheses about the psychological states of the speakers who produce them. Rather, they appeal directly to examples in which there is a dissociation between the descriptive content that a speaker has in mind in using an expression and the semantic meaning of the expression. Thus, as the reader will recall, Putnam begins one of his arguments by informing the reader that he lacks sufficient descriptive knowledge to distinguish elms from beeches, and would therefore need to obtain information from an expert in order to apply the terms “elm” and “beech” appropriately to specific trees.33 Putnam, “The Meaning of ‘Meaning’ ”. He then asks whether this fact precludes our crediting him with the ability to make true and false claims containing the words “elm” and “beech.” If, for example, he was to assert “Elms are trees,” would we deny that he had expressed a true proposition? More generally, does his lack of relevant descriptive knowledge preclude our crediting him with the ability to use “elm” and “beech” with their normal meanings? He predicts that his readers will answer “no” to both of these questions, and in most cases, this prediction has turned out to be correct. Putnam's other main arguments are also based on examples of psychological–semantic dissociations, though they tend to involve thought experiments rather than actual cases. The same is true of Kripke and Burge: They describe examples in which descriptive knowledge and meaning come apart, and then point out that these examples mandate an externalist account of meaning. It is clear that Goldberg's strategy for establishing semantic externalism is quite different than this one. Goldberg's line of thought is top-down: It relies on theoretical considerations about testimony and the elusiveness of Gricean speaker meanings. The standard arguments are bottom-up: They begin with intuitions about possible cases, and proceed from them to general morals. Goldberg's line of thought is interesting and suggestive. It deserves more attention than I can give it here. It seems unlikely, however, that it is comparable to the familiar arguments for externalism in strength and scope, even when various supplementary considerations that I have not described are added in. Among other things, the familiar arguments imply that semantic competence has an important social dimension, and more particularly, that individual agents can be credited with a grasp of the meanings of words on the of dispositions to to other of their linguistic As as I can Goldberg's line of thought to these fourth that agents must on public linguistic norms in recovering propositions from but this is about the of public linguistic norms, or what is involved in to them. As as I can the is with the that agents must have a and autonomous of all of the that would be involved in the norms and also with the that agents must have a and autonomous of the meanings that linguistic norms assign to point to the fifth It that it is possible for linguistic norms to vary while purely psychological dispositions but this is with the that agents must have a and autonomous grasp of the meanings that norms with fifth claims only that different can be is with the that agents must have an autonomous grasp of all of the meanings that in any of the In view of these it seems that the conclusion of the argument is than the conclusions of the familiar arguments for externalism. It be that there is in chapter 4 that Goldberg for his argument to have the same implications as the familiar and that, there are for the of the is a great deal of to these Goldberg to a fairly of semantic externalism. is the third that I this observation a into Goldberg must between the argument as justice to the of his and the argument in such a way that it and the familiar about the of for a grasp of and the that an agent's grasp of meanings can be and If he the argument as it Goldberg can to be a for semantic externalism that is independent of the familiar but he must also that his conclusion is than the conclusions of the familiar and also than he would And if he the argument by premises that allow for a and grasp of linguistic norms, then he can that it a fairly but the argument will no be independent of the familiar It will have to make an appeal to the that is sufficient for a grasp of and Goldberg will have to this by examples the ones by Kripke, Putnam, and Thus, no which he Goldberg will be unable to one of his main it I to of the which is to epistemological issues. As I him, Goldberg is there concerned to the following four is in the epistemic to is epistemically to testimony as there are no reasons not to the (p. as in a piece of having the epistemic or to the piece of on one of social (p. complete account of the of belief that on one of social (p. The process that in of testimony in the testimonial extends to of the social the that is for by (p. According to there is a epistemic in of testimony is the is It does not depend on having reasons to the testimony of a speaker, on having general for all speakers. The justification to which is in and are in that they both imply that the epistemic status of a testimonial belief has an important social dimension, and more particularly, that it on the reliability of the who is the of the They only in that is concerned with epistemic while is concerned with epistemic rationality. is therefore the more radical many who externalist about justification are to an internalist As for it is of a of that can be by that the testimonial beliefs of often count as knowledge, and second, that they do in forming the relevant beliefs, are from unreliable testimony by the of their and other emphasizes the most of this the that in the case of the procedures for whether testimony is called into question by are not internal to but are rather in the social of these claims are and Goldberg's of them are quite In case, his of I will not be to all of the claims here. they to the most questions, I will on and a initial but it is not to of that it into that has had to the of and is not in a to whether the is about issues. that at a and that the to begins to on other things, that there is reason to doubt that have a theory of that in to being about the and of has no for an about the of in the general And that, most is not in of a that can be to determine whether is and of can determine whether an presents the of a is in this than the to the to It would appear not. In it seems that it would be quite for to him, as it would be quite to place in a of that one on a Now, Goldberg is of examples of this and is also that, in the view of many they to the following which is of with is not epistemically in not have the epistemic to is not epistemically to testimony has or a not based on testimony, for the testimony as (p. Goldberg does not this but he its as the reader will from summary of chapter 1, Goldberg is to the following If an agent's testimonial belief is to count as knowledge, then the agent's acceptance of the relevant piece of testimony must have by a reliable capacity to distinguish reliable from unreliable (p. the one and the of principles and can it be to Goldberg's answer is that is the of two The first is that testimonial processes it is to enable the to in which is to be from in which it is not. And the second is that these processes are in fact reliable in that, in and they in acceptance in most of the the testimony is both true and and they in in most the testimony is false or (p. In then, Goldberg by it with two other The to testimony, but only in in which the following conditions are the relevant agent has no reasons for the the agent is equipped with appropriately procedures that enable him to such reasons when they and the procedures in question are reliable. The defense is at least Thus, while is a to the of is not a to the present more is not equipped with procedures for when factors are present that the of testimony into at has no such procedures that can be applied to who At first the new of is a between the of and the The new from the one in that it a on the of that can be extended to a piece of testimony only it has to a that is and that testimony can be it must a that is of with this however, seems more than In that there is a of of any for it that is only when an agent has an or a reason for that a piece of testimony is correct. who are familiar with the on testimony will as a of the view that acceptance of testimony is an autonomous and way of knowledge. According to this the of testimony does not in any way depend considerations from such of knowledge as and will also that is a of the as of these wide and are in the literature. In then, in to a between and Goldberg is to a the of the have the to be be Goldberg's it more for to their as we it the one is a significant It is from however, whether the the foundation for a general defense of Goldberg by procedures for but he about the of the What of must an agent be to And what of knowledge must an agent In the of to these questions, it is to whether the will in all of the that it is not to say, as I have that the procedures must be or strong having knowledge of the to which a piece of testimony and being to whether the testimony with that knowledge. If then Goldberg's of would be at of into for it would in effect that an agent be in a to testimony for truth, and this would tend to the that testimony can be to be I to Goldberg's of the he it as The process that in of testimony in the testimonial extends to of the social the that is for by (p. the point of in different if the testimonial beliefs of count as and therefore as knowledge, it is the in question are from unreliable testimony by the of their and other Goldberg a of interesting in this important I will however, on what he in defense of The defense begins with the following acquire knowledge (p. Knowledge a of the (p. In interpreting Goldberg's to the in we I his that if a testimonial belief is to count as knowledge, the testimony must have by a for that is both and reliable. The defense of also use of a third cannot be described as the (p. It is not to the line of thought leading from and to In it acquire knowledge testimony, the processes that produce their testimonial beliefs must at some point involve an agent who is to factors that would the reliability of testimony into are and are therefore unable to such factors on their It must be the case, therefore, that other agents an in the processes that produce testimonial beliefs in It must also be the case that these other agents have the ability to and the will to from testimony, by them from it or by them from it at that there are in fact agents who these and other who assume for the of The conclusion of the argument, is and as Goldberg is The I is that it for a between the that are for testimonial beliefs and the that are for testimony for of According to the and the may be in different It follows that it is in some for an agent to be in beliefs than has It has thought that the must have some of to and must be to to in an this allowing that agents may to the conditions concerning that there are factors in the external In to by that it is not agents but rather of agents that are the of epistemic is a and radical Although Goldberg is to be for this interesting I that, in the it is to be I we can its by our attention from and on an who testimony that is an who has in or for the of He many beliefs about what is in his but he does only on the of assertions by his is at to his on of also that the has no reason to his Rather, his is He the testimony concerning conditions the and he has no reason to that that has no sense of the factors that are reliably with testimony that is or That is to say, he has no ability to at least with to testimony about the of Now, as I Goldberg, is from by the and of his he must that has knowledge of conditions in his and also that beliefs about conditions are epistemically it seems to that these propositions are quite is an it is to him as a of epistemic He lacks even a ability to and can he has no in such an In view of these it seems to that has knowledge of the conditions in his in some sense that is quite different than the one that we have in As I it, any theory that would to beliefs, by knowledge and justification from epistemic is however, is to the question of it is possible for premises as as and to to an answer is that when we reflect on the we find that they are plausible than they at first More that both and are acquire knowledge cannot be described as the Goldberg thinks that strong from In however, it seems that our intuitions may only the that can acquire knowledge from and other in they can appropriately place a the of this that a a but in the an the with a of information about the they in that was on their and that as they the is at to make that the no from other Thus, he a who is about his the when they are a that the of and I find that I have to that this is knowledge from the The reason is that the is a and the has as no ability to distinguish between who are and who are if the beliefs that the from the are this is a though does not the is dependent on epistemic It would be to knowledge or epistemic justification to of them. I also have a about It seems to that often have a of concerning the reliability of their and other familiar in of is the I have some are in the come to be in a In a of cases, of this are by of the As a have of that their are reliable sources of information about it is for to the testimony of their concerning such their for the view that their are at least with to that their of I are to the testimony of their concerning in the great are reliable about external than ones, but it seems that the to it for them to a general of is of as they In Goldberg not to make his to the of including some fairly important ones, are by of as the of a public linguistic are the however, I the book it arguments for many significant it several important that are and it attention to a of that to be I it quite having about a of its main I a great deal from I to Goldberg's book.
It is unlikely that Standard Afrikaans has been based on one relatively uniform vernacular. Ana Deumert has convincingly argued that what we recognise as Standard Afrikaans today is a construction to be attributed to language entrepreneurs who strove for a unique South African identity towards the end of the nineteenth and early 20th century. This led to deliberately discarding some of the then metropolitan Dutch linguistic norms. The Afrikaans negative and diminutive systems will be shown to be the linguistic outcome of these conceptions of identity and purity.
We study the influence that image features may have on music tension and liveliness perception. 72 music excerpts from different genres and periods were selected, and 72 still shots were taken from different animation features little known to the subjects. 62 subjects rated the isolated images for tension and liveliness, 37 subjects rated the isolated music excerpts for tension and liveliness, and 153 subjects rated the music excerpts combined with the images for music tension and liveliness, and for music-image congruence. There is a significant variation of tension and liveliness of the music as a function of the tension and liveliness of the pairing image, showing a transfer of mood from image to music. The significance of ANOVA tests showed that 40% of music excerpts were image-sensitive for liveliness and 32% for tension. The transfer of mood was dependent on congruence: music excerpts with high congruence with the image had a higher correlation in tension and liveliness rating deviations with the image ratings. For low congruence, the liveliness correlation was not significant and the tension deviation was negatively correlated with the image tension. Feature transfer from image to music depends on the image-music congruence rated by each subject.
The basic concept of semantic Web,ontology and semantic annotation are described.Then the semantic annotation technology and tool today are introduced and analyzed,and a way of automatic semantic annotation based on HTML documents that contain rich semantic data on the Web is presented.This method couples structural analysis of documents with semantic analysis incorporating domain ontologies and lexical database Hownet,discovers the semantic partition tree corresponding to documents,and annotates HTML documents with semantic lables.The experiment is based on the HTML documents of electronic products,the result shows the method is feasible.
This paper calls for a broadening of the discussion of English language teaching (ELT) practices in Japan. We review issues associated with the global spread of English and link this discussion to the present “standard” English model of ELT in Japan. We propose three major benefits that would follow from an inclusion of non-“standard” (i.e., non American/British) Englishes in Japanese EFL classrooms. First, familiarity with different varieties could increase learners’ confidence when interacting with other non-native speakers (NNSs). Second, we review literature that shows that NNS-NNS interactions actually help learners improve their language skills. Finally, recognition of non-“standard” varieties of English would help Japanese learners challenge monolithic western-centric worldviews that marginalize regional, cultural, and linguistic norms and values. We connect this theory to practice by suggesting some possible changes to ELT in Japan. 本稿では、英語・米語に代表されるいわゆる標準英語の社会的文化的な影響について指摘し、日本英語教育において標準英語に対抗すべく多様な「非標準」英語の教育的可能性を探るものである。著者それぞれの研究を踏まえ、英米語に加え「非標準」英語を日本の英語教育現場で積極的に活用することで期待できる利点を三つ提唱する。第一に「非標準」英語に親しみを持つことにより、ノンネイティブ話者同士の対話に自信が持てるようになる。第二にノンネイティブ話者同士による対話活動は実際に第二言語習得に効果的である。第三に、「非標準」英語に触れることが、西洋的視点に偏りがちな日本人の世界観を省みる機会となり、多様な文化、言語に対する認識の向上が期待できる。以上の点を考察した上で、最後に英語教育現場における「非」標準英語の具体的な導入法ついて提案する。
Increasing the domain of locality by using tree-adjoining-grammars (TAG) encourages some researchers to use it as a modeling formalism in their language application. But parsing with a rich grammar like TAG faces two main obstacles: low parsing speed and a lot of ambiguous syntactical parses. We uses an idea of the shallow parsing based on a statistical approach in TAG formalism, named supertagging, which enhanced the standard POS tags in order to employ the syntactical information about the sentence. In this paper, an error-driven method in order to approaching a full parse from the partial parses based on TAG formalism is presented. These partial parses are basically resulted from supertagger which is followed by a simple heuristic based light parser named light weight dependency analyzer (LDA). Like other error driven methods, the process of generation the deep parses can be divided into two different phases: error detection and error correction, which in each phase, different completion heuristics applied on the partial parses. The experiments on Penn Treebank show considerable improvements in the parsing time and disambiguation process.
Az eladasban a Szeged Treebank fuggsegi fa formatumra torten atalakitasanak folyamatat mutatjuk be. Az eredetileg frazisstrukturalt treebankbl automatikus konverzio eredmenyekeppen letrejott fuggsegi fakat kezi uton ellenriztuk es javitottuk, letrehozva ezzel az els magyar nyelv kezzel annotalt dependenciakorpuszt. Jelenleg az uzleti hireket, ujsaghireket es jogi szovegeket tartalmazo alkorpuszok annotacioja fejezdott be, de terveink kozott szerepel a teljes korpusz atalakitasa fuggsegi fa formatumra. Az elkeszult adatbazis hasznosithato tobbek kozott az informaciokinyeresben es a gepi forditasban is.
The paper describes an auditory experiment aimed at testing whether the intrinsic loudness of a stimulus with a given voice quality influences the way in which it signals affect. Synthesised voice quality stimuli in which intrinsic loudness \nwas systematically manipulated were presented to listeners to test the effect of this manipulation on the affective colouring \nof the stimuli. The results showed that even when devoid of intrinsic loudness variation, non-modal voice quality stimuli \nwere capable of communicating affect. However, changing the loudness of a non-modal voice quality stimulus towards its \nintrinsic loudness resulted in the increase of affective ratings.
A known problem of WordNet is that it is too ne-grained in its sense denitions. For instance, it does not distinguish between homographs and polysemes. This distinction is crucial in many natural language processing tasks. In this paper we propose to distinguish only between homographs withinWordNet data while merging all polysemous senses. The ultimate goal of this exercise is to compute a more coarsegrained version of linguistic database. In order to achieve this task we propose to merge all polysemous senses according to similarity scores computed by a hybrid algorithm. The key idea of the algorithm is to combine the similarity scores produced by diverse semantic similarity algorithms. We implemented the algorithm and evaluated it on the dataset extracted from the WordNet. The evaluation results are promising in comparison to the other state of the art approaches.
The paper draws attention to the discrepancy between norms described in linguistic books (grammar books, dictionaries and language consulting books) and usage.Part of the reason for this discrepancy is the increasingly stronger influence of media on the speech of young people (and other speakers of the Croatian language), but other influences are discussed too.Among them, the problem of the non-distinction of functional styles in the concrete speech situation is isolated.The analysis is based on the materials gathered in the last ten years, and examples are presented at all levels of language.Key words: the Croatian language, standard Croatian language, usage, functional styles, aberrations from the linguistic norm Govorimo hrvatski (We Speak Croatian), a daily radio show which has been broadcasted for several years now, often features a number of queries related to differences in language realizations in different social situations.It is not necessary to be a linguist to observe that discrepancy.In fact, it is not even necessary to be a native speaker of the Croatian language, because even a non-native speaker proficient in the Croatian language can see that in everyday speech, native speakers do not necessary follow linguistic rules.There are differences in pronunciation, linguistic constructions, lexicon, etc.Even the first generations of Croatian immigrants, who use Croatian when talking to each other, see the influence of their new language, for example English, which they use in everyday situations.There are many reasons for these linguistic differences.To explain them, one should understand the meaning of terms such as language, standard language, usage, dialect and functional style.
Bond ratings on state‐issued debt provide a signal to credit markets that help them charge an appropriate interest rate, based on the risk of payment default. Though actual default may occur only in extreme circumstances, observed differences in ratings and interest costs across states and time demonstrate that a sound economy, strong financials, and stable policies matter. When data on the factors that presumably affect ratings is public and easily accessible, making sense of differences of opinion between bond rating agencies is difficult. We suggest that such differences—observed as so‐called split bond ratings—are often ephemeral. Utilizing a simulation method to uncover the latent credit risk presented by each state, we show that split ratings on state bonds are often due to the fact that presumed category overlap between rating agencies is absent when evaluated on a common latent scale. Most observed state bond rating splits from 1997 through 2006 can be explained by this category mismatch. Our approach has broad implications for pricing state debt, as well as pricing rated debt in other capital market sectors.
Because of the joining behavior of Persian script and its orthographic variation, the morphological and syntactic annotations of multi-token units meet various issues. By the analysis of Perso-Arabic script and its problems, the various collocation types of the tokens including the compositional, non-compositional and the new semicompositional constructions are described in the present paper. Then, to illustrate these constructions, the static and dynamic multi-token units will be presented for the generative and non-generative structures of the main categories including the verbs, infinitives, prepositions, conjunctions, adverbs, adjectives and nouns. Defining the multi-token unit templates for these categories is one of the important results of this research. The findings can be input to the segmentation module of the Persian Treebank generator system. The other usage of the present research is in the design and implementation of the morphological analyzers and syntactical parsers.
Computer-aided Acquisition of Semantic Knowledge (CASK) is aimed at describing a number of semantic fields of a few European languages using data mining techniques elaborated within the framework of the new paradigm of computation known as Knowledge Discovery in Databases (KDD). CASK's motivation is to dig deeper in order to find building blocks which could be used in various sophisticated ways. The project is interdisciplinary involving scientific cooperation of experts in linguistics with information engineers. The task of linguists consists in an interactive (computer-aided) discovery of ontology-based definitions of feature structures using the SEMANA (Semantic Analyser) software which was designed especially in order to build linguistic databases with semantic knowledge.
The referential expressions of online sales clothing generally consist of the central word and multiple attributives.Central words are always created in recent years.Multiple attributives are rich in content and varied in word order.Referential expressions of online sales clothing reflect some cultural trends in our society.However,some of the expressions deviate from linguistic norms and call for standardization.
The part-of-speech determination is necessary for resolving the part-of-speech ambiguity in English-Korean machine translation. The part-of-speech ambiguity causes high parsing complexity and makes the accurate translation difficult. In order to solve the problem, the resolution of the part-of-speech ambiguity must be performed after the lexical analysis and before the parsing. This paper proposes the CatAmRes model, which resolves the part-of-speech ambiguity, and compares the performance with that of other part-of-speech tagging methods. CatAmRes model determines the part-of-speech using the probability distribution from Bayesian network training and the statistical information, which are based on the Penn Treebank corpus. The proposed CatAmRes model consists of Calculator and POSDeterminer. Calculator calculates the degree of appropriateness of the partof-speech, and POSDeterminer determines the part-of-speech of the word based on the calculated values. In the experiment, we measure the performance using sentences from WSJ, Brown, IBM corpus.
Pp. xiii, 265, Cambridge, Cambridge University Press, 2007, $90.00. This book is not gracefully written, but it is worth penetrating its stylistic carapace if one values tough argument. When I first read the title, the naughty thought occurred to me, what a pleasure it would be to be relieved of one's individual epistemic responsibilities; of ever again having to assume the burden of finding anything out for oneself regarding matters either of fact or of value. But how could one bring off such a feat? Part I, on semantic anti-individualism, begins with an account of the communication of knowledge, and makes a case for the existence of public linguistic norms from the occurrence of successful communication on the one hand, and the existence of misunderstandings on the other. Then he argues from the reality of public linguistic norms to anti-individualism with regard to the language of thought. Part II is concerned with epistemic anti-individualism, and starts applying it to the epistemic dimension of knowledge communication. Objections are mounted which take into account the phenomena of gullibility and rationality. A final chapter recommends what the author calls ‘an ‘active’ epistemic anti-individualism’; in which the reader is given instruction about ‘nearby possible worlds’ in which someone ‘forms the testimonial belief that there is milk in the fridge, under conditions in which there is no milk in the fridge’ (p. 214). Goldberg has a good deal to say on what he calls ‘the consumption of testimony’ (a phrase I find curious, though he frequently uses it). On the acquisition by children of beliefs based on testimony, we seem to have intuitions of which the implications conflict with one another. It seems perverse to deny, on the grounds of her cognitive immaturity, that three-year-old Sally knows that her mother has just bought some ice-cream for dinner, on the basis of what she has been told by her uncle, when in fact her mother has done this. And yet we are also inclined to say that the cognitive immaturity of children of this age makes it impossible for them to have adequate grounds for believing that such testimony is credible, and therefore for knowing the fact in question. Goldberg displays an impressive mastery of the evidence amassed on these matters by empirical psychologists. One of these argues that, up to a certain age, one can indeed be properly said to know through what one is told by another person, even when one has not acquired the mental capacities necessary adequately to evaluate such information. (It appears to me that it is superstitious to believe that there is a ‘yes’ or ‘no’ answer to the question whether Sally knows about the ice-cream or not; in a sense she does, in a sense she doesn't.) The truth on the central topic under consideration, I believe, may properly be summarized something like this. By attending to the testimony of others, we get the hang of a process which is essentially private to each one of us - using our minds to attend to phenomena of sensation or feeling, to hypothesize more or less intelligently, to judge more or less reasonably what is so, and to make more or less responsible decisions accordingly. Some would say that these activities were not essentially private in that, if we had devices for inspecting the interiors of others' brains, we could observe them directly; but I remain unconvinced. Wittgenstein and his followers have demonstrated, I would concede, that if there were not public criteria for the occurrence of private mental acts, we could not talk of them, or probably even undergo or engage in them. But there are such criteria; we know what it is for people to look and sound as though they had just made an observation, or were trying to puzzle something out, or had just come to a decision after moments or months of hesitation. Having once gained the use of our mental faculties via these criteria, we can use them for ourselves, as even the most insensitive or stupid do to some extent, and persons of genius do to an exceptional degree. Our mental performances are nonetheless essentially private acts, of which we are directly aware, and of which we can enhance our awareness by suitably directed attention. The moral is, that the acquisition and cultivation of our capacity to gain knowledge, whether by testimony or otherwise, is a matter of both ‘public’ social influence and ‘private’ individual practice. We must apply this capacity to some extent for ourselves, as individuals, if we are to live reasonably and responsibly. It will not do to deny individualism so thoroughly as to imply the negation of this enormously important fact. There are times when it is proper to be Athanasius contra mundum. The balance of the public and private, the social and individual, contribution to knowledge, is of the essence. If you tip the balance too far in the direction of the public and social, as I think Goldberg's account might lead you to do, you bid fair to cut off at the root all original creativity in science, morality, or the arts.
Abstract In the preceding chapter we have delineated the boundaries of our inquiry by defining a cross-linguistically applicable domain of predicative (alienable) possession. On the basis of this definition we can now proceed to build a data base, which comprises the relevant linguistic material from the languages in the sample. Once this task has been completed (and we will assume here that it has) our next step is to construct a typology of predicative possession, on the basis of observable similarities and differences among the constructions included in the cross-linguistic database.
Multi-word Lexical Units (MWLU) are of great importance in language in general, and in Natural Language Processing in particular, since they are not governed by the free rules of the system. In this article, we give an overview of the different types of phraseological units, explaining briefly each one's features. Our priority being to process idioms automatically in Basque texts, we concisely analyze several approaches for the inflectional description of MWLUs, and then, we explain the system we have developed for Basque: (i) a general representation for describing MWLUs in the lexical database for Basque (EDBL), (ii) HABIL, a tool capable of detecting and analyzing them based on the features described in the database, and (iii) a constraint grammar for disambiguating ambiguous MWLUs.
We review lexical Association Measures (AMs) that have been employed by past work in extracting multiword expressions. Our work contributes to the understanding of these AMs by categorizing them into two groups and suggesting the use of rank equivalence to group AMs with the same ranking performance. We also examine how existing AMs can be adapted to better rank English verb particle constructions and light verb constructions. Specifically, we suggest normalizing (Pointwise) Mutual Information and using marginal frequencies to construct penalization terms. We empirically validate the effectiveness of these modified AMs in detection tasks in English, performed on the Penn Treebank, which shows significant improvement over the original AMs. 1
Business English is characterized by a specialized vocabulary,polysemy,stylistic norms of formal,nicety,preciseness,concision and emerging new words.It can improve the study effectiveness to paying attention to the accumulation of professional knowledge,to understand vocabulary through context clues,and to learn new words by chunk approach and concern about the latest business information.
Norms are essential to the human condition. Whether in the guise of tradition, culture, canon or rules, norms are therefore central to studies in the humanities. This book focuses on Russian language culture of the post-revolutionary and post-Soviet periods, times when norms — linguistic and otherwise — have been eagerly debated, challenged, broken and redefined. Exploring the intersections between linguistic authority and creative response, an international team of scholars examines different realms of linguistic practice (literary fiction, internet slang, literary criticism and aesthetics, writers’ blogs, linguistic play) and various arenas for “talk about talk” (the classroom, blogs, the media, or the courtroom). By combining various approaches and disciplines — linguistics, literary criticism, new media studies — the book as a whole explores the multiplicity of meanings that are accorded to the notion of linguistic norms in the Russian community. The result is both a broad and a detailed picture of important trends in modern Russian language culture.
Any work of art is an original bearer of information. At the same time every cultural phenomenon codes a message through the linguistic means reflecting the specific character of the given work of art. In connection with the individualisation of expressive means in the 20th century the tendency to informational isolation appears in all kinds of art the so-called ciphering of the sense that is not lying on the surface. Not only the quantity of breaches of those linguistic norms, which a composer, transferring a message to the listener, uses in his work, is of primary importance for the composer, but to what extent they are important for the listener and aimed at him. Depending on various circumstances (socio-cultural, moral, aesthetic, etc.) the listener can form his own hypotheses concerning deciphering of the concrete text, revealing his cross-initiatives. Resting upon the semiotic investigation, the author of the article tried to mark out the stages of the listener's perception of a 20th-century musical composition through the definitions of codes and subcodes.
Coordinating structure is a special phenomena with high frequence in human languages. It enjoys a special status in linguistic theories especially in dependency grammars, which require the exceptional treatment. This paper proposes three schemes to analyse coordinating structure and builds dependency treebanks for automatic parsing. We find that different schemes to annotate and parse the coordinating structure bring some distinctions in precision and other aspects. The factors, such as the part-of-speech categories matching with their syntactic functions, may play an important role in dependency parsing. We also calculate the mean dependency distance of three schemes, and the result shows that single-directional and smaller mean dependency distance is helpful for better machine learning and higher precision.
In this thesis, I present arguments for a model of language acquisition with three characteristics. These are (1) Continuity in the abstract principles of Universal Grammar; (2) Lexical Learning, or the setting of syntactic parameters based upon the acquisition of morphology; and (3) Morpholexical Learning, which is the abstraction of morphological patterns and generalizations from a lexical database. Continuity accounts for what is invariant in language development. Lexical Learning accounts for what is languageparticular, and which therefore must be learned. Morpholexical Learning accounts for the sequence of developmental stages observed in child language data. The main goal of this thesis is to demonstrate that Morpholexical Learning, in conjunction with paradigmatic structure in the lexicon, provides a model for the acquisition of inflectional morphology. I demonstrate this proposal with data on the acquisition of subject-verb agreement morphology in German. In Chapter One, I present an introduction to the concerns and main proposals of this thesis. In Chapter Two, I motivate the existence of paradigmatic structure with three diachronic case studies. In Chapter Three, I return to the acquisitional debates introduced in Chapter One. I argue that the Continuity Hypothesis represents a preferable alternative to the Maturational Hypothesis. Next, I show that Lexical Learning of clausal representations is superior to the Lexical Projection Hypothesis and the Full Competence Hypothesis. I argue that Morpholexical Learning provides an answer to the Developmental Problem which Continuity and Lexical Learning create. In Chapter Four, I provide an extended case study of the acquisition of subjectverb agreement in German. The construction of word-specific paradigms during the stages under examination accounts for the pattern of agreement errors which German children produce. In Chapter Five, I continue the analysis of paradigm mixture begun in Chapter Two. The patterns of paradigm mixture attested Latin, German, and Icelandic are in essence identical to one another, which suggests universal principles of inflectional organization. In Chapter Six, I conclude the thesis with a sketch of how children develop from the "word-specific paradigm" stage to the "general paradigm stage".
espanolComo es bien sabido, aunque para los hablantes de una lengua las variedades dialectales resulten mas evidentes en los planos lexico, fonetico o fonologico, ellas se advierten en todos los niveles del lenguaje, orbita de la que, por supuesto, no escapa la sintaxis. Asi, en el caso particular del espanol de Buenos Aires, el uso del Preterito Perfecto Compuesto del Modo Indicativo difiere sensiblemente de la norma castellana, a la vez que la conciencia de los hablantes de la lengua respecto de el es practicamente nula: o lo niegan por completo, alegando que prefieren siempre el Preterito Perfecto Simple, o bien aducen que lo emplean segun la norma de Madrid; lo cual, como se vera a lo largo de nuestro trabajo, no resulta de ese modo en ninguno de los dos casos. Asi pues, intentaremos problematizar las cuestiones de norma y uso, en relacion con la conciencia de los hablantes portenos respecto de su empleo de los tiempos pasados. Para ello, partiremos de un trabajo de campo que hemos realizado y que nos ha permitido esbozar algunos matices caracteristicos del uso del tiempo verbal que nos ocupa, es decir, el Preterito Perfecto Compuesto del Modo Indicativo del dialecto rioplatense. EnglishIt is well known that dialectal language variations appear at every level of language including syntax. However, speakers are usually aware of lexical, phonetics, and phonological variations only. In this particular case, as expected, the use of perfect tenses in Buenos Aires (Argentina) is very different from that of Madrid (Spain). The problem is that most Argentinean speakers know how to use the Present Perfect according to Spanish rules they have learned in school, but their speech do not matches their learning. Most Argentinean speakers would say (and they believe) that they do not use the Present Perfect in everyday life, when they actually do, albeit in a different way. That is why I conducted a survey among speakers of all kind of age, in order to distinguish some specific characteristics of the Present Perfect use in rioplatense dialect. Finally, I intend to discuss the concept of language norm and use related to speakers' awareness in Buenos Aires.
This paper presents a comparative study of Judgment and Assessing frames in English and Portuguese. The aim is to verify the possibility of using the FrameNet frames to construct a lexical database for Brazilian Portuguese. The research corpus is composed by 50 legal documents, totalizing 1.055,535 tokens and 39,108 types. Through a contrastive method the Judgment and Assessing frames were selected and translation equivalents for the English lexical units were established. The points considered in this research were the polysemy and the semantic relations of words. The polysemy is the main difficulty in applying FrameNet frames for Portuguese description.
Parsing is an important process of Natural Language Processing (NLP) and Computational Linguistics which is used to understand the syntax and semantics of natural language sentences confined to the grammar. Parsing models need syntax and semantic coverage for better interpretation of natural language sentences. Though statistical parsing with trigram language models gives better performance through tri-gram probabilities and large vocabulary size, it has some disadvantages like lack of support in syntax, free ordering of words and long distance relationship which are the challenging features of the Tamil language. Grammar based structural parsing provides solutions to some extent. To overcome these disadvantages, structural component is to be involved in statistical approach which results in hybrid models like phrase and dependency models. To add the structural component, balance the vocabulary size and meet the challenging features, lexicalized and statistical parsing (LSP) is to be employed with the assistance of hybrid models. To incorporate all the features in complex and large sentences, phrase structure model may not be suitable to a larger extent. When dependency relations are applied among words, direct relationships can be established. Lexicalized and statistical parsing of natural language text in Tamil language using dependency model will give better performance than using phrase structure model. New part of speech (POS) and dependency tag sets for Tamil language have been Treebank has been developed with 326 sentences which comprises more than 5000 words with manual annotation. It has been extended to 1000 sentences using bootstrapping and manual correction and used to train the dependency model. This LSP with dependency model provides better results and covers all the features of Tamil language.
In the article we compare the role of the dictionary and the lexical database, and address the issue of language register and correctness in dictionaries. We then deal with various types of sense distribution in dictionaries, the history of the word, and the principles of selection of dictionary headwords. We cite the corpus as an essential source for the treatment of meaning, collocation and syntagmatics, and investigate ways of interpreting corpus data – corpus profiling of headwords. We conclude with the thought that a dictionary represents the central language standard, whereby all of the expressed linguistic opinions contained in it must be based on corpus evidence.
Chemical sensitivity (CS) is common in the adult population and implies negative effects (e.g. physical symptoms, negative affects and behavioral disruptions) of odorous chemical substances. In Study 1, relations between self-reported CS, negative affect and neuroticism were investigated among Swedish university students (n = 103). CS and neuroticism were positively correlated, suggesting that highly neurotic persons, compared to low- neurotic persons, more easily respond negatively to environmental odors. In Study 2 (n = 40), relations between CS, self-reported noise sensitivity, physical symptoms and odor perception (pleasantness, intensity and familiarity) were examined in a sub-sample of high (HCS) and low (LCS) chemical sensitivity. The HCS group reported higher odor intensity and noise sensitivity. No differences were found in odor pleasantness and familiarity ratings, and in physical symptoms. Overall, the results suggest that normal variation in CS has a general, sensory-unspecific basis.
BACKGROUND: Memory impairment and verbal learning are the most common cognitive deficits associated with schizophrenia. Hopkins Verbal Learning Test (HVLT) is considered to be the most reliable test to asses memory and verbal learning in this mental illness. AIMS: to create one form of the HVLT which would suit our linguistic and cultural context and to study the characteristics of this test in a group of healthy subjects. METHODS: The HVLT consists of a list of 12 words belonging to 3 semantic categories and which are read orally to the subject with an immediate and differed recall. The first part of this work was to select words from a lexical database in order to create the list of the HVLT. The test was then administered to 103 subjects aged from 17- to 45-years-old (mean=27,4; SD =7,3) and having between 1 and 20 years of education ( mean=12,2; SD=5,3). RESULTS: No statistical difference was found within performances of the HVLT across gender and sex. Whereas, years of education was found to have an impact on performances. Although statistically difference was found across level of education. CONCLUSION: Our study permitted us to create one form of the HVLT which well suits our Tunisian context and which we could use to evaluate memory functions among people suffering from schizophrenia.
Un assunto sempre più condiviso nell’ambito degli studi sull’acquisizione sia di L1 che di L2 è che l’evidenza empirica privilegiata debba essere rappresentata da corpora di produzioni scritte o orali degli apprendenti, estensivamente annotate a molteplici livelli di rappresentazione linguistica. Più in generale, corpora lemmatizzati e annotati a livello morfosintattico fanno ormai parte dello strumentario comune del linguista. Accanto ad essi, si fa però strada l’esigenza di disporre di risorse testuali più sofisticate dal punto di vista delle modalità di esplorazione linguistica, come ad esempio corpora annotati a livello sintattico (le cosiddette treebank). Questi consentono infatti di osservare i processi di convergenza degli apprendenti verso la lingua “obiettivo” anche a livello di specifici tratti grammaticali astratti o di macro-strutture linguistiche.
This paper presents a simple and effective approach to improve dependency parsing by using subtrees from auto-parsed data. First, we use a baseline parser to parse large-scale unannotated data. Then we extract subtrees from dependency parse trees in the auto-parsed data. Finally, we construct new subtree-based features for parsing algorithms. To demonstrate the effectiveness of our proposed approach, we present the experimental results on the English Penn Treebank and the Chinese Penn Treebank. These results show that our approach significantly outperforms baseline systems. And, it achieves the best accuracy for the Chinese data and an accuracy which is competitive with the best known systems for the English data.
The rise of a standard language is inextricably connected to value judgements about linguistic variants. During standardisation processes certain linguistic expressions are marked as ‘correct’ and prestigious and subsequently selected as a standard form whereas other linguistic features are labelled as ‘bad’ and corrupt use of language. As Haugen rightly points out, ‘[w]here a norm is to be established, the problem will be as complex as the sociolinguistic structure of the people involved’ (1997, p. 349). It is, after all, the socio-political context that influences the evaluation of the language usage. An established norm, in turn, has many socio-political consequences. This work will be concerned with the establishment of linguistic norms in the history of specific languages and the socio-political contexts in which these norms arose as well as their influence on actual usage. More precisely, this study seeks to trace the development of the subjunctive mood in English and German, with a special focus on the Austrian variety,1 during part of their standardisation processes, namely the eighteenth century. As grammarians were attempting to shape and codify a prestige variety during this period, the question arises whether and to what extent these normative grammarians influenced the development of the inflectional subjunctive. After all, the subjunctive mood has been claimed to have been on the decline in both English and German in the eighteenth century (cf. for English: Strang, 1970, p. 209; Turner, 1980, p. 272; Görlach, 2001, p. 122; for German: von Polenz, 1994, pp. 261–263).
Since the inception of the Senseval series there has been a great deal of debate in the word sense disambiguation (WSD) community on what the right sense distinctions are for evaluation, with the consensus of opinion being that the distinctions should be relevant to the intended application. A solution to the above issue is lexical substitution, i.e. the replacement of a target word in context with a suitable alternative substitute. In this paper, we describe the English lexical substitution task and report an exhaustive evaluation of the systems participating in the task organized at SemEval-2007. The aim of this task is to provide an evaluation where the sense inventory is not predefined and where performance on the task would bode well for applications. The task not only reflects WSD capabilities, but also can be used to compare lexical resources, whether man-made or automatically created, and has the potential to benefit several natural-language applications.
Virtual reality technology is argued to be suitable to the simulation study of mass evacuation behavior, because of the practical and ethical constraints in researching this field. This article describes three studies in which a new virtual reality paradigm was used, in which participants had to escape from a burning underground rail station. Study 1 was carried out in an immersion laboratory and demonstrated that collective identification in the crowd was enhanced by the (shared) threat embodied in emergency itself. In Study 2, high-identification participants were more helpful and pushed less than did low-identification participants. In Study 3, identification and group size were experimentally manipulated, and similar results were obtained. These results support a hypothesis according to which (emergent) collective identity motivates solidarity with strangers. It is concluded that the virtual reality technology developed here represents a promising start, although more can be done to embed it in a traditional psychology laboratory setting.
Grammar extraction in deep formalisms has received remarkable attention in recent years. We recognise its value, but try to create a more precision-oriented grammar, by hand-crafting a core grammar, and learning lexical types and lexical items from a treebank. The study we performed focused on German, and we used the Tiger treebank as our resource. A completely hand-written grammar in the framework of HPSG forms the inspiration for our core grammar, and is also our frame of reference for evaluation.
Cette etude rend compte des mecanismes semantico-discursifs de construction du sens lexical, mis en œuvre lors de la realisation de dialogues argumentatifs, realises en francais en Cote d’Ivoire. Notre objectif est d’observer la maniere dont les dialogues argumentatifs peuvent servir de lieu de construction des representations sociales, par le jeu de la construction du sens des unites lexicales, afin de saisir les praxis linguistiques et les normes socio-discursives sous-jacentes. Pour ce faire, nous nous appuyons sur l’analyse d’un corpus de debats televisuels, enregistres en Cote d’Ivoire, nous permettant de decrire les processus constitutifs du sens des objets discursifs amitie et homme politique. L’etude de ces expressions nous a permis de conclure a leur dynamisme semantique, en mettant en valeur la stabilite, l’instabilite et la plasticite de leur sens lexical dans le dialogue argumentatif, en fonction de leur contexte social de production. En outre, nous avons pu etablir des correspondances entre les positions enonciatives et les formes linguistiques nous permettant de saisir les doxas constitutives des deux dialogues argumentatifs, ces phenomenes etant lies a des enjeux symboliques, voire ideologiques.
The Romance languages derive, via Latin, from the Italic branch of Indo-European. Their modern distribution is the product of two major phases of conquest and colonisation. The first, between c. 240 BC and c. AD 100 brought the whole Mediterranean basin under Roman control; the second, beginning in the sixteenth century, annexed the greater part of the Americas and sub-Saharan Africa to Romance-speaking European powers. Today, some 665 million people speak, as their first or only language, one that is genetically related to Latin. Although for historical and cultural reasons preeminence is usually accorded to European Romance, it must not be forgotten that European speakers are now outnumbered by non-Europeans by a factor of nearly three to one. The principal modern varieties of European Romance are indicated on Map 8.1. Nouniformly acceptable nomenclature has been devised for Romance and the choice of term to designate a particular variety can often be politically charged. The Romance area is not exceptional in according or withholding the status of ‘language’ (in contradistinction to ‘dialect’ or ‘patois’) on sociopolitical rather than linguistic criteria, but additional relevant factors in Romance may be cultural allegiance and length of literary tradition. Five national standard languages are recognised: Portuguese, Spanish, French, Italian and Rumanian (each treated in an individual chapter below). ‘Language’ status is usually also accorded on cultural/literary grounds to Catalan and Occitan, though most of their speakers are bilingual in Spanish and French respectively, and the ‘literary tradition’ of Occitan refers primarily to medieval Provençal, whose modern manifestation is properly considered a constituent dialect of Occitan. On linguistic grounds, Sardinian too is often described as a language, despite its internal heterogeneity. Purely linguistic criteria are difficult to apply systematically: Sicilian, which shares many features with southern Italian dialects, is not usually classed as an independent language, though its linguistic distance from standard Italian is no less than that separating Spanish from Portuguese. ‘Rhaeto-Romance’ is nowadays used as a cover term for a number of varieties spoken in southern Switzerland (principally Engadinish, Romansh and Surselvan) and in the Dolomites, but it is no longer taken to subsume Friulian. Romansh (local form romontsch) enjoys an official status for cantonal administration and so perhaps fulfils the requirements of a language. Another special case is Galician, located in Spainbut genetically and typologically very close to Portuguese; in the wake of political autonomy, galego/gallego now enjoys protected ‘language’ status in Spain, although elsewhere it continues to be thought of (erroneously) as a regional dialect of Spanish. Corsican, which clearly belongs to the Italo-Romance group, would be in a similar position if the separatist movement gained autonomy or independence from France. Outside Europe, Spanish, Portuguese and French, in descending order of native speakers,have achieved widest currency, though many other varieties are represented in localised immigrant communities, such as Sicilian in New York, Rumanian in Melbourne, Sephardic Spanish in Seattle and Buenos Aires. In addition, the colonial era gave rise to a number of creoles, of which those with lexical affinities to French are now the most vigorous, claiming some ten million speakers. In general, European variants are designated by their geographical location; ‘Latin’,as a term for the vernacular, has survived only for some subvarieties of Rhaeto-Romance (ladin) and for Biblical translations into Judaeo-Spanish (ladino). ‘Romance’ derives, through Spanish and French, from ROMĀNICĒ ‘in the Roman fashion’ but also ‘candidly, straightforwardly’, a sense well attested in early Spanish. The terminological distinction may reflect early awareness of register differentiation within the language, with ‘Latin’ reserved at first for formal styles and later for written language and Christian liturgy. The idea, once widely accepted, that Latin and Romance coexisted for centuries as natural, spoken languages, is now considered implausible. Among the chief concerns of Romance linguists have always been: the unity orotherwise of the proto-language, the causes and date of dialect differentiation and the classification of the modern variants. Plainly, Romance does not derive from the polished literary models of Classical Latin. Alternative attestations are quite plentiful, but difficultin may be stylistic artifice; inscriptional evidence is formulaic; the abundant Pompeian graffiti may be dialectal, and so on. Little is known of Roman linguistic policy or of the rate of assimilation of new conquests. We may however surmise that a vast territory, populated by widely different ethnic groups, annexed over a period exceeding three centuries, conquered by legionaries and first colonised by settlers who were probably not native speakers of Latin, and never enjoying easy or mass communications, could scarcely have possessed a single homogeneous language. The social conditions which must have accompanied latinisation – including slaveryand enforced population movements – have led some linguists to postulate a stage of creolisation, from which Latin slowly decreolised towards a spoken norm in the regions most exposed to metropolitan influences. Subsequent differentiation would then be due to the loss of administrative cohesion at the break-up of the Empire and the slow emergence of local centres of prestige whose innovations, whether internal or induced by adstrate languages, were largely resisted by neighbouring territories. Awareness of the extent of differentiation seems to have come very slowly, probably stimulated in the west by Carolingian reforms of the liturgical language, which sought to achieve a uniform pronunciation of Church Latin at the cost of rendering it incomprehensible to uneducated churchgoers. Sporadic attestations of Romance, mainly glosses and interlinear translations in religious and legal documents, begin in the eighth century. The earliest continuous texts which are indisputably Romance are dated: for French, ninth century; for Spanish and Italian, tenth; for Sardinian, eleventh; for Occitan (Provençal), Portuguese and Rhaeto-Romance, twelfth; for Catalan, thirteenth; for Dalmatian (now extinct), fourteenth; and for Rumanian, well into the sixteenth century. Most classifications of Romance give precedence, explicitly or implicitly, to histor-ical and areal factors. The traditional ‘first split’ is between East and West, located on a line running across northern Italy between La Spezia and Rimini. Varieties to the northwest are often portrayed as innovating, versus the conservative south-east. For instance, West Romance voices and weakens intervocalic plosives: SAPŌNE ‘soap’ > Ptg. sabão, Sp. jabón, Fr. savon, but It./Sard. sapone, Rum. sa˘pun; RŌTA ‘wheel’ > Ptg. roda, Sp. rueda, Cat. roda, Fr. roue, but It./Sard. rota, Rum. roata˘; URTĪCA ‘nettle’> Ptg./Sp./ Cat. ortiga, Fr. ortie, but Sard. urtica, It. ortica, Rum. urzica˘. The West also generalises /-s/ as a plural marker, while the East uses vocalic alternations: Ptg. as cabras ‘the goats’, Cat. les cabres, Romansh las chavras, contrast with It. le capre and Rum. caprele. In vocabulary, we could cite the verb ‘to weep’, where the older Latin word PLANGE˘RE survives in the East (Sard. pranghere, It. piangere, Rum. a plînge) but is completely replaced in the West by reflexes of PLORĀRE (Ptg. chorar, Sp. llorar, Cat. plorar, Oc. plourà, Fr. pleurer). In this classification, each major group splits into two subgroups: ‘East’ into Balkan-Romance and Italo-Romance, ‘West’ into Gallo-Romance and IberoRomance. The result is not entirely satisfactory. While, for example, Arumanian dialects and Istro-Rumanian group quite well with Balkan-Romance, our scant evidence of Dalmatian suggests it shared as many features with Italo-Romance as with the Balkan group. Catalan is a notorious difficulty, having been subject for centuries to alternating Occitan and Spanish influences. The unity of ‘Rhaeto-Romance’ also fails to survive closer scrutiny: Ladin groups fairly well with Friulian as part of Italo-Romance, but southern Swiss varieties share many features with eastern French dialects. ‘Family-tree’ classifications, in which variants are each assigned to a single node,give only a crude indication of relationships in Romance and tend to obscure theLatin and patterns of contact. This is readily illustrated from the lexicon. The PLANGE˘RE/ PLORĀRE example, though supportive of the East-West split, is in fact rather atypical. More common are innovations spreading from central areas but failing to reach the periphery. ‘To boil’ is Ptg. ferver, Sp. hervir, Rum. a fierbe (< FERVĒRE/FERVE˘RE), but Cat. bullir, Oc. boulí, Fr. boullir, It. bollire (< BULLĪRE, originally ‘to bubble’); ‘to request’ is Ptg./Sp. rogar, Rum. a ruga (< ROGĀRE), but Cat. pregar, Oc. pregá, Fr. prier, It. pregare (< PRECĀRE, originally ‘to pray’); ‘to find’ is Ptg. achar, Sp. hallar, Rum. a afla, but Cat. trobar, Oc. trobà, Fr. trouver, It. trovare (both forms are metaphorical – classical INVENĪRE and REPERĪRE do not survive). Among nouns, we may cite ‘bird’: Ptg. pássaro, Sp. pájaro, Rum. pasa˘re (< *PASSARE), versus Oc. aucèu, Fr. oiseau, Romansh utschè, It. uccello (< AUCELLU); and ‘cheese’: Ptg. queijo, Sp. queso, Rum. cas¸ (< CĀSEU), versus Cat. formatge, Oc. froumage, Fr. fromage, It. formaggio (< [CASEU] FORMATICU ‘moulded [cheese]’). Almost the same distribution is found in a morphosyntactic innovation: the Latin synthetic comparative in -IŌRE nowhere survives as a productive form, but peripheral areas have MAGIS as the analytic replacement (‘higher’ is Ptg. mais alto, Rum. mai înalt) whereas the centre prefers PLŪS (Fr. plus haut, It. più alto). Despite this differential diffusion and the divergences created by localised borrowingfrom adstrate languages (notably from Arabic into Portuguese and Spanish, from Germanic into northern French, from Slavonic into Rumanian), the modern Romance languages have a high degree of lexical overlap. Cognacy is about 40 per cent for all major variants using the standard lexicostatistical 100-word list. For some language pairs it is much higher: 65 per cent for French-Spanish (slightly higher if suffixal derivation is disregarded), 90 per cent for Spanish-Portuguese. This is not, of course, a guarantee of mutual comprehensibility (untrained observers are unlikely to recognise the historical relationship of Sp. /oxa/ ‘leaf’ to Fr. /fœj/), but a high rate of cognacy does increase the chances of correct identification of phonological correspondences. Intercomprehensibility is also good in technical and formal registers, owing to extensive borrowing from Latin, whether of ready-made lexemes (abstract nouns are a favoured category) or of roots recombined in the naming of a new concept, like Fr. constitutionnel, émetteur, exportation, ventilateur, etc. Indirectly, coinings like these have fed the existing propensity of all Romance languages for enriching their word stock by suffixal derivation. Turning to morphosyntax, we find that all modern Romance is VO in its basic wordorder, though southern varieties generally admit some flexibility of subject position. A much reduced suffixal case system survives in Rumanian, but has been eliminated everywhere else, with internominal relations now expressed exclusively by prepositions. All variants have developed articles, the definite ones deriving overwhelmingly from the demonstrative ILLE/ILLA (though Sardinian and Balearic Catalan use IPSE/ IPSA), the indefinite from the numeral ŪNU/ŪNA. Articles, which precede their head noun everywhere except in Rumanian where they are enclitic, are often obligatory in subject position. Concord continues to operate throughout noun phrases and between subject and verb, though its range of exponents has diminished with the loss of nominal case. French is eccentric in virtually confining plural marking to the determiner, though substantives still show number in the written language. Parallel to the definite articles, most varieties have developed deictic object pronouns from demonstratives. These, like the personal pronouns, often occur in two sets, one free and capable of taking stress, the other cliticised to the verb. There is some evidence of the grammaticalisation of anclitic and in the marking of specific animate objects. This latter is widespread (using a in West Romance and pe in Rumanian) but not found in standard French or Italian. Suffixal inflection remains vigorous in the common verb paradigms everywhere butin French. Compound tense forms everywhere supplement the basic set, though the auxiliaries vary: for perfectives, HABĒRE is most common: ‘I have sung’ is Fr. j’ai chanté, It. ho cantato, Sp. he cantado but Ptg. tenho cantado (< TENĒRE originally ‘to hold’) and Cat. vaig cantar (< VĀDO CANTĀRE) – an eccentric outcome for a combination that would be interpreted elsewhere as a periphrastic future (‘I am going to sing’). Most Romance varieties have a basic imperfective/perfective aspectual opposition, supplemented by one or more of punctual, progressive and stative. The synthetic passive has given way to a historically reflexive medio-passive which coexists uneasily with a reconstituted analytic passive based on the copula and past participle. The replacement of the Latin future indicative by a periphrasis expressing volition or mild obligation (HABĒRE is again the most widespread auxiliary, but deppo ‘I ought’ is found in Sardinian and voi ‘I wish’ in Rumanian) provided the model for two new synthetic paradigms, the future itself and the conditional, which has taken over a number of functions from the subjunctive. The subjunctive has also been affected by changes in complementation patterns, but a few new uses have evolved during the documented period of Romance, and its morphological structure, though drastically reduced in spoken French, remains largely intact. In phonology, it is more difficult to make generalisations (see the individual languagesections below and, for the development from Latin to Proto-Romance, pages 146150). We can, however, detect some shared tendencies. The rhythmic structure is predominantly syllable-timed. Stress is dynamic rather than tonal and, on the whole, rather weak – certainly more so than in Germanic; some variants, notably Italian, do use higher tones as a concomitant of intensity, but none rely on melody alone. The loss of many intertonic and post-tonic syllables suggests that stress may previously have been stronger, witness IŪDI˘CE[’i˘u-di-ke] ‘judge’ > Ptg. juiz, Sp. juez, Cat. jutge, Fr. juge; CUBI˘TU [’ku-bi-tu] ‘elbow’ > Sp. codo, Fr. coude, Rum. cot. The elimination of phonemic length from the Latin vowel system has been maintained with only minor exceptions. A strong tendency in early Romance towards diphthongisation of stressed mid vowels has given very varied results, depending on whether both higher and lower mid vowels were affected, in both open and closed syllables, and on whether the diphthong was later levelled. Romance now exhibits a wide range of vowel systems, but those of the south-central group are noticeably simpler than those of the periphery: phonemic nasals are found only in French and Portuguese, high central vowels only in Rumanian, and phonemic front rounded vowels only in French, some Rhaeto-Romance and north Italian varieties and São Miguel Portuguese. Among consonantal developments, we have already mentioned lenition, which led to wholesale reduction and syllable loss in northern French dialects. Latin geminates generally survive only in Italo-Romance, and many other medial clusters are simplified (though new ones are created by various vocalic changes). Although Latin is in Indo-European terms a centum language (with k for PIE k), one of the earliest and most far-reaching Romance changes is the palatalisation, and later affrication, of velar and dental consonants before front vowels. Only the most conservative dialect of Sardinian fails to palatalise (witness kenapura ‘Holy supper = Friday’), and the process itself has elsewhere often proved cyclic.
The Columbia Arabic Treebank (CATiB) is a database of syntactic analyses of Arabic sentences. CATiB contrasts with previous approaches to Arabic treebanking in its emphasis on speed with some constraints on linguistic richness. Two basic ideas inspire the CATiB approach: no annotation of redundant information and using representations and terminology inspired by traditional Arabic syntax. We describe CATiB's representation and annotation procedure, and report on inter-annotator agreement and speed.