Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Larysa Kolibaba PhD in Philology, Senior Research Scientist of the Department of Grammar and Scientific Terminology, Institute of the Ukrainian Language of National Academy of Sciences of Ukraine 4 Hrushevskyi St., Kyiv 01001, Ukraine Е-mail: kolibaba.lm@meta.ua Heading: Researches Language: Ukrainian Abstract: In this article the problem of fixing of morphological forms of nouns in the Ukrainian dictionaries of different time and its […]
Annotation corpus for discourse relations benefits NLP tasks such as machine translation and question answering. In this paper, we present SciDTB, a domainspecific discourse treebank annotated on scientific articles. Different from widelyused RST-DT and PDTB, SciDTB uses dependency trees to represent discourse structure, which is flexible and simplified to some extent but do not sacrifice structural integrity. We discuss the labeling framework, annotation workflow and some statistics about SciDTB. Furthermore, our treebank is made as a benchmark for evaluating discourse dependency parsers, on which we provide several baselines as fundamental work.
The Effective Set-Size model has been used to describe uncertainty in various signal detection experiments. The model regards images as if they were an effective number (M*) of searchable locations, where the observer treats each location as a location-known-exactly detection task with signals having average detectability d'. The model assumes a rational observer behaves as if he searches an effective number of independent locations and follows signal detection theory at each location. Thus the location-known-exactly detectability (d') and the effective number of independent locations M* fully characterize search performance. In this model the image rating in a single-response task is assumed to be the maximum response that the observer would assign to these many locations. The model has been used by a number of other researchers, and is well corroborated. We examine this model as a way of differentiating imaging tasks that radiologists perform. Tasks involving more searching or location uncertainty may have higher estimated M* values. In this work we applied the Effective Set-Size model to a number of medical imaging data sets. The data sets include radiologists reading screening and diagnostic mammography with and without computer-aided diagnosis (CAD), and breast tomosynthesis. We developed an algorithm to fit the model parameters using two-sample maximum-likelihood ordinal regression, similar to the classic bi-normal model. The resulting model ROC curves are rational and fit the observed data well. We find that the distributions of M* and d' differ significantly among these data sets, and differ between pairs of imaging systems within studies. For example, on average tomosynthesis increased readers’ d' values, while CAD reduced the M* parameters. We demonstrate that the model parameters M* and d' are correlated. We conclude that the Effective Set-Size model may be a useful way of differentiating location uncertainty from the diagnostic uncertainty in medical imaging tasks.
Introduction. The article explores the impact of various types of verbal representation of ethnic stereotypes in the framework of a polyethnical academic community, i.e. educational environment in modern international university. Although the educational process with subjects of different cultural backgrounds plays a crucial role in conveying world views of representatives of different cultures, the research on the linguistic representation of stereotyped views on representatives of other nationalities has not been conducted yet. This aspect determines the relevance of the study. The aim of the research is to compare the impact levels of purely linguistic and speech ways of verbalising heterostereotypes by ways of employing relevant linguistic data for academic purposes during foreign language classes. Materials and Methods. The first stage of the experiment resulted in preparation of the linguistic corpus for the research: by means of comprehensive vocabulary research the lexical database with ethnonyms or ethnonym-based adjectives was compiled. To reveal the potential of their usage in the education processes, the participants were offered the preliminary and final surveys held as free associatio n experiment. Results. The influential potential for purely linguistic and speech ways of representing national stereotypes was compared to find out if they relate to the descriptors and scripts revealed through analysis of phraseological units and national anecdote respectively, while the latter was marked as a more efficient way of delivering ethnic stereotypes. The conclusions based on the analysis of the data obtained were drawn on how to use relevant linguistic material for academic purposes in order to appropriately develop attitudes to other ethnic groups. Discussion and Conclusions. The conducted research revealed more significant impact degree for ethnic anecdotes against investigation of lexical-phraseological units containing ethnonyms or ethnonym-based adjectives. It was illustrated by collection and further analysis of verbal reactions provided by students of non-linguistic departments of the modern University who took part in the preliminary and final stages which were in line with the beginning and end of the academic term accordingly. The portraits of typical national representatives made by the students at the completion of the course which included sessions on studying dictionary extracts and national anecdotes, to a greater extent conformed with the stereotypes delivered by ethnic anecdotes than the linguistic corpus of lexical-phraseological units. The research results may be considered during the development of the curriculum for foreign language courses in international universities with polyethnical academic environment.
We introduce the syntactic scaffold, an approach to incorporating syntactic information into semantic tasks. Syntactic scaffolds avoid expensive syntactic processing at runtime, only making use of a treebank during training, through a multitask objective. We improve over strong baselines on PropBank semantics, frame semantics, and coreference resolution, achieving competitive performance on all three tasks. them he After encouraging them, told goodbye and left for STIMULATE _EMOTION
Socioemotional Selectivity Theory posits that as person progresses through the life cycle, he or she makes concerted steps to maximize social and emotional wellbeing through selective patterns of emotional processing (Carstensen, 1995). The temporal positioning of an individual’s goals, motivations, and social orientations, or future time perspective (FTP), drives this change in emotional processing. The association between FTP and age is a naturally-occurring phenomenon as FTP becomes more limited as a person ages; however, it is believed that the construct of future time perspective is sufficiently malleable to be experimentally manipulated (Carstensen, 2003). The current study assessed the effects of a future time perspective manipulation on the emotional processing of positively and negatively valenced IAPS images in a college sample. Emotional processing was indexed by heart rate variability (HRV), skin conductance, memory recall, and eye-tracking. Young adult volunteers (N=22) were randomly assigned to one of two experimental conditions, wherein their future time perspective was manipulated to become either more limited or more expansive. Participants viewed a series of positive and negative cues followed by corresponding valenced images before and after the future time perspective manipulation. Preliminary results of the current study suggest the imagery task had no significant effect on FTP. Due to limitations from the small sample size, a larger sample size will be needed to conduct valid group comparisons to sufficiently test the effectiveness of the manipulation. Results of this study show participants with a more limited FTP had lower LF and greater HF HRV, indicating greater emotional regulation of arousal during the task. Interestingly, our results also indicate that positive affect ratings on the PANAS were related to avoiding negative emotional content (lower fixation percentage for negative images and cues), remembering more positive information (greater positive memory recall), \ndetecting a greater saliency for positive information (longer skin conductance rec t/2), lower sympathetic activity (lower posttest LF and SCL) and greater parasympathetic activity (greater posttest HF and RMSSD). These data suggest that reports of affect might provide more sensitive indication of emotional processing than future time perspective.
Treebank conversion is a straightforward and effective way to exploit various heterogeneous treebanks for boosting parsing accuracy. However, previous work mainly focuses on unsupervised treebank conversion and makes little progress due to the lack of manually labeled data where each sentence has two syntactic trees complying with two different guidelines at the same time, referred as bi-tree aligned data.
High quality communication between health care providers (HCPs) and adolescents and young adults (AYAs) with type 1 diabetes (T1D) may contribute to better diabetes self-care and health outcomes. Health communication reflects both informational content and how information is conveyed, including affect and tone. The aim of this study was to assess HCP affective communication and the relationship between HCP affective communication and glycemic control in AYAs with T1D. As part of a larger study of AYA-HCP health communication, routine clinic visits for 69 AYAs with T1D (M age 17.81 years; 56.5% female) and 8 HCPs (88% female) were audiorecorded. Clinic visits were coded using the Roter Interaction Analysis System (RIAS), a validated coding structure assessing verbal and non-verbal exchanges in a medical encounter. HCP global affective ratings were used to create two composite variables—positive HCP affect (e.g., attentiveness; respectfulness; Cronbach’s a = 0.82) and negative HCP affect (e.g., anger; dominance; Cronbach’s a = 0.75). Hemoglobin A1c (A1c) was taken from the medical chart. The mean A1c was 8.97% (±2.30). Descriptive analyses of positive and negative HCP affect indicated that HCPs expressed a high level of positive affect (M = 4.21) and a relatively low level of negative affect (M = 2.80). Negative affect was positively associated with HbA1c. After controlling for salient covariates (e.g., HCP, race, regimen), A1c accounted for a significant portion of the variance in negative affect during the clinic visit (Adj R2 =.36, ß = 0.57, p < 0.001). This sample of HCPs predominantly exhibited positive affect during routine T1D visits. Glycemic control was not associated with positive affect, but higher A1c was associated with more negative affect. This finding suggests elevated A1c levels may elicit more negative affect in routine diabetes care. Future research should examine these associations over time, including how AYA-HCP health communication quality predicts long-term glycemic control. Disclosure K. Homma: None. F.R. Cogen: None. R. Streisand: None. M. Monaghan: Research Support; Self; American Diabetes Association, National Institutes of Health.
In the modern world it can be easy to encounter advertising slogans that are supposed to grab the attention of a potential recipient as well as affect their imagination and sensitivity. This communication tends to convince the client that buying the advertised product would make them feel exceptional. This article focuses on linguistic gimmicks that are used by marketing experts in order to convince the client to a certain product. The analysis in based on the cosmetics sector. The lexical database consists of websites, television commercials and Internet advertisements.
Let me begin by thanking the Association for Computational Linguistics and its Executive Committee for conferring on me the great honor of their Lifetime Achievement Award for 2018, which of course I share with all the wonderful students and colleagues that have made many essential contributions to this work over many years.At the heart of the work that I have been pursuing over my research lifetime so far, whether in parsing and sentence processing, spoken language understanding, semantics, or even in musical understanding by machine, there lies a theory of natural language grammar that brings parsing, compositional semantics, statistical modeling, and logical inference into the closest possible relation. This theory of grammar is combinatory, in the sense that its operations are type-dependent and restricted to strictly string-adjacent phonologically or graphologically-realized inputs, and categorial, in the sense that those operands pair a syntactic type with a type-transparent semantic representation or logical form.I'd like to use this opportunity to briefly address three questions that revolve around the theory of grammar, both combinatory and otherwise. The first question concerns the way that Combinatory Categorial Grammar (CCG) was developed with a number of colleagues, over a number of stages and in slightly different forms. The second is an essentially evolutionary question of why natural language grammar should take a combinatory form. The third question is that of what the future holds for CCG and other structural theories of grammar in computational linguistics and NLP in the age of deep learning.I have called this talk "The Lost Combinator" in homage to the Victorian era poem "The Lost Chord," in the hope of suggesting that the theoretical development of CCG has always been empirical, rather than axiomatic, in search of the simplest explanation of the facts of language, rather than for confirmation of linguistic received opinion, however intuitively salient.In the late 1960s (when I was a psychology undergraduate at the University of Sussex under Stuart Sutherland, and then started as a graduate student in artificial intelligence at Edinburgh under Christopher Longuet-Higgins), a broad community of theoretical linguists, psychologists, and computational linguists saw themselves as all working on the same problem, under the definition provided by the "transformational" theory of grammar proposed by Chomsky (1957, 1965), using theories of psycholinguistic processing, language acquisition, and language evolution proposed by Lashley (1951), Miller, Galanter, and Pribram (1960), Miller (1967), and Lenneberg (1967), theories of natural language semantics proposed by Carnap (1956), Montague (1970), and Lewis (1970), and computational models of parsing such as those proposed by Thorne, Bratley, and Dewar (1968) and Woods (1970). (I myself was so convinced that this program would succeed that I believed it was time to apply the same methods to other cognitive faculties, taking as my research project for Ph.D. their application to the interpretation of music by machine, following the lead of Max Clowes [1971] in machine vision.)Almost immediately, this consensus fell apart. First, Chomsky himself was among the first (1965) to recognize that transformational rules, though descriptively revealing, were so expressive as to have little explanatory force, and required many apparently arbitrary constraints (Ross 1967). Second, psychologists realized that psycholinguistic measures of processing difficulty of sentences bore almost no relation to their transformational derivational complexity (Marslen-Wilson 1973; Fodor, Bever, and Garrett 1974). Finally, computational linguists attempting to implement transformational grammars as parsers realized that they were spending all their time implementing even more constraints on rules, in order to limit search arising from overgeneration (Friedman 1971; Gross 1978). (Meanwhile, I realized that the problem had not in fact been solved, and returned to natural language processing, thanks to a postdoc at Sussex with Philip Johnson-Laird.)This disillusion wasn't just a case of internal academic squabbling. There were also a couple of influential reports commissioned by the U.S. and UK governments that ended funding for machine translation (MT) and artificial intelligence (AI) (Pierce et al. 1966; Lighthill 1973). As a result of the second of these reports, which determined that AI was never going to work, PhDs in artificial intelligence like my classmate Geoff Hinton and myself spent ten years or so after graduation in psychology departments (in my case, at the Universities of Sussex and Warwick), until yet another report said AI was working after all and that Britain and the U.S. were falling behind Japan in this vital area. As a result, I could get hired again in computer science, first briefly back at Edinburgh, and then at the University of Pennsylvania (I learned a lesson from this odyssey that I have tried to remember whenever I have been appointed to a committee to report on anything, which is that while reports very rarely do any good, they can very easily do a great deal of harm.)Meanwhile, as a result of these conflicts, the scientific study of language fragmented. The linguists swiftly abjured any responsibility for their grammars ("Competence") bearing any relation to processing ("Performance"). Because the psychologists could hardly abandon Performance, they in turn became agnostic about grammar, retreating to context-free surface grammar (which they tended to refer to as "parsing strategies"), or a touchingly optimistic belief in its emergence from neural models. Meanwhile, the computational linguists (whose machines were growing exponentially in size and speed from the 16K byte core of the machine that supported the whole group when I started my graduate studies, on to levels that would soon permit parsing the entire contents of the then embrionic Web) similarly found that very little of what the linguists and psychologists cared about was usable at scale, and that none of it significantly improved overall performance over very much simpler context-free or even finite-state methods that the linguists had shown to be incomplete. The reason of course was Zipf's law, which means that the events with respect to which the low-level methods are incomplete are off in the long tail.It also became apparent to a few computationalists working on speech, MT, and information retrieval that the real problem was not grammar but ambiguity and its resolution by world-knowledge, and that the solution lay in probabilistic models (Bar-Hillel 1960/1964; Spärck Jones 1964/1986; Wilks 1975; Jelinek and Lafferty 1991) (although it was not immediately apparent how to combine statistical models with grammar-based systems without making obviously false independence assumptions).Nevertheless, as any red-blooded psychologist had always insisted, the divorce between competence and performance that everyone else had accepted did not make any sense. The grammar and the processor had to have evolved in lock-step, as a package deal, for what could be the evolutionary selective advantage of a grammar that you cannot process, or a parser without a grammar?It seemed equally obvious that surface syntax and the underlying semantic or conceptual representation must also be closely related, since the only reasonable basis for child language acquisition that has ever been on offer is that the child attaches language-specific grammar to a universal conceptual relation or "language of mind" (Miller 1967; Bowerman 1973; Wexler and Culicover 1980). It seemed to follow that radically new theories of grammar were needed.Theoretical linguists agree that the central problem for the theory of grammar is discontinuity or non-adjacent dependency between predicates and their arguments:Chomsky described discontinuity in terms of movement, which was known to be formally very unconstrained. By contrast, the ATN parser used in the LUNAR project (Woods, Kaplan, and Nash-Webber 1972) reduced all discontinuity to local operations on registers (Thorne, Bratley, and Dewar 1968; Bobrow and Fraser 1969; Woods 1970).In particular, unbounded wh-dependencies like the above were handled by: (a) putting a pointer into a * or HOLD register as soon as the "which" was encountered without regard to where it would end up; and (b) retrieving the pointer from HOLD when the verb needing an object "had" was encountered without regard to where it had started out. (It also included an ingenious mechanism for coordination called SYSCONJ, which one finds even now being reinvented on an almost yearly basis—cf. Woods [2010].) A * register was also used for wh-constructions within a systemic grammar framework by Winograd (1972, pages 52–53) in his inspiring conversational program SHRDLU.However, it was unclear how to generalize the HOLD register to handle the multiple long-range dependencies, including crossing dependencies, that are found in many other languages. In particular, if the HOLD register were assumed to be a stack, then the ATN becomes a two-stack machine (since we are already implicitly using one stack as a PDA to parse the context-free core grammar).On the computational side at least, the reaction to this impass took two distinct forms. Both reactions took the form of trying to reduce the two major operators of the transformation theory, substitution of immediate constituents, or what is nowadays called "Merge," and "Move," or displacement of non-immediate constituents, to one. On the one hand, Lexical Functional Grammar (Bresnan and Kaplan 1982) and Head-driven Phrase Structure Grammar (Pollard and Sag 1994) followed Kay (1979) in making unification the basis of movement and merger. Because unification can pass information across unbounded structures, this can be thought of as reducing Merge to Move.On the other hand, Generalized Phrase Structure Grammar (Gazdar 1981), Tree Adjoining Grammar (TAG; Joshi and Levy 1982), and Combinatory Categorial Grammar (CCG, Ades and Steedman, 1982) sought to reduce Move to various forms of local merger. In particular, the latter authors suggested that the same stack could be used to capture both long-range dependency and recursion in CCG.1Natural language grammar exhibits discontinuity because semantically language is an applicative system. Applicative systems (such as programming languages) support the twin notions of: (a) Application of a function/concept to an argument/entity; and (b) Abstraction, or the definition of a new function/concept in terms of existing ones.Language is in that sense inherently computational. It seems to follow that linguistics is (or should be) inherently computational as well. (Of course, it does not follow that computationalists have nothing to learn from linguistics.)There are two ways of modeling abstraction in applicative systems: Taking abstraction itself as a primitive operation (λ-calculus, LISP):(2)a.fatherEsau⇒Isaacb.grandfather=λx.father(fatherx)c.grandfatherEsau⇒Abrahamor Defining abstraction in terms of a collection of operators on strictly adjacent terms aka Combinators, such as function composition (Combinatory Calculus, MIRANDA).(3)b′.grandfather=BfatherfatherThe latter does the work of the λ-calculus without using any variables.Despite the resemblance of the "traces" (or copies) and "operators" (or complementizer positions) of the transformational theory to the λ-operators and variables of applicative systems of the first kind, natural language actually seems to be a system of the second, combinatory kind. The evidence stems from the fact that natural language deals with all sorts of fragments that linguists do not normally think of as semantically typable constituents, without the use of any phonologically realized equivalent of variables, such as pronouns:(4)a.Give[Anna books]?and[Manny records]?b.(Mother to child): There's adoggie![Youlike]?#the doggie.c.Food that you must[washVP/NP[before eating](VP∖VP)/NP]?.d.ik denk dat ik1Henk2Cecilia3[zag1leren2zingen3]?These fragments are diagnostic of a Combinatory Calculus based on Bn, T, and the "duplicator" Sn, plus application (Steedman 1987; Szabolcsi 1989; Steedman and Baldridge 2011).2CCG lexicalizes all bounded dependencies, such as passive, raising, control, exceptional case-marking, and so forth, via lexical logical form. All syntactic rules are Combinatory—that is, binary operators over contiguous phonologically realized categories and their logical forms. These rules are restricted by a Combinatory Projection Principle, which in essence says they cannot override the decisions already taken in the language-specific lexicon, but must be consistent with and project unchanged the directionality specified there. All such language-specific information is specified in the lexicon: The combinatory rules like composition are free and universal. All arguments, such as subjects and objects, are lexically type-raised to be functions over the predicate, as if they were morphologically cased as in Latin, exchanging the roles of predicate and argument.All long-range dependencies are established by contiguous reduction of a wh-element, such as (N∖N)/(S/NP), with an adjacent non-standard constituent with category S/NP, formed by rules of function composition.The combinatory rules synchronize composition of the syntactic types shown here with corresponding composition of logical forms (suppressed in the derivations above), to yield the logical forms shown as λ-terms for the resulting nouns N.To capture the construction in Example (4c), whose syntactic derivation we pass over here, we also need rules based on the duplicator S:(7)a."wash X before eating X″b.VP/NP:λx.before(eatx)(washx)≡S(Bbeforeeat)washTo capture constructions like Example (4d) (whose syntactic derivation is similarly suppressed), we also need rules based on second-order composition B2:(8)a."Y saw X teach W to sing.″b.((S∖NP)∖NP)∖NP:λwλxλy.help(teach(singw)wx)xy≡B2sees(Bteachsing)CCG thus reduces the operator move of transformational theory to applications of purely adjacent operators—that is, to recursive combinatory merge.Interestingly, the latest "minimalist" form of the transformational theory has also proposed that Move should be relabeled as an "internal" form of standard or "external" Merge (Chomsky 2001/2004, page 110), though without providing any formal basis for the reduction other than identifying internal Merge as "a grammatical transformation." (If anything deserved the soubriquet "the lost combinator," it would be this notional unitary combination of application and abstraction in a single perfect operator, linking or merging all types, as in the epigraph to this article.)B2 rules allow us to "grow" categories of arbitrarily high valency, such as ((S∖NPy)∖NPx)∖NPw. As we saw earlier, in some Germanic languages like Dutch, Swiss German, and West-Flemish, serial verbs are linearized using such rules to require crossing discontinuous dependencies. Thus, B2 rules give CCG slightly greater than context-free power.Nevertheless, CCG is still not as expressive as movement. In particular, we can only capture permutations that are what is called "separable," where separability is related to the idea of obtaining the permutations by rebracketing and rotating sister nodes (Steedman 2018).For example, for the categories of the form A|B, B|C, C|D, and D, it is obvious by inspection that we cannot recognize the following permutations:(9)i.*B|CDA|BC|Dii.*C|DA|BDB|CThis generalizatiion appears likely to be true cross-linguistically for the components of this form for the NP "These five young boys":(10)i.*Five boys these youngii.*Young these boys fiveTwenty-one of the 22 separable permutations of "These five young boys" are attested (Cinque 2005; Nchare 2012). The two forbidden orders are among the unattested three.3The probability of this happening by chance is the probability of the is the of the number of ways of two of three unattested by the is, about one in a (If the order that unattested so were to be this chance would to about one in number of separable permutations much more in than the number of all example, for around of the permutations are There are obvious for the problem of in machine translation and neural semantic parsing, to which we I to and at first as a still under least, that was what Joshi and then as a of where much of the development of CCG was out. students in a of and Joshi 1987; and that the of Ades and Steedman was by that both CCG and were equivalent to Grammar (Gazdar a new of the by the and of languages by these fell within the of what Joshi called which proposed as a for what could as a theory of natural that they are and and some limit on crossing dependencies. the is much much than the including the multiple free languages and even the languages of so it seems to and CCG as with to the its CCG was assumed to be as a grammar for parsing, because of the derivational ambiguity by and the combinatory rules, these also under and as any grammar with the same as CCG the same of in the parser it is there in the In this is just another in the of derivational ambiguity that all natural language and can be handled by the same statistical models as other particular, the dependency models by and are and Steedman and CCG is also to parsing with which can be using and long and Steedman and is now used in those that for between semantic and syntactic processing, such as machine translation and and machine and parsing and Steedman 1987; and et al. et al. et al. and semantic parser and 2005; et al. et al. of my work with the same has returned to their application in musical and and Steedman and shown that CCG grammars of the same and parsing models of the same statistical are required there as It is only in the of their compositional semantics that music and language very than on this of work in like to by two more The first is an evolutionary should natural language be a combinatory in the first The second is a question about the future development of CCG and other grammar-based theories to be to NLP in the age of deep and recursive neural take these questions in order in the two language like a combinatory applicative system because T, and evolved in to support of before there was any language (Steedman like need and are for of need to to make can form the that you need to before is to that they allow and some other can with like B2 are to make with arbitrary of including and including other whose is yet to be to be to do the the and problem, using like to in order to that are of to of and so such as of and can apply to a and much work that the and other have to be there in the already for the to be to with was to that from the even if had the problem of can be as the problem of search for a of in a or of possible such search has the same recursive as parser example, there are both and for the The latter a more in evolutionary as a mechanism that is to both semantic interpretation and parsing, rather than the evolution of like the the for linguistic as as the operators for competence grammar, the two to as what was to above as an evolutionary hope to have convinced you that CCG grammars are both and as as semantic parsers that are to parse the like other CCG parsers are by the of parsing models based on only a of no how we using and the constructions they are dependencies and so as we have and off in the long As a parsers are to on performance overall by using models et al. says more about the of grammars and parsers than about the of deep in the of semantic parser for arbitrary such as and Steedman I would that CCG and other grammar-based parsers have already been by of deep neural and and the question of whether models for applications in are because we have to the universal semantic that allow the child to CCG for natural languages and that we to be using both in semantic parser and in Because the language is any of linguistic logical like the universal language of it be more with and to semantic parsers for by deep neural force, rather than by CCG semantic parser is it possible that the problem of parsing could be by neural by stack et al. et al. actually learn as has been seems likely that semantic parsers and neural machine translation to have difficulty with long-range because the evidence for their is so example, both and as a case, verb categories or a complementizer I is actually by which says that if you are an language with like and then you not in be to in languages like or languages like and German, you be to both subjects and objects, or at the time of a translation system no of learned these syntactic from a we get an sentence whose to translation means is the that the us the a is the that the us to the if we with we (which back into the a is the that the us holds the a we with using a the translation again an that is is the that they said had the is the they said they the contrast, CCG parsers do rather on and Steedman Steedman, and which is a construction that to be rather determined by the parsing methods are similarly when with long-range even when the sentence is in this think that the the and that it is et think that the the and it is in at least, these constructions could be learned by grammar-based semantic parser from using the methods of et al. and et al. make it likely that there be a need for in like where long-range dependencies like deep and are here to The future in parsing for such lies with systems using neural for and grammars for problem in NLP the fact that natural language understanding inference as as semantics, and we have no idea of the representation question is the almost it many it is almost equally to the information in a form that is not immediately with the form of the example, sentences like the following a different a rather than a an an a a and at a representation language that is we are using CCG parsers to the for between in order to consistent of between over of the same types, using over then an it and it under such as et al. of in the then that can be to a single relation and Steedman can be across from multiple languages and Steedman can then the semantics for relation with the and the entire using this now both and semantic an with the as and the as questions the in this we parse questions into the same semantics which is now the language of the the we use the and the and the following anything that in the then the the of anything is in the in the this to work, we need to be nodes in their in both and then be to like in of a semantic in the language of the need to learn between semantic and the language of the this project is the the function of a of the semantic in semantic like those of and and Steedman and while the form a similarly of the of Carnap and Fodor, Fodor, and Garrett semantic are essentially but with the advantage that they can be with logical operators such as and for the of semantic underlying natural language semantics in to another different to semantics that to use reduced of to using operations such as and and in of compositional It is an question whether can be with to of a and et al. It is likely that some of be here to the of the of the long like and work in do they work in In particular, can they learn all the syntactic in the long like and crossing in a way that support semantic they are not actually but are a finite-state or a then by on as for natural language processing, we are in of of the computational linguistic project of also providing computational of language and if we that like is a real and that learn their first language by of the sentences of their language the of the universal language of we still the of what that universal semantic language not get an to that question we can above using and and such as as for the language of to use machine for what it is such variables and their for use in a natural language work was supported in by a Award and a a University of Edinburgh and my and the and all my and students over many
Blueberries have been reported to possess several anti-inflammatory properties. Previous studies examining the anti-inflammatory effect of blueberries on acute inflammation caused by exercise-induced muscle damage are largely inconclusive. This may be due to the dose used in these studies not accounting for an individual’s lean mass (LM), the compartment directly involved during exercise, when determining appropriate blueberry dosage. PURPOSE: To examine the effect of blueberry supplementation (BB) at a dose relative to LM on delayed onset muscle soreness (DOMS) and recovery. METHODS: Fourteen recreationally active women (age: 21±1yr; body fat: 24.8±4.5%) participated in this double blind, matched-pairs study. Participants were matched by LM and randomly assigned to either a BB or a placebo (PLA) group. Leg strength was assessed via one-repetition maximum (1RM) on a leg press. Participants consumed a daily dose of freeze-dried BB powder (1.6g BB/kgLM) or a PLA (1.6g PLA/kgLM) for 7 days prior to induction of DOMS. Participants completed 6 sets of 10 repetitions at 70% 1RM on the leg press to induce DOMS. Perceived soreness (questionnaire), pressure-pain threshold (dolorimeter), and average power (AP; Biodex™) of the right thigh muscles were assessed immediately before (PRE) and after (POST), 24, 48, and 72h post induction of DOMS. Repeated measures ANOVAs were used for analyses. Significance was set at p<0.05. RESULTS: There were no group x time interactions for perceived soreness, pressure-pain threshold, and AP, however, significant time effects were observed for these variables. When comparing pre to post 24hr (p<0.001), 48hr (p=0.001), and 72hr (p=0.011) perceived soreness of the thigh muscles significantly increased. Pressure-pain threshold of the thigh muscles decreased significantly from pre to post 24hr (p=0.023), 48hr (p=0.001), and 72hr (p=0.024. Isokinetic leg extension AP decreased from pre to post 24hr (BB: 83±17 to 76±22Nm; PLA: 85±21 to 79±26Nm; p=0.02). CONCLUSION: Consumption of BB for 7 days prior to DOMS induction on a leg press does not affect rating of perceived soreness, pain threshold, nor attenuate decreases in performance compared to a PLA in recreationally active women.
Coherent texts, whether written or spoken, are built upon discourse relations linking utterances together through causal, temporal or contrastive connections, among many other types (Mann & Thompson 1988). Different types of relations are signalled by different types of markers, although there is no one-to-one mapping. These markers often belong to the functional category of discourse-relational devices, or “connectives”, such as however, because or in fact. Writers and speakers also have the option to use other signalling devices (e.g. lexical or syntactic patterns) or even to leave a discourse relation implicit (e.g. Taboada 2009). This study focuses on another strategy for discourse marking, namely the use of underspecified connectives (Spooren 1997). More particularly, we investigate the role of the additive conjunction and to signal relations of addition but also of consequence, contrast and concession. In these cases, the discourse relation is more specific than the information strictly provided by the connective: a consequence or contrast is more informative than a mere additive relation. Despite the low informative value of and, it is quite often found in authentic contexts where such enriched interpretations were assigned to the discourse relation (6% of and express a result in the Penn Discourse TreeBank 2.0, Prasad et al. 2008), which calls for more research on the conditions under which and can be used as an underspecified connective. Our research objective is thus to compare the linguistic and contextual features of utterances linked by and which either express addition, consequence, contrast or concession. To do so, we first need corpus-based data where such discourse relations are reliably identified. Discourse relation annotation is extremely costly in time and human resources, it requires heavy training and, even so, agreement scores are often rather low (Spooren & Degand 2010). As a result, researchers have recently started to turn to crowdsourcing as an alternative method to gather discourse relation disambiguations through a low-cost, non-expert workforce (Kawahara et al. 2014; Rohde et al. 2016). Scholman & Demberg (2017) report on the results of a connective insertion task which they used as an indirect method to annotate discourse relations: the sense of a relation can be retrieved through the selection of unambiguous connectives from a list to fill in a blank between utterances, provided this task is repeated by a large number of participants (around 20). The authors discuss the validity of this method and conclude that crowdsourcing connective insertions is reliable enough as an alternative to expert annotations. For connective insertion tasks, it's important to distinguish between originally implicit vs. explicit relations. For originally implicit relations, inserting a connective resembles the approach taken in PDTB annotation (except that the choice of connectives is more restricted in the crowdsourcing step in order to allow for disambiguation of relation type). In originally explicit relations, however, the meaning and interpretation may substantially change by removing the connective, see examples (1)-(3) below. In this study, where we investigate originally explicitly marked relations with the connective “and”, we will therefore compare a crowdsourced connective insertion task with a crowdsourced connective replacement task, in which the original connective “and” is not removed from the stimulus. (1) I am not going back to Germany. Therefore, I will not eat spätzle ever again. (consequence) (2) I am not going back to Germany. In fact, I will not eat spätzle ever again. (addition) (3) I am not going back to Germany. I will not eat spätzle ever again. (?cause) Our working hypothesis is that such differences in interpretations, triggered by the connective (or absence thereof), apply to utterances containing and as well. In this respect, we challenge previous experimental research on and which showed that and has a very small informative value and little or no facilitating effect on reading times (Murray 1994) or comprehension (Cain & Nash 2011). By contrast, we expect that connective insertion tasks will be affected by the presence of and, thus supporting the claim that and does trigger enriched pragmatic inferences (Blakemore & Carston 1999). We therefore propose to use crowdsourcing for the study of underspecified and by comparing stimuli with and without the original connective. In other words, we want to test whether connective elicitations will differ across stimuli which are identical except for the presence or absence of the conjunction and. To this end, we ran two crowdsourcing experiments on the Prolific Academic online platform. In the first one, we used 83 authentic pairs of utterances originally containing and. We collected them from the Loyola Corpus of Computer-Mediated Communication (Goldstein-Stewart et al. 2008), in order to avoid the high formality of existing corpora such as the Penn Discourse Treebank (economy newspaper articles). This corpus contains blogs and chat conversations between college students about topics such as gay marriage, gender discrimination or privacy rights. This data was pre-annotated by the first author as either expressing a relation of addition, contrast, concession or consequence. The stimuli are grouped in four lists of about 20 items each, which are balanced with respect to the pre-annotated relation type (about 10 addition, 6 consequence, 1 contrast, 3 concession in each list). The participants (paid 1€ per list) can choose from a list of eight connectives to fill in a blank between the two utterances (the original and has been removed). The connectives are in addition, plus, therefore, as a result, by contrast, whereas, nevertheless and yet. The second experiment uses exactly the same lists of items, except that the stimuli now show the original and connecting the utterances, and the participants are therefore instructed to substitute this and with one of the connectives from the same list of options. In the analysis, we first compare the connectives chosen by the participants with the relation type pre-identified by the expert annotator, in order to see whether they converge (e.g. therefore or as a result selected in case of a relation of consequence). This first step provides us with a dataset of utterance pairs with their disambiguated discourse relation, without resorting to costly (and partly subjective) expert annotations. The items thus classified into one of the four categories (addition, consequence, contrast, concession) will allow us to test the effect of additional variables (e.g. register) in further studies (Crible & Demberg 2017). We will replicate this analysis with the results of the second experiment. We will then compare whether the connectives selected by the participants are the same when and is present in the stimuli and when it is not. Preliminary results show that, when the relation was pre-annotated as additive, the participants tend to equally choose a consequence or an additive connective, with no significant difference, which suggests that consequence is often interpreted in the absence of a connective. For all other relation types (i.e. consequence, concessive and contrast), the great majority of participants’ choices match the pre-annotation, thus confirming that and can be used in contexts which express more than mere addition. The results from the second experiment are still pending, and should lead to interesting comparisons on the effect of and in connective elicitation (or connective substitution). This study has a number of implications on the informative value of and, which may be higher than what previous studies have suggested, and on the use of connective elicitation with or without the original connective included in the stimuli. We argue that, when dealing with authentic corpus-based stimuli, including the original connective in the experiment is a more accurate representation of the data and better reproduces the interpretation mechanisms as they would be processed in natural conditions. The experiments reported in this paper constitute the first step of a larger project on the contextual and cognitive constraints to the production and interpretation of underspecified connectives. They also relate to ongoing crosslinguistic projects on the meaning variation of and and its use across spoken registers (Crible, in press) and in translation (Abuzcki et al. 2017).
It might surprise some young researchers that the first recipient of the ACL LifeTime Achievement award,1 Aravind Joshi, was so often compared to “Yoda,” one of the oldest and most powerful of the Jedi Masters in the Star Wars universe. But Aravind was also one of the kindest, wisest, and most justly celebrated people that one was fortunate to know.When Aravind received the award in 2002, at the age of 73, he said that he hoped his lifetime wasn’t over. Fortunately, it wasn’t: For the next 15 years, Aravind continued to enjoy time spent on research; advising students and younger colleagues; attending ACL conferences both at home in the United States and in far-flung places such as Sydney, Singapore, and Jeju Island; and enjoying the company of his extraordinary wife, the embryologist Susan Heyner, his daughters, Meera and Shyamala Joshi, and his grandchildren, Marco and Ava. Then on 31 December 2017, Aravind died peacefully at home in Philadelphia, sitting in his favorite chair, at the age of 88.Aravind Joshi was born in Pune, India, on 5 August, 1929. He sailed to the United States in 1954 to study electrical engineering (EE) at the University of Pennsylvania, after he was rejected by Harvard because his application, mailed from India, arrived a day late. While completing his M.Sc. in EE, he worked as an engineer at RCA (Camden, NJ), and then while completing his Ph.D. in EE, as a research assistant at the University of Pennsylvania’s Department of Linguistics. After being awarded his doctorate, Aravind joined the Penn faculty, remaining in EE until the brilliant and prescient Saul Gorn, who chaired Penn’s Graduate Group in Computer and Information Science, convinced the University to establish a new academic department of Computer and Information Science (CIS). Aravind joined this new department as a full professor and Chair, grateful that Saul Gorn had argued so forcefully that the new department should embrace the science of information as well as the practical study of computers and computing. Allowing for such broad intellectual content allowed the evolving CIS Department to constantly reach out to researchers in other disciplines—including those at Penn’s Wharton School of Business, as well as at the Departments of Linguistics, Psychology, Philosophy and Bioinformatics.Aravind remained Chair of CIS for an incredible 13 years, until 1985—continuing throughout this time to carry out cutting-edge research, to serve leadership roles in both ACL (as President in 1975 and then as Book Series Editor from 1982) and IJCAI (as General Chair of the 1985 IJCAI Conference), while at the same time giving generously to his students and colleagues. As General Chair of IJCAI, Aravind reached out to invite Soviet refusenik computer scientists, linguists, and mathematicians to attend. Several of the invitees had been arrested and were serving long sentences in labor camps. Although Aravind knew that they would never be able to attend, the invitation reassured them that their academic accomplishments would be recognized internationally, despite the political environment.One of Aravind’s major achievements during his time as Chair of CIS was co-founding, with psycholinguist Lila Gleitman, Penn’s famous Cognitive Science Program. Funded initially by the Sloan Foundation, the Program received further funding from the National Science Foundation in 1991, to become Penn’s world-famous Institute for Research in Cognitive Science (IRCS).2 Aravind and Lila co-directed IRCS until 2001, contributing a stream of over 100 postdocs from linguistics, psychology, computer science, philosophy, neuroscience, and mathematics that it hosted and nurtured before closing its doors in 2016.Meanwhile, Aravind’s five decades at Penn saw his research and inventions span much of what we know as computational natural language processing, including• parsing using finite state transducers, with Aravind’s early FST parser reimplemented as “A Parser from Antiquity” (Joshi and Hopely 1996). Aravind always encouraged his students and colleagues to go back and re-examine earlier work. In fact, he once suggested that students in introductory CL courses be asked to look at the literature and reconstruct some old system as a way of both shortening the period to re-discovery and giving students a better historical grounding in the field.• grammatical formalisms, most notably the development and detailed characterization of the “mildly context-sensitive” Tree Adjoining Grammar (TAG) in both its original and lexicalized forms (LTAG), providing enough power to handle the range of phenomena in human language syntax while remaining computationally tractable (Joshi and Schabes 1997).• cooperative Question Answering and the range of inference it requires (Joshi, Webber, and Weischedel 1984, 1986)• prominence in discourse (Grosz, Joshi, and Weinstein 1995; Walker, Joshi, and Prince 1998), in the form of work on “Centering,” which was meant to account for ease of inference and the use of anaphoric expressions, linked by the observation that an entity that can be accessed with an expression as small as a pronoun must also be prominent.• discourse and syntax (Webber et al. 1999, 2003), where Aravind reconceptualized discourse connectives in the framework of LTAG, culminating in development of the NSF-funded Penn Discourse TreeBank3 and similarly annotated corpora in Chinese (Zhou and Xue 2012, 2015), Hindi (Oza et al. 2009), Turkish (Zeyrek et al. 2010), and biomedicine (Prasad et al. 2011).Here it is worth saying a bit more about two aspects of Aravind’s work: His work on grammar formalisms and his work on discourse. In the early 1980s, Aravind identified a set of computational properties that provided an informal definition of a class of Mildly Context Sensitive (MCS) languages that he claimed properly included all human languages and was properly included among the much vaster class of context-sensitive languages. These properties were: (a) polynomial parsability; (b) the constant growth property (which excludes languages with unbounded gaps in the length of sentences, of which an artificial example is the indexed language a2n, made up of strings of 2n a’s); and (c) a limit on crossing dependency of the kind seen in the artificial tree-adjoining language (TAL) anbncn, and hence on permutation-completeness (Joshi and Levy 1982; Joshi 1988; Joshi, Vijay-Shanker, and Weir 1991). TAG was the first fully formalized theory of grammar to be proved to characterize only languages within the MCS class, and provided the basis for further proofs of MCS expressive power for several other constrained grammar formalisms that were developed around the same time via their (weak) equivalence to TAG.One natural generalization of TAG was to the Linear Context-free Rewriting Systems (LCFRS) or Multiple Context Free Grammars (MCFG). These are considerably more expressive than TAG, and were for a while conjectured to provide a formal definition of MCS languages. As a result, other formalisms that were considerably more expressive than TAG have laid claim to the Joshian mantle of Mild Context Sensitivity, showing that it has become a highly influential meme in the field.However, it has since been shown that the artificial permutation-complete language MIX3, consisting of all permutations over the strings of the TAL anbncn, is a Multiple Context Free Language (MCFL). Thus, the formal characterization of the MCS class is currently a matter for debate. Nevertheless, TAG itself remains among the least more expressive formalisms than CFG that is known. Now the more important question is whether TAG or one of the other weakly equivalent formalisms is expressive enough to capture the full range of phenomena actually exhibited by natural languages, as argued in Frank (2004).Aravind’s interest in discourse semantics has been long-standing, going back at least to the mid-1970s and the publication by Academic Press of Subject and Topic (Li 1976). The collection contained two articles that Aravind annotated extensively—Li and Thompson’s article on topic-prominent languages and Lehmann’s article on the history of topic-prominence in Indo-European. Aravind’s interest in the notion of topic-prominence and the discourse semantics of information structure seems to have been stimulated by his knowledge of Sanskrit and of his native language, Marathi, both of which exhibit aspects of topic-prominence and ergativity. Together with the article in the same collection by Keenan and Schieffelin, this concern with prominence in discourse seems to have fuelled his later work on Centering and prominence in discourse.Later on, when Aravind started to look at discourse connectives (Webber et al. 1999, 2003), his concerns went deeper than the obvious parallels between lexically anchored trees for sentence-level syntactic analysis and trees lexically-anchored on discourse connectives that could be used in going beyond sentences to small units of discourse. Rather, his concerns were grounded in his growing belief that too many constructions found between the start of a sentence and its final punctuation didn’t really belong to syntax. In particular, Aravind described parentheticals, epithets, extraposed predicates, and sentential relatives as “constructions that require a skilled tree surgeon to force a single tree over a sentence.” Whereas syntactic analysis may simply punt by attaching a parenthetical such as “John thinks,” in “Mary, John thinks, will win the race,” to the root node of its parse tree (just as “Mary” is attached to the root node as its subject), for Aravind, the two attachments (of the subject and of the parenthetical) were completely different, with the attachment of the parenthetical belonging to an orthogonal dimension, in a “paratactic” relation, another characteristic of many constructions in Sanskrit and Marathi. He felt the same about epithets like “damn” in “I finished the damn book”: “damn” should be attached along an orthogonal dimension because it bore a different semantic relation to “book” than say “thick book” or “book about insects.” For Aravind, discourse provided this orthogonal dimension, reducing the number of “no-win” decisions that followed from insisting on a single parse tree over a sentence.Having spent so much time with sentences from the Penn TreeBank (constructed from articles from the ACL/DCI Wall Street Journal corpus), Aravind found the examples that made the most convincing demand for distributing the burden between syntax and discourse to be attribution phrases (Dinesh et al. 2005). He felt that sometimes an attribution phrase like “the company says” belongs to sentential syntax, feeding its semantics up to that of the sentence as a whole (as in the contrast expressed in “Observers say negotiations have halted, while the company says it is talking with several prospects”). In other cases, however, such as the concession expressed in “There have been no orders for the Cray-3 so far, though the company says it is talking with several prospects,” he felt that the attribution phrase belongs to an orthogonal discourse dimension, because what is contrary to expectations associated with the lack of orders for the Cray-3 is not the company (Cray) saying something, but rather the existence of several prospects that it is talking with. Both analyses become possible if syntax and discourse can provide distinct bases for analysis. Hence Aravind’s interest in both.Although Aravind’s ingenuity was all his own, the inventions he was involved with were possible only because of his unprecedented inclusion of linguists, psychologists, philosophers, and mathematicians, as well as computer scientists and engineers, in his work. In recognition of Aravind’s inclusionary spirit and many achievements, Penn’s CIS Department hosted JoshiFest in Fall 2012, an all-day symposium of talks and encomia in his honor. The complete program of talks and presentations is available for viewing at https://www.cis.upenn.edu/about-cis/events/joshi-fest/program.php.JoshiFest was neither the first nor the last time that Aravind’s achievements were publicly recognized. Besides the ACL Lifetime Achievement Award in 2002, Aravind was the recipient of the 1997 IJCAI Award for Research Excellence; the 2003 David E. Rumelhart Prize of the Cognitive Science Society; the 2005 Benjamin Franklin Medal in Computer and Cognitive Science (awarded by the Franklin Institute in Philadelphia); and the Henry Salvatori Chair in Cognitive and Computer Science. He was elected to the National Academy of Engineering in 1999 and named a Fellow of IEEE in 1976 and the Association for Computing Machinery (ACM) in 1998. In 1990, he became a Founding Fellow of the Association for the Advancement of Artificial Intelligence (AAAI), and in 2011, a Founding Fellow of the ACL. Most recently, the Charles University (Prague) recognized Aravind’s accomplishments—including his joint work with Professor Eva Hajicova’s group at Charles University—with the award of Doctor Honoris Causa in physics and mathematics.In his obituary for Aravind Joshi in Language Log,4 Mark Liberman quoted posts from Bob Frank (whose Ph.D. thesis Aravind supervised, and who is now Chair of Linguistics at Yale) and Julia Hockenmaier (a postdoc at IRCS following her Ph.D. from the University of Edinburgh, now an Associate Professor at the University of Illinois (Champaign-Urbana). Because both posts express their author’s thoughts and feelings so well, they seem an appropriate way to close this obituary.I just heard the crushing news that Aravind Joshi passed away yesterday. It’s hard for me to overstate how profoundly Aravind influenced my career and my life, since he took me on as his PhD student 30 years ago. The content of his work laid the foundations for so much of what I have worked on over the years, and his vision of interdisciplinary interaction shaped how I see the field. I will never forget his insatiable curiosity and intellectual energy, his remarkable ability to identify good problems and insightful solutions, and his gentle kindness and humanity. And I will so much miss the boyish excitement he exuded whenever he would share his latest ideas with me. Thank you for everything, Aravind. You will be missed. [Bob Frank]I can’t begin to describe how much I owe to Aravind’s advice and mentorship, his intellect, his curiosity, his kindness, and his great sense of humor. It was such a privilege to work so closely with him, even as one of his last postdocs. His impact on our field and our community can simply not be overstated. We’ve lost one of our founding fathers. Not just because he was one of the few, or probably even the only one, still around from the very early days of NLP. We’ve also lost someone who has really shaped the intellectual and social culture of our community in fundamental ways. If you are among those that feel at home in our field because of its intellectual richness and diversity, and also because you never felt out of place because you are a woman, you should know how much you owe to Aravind and the legacy of his very many distinguished students, and the culture he and his colleagues created at Penn and in the community as a whole. Rest in peace, Aravind. In sorrow, and gratitude. [Julia Hockenmaier]In writing this article, I drew in part from text by Meera and Shyamala Joshi, John Nerbonne, Mark Liberman and Mark Steedman, all of whom have written eloquently about Aravind’s life and accomplishments. All errors, however, are my own.
In this paper, we describe the annotation and development of Telugu treebank following the Universal Dependencies framework. We manually annotated 1328 sentences from a Telugu grammar textbook and the treebank is freely available from Universal Dependencies version 2.1.1 In this paper, we discuss some language specific annotation issues and decisions; and report preliminary experiments with POS tagging and dependency parsing. To the best of our knowledge, this is the first freely accessible and open dependency treebank for Telugu.
This paper reports on the extension project “Fostering entrepreneurship through computer-aided translation and interpreting practices” carried out by undergraduate language translation students at São Paulo State University in São José do Rio Preto, São Paulo, Brazil. The extension project worked in partnership with the Company Incubator Center and the Technological Park of São José do Rio Preto in 2017. The project is comprised of a series of activities aimed to integrate the practice of Brazilian Portuguese language writing and translation into English and Spanish. The translations are assisted by translation memory systems to meet the demands of small companies and entrepreneurs in their early development stages without the financial resources to advertise their services and products to potential consumers abroad. In order to promote the products and services of these companies, a series of activities were developed for the Company Incubator Center website into English and Spanish, contributing to undergraduate translation students’ education and qualification while increasing the visibility of start-up companies both domestically and internationally. The results suggest that collaboration between internal and external university communities can create products (a trilingual website, a linguistic database, and a trilingual glossary) capable of fostering both professional and personal growth for students and businesses.
Introduction: The purpose of this investigation was to examine the independent and combined effects of caffeine (CAF) alone or as a part of a multi-ingredient pre-workout supplement (PWS) on resistance exercise performance in recreationally active males. Methods: In a single-blind, randomized, placebo (PLA) controlled, crossover design; 10 recreationally active males (20.5 ± 0.9; 178.9 ± 7.7 cm; 81.8 ± 11.5 kg) completed three laboratory visits, after determination of one repetition-maximum (1-RM) on the bench press and leg press, where they performed bench press and leg press to failure at a load of 70% 1-RM. Subjects were randomly assigned to ingest either one serving of a commercially available PWS (C4 Original, Cellucor, Bryan, TX, United States), a dosage-matched anhydrous CAF beverage (150 mg), or a taste- matched PLA beverage. Heart rate (HR), affect, rating of perceived exertion (RPE), and mood state was assessed 20 minutes pre and post-substance ingestion, and immediately after exercise.Results: Participants completed significantly more repetitions to failure (p = 0.006) and lifted significantly greater weight (p = 0.009) during Leg Press in the PWS and CAF conditions compared to the PLA condition. There was not a significant difference found between CAF and PWS trials (p > 0.05).Conclusions: This data suggests that both CAF and PWS may have a positive effect on exercise performance in leg press but was not effective in increasing muscular endurance in bench press. The commercially available PWS offered no additional ergogenic effects when compared to the CAF.
treebanks, discusses the process through which they were created, and outlines the future plans and timeline for the next improvements. Special attention is paid to the possibilities of using UD in the documentation and description of endangered languages.
In today’s increasingly divided political climate there is a need for a tool that can compare news articles and organizations so that a user can receive a wider range of views and philosophies. NewsAnalyticalToolkit allows a user to compare news sites and their political articles by coverage, mood, sentiment, and objectivity. The user can sort through the news by topic, which was determined using Natural Language Processing (NLP) and Latent Dirichlet Allocation (LDA). LDA is a probabilistic method used to discover latent topics within a series of documents and cluster them accordingly. Each news article can be considered a mix of multiple topics and LDA assigns a set of topics to each with a probability of it pertaining to that topic. For each topic, a user can then discover the coverage, mood, sentiment and objectivity expressed by each author and site. The mood was determined using IBM Watsons ToneAnalyzerV3, which uses linguistic analysis to detect emotional, social and language tones in written text. The analyzer is based on the theory of psycholinguistics, a field of research that explores the relationship between linguistic behavior and psychological theories. The sentiment and objectivity scores were determined using SentiWordNet, which is a lexical database that groups English words into sets of synonyms and assigns sentiment scores to them. The features were combined to plot an interactive graph of how opinionated versus how analytical an article is, so that the user can click through them to get a better understanding of the topic in question.
We present a progress report of the Turkish Treebank concentrating on various aspects of its design and implementation. In addition to a review of the corpus compilation process and the design of the annotation scheme, we describe the details of various pre-processing stages and the computer-assisted annotation process.
This study aims to electronically assessed (e-assessment) students’ replies in response to teachers’ question. It can be useful to systematize the question answering context regarding matching text semantically through WordNet semantic similarity techniques. WordNet is a lexical database of words’ synonyms. It uses group of synonyms called synsets for semantical operation of English text. For this purpose, a new methodology is proposed to automate e-assessment in the field of education. The collected dataset contains 210 pairs of words extracted from different undergraduate students’ replies in contradiction of teacher’s question statement. Further WordNet similarity measures i.e. Path Length, Lin, Wu &Palmer and Hirst & Onge are used to compute the semantic relatedness score. In the pilot study 42 pair of words were extracted from 8 students’ replies, which are marked using semantic similarity measures and equated with teacher’s marks. Teachers are provided with four boxes of the mark while our developed method provides a precise measure of marks. The experiment is shown with comprehensive dataset resulting with words’ frequencies in similarity measures.
Rapid advances in information technology and proliferation of social media services have caused a radical transformation of human communication. Having created a social media presence people engage in computer-mediated communication, set their own goals as well as perfect their knowledge of English as a global language. The richness and diversity of computer-mediated discourse is concentrated in multiple online experiences and therefore enables to study a great number of linguistic changes. Various studies of computer-mediated discourse analyze socio-psychological characteristics in coherent sequences of sentences, propositions, speech or turns-at-talk. The given article aims at presenting vocabulary teaching strategies to new computer-mediated language and their influence on students' acquisition. Our contribution provides an overview of the recent new entries of computer-mediated vocabulary in online crowdsourced dictionaries. The main ways of forming new words as well as wide-spread semantic changes are viewed (acronyms, compounds, suffixes, blended words, conversion, etc.) Computer-mediated vocabulary teaching in the classroom covers change, diversity, disputes in economical, political, social spheres (such as narcissistic tendencies, emotional correctness, excessive use of social media and increasing reliance on technology, equal rights movement, task-based employment, etc). As the social media universe strives for a thorough integration with the user's life language learners are expected to embrace the latest changes in linguistic norms and devise an appropriate philosophy of language management.
This paper describes our system (SLT-Interactions) for the CoNLL 2018 shared task: Multilingual Parsing from Raw Text to Universal Dependencies. Our system performs three main tasks: word segmentation (only for few treebanks), POS tagging and parsing. While segmentation is learned separately, we use neural stacking for joint learning of POS tagging and parsing tasks. For all the tasks, we employ simple neural network architectures that rely on long short-term memory (LSTM) networks for learning task-dependent features. At the basis of our parser, we use an arc-standard algorithm with Swap action for general non-projective parsing. Additionally, we use neural stacking as a knowledge transfer mechanism for cross-domain parsing of low resource domains. Our system shows substantial gains against the UDPipe baseline, with an average improvement of 4.18% in LAS across all languages. Overall, we are placed at the 12 th position on the official test sets.
Tree-structured neural network architectures for sentence encoding draw inspiration from the approach to semantic composition generally seen in formal linguistics, and have shown empirical improvements over comparable sequence models by doing so. Moreover, adding multiplicative interaction terms to the composition functions in these models can yield significant further improvements. However, existing compositional approaches that adopt such a powerful composition function scale poorly, with parameter counts exploding as model dimension or vocabulary size grows. We introduce the Lifted Matrix-Space model, which uses a global transformation to map vector word embeddings to matrices, which can then be composed via an operation based on matrix-matrix multiplication. Its composition function effectively transmits a larger number of activations across layers with relatively few model parameters. We evaluate our model on the Stanford NLI corpus, the Multi-Genre NLI corpus, and the Stanford Sentiment Treebank and find that it consistently outperforms TreeLSTM
This paper describes our system (HIT-SCIR) submitted to the CoNLL 2018 shared task on Multilingual Parsing from Raw Text to Universal Dependencies. We base our submission on Stanford's winning system for the CoNLL 2017 shared task and make two effective extensions: 1) incorporating deep contextualized word embeddings into both the part of speech tagger and parser; 2) ensembling parsers trained with different initialization. We also explore different ways of concatenating treebanks for further improvements. Experimental results on the development data show the effectiveness of our methods. In the final evaluation, our system was ranked first according to LAS (75.84%) and outperformed the other systems by a large margin.
In this paper we present the linguistic databases developed during our 8-year lexicographic research on the Modern Greek Standard (MGS) verbal system. Apart from the intermediate databases presented, the main products are (a) a new conjugation system of 385 paradigmatic models, which allows for the automatic generation of all verbal lexical morphemes and monolexical forms (b) a statistically established database of 151,536 distinctive verb-final grapheme sequences which allow for the automatic tagging of all monolexical verbal tokens without the traditional intervention of any built-in lexicon, and (c) a linear Iemmatisation morphophonological rule system accessed on the basis of the distinctive grapheme sequences identified.
Detecting lexical entailment plays a fundamental role in a variety of natural language processing tasks and is key to language understanding. Unsupervised methods still play an important role due to the lack of coverage of lexical databases in some domains and languages. Most of the previous approaches were either based on statistical hypothesis of specific entailment relations or tried to encode word relations in low-dimensional vector embeddings. This thesis builds upon one of the few approaches which intrinsically model entailment in a vector space. We then further generalize this model by introducing an alternative, distributional representations for words which harnesses tools from optimal transport to define distance or entailment measures between such representations. We evaluated the models on hypernymy detection where our distributional estimate significantly improves over the underlying model and even outperforms state-of-the-art on some datasets.
The statistical parsing of morphologically rich languages is hindered by the inability of parsers to collect solid statistics because of the large number of word types in such languages. There are however two separate but connected problems, reducing data sparsity of known words and handling rare and unknown words. Methods for tackling one problem may inadvertently negatively impact methods to handle the other. We perform a tightly controlled set of experiments to reduce data sparsity through class-based representations in combination with unknown word signatures with two PCFG-LA parsers that handle rare and unknown words differently on the German TiGer treebank. We demonstrate that methods that have improved results for other languages do not transfer directly to German, and that we can obtain better results using a simplistic model rather than a more generalized model for rare and unknown word handling.
Abstract People remember events and materials better when these are congruent with their mood at retrieval; this is known as the mood-congruent memory bias. This effect is largest when the materials are self-referential and this is known as the self-reference effect. We present two word rating studies, to create a list of self-referential valenced words that may be used as stimuli to investigate the influence of valence on cognitive processing in depressive ruminators. Words selected from the Affective Norms for English Words pool were rated by an unselected sample for self-referentiality (Study 1) and validated with ratings provided by depressive ruminators. As hypothesized, depressive ruminators rated negative words as more self-referential than an unselected sample. Using this list, valence differentiated performance between depressive ruminators and healthy controls in a working memory updating task. We thus created a list of self-referential valenced words matched on factors that influence word processing.
This study aims at exploring new norms as to the textual additions in parentheses (=TAiPs) in the translation of a Quranic text as writer-oriented devices of textuality. Coding for this sort of information could be useful in establishing an impact on any decision-making process on the TL version; such TAiPs can give a translated text of the Quran unity and purpose and distinguish it from a disconnected sequence of sentences. Six small-sized chapters of the Quran were selected as a research sample including a number of four handred forty two (442) TAiPs. Two writer-oriented kinds of textuality were found: cohesivity at the levels of grammar and lexis to be in form of recurrence, reference, substitution, ellipsis and conjunction; and relationality by coherence and intentionality to be in form of reiteration, collocation, connotation, evocation and interpretation. The study is a detailed analysis of such a severely criticized yet officially approved English interpretation of the Quran as the Hilali and Khan Translation (=HKT) against a predetermined set of text-linguistic norms. The strength or weakness of TAiPs as to how they might alleviate or aggravate the TL version is eventually identified for sake of improvement.
We propose the dense RNN, which has the fully connections from each hidden state to multiple preceding hidden states of all layers directly. As the density of the connection increases, the number of paths through which the gradient flows can be increased. It increases the magnitude of gradients, which help to prevent the vanishing gradient problem in time. Larger gradients, however, can also cause exploding gradient problem. To complement the trade-off between two problems, we propose an attention gate, which controls the amounts of gradient flows. We describe the relation between the attention gate and the gradient flows by approximation. The experiment on the language modeling using Penn Treebank corpus shows dense connections with the attention gate improve the model’s performance.
This study examines how the acoustic input (the surface form) and the abstract linguistic representation (the underlying representation) interact during spoken word recognition by investigating left-dominant tone sandhi, a tonal alternation in which the underlying tone of the first syllable spreads to the sandhi domain. We conducted an auditory-auditory priming lexical decision experiment on Shanghai left-dominant sandhi words, in which each disyllabic target ([tɕi55 dɛ31] “egg”) was preceded by monosyllabic primes either sharing the same underlying tone ([tɕi55]), surface tone ([tɕi53] “machine”), or being unrelated to the tone of the first syllable of the sandhi targets ([tɕi24] “to remember”). Results showed a surface priming effect, but not an underlying priming effect. Moreover, the surface priming did not interact with speakers’ familiarity ratings to the sandhi targets. The results are discussed in the context of how phonological opacity, productivity, and the directionality of tone sandhi patterns influence the representation of tone sandhi words as well as how the lexicality of the primes and the participants’ usage pattern of Shanghai may have influenced the results.
The Universal Dependencies project is currently comprised of 71 languages and 122 treebanks, and aims to find morphological and syntactic characteristics that can be applied to multiple languages for parallel language processing. In this paper, we introduce Universal POS, which is a morphological tagset for UD, and propose a method to automatically convert existing Korean morphological tagset into UPOS. In order to apply the UPOS tagset, which is based on refraction words such as English, to the Korean language, it is necessary to try a one-to-many mapping between the UPOS individual tag and the 21st century Sejong tag combination. (Yonsei University)
Because the most common transition systems are projective, training a transition-based dependency parser often implies to either ignore or rewrite the non-projective training examples, which has anadverse impact on accuracy. In this work, we propose a simple modification of dynamic oracles, which enables the use of non-projective data when training projective parsers. Evaluation on 73~treebanks shows that our method achieves significant gains (+2 to +7 UAS for the most non-projective languages) and consistently outperforms traditional projectivization and pseudo-projectivizationapproaches.
Polycentric Spanish Norm Towards the Polish‑Spanish Legal Translation The Spanish, being the official language of Spain and many other countries, is characterized by an important dialectal diversity that is reflected in the differences at all linguistic levels: phonetic, morphological, syntactic and lexico‑semantic, etc. All these differences raise controversies and discussions about the existence of a linguistic norm depending on the perspective that can have a monocentric or polycentric character. In this contribution we present some arguments for the second one. To this end, we rely on translations, starting simultaneously from the semasiological and onomasiological perspective, of some Polish‑Spanish legal terms in which it is essential to take into account, the diatopic variation as well as the norm whose character is polycentric.
The short note describes the chart parser for multimodal type-logical grammars which has been developed in conjunction with the type-logical treebank for French. The chart parser presents an incomplete but fast implementation of proof search for multimodal type-logical grammars using the "deductive parsing" framework. Proofs found can be transformed to natural deduction proofs.
The article deals with the issue of translation as an important means of communication between individuals who speak different languages and belong to different cultures.The article analyzes the translation as interlingual communicative phenomenon.The translation process is determined by the linguistic norms, communicative situations, functional parameters of the original text and translation norms.The role and tasks of an interpreter in the process of intercultural communication of individuals are defined.
This chapter discusses the new and changing conditions for linguistic norms in literary fiction of the post-Soviet era. In particular, it looks at how the interrelationship between the language of literature (<italic>iazyk literatury</italic>) and the standard language (<italic>literaturnyi iazyk</italic>) has been challenged by several processes of sociolinguistic change, including initiatives in language policy.
The article discusses the terms that nominate the language of written artifacts documented by the Cyrillic on the Ukrainian-Byelorussian lands in the XIV-XVI centuries; the expediency of using the notion “literary language” as to the Ukrainian literary written tradition of the XIV–XVI centuries is clarified; the content of the term “linguistic norm” is outlined, its characteristics in the investigated period are determined.
The linguistic database is also positioned as an actual way of formalizing and organizing phraseological units, terms for designating types of phraseological units. The main principle of systematization of the latter in the study is the thesaurus principle, that is the filling of the paradigm «terminological system – terminological microsystem – terminological subsystem – term», represented by a linguistic database.
Abstract The aim of the contribution is to introduce a database of linguistic forms and their functions built with the use of the multi-layer annotated corpora of Czech, the Prague Dependency Treebanks. The purpose of the Prague Database of Forms and Functions (ForFun) is to help the linguists to study the form-function relation, which we assume to be one of the principal tasks of both theoretical linguistics and natural language processing. We demonstrate possibilities of the exploitation of the ForFun database. This article is largely based on a paper presented at the 16th International Workshop on Treebanks and Linguistic Theories in Prague (Bejček et al., 2017).
Functional alterations of the default mode network (DMN) are frequently reported in psychotic disorders, but the functional role of these alterations remains poorly known. In addition to previous studies that have applied different types of tasks or recorded resting-state neuroimaging data, there has recently been more interest in the use of movie stimuli in studying brain functioning in patient populations, because this could provide a more naturalistic account of brain functioning in real life-like situations. Seventy-one first-episode psychosis (FEP) patients (mean age = 26.0 yrs, 47 (66%) males) and 57 controls (mean age = 26.86 yrs, 24 (42%) males) from the Helsinki Early Psychosis Study watched scenes from the movie Alice in Wonderland (Tim Burton, 2010) during 3 T fMRI-BOLD imaging. We used intersubject correlation (ISC) analysis, in which the correlation between voxel-wise BOLD time series in every within-group pair of subjects is calculated. In this study, time-windowed ISC was calculated with a 10-TR (time of repetition, 1.8 s) window with 1-TR steps over the fMRI time series. In each ISC window, a two-sample t test was performed to obtain a t-statistic time series of differences between the groups. An independent group of control subjects (n = 17, 10 males, mean age 26.5 yrs) rated how emotionally arousing the currently seen events of the stimulus are, producing a time-varying rating used as a regressor. General linear model was used to identify brain regions where the t-statistic time series covaries with the arousal rating. To make the interpretation of results less ambiguous, the arousal rating was divided into high and low arousal regressor by z scoring the rating and taking only the positive and negative values, respectively. Nonparametric clusterwise permutation test was used for statistical inference (cluster-defining threshold of p = 0.05, familywise error corrected threshold of p = 0.05, number of permutations = 5000). Furthermore, by using an experience-sampling setup during the same brain-scanning session, a partially overlapping sample of participants reported how emotionally aroused they were feeling during scanning. The results show significant correlation between the t-statistic time series and low arousal regressor, especially in the DMN including the anterior and posterior cingulate cortex, medial prefrontal cortex, precuneus, and bilateral lateral temporoparietal regions. Closer inspection reveals that during moments of low arousal in the movie stimulus, the ISC of healthy controls goes up but the ISC of patients does not. In the experience-sampling portion of the study, the patients reported more arousal than the control subjects. Intersubject correlation in the DMN depended differentially on arousal in FEP patients and control subjects. More specifically, during moments when the stimulus was rated less emotionally arousing, control subjects’ DMN functioning synchronized more while the patients’ did not. In connection with the difference in reported arousal during the same imaging session, our findings provide preliminary evidence for a contribution of arousal on the functional alterations of the DMN and suggest that this may be related to higher baseline arousal in the patients. Higher arousal and the related distortion of high order integrative functioning that characterizes DMN could contribute to the pathogenesis of psychosis.
Temporal models based on recurrent neural networks have proven to be quite\npowerful in a wide variety of applications. However, training these models\noften relies on back-propagation through time, which entails unfolding the\nnetwork over many time steps, making the process of conducting credit\nassignment considerably more challenging. Furthermore, the nature of\nback-propagation itself does not permit the use of non-differentiable\nactivation functions and is inherently sequential, making parallelization of\nthe underlying training process difficult. Here, we propose the Parallel\nTemporal Neural Coding Network (P-TNCN), a biologically inspired model trained\nby the learning algorithm we call Local Representation Alignment. It aims to\nresolve the difficulties and problems that plague recurrent networks trained by\nback-propagation through time. The architecture requires neither unrolling in\ntime nor the derivatives of its internal activation functions. We compare our\nmodel and learning procedure to other back-propagation through time\nalternatives (which also tend to be computationally expensive), including\nreal-time recurrent learning, echo state networks, and unbiased online\nrecurrent optimization. We show that it outperforms these on sequence modeling\nbenchmarks such as Bouncing MNIST, a new benchmark we denote as Bouncing\nNotMNIST, and Penn Treebank. Notably, our approach can in some instances\noutperform full back-propagation through time as well as variants such as\nsparse attentive back-tracking. Significantly, the hidden unit correction phase\nof P-TNCN allows it to adapt to new datasets even if its synaptic weights are\nheld fixed (zero-shot adaptation) and facilitates retention of prior generative\nknowledge when faced with a task sequence. We present results that show the\nP-TNCN's ability to conduct zero-shot adaptation and online continual sequence\nmodeling.\n
This thesis aims to examine metapragmatic discourses on linguistic politeness illustrated in Korean language how-to literature. The primary task lies in contextualizing the native awareness of ene yeycel (linguistic politeness in Korean) within the interests or values of certain social groups. The first group, South Korean government-sanctioned agencies, led a linguistic campaign promoting a new standard speech model in 1992. Language professionals, the second group of social actors, produced popular language how-to literature, especially after the establishment of the hegemonic standard speech model. Both language standardizing policy and the participants in the how-to industry represent the cultural process of constructing language and social conventions. The “normative” culture of ene yeycel can be empowered and widely circulated, gaining wider social practice. Standardization of honorification came to the surface as a public issue along with a new “cultural policy” of the Ministry of Cultural Affairs in 1990. In this cultural-political circumstance, the social meaning of standardized honorification was rediscovered as indigenous culture, a group identity shared by Korean speakers. Positively valorizing honorification as linguistic and cultural tradition, the standardized model preserves the sophisticated use of honorifics and reinforces superior-inferior relationships. However, the standard model of ene yeycel can be subjective and arbitrary. Moreover, different styles are too easily proscribed as errors made by sloppy speakers. Language how-to literature produces more diversified interpretations than the standard speech manual. As language users are confronted with the challenges of finding the proper level of honorification, language how-to manuals provide justifications to help speakers prioritize linguistic norms when internalizing social relationships. Positive valorizations of honorification derive from a speaker's respect for the interlocutor's social status or personality. Negative valorizations of honorification view deferential politeness as a kind of discriminatory behaviour indexing power-difference. The positive or negative values of honorification are based on different concepts of ene yeycel and on different identifications of social relationships. Such conceptualizations rationalize whether speakers should support honorification or not, and lead them to discuss language use in current society.
Introduction:Iliac artery endofibrosis (IAE) is an uncommon disease, poorly studied pathology with devastating effects and different therapeutic approaches affecting young people who practise intensive sports, especially cyclists. The evolution of the process not only depends on the diagnosis and therapeutic action, but also on the acceptance and attitude of the patient and subsequent professional guidance. Case description:This is the case description of a professional triathlon athlete that had one previous iliac surgical revascularization for an IAE Iliac and was admitted in our department five times with subacute lower limb ischemia affecting both legs between 2013 and 2016. Clinical findings and image tests are reported, as well as medical procedures performed. Indications based on clinical, functional and imaging ratings were clear, but his professional activity was not completely abandoned. Finally, after four endovascular procedures with good immediate results, he was warned of the seriousness of the process since the etiopathogenic reason. At the present moment patient is asymptomatic, under routine controls, working as successful triathlon coach. Discussion and conclusion:The fact that an external mechanical stress is the reason of repeated iliac artery injury suggests that an open surgical approach correcting the external muscular compression or arterial deformation should be a definitive but also aggressive solution according to literature. However, endovascular procedures and new endovascular devices are an increasingly promising option with a very low surgical risk. No matter the revascularization performed, the persistence of sports intensive practice carries a high risk of recurrence. Sport practise cessation is mandatory in some cases in order to assure revascularization long-term patency, but also a well conducted professional orientation is needed to complete the therapeutic action.