Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Introduction. Nowadays the language of modern television broadcasts’ speechesis more and more in the focus of linguistic study. Special interest of given paper is comprisedby the variability in gender categorization of nouns as it occurs in speeches of Ukrainian TVprograms anchorpersons. The object of the paper is the choice of nouns in modern TV speech,that is distinguished due to the variability of grammatical category of the gender of nouns waysof realization.Purpose of the article is to analyze the nouns that have suffered changes in thegrammatical category of gender within the current language trends being implemented inthe broadcast of Ukrainian television. In addition, the aim is to outline some of the reasonsfor the emergence of such tendencies, the relative use of these nouns, the degree of codificationin modern lexicographic sources.Methods of research. The research is grounded on descriptive method, the method ofempiric analysis, immediate constituents’ analysis, and contextual analysis.Results. Studying the language used in contemporary Ukrainian TV speeches convincinglydemonstrated how formal-grammatical indicators of the category of gender of nouns in moderntelecommunication vary (shift), implemented in the following modifications: male genius femalegenus, female genus male genus. In addition, the review of codification in the dictionariesof the analyzed noun units shows that variational changes in the morphology of the noun andin particular, in its morphological and grammatical categories nowadays cause not onlythe dislocation of the current linguistic norm, but also tend to change the morphological norms. Conclusion. Analyzing the language used in speeches of informative and entertainingUkrainian TV programs’ anchorpersons, we have concluded that they prefer to choose differentgrammatical variations in favor of a specific counterpart to the grammatical category of the gender,often a revitalized or dialectal word used.
There exist distinctive words that are used to express same semantics and as a result of this it has become hard to quantify the exact matching of words. To deal with this issue, past investigations endeavored to ascertain a likeness between distinctive pair of words. Conventional methodologies for computing word similarity are based on repositories like WordNet. It is a manually created lexical database and it processes semantic connection between various words. However, WordNet is a universally useful asset but wide range of words are not present in it and furthermore there exist an issue of identifying the meaning of words. Implication of words are diverse in WordNet when we utilize it in a textual framework. There exists a need of the refined approach that can gauge words resemblance in light of their co-occurrence. In this examination, we proposed an approach that registers likeness in text particular words, with the assistance of literary substance of various posts on StackOverflow. Our proposed strategy figures out word similarities in text by ascertaining the weighted co-occurrence in view of Computing Term Cooccurrence (CTC) and SentiWordNet. The exploratory outcome demonstrates that our system proposed an arrangement of words that are identified with text data is exceptional. Moreover, when it was compared with WordNet-based strategy named as WordNetres, it results with better outcomes.
In this work we describe the system built for the three English subtasks of\nthe SemEval 2016 Task 3 by the Department of Computer Science of the University\nof Houston (UH) and the Pattern Recognition and Human Language Technology\n(PRHLT) research center - Universitat Polit`ecnica de Val`encia: UH-PRHLT. Our\nsystem represents instances by using both lexical and semantic-based similarity\nmeasures between text pairs. Our semantic features include the use of\ndistributed representations of words, knowledge graphs generated with the\nBabelNet multilingual semantic network, and the FrameNet lexical database.\nExperimental results outperform the random and Google search engine baselines\nin the three English subtasks. Our approach obtained the highest results of\nsubtask B compared to the other task participants.\n
Abstract. Patterns of facial reactivity and attentional allocation to emotional facial expressions, and how these are moderated by gaze direction, are not clearly established. Among a sample of undergraduate university students, aged between 17 and 22 years (76% female), corrugator and zygomatic reactivity, as measured by facial electromyography, and attention allocation, as measured by the startle reflex and startle-elicited N100, was examined while viewing happy, neutral, angry and fearful facial expressions, which were presented at either 0- or 30-degree gaze. Results indicated typically observed facial mimicry to happy faces but, unexpectedly, “smiling” facial responses to fearful, and to a lesser extent, angry faces. This facial reactivity was not influenced by gaze direction. Furthermore, emotional facial expressions did not elicit increased attentional allocation. Likewise, matched facial expressions did not elicit increased attentional allocation. Rather, happy and fearful faces with direct (0°) gaze elicited increased controlled attentional allocation, and averted (30°) gaze faces, regardless of emotional expression, elicited preferential, early cortical processing. These findings suggest typical facial mimicry to happy faces, but unexpected facial reactivity to angry and fearful faces, perhaps due to an attempt to regulate social bonds during threat perception. Findings also suggest a divergence in controlled versus preferential, early cortical attentional processing for direct compared to averted gaze faces. These findings relate to young, mostly female, adults attending university. The experiment should be repeated with a larger sample drawn from the general community, with a broader age range and gender balance, and with a stimulus set with validated subjective valence and arousal ratings. This can reduce Type II error and establish normative patterns of facial reactivity and attentional processing of emotional facial expressions with different gaze directions.
Requirement is a formal expression of user’s need. It is the main foundation of any software development project. Natural language (NL) is often used to express and write system requirements specifications as well as user requirements. However, there is a very high probability that more than half natural language requirements can be ambiguous, incomplete and inaccurate. A software engineer can miss-interpret the natural language requirements and can generate an erroneous software model, which finally will lead to project failure. Earlier, we have introduced a prototype tool that provides natural language requirements authoring facilities and consistency checking to assist requirement engineers when working with informal and semi-formal requirements. However, the tool has pattern limitation to support the extraction of the essential requirements from the NL requirements. Therefore this study is aimed to enhance the accuracy and scalability of the tool to capture the essential requirements from the NL requirements. Our approach is to implement lexical analysis and embed an English lexical database where it will serve as a thesaurus in the tool. This tool is expected to be able to find the synonym of the extracted phrases (essential requirements) in the database to match it to the essential interaction pattern (phrases and expressions) in the library. Our future work will focus on the next phase of requirements engineering, which is requirements validation.
Artikkeli käsittelee suomentamiseen liittyviä ideologioita ja normeja 1800-luvun tietokirjallisuudessa. Tapaustutkimuksena on Werner Söderström Osakeyhtiön tietokirjojen suomennostoiminta 1800-luvun lopulla. Tutkimus kytkeytyy kääntämisen sosiologiaan ja historiaan, ja siinä arvioidaan myös, miten ja missä määrin historiallisia käännösprosesseja voidaan rekonstruoida. Käännösprosesseja lähestytään tarkastelemalla eri toimijoiden − kustantaja, kääntäjä, kieliasiantuntija, tekstin arvioija − osuutta käännösprosessissa. Tutkimuksen aineistona on kustantajan ja kääntäjän kääntämistä ja kielellisiä valintoja käsittelevä kirjeenvaihto, jonka avulla on mahdollista valottaa eri suunnista kääntäjän arkea, yhteisöllisiä arvoja ja normeja käännösvalintojen taustalla sekä niitä henkilökohtaisia asenteita, jotka ohjaavat kääntäjiä erilaisiin valintoihin.
 Analyysin tuloksena voi päätellä, että ammattikirjoittajina kääntäjät olivat hyvin tietoisia erilaisista kielellisistä ja kääntämiseen liittyvistä normeista. Käytännön työssä kääntäjät toimivat kuitenkin usein erilaisten normien ristipaineessa, jolloin vastakkain asettuivat esimerkiksi alkuteoksen luonteen säilyttäminen ja toisaalta sen kotouttaminen. Kääntäjät olivat myös tietoisia kielen vaihtelevista normeista, tunsivat käynnissä olevat kielikeskustelut ja mukauttivat herkästi kielenkäyttöään kulloinkin vallitsevien kirjakielen normien mukaiseksi.
 
 Norms and ideologies of translation in light of correspondence between publisher and translator in 19th-century Finland
 This article analyses the ideologies and norms that guided the translation of works of non-fiction in 19th-century Finland. As a case study the article analyses the processes involved in the publication of non-fiction at the Werner Söderström Ltd publishing house at the end of the 19th century. The research takes as its base theories examining the sociology and history of translation. It also aims to evaluate how and to what extent historical translation processes can be reconstructed. Translation is approached as a collaborative process involving various actors: publisher, translator, language editor, and expert reader. The data consists of correspondence between publisher and translator that deals with matters of translation or language. This correspondence sheds light on the everyday life of the translator and the socially accepted norms and ideologies that guide the translation process. It also reveals the stance of publishers concerning the choice of translator, a factor that can lead to very different end products.
 The analysis shows that, as professional writers, translators at the end of the 19th century were well aware of contemporary translational norms. In practice, translators were caught between various conflicting pressures – regarding, for instance, questions such as whether one should follow the original text as close as possible to preserve its unique style or assimilate the text to a Finnish context to help the reader. The data also shows that translators were well aware of linguistic norms; they were acquainted with current and past debates, and in assimilating their use of language they remained sensitive to prevailing norms.
Estimating the entropy based on data is one of the prototypical problems in distribution property testing and estimation. For estimating the Shannon entropy of a distribution on $S$ elements with independent samples, [Paninski2004] showed that the sample complexity is sublinear in $S$, and [Valiant--Valiant2011] showed that consistent estimation of Shannon entropy is possible if and only if the sample size $n$ far exceeds $\frac{S}{\log S}$. In this paper we consider the problem of estimating the entropy rate of a stationary reversible Markov chain with $S$ states from a sample path of $n$ observations. We show that: (1) As long as the Markov chain mixes not too slowly, i.e., the relaxation time is at most $O(\frac{S}{\ln^3 S})$, consistent estimation is achievable when $n \gg \frac{S^2}{\log S}$. (2) As long as the Markov chain has some slight dependency, i.e., the relaxation time is at least $1+Ω(\frac{\ln^2 S}{\sqrt{S}})$, consistent estimation is impossible when $n \lesssim \frac{S^2}{\log S}$. Under both assumptions, the optimal estimation accuracy is shown to be $Θ(\frac{S^2}{n \log S})$. In comparison, the empirical entropy rate requires at least $Ω(S^2)$ samples to be consistent, even when the Markov chain is memoryless. In addition to synthetic experiments, we also apply the estimators that achieve the optimal sample complexity to estimate the entropy rate of the English language in the Penn Treebank and the Google One Billion Words corpora, which provides a natural benchmark for language modeling and relates it directly to the widely used perplexity measure.
The article presents the results of word-formative and semantic analysis of Middle Czech verbs consisting of the prefix roz(e)- contained in Lexical database of humanistic and baroque Czech (https://madla.ujc.cas.cz). The analysis partly confirms, partly corrects the results of earlier analyses, above all, carried out by D. Šlosar (1981).
This paper aims to present the theoretical considerations and methodology used in elaborating a formal characterization of a fuzzy grammar in natural language grammars. It specifically focuses on the syntax of Spanish. However, we suggest that this methodology could be used to define a universal model for describing any kind of natural language grammar that takes into account fuzziness. Objective data based on frequencies were extracted using the Spanish Universal Dependencies Corpus Treebank and the Marsagram tool. These data allow us to describe Spanish Natural Language Grammar in terms of its constraints using the Property Grammars Theory. The work presented here could be applied in the form of an algorithm for parsing which could be of benefit to various areas of language and technology such as self-taught language learning software (in which violations and degree of violation could be tagged), data mining or human-machine interfaces.
We explore whether it is possible to build lighter parsers, that are statistically equivalent to their corresponding standard version, for a wide set of languages showing different structures and morphologies. As testbed, we use the Universal Dependencies and transition-based dependency parsers trained on feed-forward networks. For these, most existing research assumes de facto standard embedded features and relies on pre-computation tricks to obtain speed-ups. We explore how these features and their size can be reduced and whether this translates into speed-ups with a negligible impact on accuracy. The experiments show that grand-daughter features can be removed for the majority of treebanks without a significant (negative or positive) LAS difference. They also show how the size of the embeddings can be notably reduced.
Let me begin by thanking the Association for Computational Linguistics and its Executive Committee for conferring on me the great honor of their Lifetime Achievement Award for 2018, which of course I share with all the wonderful students and colleagues that have made many essential contributions to this work over many years.At the heart of the work that I have been pursuing over my research lifetime so far, whether in parsing and sentence processing, spoken language understanding, semantics, or even in musical understanding by machine, there lies a theory of natural language grammar that brings parsing, compositional semantics, statistical modeling, and logical inference into the closest possible relation. This theory of grammar is combinatory, in the sense that its operations are type-dependent and restricted to strictly string-adjacent phonologically or graphologically-realized inputs, and categorial, in the sense that those operands pair a syntactic type with a type-transparent semantic representation or logical form.I'd like to use this opportunity to briefly address three questions that revolve around the theory of grammar, both combinatory and otherwise. The first question concerns the way that Combinatory Categorial Grammar (CCG) was developed with a number of colleagues, over a number of stages and in slightly different forms. The second is an essentially evolutionary question of why natural language grammar should take a combinatory form. The third question is that of what the future holds for CCG and other structural theories of grammar in computational linguistics and NLP in the age of deep learning.I have called this talk "The Lost Combinator" in homage to the Victorian era poem "The Lost Chord," in the hope of suggesting that the theoretical development of CCG has always been empirical, rather than axiomatic, in search of the simplest explanation of the facts of language, rather than for confirmation of linguistic received opinion, however intuitively salient.In the late 1960s (when I was a psychology undergraduate at the University of Sussex under Stuart Sutherland, and then started as a graduate student in artificial intelligence at Edinburgh under Christopher Longuet-Higgins), a broad community of theoretical linguists, psychologists, and computational linguists saw themselves as all working on the same problem, under the definition provided by the "transformational" theory of grammar proposed by Chomsky (1957, 1965), using theories of psycholinguistic processing, language acquisition, and language evolution proposed by Lashley (1951), Miller, Galanter, and Pribram (1960), Miller (1967), and Lenneberg (1967), theories of natural language semantics proposed by Carnap (1956), Montague (1970), and Lewis (1970), and computational models of parsing such as those proposed by Thorne, Bratley, and Dewar (1968) and Woods (1970). (I myself was so convinced that this program would succeed that I believed it was time to apply the same methods to other cognitive faculties, taking as my research project for Ph.D. their application to the interpretation of music by machine, following the lead of Max Clowes [1971] in machine vision.)Almost immediately, this consensus fell apart. First, Chomsky himself was among the first (1965) to recognize that transformational rules, though descriptively revealing, were so expressive as to have little explanatory force, and required many apparently arbitrary constraints (Ross 1967). Second, psychologists realized that psycholinguistic measures of processing difficulty of sentences bore almost no relation to their transformational derivational complexity (Marslen-Wilson 1973; Fodor, Bever, and Garrett 1974). Finally, computational linguists attempting to implement transformational grammars as parsers realized that they were spending all their time implementing even more constraints on rules, in order to limit search arising from overgeneration (Friedman 1971; Gross 1978). (Meanwhile, I realized that the problem had not in fact been solved, and returned to natural language processing, thanks to a postdoc at Sussex with Philip Johnson-Laird.)This disillusion wasn't just a case of internal academic squabbling. There were also a couple of influential reports commissioned by the U.S. and UK governments that ended funding for machine translation (MT) and artificial intelligence (AI) (Pierce et al. 1966; Lighthill 1973). As a result of the second of these reports, which determined that AI was never going to work, PhDs in artificial intelligence like my classmate Geoff Hinton and myself spent ten years or so after graduation in psychology departments (in my case, at the Universities of Sussex and Warwick), until yet another report said AI was working after all and that Britain and the U.S. were falling behind Japan in this vital area. As a result, I could get hired again in computer science, first briefly back at Edinburgh, and then at the University of Pennsylvania (I learned a lesson from this odyssey that I have tried to remember whenever I have been appointed to a committee to report on anything, which is that while reports very rarely do any good, they can very easily do a great deal of harm.)Meanwhile, as a result of these conflicts, the scientific study of language fragmented. The linguists swiftly abjured any responsibility for their grammars ("Competence") bearing any relation to processing ("Performance"). Because the psychologists could hardly abandon Performance, they in turn became agnostic about grammar, retreating to context-free surface grammar (which they tended to refer to as "parsing strategies"), or a touchingly optimistic belief in its emergence from neural models. Meanwhile, the computational linguists (whose machines were growing exponentially in size and speed from the 16K byte core of the machine that supported the whole group when I started my graduate studies, on to levels that would soon permit parsing the entire contents of the then embrionic Web) similarly found that very little of what the linguists and psychologists cared about was usable at scale, and that none of it significantly improved overall performance over very much simpler context-free or even finite-state methods that the linguists had shown to be incomplete. The reason of course was Zipf's law, which means that the events with respect to which the low-level methods are incomplete are off in the long tail.It also became apparent to a few computationalists working on speech, MT, and information retrieval that the real problem was not grammar but ambiguity and its resolution by world-knowledge, and that the solution lay in probabilistic models (Bar-Hillel 1960/1964; Spärck Jones 1964/1986; Wilks 1975; Jelinek and Lafferty 1991) (although it was not immediately apparent how to combine statistical models with grammar-based systems without making obviously false independence assumptions).Nevertheless, as any red-blooded psychologist had always insisted, the divorce between competence and performance that everyone else had accepted did not make any sense. The grammar and the processor had to have evolved in lock-step, as a package deal, for what could be the evolutionary selective advantage of a grammar that you cannot process, or a parser without a grammar?It seemed equally obvious that surface syntax and the underlying semantic or conceptual representation must also be closely related, since the only reasonable basis for child language acquisition that has ever been on offer is that the child attaches language-specific grammar to a universal conceptual relation or "language of mind" (Miller 1967; Bowerman 1973; Wexler and Culicover 1980). It seemed to follow that radically new theories of grammar were needed.Theoretical linguists agree that the central problem for the theory of grammar is discontinuity or non-adjacent dependency between predicates and their arguments:Chomsky described discontinuity in terms of movement, which was known to be formally very unconstrained. By contrast, the ATN parser used in the LUNAR project (Woods, Kaplan, and Nash-Webber 1972) reduced all discontinuity to local operations on registers (Thorne, Bratley, and Dewar 1968; Bobrow and Fraser 1969; Woods 1970).In particular, unbounded wh-dependencies like the above were handled by: (a) putting a pointer into a * or HOLD register as soon as the "which" was encountered without regard to where it would end up; and (b) retrieving the pointer from HOLD when the verb needing an object "had" was encountered without regard to where it had started out. (It also included an ingenious mechanism for coordination called SYSCONJ, which one finds even now being reinvented on an almost yearly basis—cf. Woods [2010].) A * register was also used for wh-constructions within a systemic grammar framework by Winograd (1972, pages 52–53) in his inspiring conversational program SHRDLU.However, it was unclear how to generalize the HOLD register to handle the multiple long-range dependencies, including crossing dependencies, that are found in many other languages. In particular, if the HOLD register were assumed to be a stack, then the ATN becomes a two-stack machine (since we are already implicitly using one stack as a PDA to parse the context-free core grammar).On the computational side at least, the reaction to this impass took two distinct forms. Both reactions took the form of trying to reduce the two major operators of the transformation theory, substitution of immediate constituents, or what is nowadays called "Merge," and "Move," or displacement of non-immediate constituents, to one. On the one hand, Lexical Functional Grammar (Bresnan and Kaplan 1982) and Head-driven Phrase Structure Grammar (Pollard and Sag 1994) followed Kay (1979) in making unification the basis of movement and merger. Because unification can pass information across unbounded structures, this can be thought of as reducing Merge to Move.On the other hand, Generalized Phrase Structure Grammar (Gazdar 1981), Tree Adjoining Grammar (TAG; Joshi and Levy 1982), and Combinatory Categorial Grammar (CCG, Ades and Steedman, 1982) sought to reduce Move to various forms of local merger. In particular, the latter authors suggested that the same stack could be used to capture both long-range dependency and recursion in CCG.1Natural language grammar exhibits discontinuity because semantically language is an applicative system. Applicative systems (such as programming languages) support the twin notions of: (a) Application of a function/concept to an argument/entity; and (b) Abstraction, or the definition of a new function/concept in terms of existing ones.Language is in that sense inherently computational. It seems to follow that linguistics is (or should be) inherently computational as well. (Of course, it does not follow that computationalists have nothing to learn from linguistics.)There are two ways of modeling abstraction in applicative systems: Taking abstraction itself as a primitive operation (λ-calculus, LISP):(2)a.fatherEsau⇒Isaacb.grandfather=λx.father(fatherx)c.grandfatherEsau⇒Abrahamor Defining abstraction in terms of a collection of operators on strictly adjacent terms aka Combinators, such as function composition (Combinatory Calculus, MIRANDA).(3)b′.grandfather=BfatherfatherThe latter does the work of the λ-calculus without using any variables.Despite the resemblance of the "traces" (or copies) and "operators" (or complementizer positions) of the transformational theory to the λ-operators and variables of applicative systems of the first kind, natural language actually seems to be a system of the second, combinatory kind. The evidence stems from the fact that natural language deals with all sorts of fragments that linguists do not normally think of as semantically typable constituents, without the use of any phonologically realized equivalent of variables, such as pronouns:(4)a.Give[Anna books]?and[Manny records]?b.(Mother to child): There's adoggie![Youlike]?#the doggie.c.Food that you must[washVP/NP[before eating](VP∖VP)/NP]?.d.ik denk dat ik1Henk2Cecilia3[zag1leren2zingen3]?These fragments are diagnostic of a Combinatory Calculus based on Bn, T, and the "duplicator" Sn, plus application (Steedman 1987; Szabolcsi 1989; Steedman and Baldridge 2011).2CCG lexicalizes all bounded dependencies, such as passive, raising, control, exceptional case-marking, and so forth, via lexical logical form. All syntactic rules are Combinatory—that is, binary operators over contiguous phonologically realized categories and their logical forms. These rules are restricted by a Combinatory Projection Principle, which in essence says they cannot override the decisions already taken in the language-specific lexicon, but must be consistent with and project unchanged the directionality specified there. All such language-specific information is specified in the lexicon: The combinatory rules like composition are free and universal. All arguments, such as subjects and objects, are lexically type-raised to be functions over the predicate, as if they were morphologically cased as in Latin, exchanging the roles of predicate and argument.All long-range dependencies are established by contiguous reduction of a wh-element, such as (N∖N)/(S/NP), with an adjacent non-standard constituent with category S/NP, formed by rules of function composition.The combinatory rules synchronize composition of the syntactic types shown here with corresponding composition of logical forms (suppressed in the derivations above), to yield the logical forms shown as λ-terms for the resulting nouns N.To capture the construction in Example (4c), whose syntactic derivation we pass over here, we also need rules based on the duplicator S:(7)a."wash X before eating X″b.VP/NP:λx.before(eatx)(washx)≡S(Bbeforeeat)washTo capture constructions like Example (4d) (whose syntactic derivation is similarly suppressed), we also need rules based on second-order composition B2:(8)a."Y saw X teach W to sing.″b.((S∖NP)∖NP)∖NP:λwλxλy.help(teach(singw)wx)xy≡B2sees(Bteachsing)CCG thus reduces the operator move of transformational theory to applications of purely adjacent operators—that is, to recursive combinatory merge.Interestingly, the latest "minimalist" form of the transformational theory has also proposed that Move should be relabeled as an "internal" form of standard or "external" Merge (Chomsky 2001/2004, page 110), though without providing any formal basis for the reduction other than identifying internal Merge as "a grammatical transformation." (If anything deserved the soubriquet "the lost combinator," it would be this notional unitary combination of application and abstraction in a single perfect operator, linking or merging all types, as in the epigraph to this article.)B2 rules allow us to "grow" categories of arbitrarily high valency, such as ((S∖NPy)∖NPx)∖NPw. As we saw earlier, in some Germanic languages like Dutch, Swiss German, and West-Flemish, serial verbs are linearized using such rules to require crossing discontinuous dependencies. Thus, B2 rules give CCG slightly greater than context-free power.Nevertheless, CCG is still not as expressive as movement. In particular, we can only capture permutations that are what is called "separable," where separability is related to the idea of obtaining the permutations by rebracketing and rotating sister nodes (Steedman 2018).For example, for the categories of the form A|B, B|C, C|D, and D, it is obvious by inspection that we cannot recognize the following permutations:(9)i.*B|CDA|BC|Dii.*C|DA|BDB|CThis generalizatiion appears likely to be true cross-linguistically for the components of this form for the NP "These five young boys":(10)i.*Five boys these youngii.*Young these boys fiveTwenty-one of the 22 separable permutations of "These five young boys" are attested (Cinque 2005; Nchare 2012). The two forbidden orders are among the unattested three.3The probability of this happening by chance is the probability of the is the of the number of ways of two of three unattested by the is, about one in a (If the order that unattested so were to be this chance would to about one in number of separable permutations much more in than the number of all example, for around of the permutations are There are obvious for the problem of in machine translation and neural semantic parsing, to which we I to and at first as a still under least, that was what Joshi and then as a of where much of the development of CCG was out. students in a of and Joshi 1987; and that the of Ades and Steedman was by that both CCG and were equivalent to Grammar (Gazdar a new of the by the and of languages by these fell within the of what Joshi called which proposed as a for what could as a theory of natural that they are and and some limit on crossing dependencies. the is much much than the including the multiple free languages and even the languages of so it seems to and CCG as with to the its CCG was assumed to be as a grammar for parsing, because of the derivational ambiguity by and the combinatory rules, these also under and as any grammar with the same as CCG the same of in the parser it is there in the In this is just another in the of derivational ambiguity that all natural language and can be handled by the same statistical models as other particular, the dependency models by and are and Steedman and CCG is also to parsing with which can be using and long and Steedman and is now used in those that for between semantic and syntactic processing, such as machine translation and and machine and parsing and Steedman 1987; and et al. et al. et al. and semantic parser and 2005; et al. et al. of my work with the same has returned to their application in musical and and Steedman and shown that CCG grammars of the same and parsing models of the same statistical are required there as It is only in the of their compositional semantics that music and language very than on this of work in like to by two more The first is an evolutionary should natural language be a combinatory in the first The second is a question about the future development of CCG and other grammar-based theories to be to NLP in the age of deep and recursive neural take these questions in order in the two language like a combinatory applicative system because T, and evolved in to support of before there was any language (Steedman like need and are for of need to to make can form the that you need to before is to that they allow and some other can with like B2 are to make with arbitrary of including and including other whose is yet to be to be to do the the and problem, using like to in order to that are of to of and so such as of and can apply to a and much work that the and other have to be there in the already for the to be to with was to that from the even if had the problem of can be as the problem of search for a of in a or of possible such search has the same recursive as parser example, there are both and for the The latter a more in evolutionary as a mechanism that is to both semantic interpretation and parsing, rather than the evolution of like the the for linguistic as as the operators for competence grammar, the two to as what was to above as an evolutionary hope to have convinced you that CCG grammars are both and as as semantic parsers that are to parse the like other CCG parsers are by the of parsing models based on only a of no how we using and the constructions they are dependencies and so as we have and off in the long As a parsers are to on performance overall by using models et al. says more about the of grammars and parsers than about the of deep in the of semantic parser for arbitrary such as and Steedman I would that CCG and other grammar-based parsers have already been by of deep neural and and the question of whether models for applications in are because we have to the universal semantic that allow the child to CCG for natural languages and that we to be using both in semantic parser and in Because the language is any of linguistic logical like the universal language of it be more with and to semantic parsers for by deep neural force, rather than by CCG semantic parser is it possible that the problem of parsing could be by neural by stack et al. et al. actually learn as has been seems likely that semantic parsers and neural machine translation to have difficulty with long-range because the evidence for their is so example, both and as a case, verb categories or a complementizer I is actually by which says that if you are an language with like and then you not in be to in languages like or languages like and German, you be to both subjects and objects, or at the time of a translation system no of learned these syntactic from a we get an sentence whose to translation means is the that the us the a is the that the us to the if we with we (which back into the a is the that the us holds the a we with using a the translation again an that is is the that they said had the is the they said they the contrast, CCG parsers do rather on and Steedman Steedman, and which is a construction that to be rather determined by the parsing methods are similarly when with long-range even when the sentence is in this think that the the and that it is et think that the the and it is in at least, these constructions could be learned by grammar-based semantic parser from using the methods of et al. and et al. make it likely that there be a need for in like where long-range dependencies like deep and are here to The future in parsing for such lies with systems using neural for and grammars for problem in NLP the fact that natural language understanding inference as as semantics, and we have no idea of the representation question is the almost it many it is almost equally to the information in a form that is not immediately with the form of the example, sentences like the following a different a rather than a an an a a and at a representation language that is we are using CCG parsers to the for between in order to consistent of between over of the same types, using over then an it and it under such as et al. of in the then that can be to a single relation and Steedman can be across from multiple languages and Steedman can then the semantics for relation with the and the entire using this now both and semantic an with the as and the as questions the in this we parse questions into the same semantics which is now the language of the the we use the and the and the following anything that in the then the the of anything is in the in the this to work, we need to be nodes in their in both and then be to like in of a semantic in the language of the need to learn between semantic and the language of the this project is the the function of a of the semantic in semantic like those of and and Steedman and while the form a similarly of the of Carnap and Fodor, Fodor, and Garrett semantic are essentially but with the advantage that they can be with logical operators such as and for the of semantic underlying natural language semantics in to another different to semantics that to use reduced of to using operations such as and and in of compositional It is an question whether can be with to of a and et al. It is likely that some of be here to the of the of the long like and work in do they work in In particular, can they learn all the syntactic in the long like and crossing in a way that support semantic they are not actually but are a finite-state or a then by on as for natural language processing, we are in of of the computational linguistic project of also providing computational of language and if we that like is a real and that learn their first language by of the sentences of their language the of the universal language of we still the of what that universal semantic language not get an to that question we can above using and and such as as for the language of to use machine for what it is such variables and their for use in a natural language work was supported in by a Award and a a University of Edinburgh and my and the and all my and students over many
We introduce a method to reduce constituent parsing to sequence labeling. For each word w_t, it generates a label that encodes: (1) the number of ancestors in the tree that the words w_t and w_{t+1} have in common, and (2) the nonterminal symbol at the lowest common ancestor. We first prove that the proposed encoding function is injective for any tree without unary branches. In practice, the approach is made extensible to all constituency trees by collapsing unary branches. We then use the PTB and CTB treebanks as testbeds and propose a set of fast baselines. We achieve 90.7% F-score on the PTB test set, outperforming the Vinyals et al. (2015) sequence-to-sequence parser. In addition, sacrificing some accuracy, our approach achieves the fastest constituent parsing speeds reported to date on PTB by a wide margin.
In recent years, with increase in the use of internet the multimedia contents on it have rapidly increased. Users may need to go through a video in a top down manner i.e. browsing the videos, or in bottom up manner i.e. retrieving specific information from videos. They may also want to go through the summary or through the highlights of the videos. This has necessitated the need to handle multimedia resources effectively. This paper proposes an automatic method for aligning scripts of lecture videos with captions. Alignment is needed to extract time information from captions and insert it in the scripts, to create index of the videos. No alignment work has been previously done in lecture videos domain. Alignment methods proposed for other type of videos are not applicable for lecture videos because, different similarity techniques behave differently on different types of datasets. The proposed method uses transcripts of lecture videos, SRT file of captions available along with lecture videos and captions generated from auto-caption generation feature of YouTube. The captions and scripts are then aligned using a dynamic programming technique. No such work has been previously done for lecture videos. Most important aspect of alignment is similarity measure. In the proposed work we have used three similarity measures cosine, jaccard, and dice. A comparative analysis of these measures is given in the paper. We also use a large lexical database of English words known as WordNet for word-to-word similarity. The experimental result shows comparison of various similarity techniques and YouTube captions.
Hoarding is a mental and public health problem stemming from difficulty associated with discarding one's possessions and resulting clutter. In the last decade, a visual method, called "Clutter Image Rating" (CIR), has been developed for the assessment of hoarding severity. It involves rating clutter in patient's home on the CIR scale from 1 to 9 using a set of reference images. Such assessment, however, is time-consuming, subjective, and may be non-repeatable. In this paper, we propose a new automatic clutter assessment method from images, according to the CIR scale, based on deep learning. While, ideally, the goal is to perfectly classify clutter, trained professionals admit assigning CIR values within ±1. Therefore, we study two loss functions for our network: one that aims to precisely assign a CIR value and one that aims to do so within ±1. We also propose a weighted combination of these loss functions that, as a byproduct, allows us to control the CIR mean absolute error (MAE). On a recently-collected dataset, we achieved ±1 accuracy of 82% and MAE of 0.88, significantly outperforming our previous results of 60% and 1.58, respectively.
STUDY DESIGN: Descriptive analysis using publicly available data. OBJECTIVES: The purpose of this study was 2-fold: to assess patient-rated trustworthiness of spine surgeons as a whole and to assess if academic proclivity, region of practice, or physician sex affects ratings of patient perceived trust. METHODS: Orthopedic spine surgeons were randomly selected from the North American Spine Society directory. Surgeon profiles on 3 online physician rating websites, HealthGrades, Vitals, and RateMDs were analyzed for patient-reported trustworthiness. Whether or not the surgeon had published a PubMed-indexed paper in 2016 was assessed with regard to trustworthiness scores. Total number of publications was also assessed. Individuals with >300 publications were excluded due to the likelihood of repeat names. RESULTS: Recent publication and total number of publications has no relationship with online patient ratings of trustworthiness across all surgeons in this study. Region of practice likewise has no influence on mean trust ratings, yet varied levels of correlation are observed. Furthermore, there was no difference in trust scores between male and female surgeons. CONCLUSION: Total academic proclivity via indexed publications does not correlate with patient perceived physician trustworthiness among spine surgeons as reported on physician review websites. Furthermore, region of practice within the United States does not have an influence on these trust scores. Likewise, there is no difference in trust score between female and male spine surgeons. This study also highlights an increasing utility for physician rating websites in spine surgery for evaluating and monitoring patient perception.
Based on research linking depressive symptoms and intimate partner aggression perpetration with negatively biased perception of social stimuli, the present authors examined biased perception of emotional expressions as a mechanism in the frequently observed relationship between depression and psychological aggression perpetration.In all, 30 university students made valence ratings (negative to positive) of emotional facial expressions and completed measures of depressive symptoms and psychological aggression perpetration.As expected, depressive symptoms were positively associated with psychological aggression perpetration in an individual's current relationship, and this relationship was mediated by ratings of negative emotional expressions.These findings suggest that negatively biased perception of emotional expressions within the context of elevated depressive symptoms may represent an early stage of information processing that leads to aggressive relationship behaviors.
In Russom (2011), I defended a universalist hypothesis that the constituents of poetic form are abstracted from natural linguistic constituents: metrical positions from phonological constituents, usually syllables; metrical feet from morphological constituents, usually words; and metrical lines from syntactic constituents, usually sentences. An important corollary to this hypothesis is that norms for realization of a metrical constituent are based on norms for the corresponding linguistic constituent. Optimality Theory provides a universalist account of relevant linguistic norms and deals effectively with situations in which norms conflict, employing ranked violable rules. Language Typology provides a universalist account of relevant syntactic norms. In this paper I integrate these independently grounded methodologies and use them to explain the distribution of constituents within the line, identifying a variety of important facts that seem to have escaped previous notice. Universalist claims are tested against meters from each of the major language types: subject-verb-object (SVO), subject-object-verb (SOV) and verb-subject-object (VSO). My findings are incompatible with the claim that “lines are sequences of syllables, rather than of words or phrases” (Fabb, Halle 2008: 11).
We investigated user experiences from 117 Finnish children aged between 8 and 12 years in a trial of an English language learning programme that used automatic speech recognition (ASR). We used measures that encompassed both affective reactions and questions tapping into the children' sense of pedagogical utility. We also tested their perception of sound quality and compared reactions of game and nongame-based versions of the application. Results showed that children expressed higher affective ratings for the game compared to nongame version of the application. Children also expressed a preference to play with a friend compared to playing alone or playing within a group. They found that assessment of their speech is useful although they did not necessarily enjoy hearing their own voices. The results are discussed in terms of the implications for user interface (UI) design in speech learning applications for children.
This paper aims to observe, describe and explain the translation of film titles with the Chinese character "Xia". In this paper, a small parallel corpus is firstly constructed and annotated to analyze the basic information, language features, translation principles, strategies, methods and techniques of film titles. And then, based on Holmes' methodology of Descriptive Translation Studies, some descriptive corpus is analyzed and some assumptions are predicated. Finally, these assumptions are divided into three preliminary norms, two initial norms, two matrix norms and three textual-linguistic norms under the guidance of Toury's translation norm system with a view to provide some references and inspiration for the standardization of film title translation.
Many of the leading approaches in language modeling introduce novel, complex and specialized architectures. We take existing state-of-the-art word level language models based on LSTMs and QRNNs and extend them to both larger vocabularies as well as character-level granularity. When properly tuned, LSTMs and QRNNs achieve state-of-the-art results on character-level (Penn Treebank, enwik8) and word-level (WikiText-103) datasets, respectively. Results are obtained in only 12 hours (WikiText-103) to 2 days (enwik8) using a single modern GPU.
This artifact depicts the depth map of the Rosetta stone, which was algorithmically generated in 2018 as part of the Digital Rosetta Stone project. The Digital Rosetta Stone is a project developed at Leipzig University by the Chair of Digital Humanities and the Egyptological Institute/Egyptian Museum Georg Steindorff in collaboration with the British Museum and the Digital Epigraphy and Archaeology Project of the University of Florida. The aims of the project are to produce a collaborative digital edition of the Rosetta Stone, address standardization and customization issues for the scholarly community, create data that can be used by students to understand the document in terms of language and content, and produce a high-resolution 3D model of the inscription. The three versions of the text were transcribed and outputted in XML, according to the EpiDoc guidelines. Next, the versions were aligned with the Ugarit iAligner tool that supports the alignment of ancient texts with modern languages, such as English and German. All three texts were then parsed syntactically and morphologically through Treebank annotation. Finally, the project explored new 3D-digitization methodologies of the Rosetta Stone in the British Museum that enhances traditional archaeological methods and facilitates the study of the artifact. The results of this work were used in different courses in Digital Humanities, Digital Philology, and Egyptology.
Neural network-based language models deal with data sparsity problems by mapping the large discrete space of words into a smaller continuous space of real-valued vectors. By learning distributed vector representations for words, each training sample informs the neural network model about a combinatorial number of other patterns. In this paper, we exploit the sparsity in natural language even further by encoding each unique input word using a fixed sparse random representation. These sparse codes are then projected onto a smaller embedding space which allows for the encoding of word occurrences from a possibly unknown vocabulary, along with the creation of more compact language models using a reduced number of parameters. We investigate the properties of our encoding mechanism empirically, by evaluating its performance on the widely used Penn Treebank corpus. We show that guaranteeing approximately equidistant (nearly orthogonal) vector representations for unique discrete inputs is enough to provide the neural network model with enough information to learn --and make use-- of distributed representations for these inputs.
We introduce a method to reduce constituent parsing to sequence labeling. For each word w t, it generates a label that encodes: (1) the number of ancestors in the tree that the words w t and w t+1 have in common, and (2) the nonterminal symbol at the lowest common ancestor. We first prove that the proposed encoding function is injective for any tree without unary branches. In practice, the approach is made extensible to all constituency trees by collapsing unary branches. We then use the PTB and CTB treebanks as testbeds and propose a set of fast baselines. We achieve 90.7% F-score on the PTB test set, outperforming the Vinyals et al. ( In addition, sacrificing some accuracy, our approach achieves the fastest constituent parsing speeds reported to date on PTB by a wide margin. 1
Recently, span-based constituency parsing has achieved competitive accuracies with extremely simple models by using bidirectional RNNs to model "spans". However, the minimal span parser of Stern et al (2017a) which holds the current state of the art accuracy is a chart parser running in cubic time, $O(n^3)$, which is too slow for longer sentences and for applications beyond sentence boundaries such as end-to-end discourse parsing and joint sentence boundary detection and parsing. We propose a linear-time constituency parser with RNNs and dynamic programming using graph-structured stack and beam search, which runs in time $O(n b^2)$ where $b$ is the beam size. We further speed this up to $O(n b\log b)$ by integrating cube pruning. Compared with chart parsing baselines, this linear-time parser is substantially faster for long sentences on the Penn Treebank and orders of magnitude faster for discourse parsing, and achieves the highest F1 accuracy on the Penn Treebank among single model end-to-end systems.
Italian unification aimed to ‘make Italian women’, as an equally important and complementary goal to that of forging men ‘of character’, worthy of being citizens of the new Italian nation. In the process of nation-building, women were entrusted with an essential role as educators and for this reason were subject to standardizing pressure by the cultural industry. There was a marked growth in the number of series of novels which served to instil moral and linguistic norms. The serial nature of these publications, the stereotypical social roles they depicted, and the paradigmatic nature of their stories — all transmitted in an accessible, Tuscanized Italian — ensured that a uniform model was presented to unmarried and married women. This article provides an analysis of a representative corpus of various genres of novel (e.g. moral novels, protest novels, romances) which reveals how women evolved from being passive and silent readers to authors of female-centred but not feminist novels, thus taking on a socio-political role, despite lacking the right to vote.
There is general agreement that the body expresses emotion, and while there is research on the expression of emotion in so-called “body language” (e.g. facial expression, posture), only little research exists on the expression of emotion in hand gestures. Two previous studies (Casasanto & Jasmin, 2010; Kipp & Martin, 2009) found opposing patterns of correlation between the gesturing hand and the valence of the co-occurring speech, and this thesis aims to help fill the research gap with the help of two empirical studies, the first on gesture production and the second on gesture perception, both in relation to valence. The first study analyses the handedness of gestures and valence of speech produced by three guests on the Tavis Smiley Show. For two of the three speakers it was found that gestures produced with the dominant hand tended to co-occur with positive speech, while gestures produced with the non-dominant hand tended to co-occur with negative speech, a pattern similar to that found by Casasanto and Jasmin (2010). No pattern of correlation was found for the third speaker. The second study used an experiment to explore whether the handedness of gesture affect the valence ratings of co-occurring speech, using systematically varied video data of one-word utterances and pragmatic gestures. This study found some evidence that words co-occurring with gestures performed with the right hand receive higher valence ratings, while words co-occurring with gestures performed with the left hand receive lower valence ratings, with a larger effect for negative words than positive words. In sum, this thesis supports previous findings of correlations between the handedness of gestures and the valence of speech in production, but also shows inter-individual variation and points out the need to consider demographic factors in future research. It found some evidence of the effect of handedness of gesture on the perceived valence of speech, but more research is needed, and should include left-handed participants and speakers to determine the pattern of correlation between handedness of gesture and speech valence. (Less)
Affective processing appears to be altered in tinnitus, and the condition is to a large extent characterized by the emotional reaction to the phantom sound. Psychophysiological models of tinnitus and supporting brain imaging studies have suggested a role for the limbic system in the emergence and maintenance of tinnitus. It is not clear whether the tinnitus-related changes in these systems are specific for tinnitus only, or whether they affect emotional processing more generally. In this study, we aimed to quantify possible deviations in affective processing in tinnitus patients by behavioral and physiological measures. Tinnitus patients rated the valence and arousal of sounds from the International Affective Digitized Sounds database. Sounds were chosen based on the normative valence ratings, that is, negative, neutral, or positive. The individual autonomic response was measured simultaneously with pupillometry. We found that the subjective ratings of the sounds by tinnitus patients differed significantly from the normative ratings. The difference was most pronounced for positive sounds, where sounds were rated lower on both valence and arousal scales. Negative and neutral sounds were rated differently only for arousal. Pupil measurements paralleled the behavioral results, showing a dampened response to positive sounds. Taken together, our findings suggest that affective processing is altered in tinnitus patients. The results are in line with earlier studies in depressed patients, which have provided evidence in favor of the so-called positive attenuation hypothesis of depression. Thus, the current results highlight the close link between tinnitus and depression.
Researchers increasingly recognize that exercise and physical activity may be determined by interacting explicit and implicit processes. While this trend has led to a substantial increase in the use of measures of implicit processes within exercise psychology, such measures are typically selected without a supporting rationale. To facilitate the refinement of theoretical models and the testing of interventions, investigating the internal consistency and validity of measures of implicit processes, such as automatic exercise associations, is crucial. Objectives: To assess the internal consistency and validity of nine measures of automatic exercise associations. Method: Participants (N = 95) completed an exercise session at the intensity of the ventilatory threshold, intended to generate heterogeneous affective responses. One week later, they also completed nine randomly ordered measures of automatic exercise associations. The slope of affect ratings during exercise, affect ratings at the end of exercise, recalled affect, explicit affective attitude, self-reported exercise, and situated decisions to exercise served as validation criteria. Results: Three of the nine measures exhibited acceptable to good internal consistency. Only the Approach-Avoidance Task was significantly and meaningfully related to any of the validation criteria (i.e., self-reported exercise and situated decisions to exercise). Conclusions: Most measures of automatic exercise associations exhibited unsatisfactory internal consistency and were unrelated to the validity criteria. For research on the role of implicit processes in exercise and physical activity to advance, further psychometric evaluation and refinement of measures is needed.
Automatically captioning images with natural language sentences is an important research topic. State of the art models are able to produce human-like sentences. These models typically describe the depicted scene as a whole and do not target specific objects of interest or emotional relationships between these objects in the image. However, marketing companies require to describe these important attributes of a given scene. In our case, objects of interest are consumer goods, which are usually identifiable by a product logo and are associated with certain brands. From a marketing point of view, it is desirable to also evaluate the emotional context of a trademarked product, i.e., whether it appears in a positive or a negative connotation. We address the problem of finding brands in images and deriving corresponding captions by introducing a modified image captioning network. We also add a third output modality, which simultaneously produces real-valued image ratings. Our network is trained using a classification-aware loss function in order to stimulate the generation of sentences with an emphasis on words identifying the brand of a product. We evaluate our model on a dataset of images depicting interactions between humans and branded products. The introduced network improves mean class accuracy by 24.5 percent. Thanks to adding the third output modality, it also considerably improves the quality of generated captions for images depicting branded products.
Based on dependency syntactic treebanks, the present study focuses on relative clauses (RCs) and explores their relations between dependency distance (DD), dependency direction, the embedding position, and the length of RC. It was found that: (1) the probability distributions of the DD of RCs aren't influenced by their embedding positions; (2) the DD of RC embedded in the object position is relevant to the length of the RC; (3) the longer the RC is (≥10), the more likely the dependency relations within the RC present a head-initial tendency; (4) the embedding positions of RCs in main clauses don't influence their dependency directions and dependency directions of RCESs and RCEOs are almost head-initial; (5) the changes of dependency relations of main clauses are related to the embedding position of relative clauses; (6) the embedded RCs increase the MDDs of their main clauses significantly.
This paper proposes a state-of-the-art recurrent neural network (RNN) language model that combines probability distributions computed not only from a final RNN layer but also from middle layers. Our proposed method raises the expressive power of a language model based on the matrix factorization interpretation of language modeling introduced by Yang et al. (2018). The proposed method improves the current state-of-the-art language model and achieves the best score on the Penn Treebank and WikiText-2, which are the standard benchmark datasets. Moreover, we indicate our proposed method contributes to two application tasks: machine translation and headline generation. Our code is publicly available at: https://github.com/nttcslab-nlp/doc_lm.
Enriched discourse annotation of a subset of the Prague Discourse Treebank, adding implicit relations, entity based relations, question-answer relations and other discourse structuring phenomena.
Social anxiety disorder affects approximately 7% of the adult population in the U.S., yet a vast majority of these individuals do not seek treatment. Thus, it is critical to examine models that deliver treatment to them. Computerized Cognitive Bias Modification (CBM) training programs can be effective in targeting interpretation bias, a key cognitive mechanism underlying social anxiety, and have potential for widespread dissemination, especially if they can be delivered via smart phones, which are becoming ubiquitous. However, the efficacy of CBM interpretation training paradigms that are adapted to and delivered via smart phones remains unknown. We present a pilot study to investigate if physiologic data can be used to track the changes over a smartphone-based CBM intervention for social anxiety. In a 3-week open trial, pilot study involving 20 high socially anxious participants, self-report affect ratings, heart rate and accelerometer data were collected using a smartphone and smartwatch before, after, and during the CBM intervention. The study focused on the relationship between accelerometer and heart rate to track change following the intervention. Results provide preliminary evidence for the viability of using physiological data to identify the change in mental state influenced by CBM interventions.
The widely deployed and easy-to-use Linguistic Inquiry and Word Count (LIWC) tool is the gold standard for many computerized text analysis tasks for many medical applications such as patient sentiment analysis, depression detection, and ADHD detection. Compared to most other natural language processing (NLP) tasks, in the medical field it is often very difficult to obtain large-scale data sets, making effective automatic representation learning from complex text patterns (e.g., using a deep auto-encoder) challenging. LIWC can solve this problem by using a human-designed dictionary as a substitution of a machine learning model to convert text into a concise and effective vector representation. However, while LIWC's dictionary is large, some potentially informative words might still be neglected due to the knowledge constraint of the dictionary editors. This problem is particularly conspicuous when the analyzed text is not a formal language (e.g., dialect, slang, or cyber words). To address this problem, we propose a new matching scheme that does not require an exact word match, but instead counts all words that are similar to a key in the LIWC dictionary. This scheme is implemented using WordNet, a large lexical database, and Word2Vec, a machine learning based word embedding technology. The output of the proposed method is in the exact same format as LIWC's output, thereby maintaining the usability. Similar to previous work, the proposed method can be viewed as a combination of human domain knowledge and machine learning for text representation encoding.
Abstract Non-configurationality is a linguistic property associated with free word order, discontinuous constituents, including NPs, and null anaphora of referential arguments. Quantitative metrics, based both on local networks (syntactic trees and word order within sentences) and on global networks (incorporating the relations within a whole treebank into a shared graph), can reveal correlations among these features. Using treebanks we focus on diachronic varieties of Ancient Greek and Latin, in which non-configurationality tapered off over time, leading to the largely configurational nature of the Romance languages and of Modern Greek. A property of global networks (density of their spectra around zero eigenvalues) measuring the regularity in word order is shown to be strengthened from classical to late varieties. Discontinuous NPs are traced by counting the words creating non-projectivity in dependency trees: these drop dramatically in late varieties. Finally, developments in the use of null referential direct objects are gauged by assessing the percentage of third-person personal pronouns among verb objects. All three features turn out to change over time due to the decay of non-configurationality. Evaluation of the strength of their pairwise correlation shows that null direct objects and discontinuous NPs are deeply intertwined.
Language is integral to educational processes because it forms the basis for classroom communication and the medium for knowledge transfer. However, language is imbued with race- and class-related ideologies: ideas about "proper" and "educated" uses of language. Language ideologies are shaped by the linguistic norms of powerful groups and are based on political rather than linguistic factors. In this paper, I explore how language ideologies operated in three educational sites on the Cape Flats. Multisite ethnography was used to research language ideologies in classrooms, amongst a hip-hop group, and at a youth radio show. Participants in the study spoke a variety of Afrikaans known as Kaapse Afrikaans, which differs from the standard Afrikaans inscribed in the school curriculum. The research showed that language ideologies were perpetuated through semiotic processes known as iconicity, recursiveness, and erasure. Through iconicity, Rosemary Gardens youths' language was inextricably linked to colouredness-a mixed race and language with low status attributed to both. Whereas standard Afrikaans was described as "pure, high, proper, and real," Kaapse Afrikaans was recursively depicted as "low, deficient and slang." These semiotic processes functioned to erase young people's use of language at schools, particularly repressing Kaapse Afrikaans in its written form. On certain occasions, the hip-hop group used language freely as they commented on their local environments. Powerful linguistic ideologies will continue to denigrate marginalised youth, even if radical teachers and hip-hop culture dismiss them. Educators should, therefore, both endorse the linguistic resources youth bring to classrooms and arm them with powerful forms of language and knowledge that hold power elsewhere.
This paper describes CzEngClass, a bilingual lexical resource being built to investigate verbal synonymy in bilingual context and to relate semantic roles common to one synonym class to verb arguments (verb valency). In addition, the resource is linked to existing resources with the same of a similar aim: English and Czech WordNet, FrameNet, PropBank, VerbNet (SemLink), and valency lexicons for Czech and English (PDT-Vallex, Vallex, and EngVallex). There are several goals of this work and resource: (a) to provide gold standard data for automatic experiments in the future (such as automatic discovery of synonym classes, word sense disambiguation, assignment of classes to occurrences of verbs in text, coreferential linking of verb and event arguments in text, etc.), (b) to build a core (bilingual) lexicon linked to existing resources, for comparative studies and possibly for training automatic tools, and (c) to enrich the annotation of a parallel treebank, the Prague Czech English Dependency Treebank, which so far contained valency annotation but has not linked synonymous senses of verbs together. The method used for extracting the synonym classes is a semi-automatic process with a substantial amount of manual work during filtering, role assignment to classes and individual Class members’ arguments, and linking to the external lexical resources. We present the first version with 200 classes (about 1800 verbs) and evaluate interannotator agreement using several metrics.
We report our work on building linguistic resources and data-driven parsers in the grammatical relation (GR) analysis for Mandarin Chinese. Chinese, as an analytic language, encodes grammatical information in a highly configurational rather than morphological way. Accordingly, it is possible and reasonable to represent almost all grammatical relations as bilexical dependencies. In this work, we propose to represent grammatical information using general directed dependency graphs. Both only-local and rich long-distance dependencies are explicitly represented. To create high-quality annotations, we take advantage of an existing TreeBank, namely, Chinese TreeBank (CTB), which is grounded on the Government and Binding theory. We define a set of linguistic rules to explore CTB’s implicit phrase structural information and build deep dependency graphs. The reliability of this linguistically motivated GR extraction procedure is highlighted by manual evaluation. Based on the converted corpus, data-driven, including graph- and transition-based, models are explored for Chinese GR parsing. For graph-based parsing, a new perspective, graph merging, is proposed for building flexible dependency graphs: constructing complex graphs via constructing simple subgraphs. Two key problems are discussed in this perspective: (1) how to decompose a complex graph into simple subgraphs, and (2) how to combine subgraphs into a coherent complex graph. For transition-based parsing, we introduce a neural parser based on a list-based transition system. We also discuss several other key problems, including dynamic oracle and beam search for neural transition-based parsing. Evaluation gauges how successful GR parsing for Chinese can be by applying data-driven models. The empirical analysis suggests several directions for future study.
Despite the well-established benefits of regular participation in physical activity, many Australians still fail to maintain sufficient levels. More self-determined types of motivation and more positive affect during activity have been found to be associated with the maintenance of physical activity behaviour over time. Need-supportive approaches to physical activity behaviour change have previously been shown to improve quality of motivation and psychological well-being. This paper outlines the development of a need-supportive, person-centred physical activity program for frontline aged-care workers. The program emphasises the use of self-determined methods of regulating activity intensity (affect, rating of perceived exertion and self-pacing) and is aimed at increasing physical activity behaviour and psychological well-being. The development process was undertaken in six steps using guidance from the Intervention Mapping framework: (i) an in-depth needs assessment (including qualitative interviews where information was gathered from members of the target population); (ii) formation of change objectives; (iii) selecting theory-informed and evidence-based intervention methods and planning their practical application; (iv) producing program components and materials; (v) planning program adoption and implementation, and (vi) planning for evaluation. The program is based in Self-Determination Theory (SDT) and provides tools and elements to support autonomy (the use of a collaboratively developed activity plan and participant choice in activity types), competence (action/coping planning, goal-setting and pedometers), and relatedness (the use of a motivational interviewing-inspired appointment and ongoing support in activity).
We perform a fine-grained large-scale analysis of coreference projection. By projecting gold coreference from Czech to English and vice versa on Prague Czech-English Dependency Treebank 2.0 Coref, we set an upper bound of a proposed projection approach for these two languages. We undertake a detailed thorough analysis that combines the analysis of projection's subtasks with analysis of performance on individual mention types. The findings are accompanied with examples from the corpus.
We aimed to examine the effects of contextual factors (ie, observers' training background and priming texts) on decoding facial pain expressions of younger and older adults. A total of 165 participants (82 nursing students and 83 nonhealth professionals) were randomly assigned to one of 3 priming conditions: (1) information about the possibility of secondary gain (misuse); (2) information about the frequency and undertreatment of pain in the older adult (undertreatment); or (3) neutral information (control). Subsequently, participants viewed 8 videos of older adults and 8 videos of younger adults undergoing a discomforting physical therapy examination. Participants rated their perception of each patient's pain intensity, unpleasantness, and condition severity. They also rated their willingness to help, sympathy level, patient deservingness of financial compensation, and how negatively/positively they feel towards the patient (ie, valence). Results demonstrated that observers ascribed greater levels of pain and other indicators (eg, sympathy and help) to older compared with younger patients. An interaction between observer type and patient age demonstrated that nursing students endorsed higher ratings of younger adults' pain compared with other students. In addition, observers in the undertreatment priming condition reported more positive valence towards older patients. By contrast, priming observers with the misuse text attenuated their valence ratings towards younger patients. Finally, the undertreatment prime influenced observers' pain estimates indirectly through observers' valence towards patients. In summary, results add specificity to the theoretical formulations of pain by demonstrating the influence of patient and observer characteristics, as well as informational primes, on decoding pain expressions.