Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The article presents the results of word-formative and semantic analysis of Middle Czech verbs consisting of the prefix roz(e)- contained in Lexical database of humanistic and baroque Czech (https://madla.ujc.cas.cz). The analysis partly confirms, partly corrects the results of earlier analyses, above all, carried out by D. Šlosar (1981).
Estimating the entropy based on data is one of the prototypical problems in distribution property testing and estimation. For estimating the Shannon entropy of a distribution on $S$ elements with independent samples, [Paninski2004] showed that the sample complexity is sublinear in $S$, and [Valiant--Valiant2011] showed that consistent estimation of Shannon entropy is possible if and only if the sample size $n$ far exceeds $\frac{S}{\log S}$. In this paper we consider the problem of estimating the entropy rate of a stationary reversible Markov chain with $S$ states from a sample path of $n$ observations. We show that: (1) As long as the Markov chain mixes not too slowly, i.e., the relaxation time is at most $O(\frac{S}{\ln^3 S})$, consistent estimation is achievable when $n \gg \frac{S^2}{\log S}$. (2) As long as the Markov chain has some slight dependency, i.e., the relaxation time is at least $1+Ω(\frac{\ln^2 S}{\sqrt{S}})$, consistent estimation is impossible when $n \lesssim \frac{S^2}{\log S}$. Under both assumptions, the optimal estimation accuracy is shown to be $Θ(\frac{S^2}{n \log S})$. In comparison, the empirical entropy rate requires at least $Ω(S^2)$ samples to be consistent, even when the Markov chain is memoryless. In addition to synthetic experiments, we also apply the estimators that achieve the optimal sample complexity to estimate the entropy rate of the English language in the Penn Treebank and the Google One Billion Words corpora, which provides a natural benchmark for language modeling and relates it directly to the widely used perplexity measure.
Artikkeli käsittelee suomentamiseen liittyviä ideologioita ja normeja 1800-luvun tietokirjallisuudessa. Tapaustutkimuksena on Werner Söderström Osakeyhtiön tietokirjojen suomennostoiminta 1800-luvun lopulla. Tutkimus kytkeytyy kääntämisen sosiologiaan ja historiaan, ja siinä arvioidaan myös, miten ja missä määrin historiallisia käännösprosesseja voidaan rekonstruoida. Käännösprosesseja lähestytään tarkastelemalla eri toimijoiden − kustantaja, kääntäjä, kieliasiantuntija, tekstin arvioija − osuutta käännösprosessissa. Tutkimuksen aineistona on kustantajan ja kääntäjän kääntämistä ja kielellisiä valintoja käsittelevä kirjeenvaihto, jonka avulla on mahdollista valottaa eri suunnista kääntäjän arkea, yhteisöllisiä arvoja ja normeja käännösvalintojen taustalla sekä niitä henkilökohtaisia asenteita, jotka ohjaavat kääntäjiä erilaisiin valintoihin.
 Analyysin tuloksena voi päätellä, että ammattikirjoittajina kääntäjät olivat hyvin tietoisia erilaisista kielellisistä ja kääntämiseen liittyvistä normeista. Käytännön työssä kääntäjät toimivat kuitenkin usein erilaisten normien ristipaineessa, jolloin vastakkain asettuivat esimerkiksi alkuteoksen luonteen säilyttäminen ja toisaalta sen kotouttaminen. Kääntäjät olivat myös tietoisia kielen vaihtelevista normeista, tunsivat käynnissä olevat kielikeskustelut ja mukauttivat herkästi kielenkäyttöään kulloinkin vallitsevien kirjakielen normien mukaiseksi.
 
 Norms and ideologies of translation in light of correspondence between publisher and translator in 19th-century Finland
 This article analyses the ideologies and norms that guided the translation of works of non-fiction in 19th-century Finland. As a case study the article analyses the processes involved in the publication of non-fiction at the Werner Söderström Ltd publishing house at the end of the 19th century. The research takes as its base theories examining the sociology and history of translation. It also aims to evaluate how and to what extent historical translation processes can be reconstructed. Translation is approached as a collaborative process involving various actors: publisher, translator, language editor, and expert reader. The data consists of correspondence between publisher and translator that deals with matters of translation or language. This correspondence sheds light on the everyday life of the translator and the socially accepted norms and ideologies that guide the translation process. It also reveals the stance of publishers concerning the choice of translator, a factor that can lead to very different end products.
 The analysis shows that, as professional writers, translators at the end of the 19th century were well aware of contemporary translational norms. In practice, translators were caught between various conflicting pressures – regarding, for instance, questions such as whether one should follow the original text as close as possible to preserve its unique style or assimilate the text to a Finnish context to help the reader. The data also shows that translators were well aware of linguistic norms; they were acquainted with current and past debates, and in assimilating their use of language they remained sensitive to prevailing norms.
Requirement is a formal expression of user’s need. It is the main foundation of any software development project. Natural language (NL) is often used to express and write system requirements specifications as well as user requirements. However, there is a very high probability that more than half natural language requirements can be ambiguous, incomplete and inaccurate. A software engineer can miss-interpret the natural language requirements and can generate an erroneous software model, which finally will lead to project failure. Earlier, we have introduced a prototype tool that provides natural language requirements authoring facilities and consistency checking to assist requirement engineers when working with informal and semi-formal requirements. However, the tool has pattern limitation to support the extraction of the essential requirements from the NL requirements. Therefore this study is aimed to enhance the accuracy and scalability of the tool to capture the essential requirements from the NL requirements. Our approach is to implement lexical analysis and embed an English lexical database where it will serve as a thesaurus in the tool. This tool is expected to be able to find the synonym of the extracted phrases (essential requirements) in the database to match it to the essential interaction pattern (phrases and expressions) in the library. Our future work will focus on the next phase of requirements engineering, which is requirements validation.
In this work we describe the system built for the three English subtasks of\nthe SemEval 2016 Task 3 by the Department of Computer Science of the University\nof Houston (UH) and the Pattern Recognition and Human Language Technology\n(PRHLT) research center - Universitat Polit`ecnica de Val`encia: UH-PRHLT. Our\nsystem represents instances by using both lexical and semantic-based similarity\nmeasures between text pairs. Our semantic features include the use of\ndistributed representations of words, knowledge graphs generated with the\nBabelNet multilingual semantic network, and the FrameNet lexical database.\nExperimental results outperform the random and Google search engine baselines\nin the three English subtasks. Our approach obtained the highest results of\nsubtask B compared to the other task participants.\n
There exist distinctive words that are used to express same semantics and as a result of this it has become hard to quantify the exact matching of words. To deal with this issue, past investigations endeavored to ascertain a likeness between distinctive pair of words. Conventional methodologies for computing word similarity are based on repositories like WordNet. It is a manually created lexical database and it processes semantic connection between various words. However, WordNet is a universally useful asset but wide range of words are not present in it and furthermore there exist an issue of identifying the meaning of words. Implication of words are diverse in WordNet when we utilize it in a textual framework. There exists a need of the refined approach that can gauge words resemblance in light of their co-occurrence. In this examination, we proposed an approach that registers likeness in text particular words, with the assistance of literary substance of various posts on StackOverflow. Our proposed strategy figures out word similarities in text by ascertaining the weighted co-occurrence in view of Computing Term Cooccurrence (CTC) and SentiWordNet. The exploratory outcome demonstrates that our system proposed an arrangement of words that are identified with text data is exceptional. Moreover, when it was compared with WordNet-based strategy named as WordNetres, it results with better outcomes.
Introduction. Nowadays the language of modern television broadcasts’ speechesis more and more in the focus of linguistic study. Special interest of given paper is comprisedby the variability in gender categorization of nouns as it occurs in speeches of Ukrainian TVprograms anchorpersons. The object of the paper is the choice of nouns in modern TV speech,that is distinguished due to the variability of grammatical category of the gender of nouns waysof realization.Purpose of the article is to analyze the nouns that have suffered changes in thegrammatical category of gender within the current language trends being implemented inthe broadcast of Ukrainian television. In addition, the aim is to outline some of the reasonsfor the emergence of such tendencies, the relative use of these nouns, the degree of codificationin modern lexicographic sources.Methods of research. The research is grounded on descriptive method, the method ofempiric analysis, immediate constituents’ analysis, and contextual analysis.Results. Studying the language used in contemporary Ukrainian TV speeches convincinglydemonstrated how formal-grammatical indicators of the category of gender of nouns in moderntelecommunication vary (shift), implemented in the following modifications: male genius femalegenus, female genus male genus. In addition, the review of codification in the dictionariesof the analyzed noun units shows that variational changes in the morphology of the noun andin particular, in its morphological and grammatical categories nowadays cause not onlythe dislocation of the current linguistic norm, but also tend to change the morphological norms. Conclusion. Analyzing the language used in speeches of informative and entertainingUkrainian TV programs’ anchorpersons, we have concluded that they prefer to choose differentgrammatical variations in favor of a specific counterpart to the grammatical category of the gender,often a revitalized or dialectal word used.
This paper explains transition dependency parsing approaches to build a dependency parser for Telugu language. Telugu treebank is given as an input to transition dependency parsers. One of the best transition dependency parser is the Malt parser. It is an independent system and it has nine methods to parse a sentence of any language. We have applied the treebank on all the methods of a Malt parser among which Arc-eager parser produces state-of-art results for Telugu language. Arc-eager method was produced LA (Label Accuracy) of 63%, UAS (Unlabeled Attachment Score) of 88.1% and LAS (Label Attachment Score) of 62.3%. In this paper we discuss a brief introduction of all Malt Parsing methods and an in detail explanation of Arc-eager dependency parsing.
The paper tackles the question of what the dynamics of wordplay mean for Early Modern language philosophy and what function wordplay fulfills at a time when linguistic norms and cultural values of a particular language are being sought. In Part 1, the current definition of wordplay suggested in In Part 2, we give a brief sketch of the main features of Early Modern linguistic thought with a particular focus on the concepts of play and wordplay. As one of the language theorists of 17th century Germany, Georg Philipp Harsdrffer (1607-1658) is widely known for the sophisticated integration of these concepts into his "linguistic" oeuvre, and this will determine the main focus of the current article. Two of Harsdrffer's works will be the center of attention: the Frauenzimmer Gesprchspiele (FZG), published 1643-1649 in Nuremberg, an eight-volume series of dialogues about social, poetic and scientific matters, which incorporates much of Harsdrffer's thoughts on language and one of the best-sellers of the 17th century, and the Delitiae Mathematicae et Physicae (DMP), a three-volume scientific work, to which Harsdrffer added the last two of the three volumes (1651( -1653, Nuremberg), Nuremberg). Based on the study of various subtypes of wordplay with letters in Part 3, we shall argue that in the context of baroque linguistic ideas wordplay should be defined in a broader sense. It is deeply rooted in a particular view of language peculiar to European baroque culture that provided a conceptual background not only for language "theories", poetry, education and standards of knowledge but also for the role and functions of wordplay. As Harsdrffer found his inspiration in and was strongly influenced by similar ideas of other scientists, particularly in Italy and France, the results of the analysis of the German baroque sources allow for more general assumptions that are not restricted to one language only.
Temporal models based on recurrent neural networks have proven to be quite powerful in a wide variety of applications, including language modeling and speech processing. However, to train these models, one relies on back-propagation through time, which entails unfolding the network over many time steps, making the process of conducting credit assignment considerably more challenging. Furthermore, the nature of back-propagation itself does not permit the use of non-differentiable activation functions and is inherently sequential, making parallelization of the underlying training process very difficult. In this work, we propose the Parallel Temporal Neural Coding Network, a biologically inspired model trained by the local learning algorithm known as Local Representation Alignment, that aims to resolve the difficulties and problems that plague recurrent networks trained by back-propagation through time. Most notably, this architecture requires neither unrolling nor the derivatives of its internal activation functions. We compare our model and learning procedure to other online back-propagation-through-time alternatives (which also tend to be computationally expensive), including real-time recurrent learning, echo state networks, and unbiased online recurrent optimization, and show that it outperforms them on sequence modeling benchmarks such as Bouncing MNIST, a new benchmark we call Bouncing NotMNIST, and Penn Treebank. Notably, our approach can, in some instances, even outperform full back-propagation through time itself as well as variants such as sparse attentive back-tracking. Furthermore, we present promising experimental results that demonstrate our model's ability to conduct zero-shot adaptation.
Introduction:Iliac artery endofibrosis (IAE) is an uncommon disease, poorly studied pathology with devastating effects and different therapeutic approaches affecting young people who practise intensive sports, especially cyclists. The evolution of the process not only depends on the diagnosis and therapeutic action, but also on the acceptance and attitude of the patient and subsequent professional guidance. Case description:This is the case description of a professional triathlon athlete that had one previous iliac surgical revascularization for an IAE Iliac and was admitted in our department five times with subacute lower limb ischemia affecting both legs between 2013 and 2016. Clinical findings and image tests are reported, as well as medical procedures performed. Indications based on clinical, functional and imaging ratings were clear, but his professional activity was not completely abandoned. Finally, after four endovascular procedures with good immediate results, he was warned of the seriousness of the process since the etiopathogenic reason. At the present moment patient is asymptomatic, under routine controls, working as successful triathlon coach. Discussion and conclusion:The fact that an external mechanical stress is the reason of repeated iliac artery injury suggests that an open surgical approach correcting the external muscular compression or arterial deformation should be a definitive but also aggressive solution according to literature. However, endovascular procedures and new endovascular devices are an increasingly promising option with a very low surgical risk. No matter the revascularization performed, the persistence of sports intensive practice carries a high risk of recurrence. Sport practise cessation is mandatory in some cases in order to assure revascularization long-term patency, but also a well conducted professional orientation is needed to complete the therapeutic action.
This paper is an analysis of the Hill‟s Strategy Development Framework and application of the framework to a Fast Food Restaurant business which is operating in Brunei Darussalam.The study of this article concentrates on corporate objectives, marketing strategy, order qualifiers, order winners, and the operations strategy within the company.A review of the relevant literature conducted on corporate goals, competitive priorities, Order Qualifier and Order Winner. The methodology based on a desk review of secondary data and non-participant observations research approach.The findings demonstrated that Fast Food Restaurant Business in Brunei Darussalam required more considerable attention to focus on improving the Customer Service Relationship (CSR) value.It is crucial to concentrate on CSR that would enhance the business brand value, a better image rating and thereby contributing to company‟s sales and gaining a new customer.The study recommended that Fast Food Restaurant Business create a new market such as selling the frozen product.Fast Food Restaurant Business potentially can sell this in their store or export their patent right on the product internationally.The contribution of this paper is to provide provides a review of the application of the strategic framework to uncover the potential customer benefits package of the strategy.
This thesis aims to examine metapragmatic discourses on linguistic politeness illustrated in Korean language how-to literature. The primary task lies in contextualizing the native awareness of ene yeycel (linguistic politeness in Korean) within the interests or values of certain social groups. The first group, South Korean government-sanctioned agencies, led a linguistic campaign promoting a new standard speech model in 1992. Language professionals, the second group of social actors, produced popular language how-to literature, especially after the establishment of the hegemonic standard speech model. Both language standardizing policy and the participants in the how-to industry represent the cultural process of constructing language and social conventions. The “normative” culture of ene yeycel can be empowered and widely circulated, gaining wider social practice. Standardization of honorification came to the surface as a public issue along with a new “cultural policy” of the Ministry of Cultural Affairs in 1990. In this cultural-political circumstance, the social meaning of standardized honorification was rediscovered as indigenous culture, a group identity shared by Korean speakers. Positively valorizing honorification as linguistic and cultural tradition, the standardized model preserves the sophisticated use of honorifics and reinforces superior-inferior relationships. However, the standard model of ene yeycel can be subjective and arbitrary. Moreover, different styles are too easily proscribed as errors made by sloppy speakers. Language how-to literature produces more diversified interpretations than the standard speech manual. As language users are confronted with the challenges of finding the proper level of honorification, language how-to manuals provide justifications to help speakers prioritize linguistic norms when internalizing social relationships. Positive valorizations of honorification derive from a speaker's respect for the interlocutor's social status or personality. Negative valorizations of honorification view deferential politeness as a kind of discriminatory behaviour indexing power-difference. The positive or negative values of honorification are based on different concepts of ene yeycel and on different identifications of social relationships. Such conceptualizations rationalize whether speakers should support honorification or not, and lead them to discuss language use in current society.
Temporal models based on recurrent neural networks have proven to be quite\npowerful in a wide variety of applications. However, training these models\noften relies on back-propagation through time, which entails unfolding the\nnetwork over many time steps, making the process of conducting credit\nassignment considerably more challenging. Furthermore, the nature of\nback-propagation itself does not permit the use of non-differentiable\nactivation functions and is inherently sequential, making parallelization of\nthe underlying training process difficult. Here, we propose the Parallel\nTemporal Neural Coding Network (P-TNCN), a biologically inspired model trained\nby the learning algorithm we call Local Representation Alignment. It aims to\nresolve the difficulties and problems that plague recurrent networks trained by\nback-propagation through time. The architecture requires neither unrolling in\ntime nor the derivatives of its internal activation functions. We compare our\nmodel and learning procedure to other back-propagation through time\nalternatives (which also tend to be computationally expensive), including\nreal-time recurrent learning, echo state networks, and unbiased online\nrecurrent optimization. We show that it outperforms these on sequence modeling\nbenchmarks such as Bouncing MNIST, a new benchmark we denote as Bouncing\nNotMNIST, and Penn Treebank. Notably, our approach can in some instances\noutperform full back-propagation through time as well as variants such as\nsparse attentive back-tracking. Significantly, the hidden unit correction phase\nof P-TNCN allows it to adapt to new datasets even if its synaptic weights are\nheld fixed (zero-shot adaptation) and facilitates retention of prior generative\nknowledge when faced with a task sequence. We present results that show the\nP-TNCN's ability to conduct zero-shot adaptation and online continual sequence\nmodeling.\n
Prestige and dominance are thought to be two evolutionarily distinct routes to gaining status and influence in human social hierarchies. Prestige is attained by having specialist knowledge or skills that others wish to learn, whereas dominant individuals use threat or fear to gain influence over others. Previous studies with groups of unacquainted students have found prestige and dominance to be two independent avenues of gaining influence within groups. We tested whether this result extends to naturally-occurring social groups. We ran an experiment with 30 groups of 5 people from Cornwall, UK (n=150). Participants answered general knowledge questions individually and as a group, and subsequently nominated a team representative to answer bonus questions to win money on behalf of the team. Participants then rated all other team-mates anonymously on scales of prestige, dominance, likeability and influence on the task. Using a model comparison approach with Bayesian multi-level models, we found that prestige and dominance ratings were predicted by influence ratings on the task, replicating previous studies. However, prestige and dominance ratings did not predict who was nominated as group representative. Instead, participants nominated team members with the highest individual quiz scores, despite this information being unavailable to them. Interestingly, team members who were initially rated as being high status in the group, such as a team captain or group administrator, had higher ratings of both dominance and prestige than other group members. In contrast, those who were initially rated as someone from whom group members would like to learn had higher prestige ratings, but not higher dominance ratings, supporting the claim that prestige reflects social learning opportunities. Our results suggest that prestige and dominance hierarchies do become established in naturally occurring human social groups, but that these hierarchies may be more domain-specific and less flexible than we anticipated.
Annotation Guidelines for Text Analytics in Social Media A person's language use reveals much about their profile, however, research in author profiling has always been constrained by the limited availability of training data, since collecting textual data with the appropriate meta-data requires a large collection and annotation effort (Maamouri et al. 2010; Diab et al. 2008; Hawwari et al. 2013).For every text, the characteristics of the author have to be known in order to successfully profile the author. Moreover, when the text is written in a dialectal variety such as the Arabic text found online in social media a representative dataset need to be available for each dialectal variety (Zaghouani et al. 2012; Zaghouani et al. 2016).The existing Arabic dialects are historically related to the classical Arabic and they co-exist with the Modern Standard Arabic in a diglossic relation. While the standard Arabic, has a clearly defined set of orthographic standards, the various Arabic dialects have no official orthographies and a given word could be written in multiple ways in different Arabic dialects (Maamouri et al. 2012; Jeblee et al. 2014).This abstract presents the guidelines and annotation work carried out within the framework of the Arabic Author profiling project (ARAP), a project that aims at developing author profiling resources and tools for a set of 12 regional Arabic dialects. We harvested our data from social media which reflect a natural and spontaneous writing style in dialectal Arabic from users in different regions of the Arabworld.For the Arabic language and its dialectal varieties as foundin social media, to the best of our knowledge, there is nocorpus available for the detection of age, gender, nativelanguage and dialectal variety. Most of the existingresources are available for English or other Europeanlanguages. Having a large amount of annotated data remains the key to reliable results in the taskof author profiling. In order to start the annotation process, we createdguidelines for the annotation of the Tweets according totheir dialectal variety, their native language, the gender of the user and the age. Before starting theannotation process, we hired and trained a group of annotators and we implemented a smooth annotation pipeline to optimize the annotation task. Finally, we followed a consistent annotation evaluation protocol to ensure a high inter-annotator agreement.The Annotations were done by carefully analyzing each ofthe user's profiles, their tweets, and when possible, weinstructed the annotators to use external resources such asLinkedIn or Facebook. We created a general profilesvalidation guidelines and task-specific guidelines toannotate the users according to their gender, age, dialectand their native language. For some accounts, the annotators were not able to identifythe gender as this was based in most of the cases on thename of the person or his profile photo and in some casesby their biography or profile description. In case thisinformation is not available, we instructed the annotators toread the user posts and find linguistic indicators of thegender of the user.Like many other languages, Arabic conjugates verbsthrough numerous prefixes and suffixes and the gender issometimes clearly marked such as in the case of the verbsending in taa marbuTa which is usually of femininegender.In order to annotate the users for their age, we used threecategories: under 20 years, between 20 years and 40 years,and 40 years and up.In our guidelines, we asked our annotators to try their bestto annotate the exact age, for example, they can check theeducation history of the users in LinkedIn and Facebookprofile and find when the graduated from high school forexample in order to guess the age of the users. As the dialect and the regions are known in advance to theannotators, we instructed them to double check and markthe cases when the profile appears to be from a differentdialect group. This is possible despite our initial filteringbased on distinctive regional keywords. We noticed that inmore than 90% the profiles selected belong to the specifieddialect group. Moreover, we asked the annotators to mark and identifyTwitter profiles with a native language other than Arabic,so they are considered as Arabic L2 speakers. In order tohelp the annotators identify those, we instructed them tolook for various cues such as the writing style, the sentence structure, the word order and the spelling errors.AcknowledgementsThis publication was made possible by NPRP grant #9-175-1-033 from the Qatar National Research Fund (a member ofQatar Foundation). The statements made herein are solelythe responsibility of the authors. ReferencesDiab Mona, Aous Mansouri, Martha Palmer, Olga Babko-Malaya, Wajdi Zaghouani, Ann Bies, Mohammed Maamouri. A Pilot Arabic Propbank; LREC 2008, Marrakech, Morocco, May 28-30, 2008.Hawwari, A.; Zaghouani, W.; O»Gorman, T.; Badran, A.; Diab, M., «Building a Lexical Semantic Resource for Arabic Morphological Patterns,» Communications, Signal Processing, and their Applications (ICCSPA), 2013, vol., no., pp.1,6, 12-14 Feb. 2013. Jeblee Serena; Houda Bouamor; Wajdi Zaghouani; Kemal Oflazer. CMUQ@QALB-2014: An SMT-based System for Automatic Arabic Error Correction. In Proceedings of the EMNLP 2014 Workshop on Arabic Natural Language Processing (ANLP), Doha, Qatar, October 2014.Maamouri Mohamed, Ann Bies, Seth Kulick, Wajdi Zaghouani, Dave Graff and Mike Ciul. 2010. From Speech to Trees: Applying Treebank Annotation to Arabic Broadcast News. In Proceedings of LREC 2010, Valetta, Malta, May 17-23, 2010.Maamouri Mohammed, Wajdi Zaghouani, Violetta Cavalli-Sforza, Dave Graff and Mike Ciul. 2012. Developing ARET: An NLP-based Educational Tool Set for Arabic Reading Enhancement. In Proceedings of The 7th Workshop on Innovative Use of NLP for Building Educational Applications, NAACL-HLT 2012, Montreal, Canada.Obeid Ossama, Wajdi Zaghouani, Behrang Mohit, Nizar Habash, Kemal Oflazer and Nadi Tomeh. A Web-based Annotation Framework For Large- Scale Text Correction. In Proceedings of IJCNLP'2013, Nagoya, Japan.Zaghouani Wajdi, Nizar Habash, Ossama Obeid, Behrang Mohit, Houda Bouamor, Kemal Oflazer. 2016. Building an arabic machine translation post-edited corpus: Guidelines and annotation. In Proceedings of the International Conference on Language Resources and Evaluation (LREC»2016).Zaghouani Wajdi, Abdelati Hawwari and Mona Diab. 2012. A Pilot PropBank Annotation for Quranic Arabic. In Proceedings of the first workshop on Computational Linguistics for Literature, NAACL-HLT 2012, Montreal, Canada.
Currently, the biaffine classifier has been attracting attention as a method to introduce an attention mechanism into the modeling of binary relations. For instance, in the field of dependency parsing, the Deep Biaffine Parser by Dozat and Manning has achieved state-of-the-art performance as a graph-based dependency parser on the English Penn Treebank and CoNLL 2017 shared task. On the other hand, it is reported that parameter redundancy in the weight matrix in biaffine classifiers, which has O(n^2) parameters, results in overfitting (n is the number of dimensions). In this paper, we attempted to reduce the parameter redundancy by assuming either symmetry or circularity of weight matrices. In our experiments on the CoNLL 2017 shared task dataset, our model achieved better or comparable accuracy on most of the treebanks with more than 16% parameter reduction.
Emotional imagery is a common induction technique used in the laboratory and also employed in various exposure therapy treatments across the anxiety spectrum (e.g., specific and social phobias). Despite its clinical uses, there is a surprising dearth of literature regarding the basic central neural processes underlying emotional imagery, though other peripheral physiological processes have been investigated extensively using heart rate, skin conductance, and startle-blink responses. One imagery study that used a central nervous system psychophysiological measure -event specific brainwave or the event-related potential (ERP) technique- suggests the late positive potential (LPP) of the ERP is larger for unpleasant versus neutral stimuli, implying this ERP may index emotional engagement during imagery. This effect is consistent with the visual perception literature of emotion; however, the visual perception literature also indicates that the LPP is larger for pleasant stimuli versus neutral stimuli and positively correlated with subjective emotional arousal ratings. Using script-driven emotional imagery, we will extend research on the LPP to establish whether 1) the LPP is larger for both pleasant and unpleasant scripts relative to neutral ones and 2) this LPP effect is positively correlated with emotional arousal ratings. Fifty-five participants will make subjective ratings of the scripts and then imagine the scripts while electroencephalographic data are, recorded. Upon demonstrating the LPP is larger for emotional (both pleasant and unpleasant) scripts than neutral ones, this study will lay the foundation for future work aimed at determining whether LPP effects are hyper- or hypo-active for socially anxious participants.
To reinforce sports informatization management and the sports service quality on college campuses, this paper researches combined multi-agent technology, and structures undergraduate sports service system framework. It has elaborated functions of every feature and workflow of the system, and puts forward agent design procedure based on JADE. What is more, it also adopts FIPAACl linguistic norms between Agents communication and proposes some advices which based on the fundamental of Agent.
The authors try to answer two questions: 1. How the Polish philology students understand the concept of linguistic norm? and 2. When, according to those students, people should follow it? The article presents the results of the survey conducted among 200 respondents. It turns out that the students understand the concept of norm well, usually as a set of rulles, which are established by linguists or/and accepted by society. They also think that the respect of rules takes effect in varying degrees in different communication situations.
The short note describes the chart parser for multimodal type-logical grammars which has been developed in conjunction with the type-logical treebank for French. The chart parser presents an incomplete but fast implementation of proof search for multimodal type-logical grammars using the "deductive parsing" framework. Proofs found can be transformed to natural deduction proofs.
Polycentric Spanish Norm Towards the Polish‑Spanish Legal Translation The Spanish, being the official language of Spain and many other countries, is characterized by an important dialectal diversity that is reflected in the differences at all linguistic levels: phonetic, morphological, syntactic and lexico‑semantic, etc. All these differences raise controversies and discussions about the existence of a linguistic norm depending on the perspective that can have a monocentric or polycentric character. In this contribution we present some arguments for the second one. To this end, we rely on translations, starting simultaneously from the semasiological and onomasiological perspective, of some Polish‑Spanish legal terms in which it is essential to take into account, the diatopic variation as well as the norm whose character is polycentric.
Because the most common transition systems are projective, training a transition-based dependency parser often implies to either ignore or rewrite the non-projective training examples, which has anadverse impact on accuracy. In this work, we propose a simple modification of dynamic oracles, which enables the use of non-projective data when training projective parsers. Evaluation on 73~treebanks shows that our method achieves significant gains (+2 to +7 UAS for the most non-projective languages) and consistently outperforms traditional projectivization and pseudo-projectivizationapproaches.
The Universal Dependencies project is currently comprised of 71 languages and 122 treebanks, and aims to find morphological and syntactic characteristics that can be applied to multiple languages for parallel language processing. In this paper, we introduce Universal POS, which is a morphological tagset for UD, and propose a method to automatically convert existing Korean morphological tagset into UPOS. In order to apply the UPOS tagset, which is based on refraction words such as English, to the Korean language, it is necessary to try a one-to-many mapping between the UPOS individual tag and the 21st century Sejong tag combination. (Yonsei University)
This study examines how the acoustic input (the surface form) and the abstract linguistic representation (the underlying representation) interact during spoken word recognition by investigating left-dominant tone sandhi, a tonal alternation in which the underlying tone of the first syllable spreads to the sandhi domain. We conducted an auditory-auditory priming lexical decision experiment on Shanghai left-dominant sandhi words, in which each disyllabic target ([tɕi55 dɛ31] “egg”) was preceded by monosyllabic primes either sharing the same underlying tone ([tɕi55]), surface tone ([tɕi53] “machine”), or being unrelated to the tone of the first syllable of the sandhi targets ([tɕi24] “to remember”). Results showed a surface priming effect, but not an underlying priming effect. Moreover, the surface priming did not interact with speakers’ familiarity ratings to the sandhi targets. The results are discussed in the context of how phonological opacity, productivity, and the directionality of tone sandhi patterns influence the representation of tone sandhi words as well as how the lexicality of the primes and the participants’ usage pattern of Shanghai may have influenced the results.
The repertoire of forms of address can be considered as one of the determinants of the discourse genre, which makes it possible to capture its evolution and cultural variations. From such comparative, intra- and intercultural perspective, adopting an interactive approach in the analysis of political discourse, we will look at the practice of addressing one another in the French and Polish politicalmedia discourse. While in both languages the linguistic norm recommends the use of the polite forms of address in official situations, the cases of the use of the familiar pronoun tu / ty in media interactions between politicians are not rare at all. Whether it is an informal talk of politicians caught by the media, a television pre-election debate, or a meeting of the heads of state, addressing the other person by the familiar forms is a manifestation of a deliberate blurring of the boundaries between the front-stage and backstage in political discourse in order to create the impression of intimacy andequality between the interlocutors.
The first edition of one of the most important and mysterious novels of the 20th century appeared more than fifty years ago. Despite the passage of time The Master and Margarita still enjoys popularity; it also intrigues and inspires. Until now five Polish translations of Bulgakov’s novel have appeared. It is known that the interpretation of the original might be expressed in the form of many potential texts that are communicatively equivalent. There is no doubt that it is the translator who plays a vital role in any translation; her/his personality, life experience, knowledge, skills, and also the times s/he lives in regulate the target text. That is why, no matter how many times a text is translated, the final product will always be different. Taking this into consideration, the author will compare the three Polish translations of Bulgakov’s Master and Margarita, paying attention to the diachronic perspective as far as linguistic norms are concerned, the modernity of language, and the way the anthroponyms are expressed.
This study aims at exploring new norms as to the textual additions in parentheses (=TAiPs) in the translation of a Quranic text as writer-oriented devices of textuality. Coding for this sort of information could be useful in establishing an impact on any decision-making process on the TL version; such TAiPs can give a translated text of the Quran unity and purpose and distinguish it from a disconnected sequence of sentences. Six small-sized chapters of the Quran were selected as a research sample including a number of four handred forty two (442) TAiPs. Two writer-oriented kinds of textuality were found: cohesivity at the levels of grammar and lexis to be in form of recurrence, reference, substitution, ellipsis and conjunction; and relationality by coherence and intentionality to be in form of reiteration, collocation, connotation, evocation and interpretation. The study is a detailed analysis of such a severely criticized yet officially approved English interpretation of the Quran as the Hilali and Khan Translation (=HKT) against a predetermined set of text-linguistic norms. The strength or weakness of TAiPs as to how they might alleviate or aggravate the TL version is eventually identified for sake of improvement.
Abstract People remember events and materials better when these are congruent with their mood at retrieval; this is known as the mood-congruent memory bias. This effect is largest when the materials are self-referential and this is known as the self-reference effect. We present two word rating studies, to create a list of self-referential valenced words that may be used as stimuli to investigate the influence of valence on cognitive processing in depressive ruminators. Words selected from the Affective Norms for English Words pool were rated by an unselected sample for self-referentiality (Study 1) and validated with ratings provided by depressive ruminators. As hypothesized, depressive ruminators rated negative words as more self-referential than an unselected sample. Using this list, valence differentiated performance between depressive ruminators and healthy controls in a working memory updating task. We thus created a list of self-referential valenced words matched on factors that influence word processing.
The focus of the current study was on idiom comprehension in younger and older adults. Due to inconsistent results in previous studies, it is unclear whether older adults may have problems understanding idioms. For the current study, I used a sentence-to-word matching task presented on an iPad with software that recorded participants’ response time and accuracy. Participants also completed a familiarity task where they rated idioms on how frequently these phrases were encountered. I predicted that older adults would have more difficulty comprehending idioms because of the context in which the idioms were embedded and the timed nature of the task. I also predicted that both age groups would rate the idioms as highly familiar because we purposefully selected these types of expressions. With respect to the sentence-to-word matching task, results showed that although older adults were slower overall, both younger and older adults showed faster response times and greater accuracy for idiomatic targets following idiomatically-biased contexts than for literal targets following literally-biased contexts. With respect to the familiarity ratings task, results showed that both age groups were very familiar with the idioms. These findings suggest that older adults are able to successfully use context to understand familiar ambiguous idioms and that they do not have difficulty comprehending idioms in a cognitively demanding timed task.
The statistical parsing of morphologically rich languages is hindered by the inability of parsers to collect solid statistics because of the large number of word types in such languages. There are however two separate but connected problems, reducing data sparsity of known words and handling rare and unknown words. Methods for tackling one problem may inadvertently negatively impact methods to handle the other. We perform a tightly controlled set of experiments to reduce data sparsity through class-based representations in combination with unknown word signatures with two PCFG-LA parsers that handle rare and unknown words differently on the German TiGer treebank. We demonstrate that methods that have improved results for other languages do not transfer directly to German, and that we can obtain better results using a simplistic model rather than a more generalized model for rare and unknown word handling.
It is no secret that people often use taboo words when speaking about persons and objects in their environment. Taboo words are charged with emotion and have observable impact on the listener as well as the speaker. The purpose of this study was to determine whether taboo words were quantitatively more offensive when used in combination with a proper name versus being used with a non-human object. We found that using taboo words to describe proper names does not cause a significant effect; however, we found that participants rated certain categories of taboo words as more offensive than other categories. In a second experiment, taboo words did affect ratings and memory for proper names and non-human objects.
Shi, Huang, and Lee (2017) obtained state-of-the-art results for English and Chinese dependency parsing by combining dynamic-programming implementations of transition-based dependency parsers with a minimal set of bidirectional LSTM features. However, their results were limited to projective parsing. In this paper, we extend their approach to support non-projectivity by providing the first practical implementation of the MH_4 algorithm, an $O(n^4)$ mildly nonprojective dynamic-programming parser with very high coverage on non-projective treebanks. To make MH_4 compatible with minimal transition-based feature sets, we introduce a transition-based interpretation of it in which parser items are mapped to sequences of transitions. We thus obtain the first implementation of global decoding for non-projective transition-based parsing, and demonstrate empirically that it is more effective than its projective counterpart in parsing a number of highly non-projective languages
Syntactic parsing plays a crucial role in improving the quality of natural language processing tasks. Although there have been several research projects on syntactic parsing in Vietnamese, the parsing quality has been far inferior than those reported in major languages, such as English and Chinese. In this work, we evaluated representative constituency parsing models on a Vietnamese Treebank to look for the most suitable parsing method for Vietnamese. We then combined the advantages of automatic and manual analysis to investigate errors produced by the experimented parsers and find the reasons for them. Our analysis focused on three possible sources of parsing errors, namely limited training data, part-of-speech (POS) tagging errors, and ambiguous constructions. As a result, we found that the last two sources, which frequently appear in Vietnamese text, significantly attributed to the poor performance of Vietnamese parsing.
Detecting lexical entailment plays a fundamental role in a variety of natural language processing tasks and is key to language understanding. Unsupervised methods still play an important role due to the lack of coverage of lexical databases in some domains and languages. Most of the previous approaches were either based on statistical hypothesis of specific entailment relations or tried to encode word relations in low-dimensional vector embeddings. This thesis builds upon one of the few approaches which intrinsically model entailment in a vector space. We then further generalize this model by introducing an alternative, distributional representations for words which harnesses tools from optimal transport to define distance or entailment measures between such representations. We evaluated the models on hypernymy detection where our distributional estimate significantly improves over the underlying model and even outperforms state-of-the-art on some datasets.
In this paper we present the linguistic databases developed during our 8-year lexicographic research on the Modern Greek Standard (MGS) verbal system. Apart from the intermediate databases presented, the main products are (a) a new conjugation system of 385 paradigmatic models, which allows for the automatic generation of all verbal lexical morphemes and monolexical forms (b) a statistically established database of 151,536 distinctive verb-final grapheme sequences which allow for the automatic tagging of all monolexical verbal tokens without the traditional intervention of any built-in lexicon, and (c) a linear Iemmatisation morphophonological rule system accessed on the basis of the distinctive grapheme sequences identified.
Tree-structured neural network architectures for sentence encoding draw inspiration from the approach to semantic composition generally seen in formal linguistics, and have shown empirical improvements over comparable sequence models by doing so. Moreover, adding multiplicative interaction terms to the composition functions in these models can yield significant further improvements. However, existing compositional approaches that adopt such a powerful composition function scale poorly, with parameter counts exploding as model dimension or vocabulary size grows. We introduce the Lifted Matrix-Space model, which uses a global transformation to map vector word embeddings to matrices, which can then be composed via an operation based on matrix-matrix multiplication. Its composition function effectively transmits a larger number of activations across layers with relatively few model parameters. We evaluate our model on the Stanford NLI corpus, the Multi-Genre NLI corpus, and the Stanford Sentiment Treebank and find that it consistently outperforms TreeLSTM
This paper describes our system (SLT-Interactions) for the CoNLL 2018 shared task: Multilingual Parsing from Raw Text to Universal Dependencies. Our system performs three main tasks: word segmentation (only for few treebanks), POS tagging and parsing. While segmentation is learned separately, we use neural stacking for joint learning of POS tagging and parsing tasks. For all the tasks, we employ simple neural network architectures that rely on long short-term memory (LSTM) networks for learning task-dependent features. At the basis of our parser, we use an arc-standard algorithm with Swap action for general non-projective parsing. Additionally, we use neural stacking as a knowledge transfer mechanism for cross-domain parsing of low resource domains. Our system shows substantial gains against the UDPipe baseline, with an average improvement of 4.18% in LAS across all languages. Overall, we are placed at the 12 th position on the official test sets.
Abstract The proposal presented in this study seeks to properly represent natural language to ontologies and vice-versa. Therefore, the semi-automatic creation of a lexical database in Brazilian Portuguese containing morphological, syntactic, and semantic information that can be read by machines was proposed, allowing the link between structured and unstructured data and its integration into an information retrieval model to improve precision. The results obtained demonstrated that the methodology can be used in the risco financeiro (financial risk) domain in Portuguese for the construction of an ontology and the lexical-semantic database and the proposal of a semantic information retrieval model. In order to evaluate the performance of the proposed model, documents containing the main definitions of the financial risk domain were selected and indexed with and without semantic annotation. To enable the comparison between the approaches, two databases were created based on the texts with the semantic annotations to represent the semantic search. The first one represents the traditional search and the second contained the index built based on the texts with the semantic annotations to represent the semantic search. The evaluation of the proposal was based on recall and precision. The queries submitted to the model showed that the semantic search outperforms the traditional search and validates the methodology used. Although more complex, the procedure proposed can be used in all kinds of domains.
We live in a society where the large majority of the population has a camera-equipped smartphone. In addition, hard drives and cloud storage are getting cheaper and cheaper, leading to a tremendous growth in stored personal photos. Unlike photo collections captured by a digital camera, which typically are pre-processed by the user who organizes them into event-related folders, smartphone pictures are automatically stored in the cloud. As a consequence, photo collections captured by a smartphone are highly unstructured and because smartphones are ubiquitous, they present a larger variability compared to pictures captured by a digital camera. To solve the need of organizing large smartphone photo collections automatically, we propose here a new methodology for hierarchical photo organization into topics and topic-related categories. Our approach successfully estimates latent topics in the pictures by applying probabilistic Latent Semantic Analysis, and automatically assigns a name to each topic by relying on a lexical database. Topic-related categories are then estimated by using a set of topic-specific Convolutional Neuronal Networks. To validate our approach, we ensemble and make public a large dataset of more than 8,000 smartphone pictures from 10 persons. Experimental results demonstrate better user satisfaction with respect to state of the art solutions in terms of organization.
We live in a society where the large majority of the population has a camera-equipped smartphone. In addition, hard drives and cloud storage are getting cheaper and cheaper, leading to a tremendous growth in stored personal photos. Unlike photo collections captured by a digital camera, which typically are pre-processed by the user who organizes them into event-related folders, smartphone pictures are automatically stored in the cloud. As a consequence, photo collections captured by a smartphone are highly unstructured and because smartphones are ubiquitous, they present a larger variability compared to pictures captured by a digital camera. To solve the need of organizing large smartphone photo collections automatically, we propose here a new methodology for hierarchical photo organization into topics and topic-related categories. Our approach successfully estimates latent topics in the pictures by applying probabilistic Latent Semantic Analysis, and automatically assigns a name to each topic by relying on a lexical database. Topic-related categories are then estimated by using a set of topic-specific Convolutional Neuronal Networks. To validate our approach, we ensemble and make public a large dataset of more than 8,000 smartphone pictures from 10 persons. Experimental results demonstrate better user satisfaction with respect to state of the art solutions in terms of organization.
Recurrent Neural Networks (RNNs) play a major role in the field of sequential\nlearning, and have outperformed traditional algorithms on many benchmarks.\nTraining deep RNNs still remains a challenge, and most of the state-of-the-art\nmodels are structured with a transition depth of 2-4 layers. Recurrent Highway\nNetworks (RHNs) were introduced in order to tackle this issue. These have\nachieved state-of-the-art performance on a few benchmarks using a depth of 10\nlayers. However, the performance of this architecture suffers from a\nbottleneck, and ceases to improve when an attempt is made to add more layers.\nIn this work, we analyze the causes for this, and postulate that the main\nsource is the way that the information flows through time. We introduce a novel\nand simple variation for the RHN cell, called Highway State Gating (HSG), which\nallows adding more layers, while continuing to improve performance. By using a\ngating mechanism for the state, we allow the net to "choose" whether to pass\ninformation directly through time, or to gate it. This mechanism also allows\nthe gradient to back-propagate directly through time and, therefore, results in\na slightly faster convergence. We use the Penn Treebank (PTB) dataset as a\nplatform for empirical proof of concept. Empirical results show that the\nimprovement due to Highway State Gating is for all depths, and as the depth\nincreases, the improvement also increases.\n
OBJECTIVE: To evaluate variables of tobacco health warnings associated with their emotional impact, the perception of smoking risks and the perceived effectiveness to avoid tobacco use. MATERIALS AND METHODS: Teenagers (151) and adults (168) evaluated 27 tobacco health warnings selected from the sets used on tobacco packages in Argentina and in other countries. A standardized affective rating-scale system and a structured questionnaire measured respectively the emotional impact (hedonic valence and emotional arousal), and the cognitive-behavioral attributions. The correlation between emotional and cognitive-behavioral evaluations was analyzed by age, sex, education level, smoker status,stage of quitting and susceptibility of non-smokers teenagers. RESULTS: Strong significant correlations between cognitivebehavioral and emotional assessments were observed. The warnings depicting graphic images of tobacco-related injuries and suffering were considered more valuable for tobacco. control, helping quitting and preventing initiation. CONCLUSIONS: Using graphic images with high emotional arousal is recommended for both adults and teenagers.
This era, in which we currently stand, is an era of public opinion and mass information. People from all around the globe are joined together through various information junctions to create a global community, where one thing from the far east reaches to the people of the far west within seconds. Nothing is hidden, everything and anything can be scrutinized to its core and through these global criticisms and mass discussions of gigantic magnitude, we have reached to the pinnacle of correct decisions and better choices. These pseudo social groups and data junctions have bombarded our society so much that they now hold the forelock of our opinions and sentiments, ergo, we reach out to these groups to achieve a better outcome. But, all this enormous data and all these opinions cannot be researched by a single person, hence, comes the need of sentiment analysis. In this paper we’ll try to accomplish this by creating a system that will enable us to fetch tweets from twitter and use those tweets against a lexical database which will create a training set and then compare it with the pre-fetched tweets. Through this we will be able to assign a polarity to all the tweets by means of which we can address them as negative, positive or neutral and this is the very foundation of sentiment analysis, so subtle yet so magnificent.
The article deals with the variant terms in normative aspect codified in Ukrainian art lexicography of the 21st century. Dictionary codification of variant terms indicates changing in the language and deliberate influence of the society on the development of terminological norm. Variation is a existence form of objects of the surrounding reality, in particular, of scientific concepts, which defines the laws of their function and interaction. The choice of sources is due to the fact that the selected dictionaries are represented modern art knowledge. Dictionaries play a significant role in the normalization of language, the spread of linguistic norms, and therefore they are a grateful and relevant material for the analysis of variation in the Ukrainian art terminology. The article focuses on the importance of the scientific philological study of art terminology – the field of knowledge, which is rapidly developing in modern conditions, acquiring new meanings and forms. The variant terms of the art terminology, codified in Ukrainian special vocabulary, are analyzed. Three types of variant terms, phonemic, derivational and morphological-phonemic units, are fixed in the Ukrainian art terminology. It was found out that among the reasons for the occurrence of phonemic variant terms of the analyzed terminology tends to facilitate articulation of the learned term; the appearance of derivation of variant terms is conditioned by the presence of various derivative models in the Ukrainian language and the search for forms of terms that correspond most closely to modern productive models of term derivation; functioning of morphological-phonemic variant terms is explained by different degrees of grammatical adaptation of foreign-language art terms. It also traces the effect of an analogy inherent to all three of the varieties mentioned. In general, the article discusses the essence of the problem of terminological variation as one of the most relevant processes in the regulation and standardization of the Ukrainian art terminology.