Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This paper focuses on the process and principle of parsing, which is an essential task for machine to understand the syntactic, semantic structure of a sentence. First, a series of machine analysis procedures such as word segmentation, part-of-speech tagging and parsing of Chinese sentences are visually represented by using Cparser, a rule-based constituency parser developed by Peking University. Next, to better understand parsing mechanism, we explain in detail how the linguistic knowledge is embodied in the lexical, syntactic and semantic component of Cparser, showing their complex interplay that allows automatic parsing. As a practical example, a Chinese textbook treebank is also constructed using Cparser. According to the theoretical and practical discussion in this paper, Peking University Cparser, which is easy to reflect and modify linguistic knowledge, is expected to be widely used as an analysis and verification tool for Chinese grammar research.
This paper presents the first gold-standard resource for Russian annotated with compositionality information of noun compounds. The compound phrases are collected from the Universal Dependency treebanks according to part of speech patterns, such as ADJ+NOUN or NOUN+NOUN, using the gold-standard annotations. Each compound phrase is annotated by two experts and a moderator according to the following schema: the phrase can be either compositional, non-compositional, or ambiguous (i.e., depending on the context it can be interpreted both as compositional or noncompositional). We conduct an experimental evaluation of models and methods for predicting compositionality of noun compounds in unsupervised and supervised setups. We show that methods from previous work evaluated on the proposed Russian-language resource achieve the performance comparable with results on English corpora.
The article refers to the concept of intelligentsia as a social group which exerts significant influence on Polish standard patterns. Although the term intelligentsia is vague and questionable, it is well-established term in Polish linguistics, especially in sociolinguistics. Author argues that science communicators (young professional researchers, science journalists, PhD students) represent the young intelligentsia, because these well-educated people pursue their intellectual development and they have sense of public duty. The article examines standard of popular science texts in Internet, new tendencies in written Polish and attitude of young intelligentsia toward traditional linguistic norm. The errors (esp. punctuation and syntax) exemplify impact of technological changes and phenomenon of secondary orality. It would be useful for science communicators to edit carefully their texts. Both researchers and journalists need to improve their writing skills permanently. Nevertheless it must be emphasized that school education and competent teachers seem to have important influence on the linguistic patterns.
The communicative role of nonlinear vocal phenomena remains poorly understood since they are difficult to manipulate or even measure with conventional tools. In this study parametric voice synthesis was employed to add pitch jumps, subharmonics/sidebands, and chaos to synthetic human nonverbal vocalizations. In Experiment 1 (86 participants, 144 sounds), chaos was associated with lower valence, and subharmonics with higher dominance. Arousal ratings were not noticeably affected by any nonlinear effects, except for a marginal effect of subharmonics. These findings were extended in Experiment 2 (83 participants, 212 sounds) using ratings on discrete emotions. Listeners associated pitch jumps, subharmonics, and especially chaos with aversive states such as fear and pain. The effects of manipulations in both experiments were particularly strong for ambiguous vocalizations, such as moans and gasps, and could not be explained by a non-specific measure of spectral noise (harmonics-to-noise ratio) – that is, they would be missed by a conventional acoustic analysis. In conclusion, listeners interpret nonlinear vocal phenomena quite flexibly, depending on their type and the kind of vocalization in which they occur. These results showcase the utility of parametric voice synthesis and highlight the need for a more fine-grained analysis of voice quality in acoustic research.
Abstract This article presents results from a study on hybrid linguistic norms in translated articles from New York Times made available on UOL website. According to Faraco (2008) and Bagno (2012), there is a difference between norma padrão (a prescriptive norm, but not based on usage) and norma culta (an alternative, usage-based norm). The first one combines normative rules that determine correct linguistic forms, but generally hard to follow by most users, while the second one brings together a set of linguistic forms frequently employed by users, because they are more intuitively accessible, although not subscribed by the conservative standard norm (norma padrão) commonly taught in grammar books and in writing style manuals. The research was meant to verify if journalistic texts translated from English have been as permeable to linguistic forms not subscribed by the prescriptive standard norm, as the ones originally written in Portuguese have proven to be.
Location: Dewberry Hall With over 34,000 students representing 123 countries, George Mason University is a vastly diverse university with students bringing different learning experiences and skill sets with them into the classroom. The study that our team conducted analyzes the essays of native English speakers, as well as the essays of students for whom English is their second language. Our objective when conducting this research was to observe the essays for signs of syntactic complexity and patterns of language errors. Specifically, we looked for subordinating clauses, transitions, subject/verb agreement, article usage, run-on sentences, and fragments. We found that some errors in L1 and L2 populations were consistent with our expectations, but others reveled a more complex understanding of the linguistic norms of the groups studied. The results of these findings will give professors of all disciplines and modalities insight to the challenges that first-year L1 and L2 students confront when faced with a writing assignment.
Contextualized embeddings, which capture appropriate word meaning depending\non context, have recently been proposed. We evaluate two meth ods for\nprecomputing such embeddings, BERT and Flair, on four Czech text processing\ntasks: part-of-speech (POS) tagging, lemmatization, dependency pars ing and\nnamed entity recognition (NER). The first three tasks, POS tagging,\nlemmatization and dependency parsing, are evaluated on two corpora: the Prague\nDependency Treebank 3.5 and the Universal Dependencies 2.3. The named entity\nrecognition (NER) is evaluated on the Czech Named Entity Corpus 1.1 and 2.0. We\nreport state-of-the-art results for the above mentioned tasks and corpora.\n
This article proposes a character-level neural language model (NLM) that is based on quantum theory. The input of the model is the character-level coding represented by the quantum semantic space model. Our model integrates a convolutional neural network (CNN) that is based on network-in-network (NIN). We assessed the effectiveness of our model through extensive experiments based on the English-language Penn Treebank dataset. The experiments results confirm that the quantum semantic inputs work well for the language models. For example, the PPL of our model is 10%–30% less than the states of the arts, while it keeps the relatively smaller number of parameters (i.e., 6 m).
Modern machine learning (ML) techniques are transforming many disciplines ranging from transportation to healthcare by uncovering patterns in data, developing autonomous systems that mimic human abilities, and supporting human decision-making. Modern ML techniques, such as deep neural networks, are fueling the rapid developments in artificial intelligence. Engineering design researchers have increasingly used and developed ML techniques to support a wide range of activities from preference modeling to uncertainty quantification in high-dimensional design optimization problems. This special issue brings together fundamental scientific contributions across these areas.The special issue consists of 24 papers spread over two issues of the Journal of Mechanical Design. The papers use various ML techniques, including artificial neural networks, Gaussian processes, reinforcement learning, clustering techniques, and natural language processing. Based on their research objective, the papers can be broadly classified into four groups: (i) ML to support surrogate modeling, design exploration, and optimization, (ii) ML for design synthesis, (iii) ML for extracting human preferences and design strategies, and (iv) comparative studies of ML techniques and research platforms to help design researchers. The papers are summarized in Secs. 1–4. An analysis of the themes covered in the special issue and the potential opportunities for future research in ML for Engineering Design are presented in Sec. 5.In the paper titled Multifidelity Physics-Constrained Neural Network and Its Application in Materials Modeling, Liu and Yang address how to incorporate multifidelity, physics-based constraints into neural network predictions. The paper contributes two key insights. First, the paper extends existing Physics-Constraints Neural Network architectures by imposing a multifidelity constraint scheme wherein an auxiliary network minimizes discrepancies between low and high fidelity models—essentially learning how to correct the low-fidelity one. Second, it proposes an adaptive weighting scheme to control the convergence of individual losses among the different fidelities. They demonstrate the impact of these improvements on several fundamental multiscale material modeling challenges including two-dimensional heat transfer, phase transition, and dendritic growth problems. On these problems, the proposed multifidelity, physics-based constraints decrease the prediction error up to order of magnitude compared with networks without such constraints. This achieves comparable accuracy to that of direct numerical solutions of the underlying equations.Sarkar et al. present a multifidelity modeling and information-theoretic sequential sampling strategy for optimization in their paper titled Multifidelity and Multiscale Bayesian Framework for High-Dimensional Engineering Design and Calibration. The approach is based on modeling of the varied fidelity information sources via Gaussian processes, augmented with efficient active learning strategies that involve sequential selection of optimal points in a multiscale architecture. The strategy is demonstrated using the design optimization of a compressor rotor and calibration of a microstructure prediction model.In the paper titled A Case Study of Deep Reinforcement Learning for Engineering Design: Application to Microfluidic Devices for Flow Sculpting, Lee et al. address how to design micro-fluidic flow sculpting devices by overcoming some of the key weaknesses of evolutionary optimization-based methods, namely, poor sample efficiency and slow optimization convergence. The paper adapts deep reinforcement learning (DRL) techniques to the flow sculpting task and also studies the effectiveness of transfer learning on accelerating the design of target flow shapes. The paper demonstrates that DRL is able to match 90% of the target flow shapes using significantly fewer sculpting pillars than comparable GA models as well as provides a means to interpret the learned model (using Principal Components) that existing approaches to fluidic sculpting do not provide.Lynch et al., in their paper Machine Learning to Aid Tuning of Numerical Parameters in Topology Optimization, present an ML-based meta-learning framework to determine tuning parameters in topology optimization. The parameters are learned from similar optimization problems carried out in the past and adjusted for the problem at hand. This helps in avoiding costly trial-and-error involved in manual parameter tuning.In the paper Data-Driven Design Space Exploration and Exploitation for Design for Additive Manufacturing, Xiong et al. present a data-driven approach for design search and optimization at successive stages in the design process. They use Bayesian network classifier in the embodiment design stage and Gaussian process regression in the detailed design phase. The approach is illustrated in the paper through the design of a customized ankle brace design.Odonkor and Lewis apply data-driven design to the design of operational strategies of complex systems, specifically distributed energy resources. The paper is titled Data-Driven Design of Control Strategies for Distributed Energy Systems. The problem of maximizing arbitrage value is formulated as an optimization problem and solved using reinforcement learning. The approach is demonstrated for shared distributed energy resources in multi-building residential clusters.In Globally Approximate Gaussian Processes for Big Data With Application to Data-Driven Metamaterials Design by Bostanabad et al., a globally approximate Gaussian process (GAGP) is introduced for the purpose of handling large datasets. A GAGP is constructed by pooling several Gaussian processes using identical hyperparameters but built from different subsets of the training data. The predictive capability of GAGPs is shown to be at least as good as state-of-the-art supervised learning methods. It is demonstrated on the unit-cell design of metamaterials through inverse optimization.Liu et al. present a method for the design for crashworthiness involving categorical multimaterial structures in their paper titled Design for Crashworthiness of Categorical Multimaterial Structures Using Cluster Analysis and Bayesian Optimization. Following a topology optimization, the dimensionality of the problem is reduced through clustering followed by a Bayesian optimization to assign a given material to a specific cluster. The approach is applied to the maximization of absorbed energy of an S-rail.Garriga et al. propose a framework to assist the optimization of aircraft systems at the early design stages. The approach in their paper titled A Machine Learning Enabled Multifidelity Platform for the Integrated Design of Aircraft Systems is based on the screening of designs using clustering followed by an identification of the best candidate on a Pareto front. The framework enables the use of models of various fidelities and is demonstrated on a primary flight control system and a landing gear.In Synthesizing Designs With Interpart Dependencies Using Hierarchical Generative Adversarial Networks, Chen and Fuge present a method for synthesizing hierarchical designs with inter-part dependencies using generative models learned from examples. The method constructs multiple generative models using generative adversarial networks (GANs) while satisfying the dependencies through part dependency graphs. The paper lays the foundation for extending the use of generative models from creative individual parts to more realistic engineering systems.The objective in Evolving a Psycho-Physical Distance Metric for Generative Design Exploration of Diverse Shapes by Khan et al. is to incorporate humans’ psychological perceptions about design into the design exploration process. A psycho-physical distance metric is proposed that enables the augmentation of CAD designs based on feedback from users. Results reveal that the proposed method generates more distinct variations of CAD designs compared with a baseline Euclidean distance method.Oh et al. in their paper titled Deep Generative Design: Integration of Topology Optimization and Generative Models present a design framework for creating diverse aesthetic designs that are optimized for engineering performance. The framework integrates topology optimization and generative adversarial networks (GANs) to generate large numbers of design options from limited previous design data. The approach is validated using a 2D wheel design problem.Deshpande and Purwar, in their paper Computational Creativity Via Assisted Variational Synthesis of Mechanisms Using Deep Generative Models, present an approach for variational synthesis of mechanisms and an End-to-End synthesis pipeline that accepts raw, high-level input from users and provides them with distinct concept solutions. The approach is based on learning the probability distribution of linkage parameters and their interdependence to perform tasks such as input conditioning, imputation, and variational synthesis. The approach is a step in the direction of enhancing users’ computational creativity for engineering design.Stump et al. in their paper, Spatial Grammar-Based Recurrent Neural Network for Design Form and Behavior Optimization, present a method for simultaneous optimization of form and behavior through a combination of physics-based models and ML techniques. Specifically, they use character-Recurrent Neural Networks to embody spatial grammars and reinforcement learning to optimize the behavior. The design of a modular multi-hull sailing craft is used as a demonstration problem.Suryadi and Kim utilize machine-learning algorithms for customer choice modeling in their paper titled A Data-Driven Methodology to Construct Customer Choice Sets Using Online Data and Customer Reviews. They present an approach that utilizes publicly available online data and customer reviews from e-commerce websites to construct customer choice sets in the absence of both an actual choice set and customer sociodemographic data. The approach consists of clustering (i) products based on their attributes and (ii) customers based on their reviews, and constructing the choice-sets based on a sampling probability scenario that relies on product and customer clusters. The approach generates choice models with higher predictive ability than randomly constructed choice sets.In their paper Extracting Customer Perceptions of Product Sustainability From Online Reviews, El Dehaibi et al. seek to extract perceived sustainable design features from online reviews. Annotators from Amazon’s Mechanical Turk are used to annotate product reviews and develop a natural language processing model that predicts the positive/negative sentiment of sustainable phrases. The results reveal that the model is more efficient at predicting positive sentiment pertaining to sustainable product features compared with negative sentiments.Raina et al. take a step toward transfer learning from human designers to computational agents in their paper titled Transferring Design Strategies From Human To Computer and Across Design Problems. They present an approach where design strategies are represented using a probabilistic model that provides a general mechanism to transfer strategies from human designers to computational design agents and to generate new designs. The approach is illustrated using a configuration design problem.The goal in Learning to Design From Humans: Imitating Human Designers Through Deep Learning by Raina et al. is to teach computational agents to generate designs without the need for explicit information about objective or performance metrics. A deep learning model is proposed that learns from historical human data and identifies the important regions of a design space. The results reveal that the machine learning agent learns to create designs that are comparable to human-generated ones, despite not having the same explicit feedback that humans do to guide them through the design exploration process.He et al. address the challenge of mining large numbers of design ideas generated from the crowd in their paper titled Mining and Representing the Concept Space of Existing Ideas for Directed Ideation. The authors use natural language processing to extract keywords as elementary concepts and represent the concepts in a way that they can be recombined to generate new ideas.In the paper titled A Data-Driven Approach to Product Usage Context Identification From Online Customer Reviews, Suryadi and Kim use machine learning and natural language processing to identify and cluster usage contexts from a large volume of customer reviews. The methodology also captures sentiments toward a particular usage context in a sentence. The methodology enables designers to effectively use online product reviews by focusing on several specific reviews regarding particular usage contexts and potentially to identify market opportunities for new products that excel in specific usage contexts.Sharpe et al. illuminate differences between Supervised Learning algorithms in terms of how and where different algorithms may apply to different Engineering Design applications. Their paper titled A Comparative Evaluation of Supervised Machine Learning Classification Techniques for Engineering Design Applications does this by comparing four common supervised learning approaches—Support Vector Machines, Random Forests, Gaussian Näive Bayes, and shallow depth Neural Networks—across six example problems that demonstrate different facets or challenges classifiers may face within the engineering design. The results from the work are multifaceted with different algorithms performing better or worse under different conditions and performance measures. However, this leads to the general notion of strong problem dependence for the classifier choice and highlights the importance of understanding appropriate benchmark problems within the engineering design that can shed light on such issues in the future.The availability of data enables not just designers but also design researchers. Rahman et al., in their paper A Computer-Aided Design Based Research Platform for Design Thinking Studies, present a research platform to support data-driven design-thinking and decision-making research. Through the use of fine-grained design action data and unsupervised clustering methods in conjunction with design process models, the authors show how the platform enables data-driven research studies on designers’ sequential decision-making behaviors.In Design Repository Effectiveness for 3D Convolutional Neural Networks: Application to Additive Manufacturing, Williams et al. address the question of whether or not a data repository is useful for training effective ML tools. The authors experimentally test the effects of changes in CAD datasets on the precision and generalizability of trained convolutional neural networks (CNNs) for additive manufacturing applications. The study sheds light on how standardization of design repositories can influence the performance of ML tools.Cunningham et al. study the construction of a performance surrogate based on 3D point cloud representations. In their paper titled An Investigation of Surrogate Models for Efficient Performance-Based Decoding of 3D Point Clouds, a radial basis function (RBF) surrogate is used to link performance and cloud representation mapped onto a latent vector. The proposed RBF-based approach was found to be more efficient and accurate than traditional neural network-based approaches.In the call for proposals for this special issue, the guest editors posed three primary questions: How to effectively use ML for new design applications that are not well-supported by existing ML practice or tools?How to leverage the unique aspects of engineering design in creating new ML approaches?How to share benchmark problems or datasets that can measure ML progress in design?Looking back at the papers collectively within this Special Issue helps shed light on the areas that are receiving significant within the design research and the areas where are opportunities for future design research a strong of using machine learning. techniques have used for data-driven techniques such as Gaussian process regression and neural networks have applications in learning complex between design and performance natural language processing used for mining customer and reinforcement learning is an part of control systems design. of these techniques are in the papers for the special issue Bostanabad et et Suryadi and Xiong et of the deep learning techniques such as convolutional neural networks (CNNs) and generative adversarial networks (GANs) are their way into engineering design research and and and et papers in this special issue address diverse applications including modeling, additive distributed energy systems, and topology optimization, synthesis, mechanism preference modeling, and learning from human of the areas where is potential for ML techniques, but are not well represented in this special issue, modeling human design of market systems, of products with humans or the design for use of data from product usage to using data from of the product or data, or across different aspects of the are opportunities for design challenges in engineering such as and and that from the products and ML can also be used to support engineering design for supporting studies and the generalizability of research terms of understanding and of machine learning methods, the papers in the special issue on the the of multifidelity or multiple data or with the common surrogate modeling and physics-based constraints into ML and ML-based models of human preferences and of the papers in the special issue present surrogate modeling this an active of research for the past two the approaches are using ML techniques. A of papers issues that ML models in multifidelity and et The availability of multifidelity models is in engineering design. In some designers’ are to construct multiple and fidelities of models or of a system as to with it at the appropriate This is not that ML systems are to In this of do multiple or important to to this special issue and is an active of of the approaches on existing with changes that constraints. of the is that many ML systems are not used to many of the problems that need for engineering design. in et al., how the between form and behavior the generated or in and Lewis the between Control and Design these of between and within Engineering and existing ML approaches do not need to for papers the of design or constraints in a constraints by or and hierarchical or et Chen and and constraints and This is important the of many ML systems relies on their to but specific of into the This is where engineering design researchers are well to are opportunities for research in engineering with ML models, such as physics-based models from or via more system models, and studies a problem and as a In the can ML be used with a between design and A is learning from multiple of a or learning among design with multiple data structures and of the papers to this special issue used supervised or reinforcement learning where data or available via used Learning clustering and dimensionality in natural and have from unsupervised learning of using deep learning or and this be a for future work in design. are opportunities for approaches to or for in uncertainty with or data. and calibration of ML models an papers on the of ML for optimization and specifically how to about generative models of two of approach to on the of the the approach Optimization as as by approaches that used or latent methods et al. and et or Optimization as such as that formulated the problem using reinforcement learning or inverse problem et are active and areas of in to for optimization that not in this special issue, such as direct inverse design papers how to best or behavior into an ML This is a are many opportunities for human information and strategies into ML models, and models of and computational or for behavior. This to approaches have used to models into the same be for of models, such as involving human or to ML models of human behavior designers or the design research on in both and and this is within the special However, are constraints that engineering design on of a notion that be to function and to that out as a of approaches to do not such aspects and this a that the Engineering Design can This also for are the fundamental of designers as conditions and does this datasets have ML research and a common for performance. collectively the the among have collectively of of They have in and This special issue set out to to new datasets for engineering design papers in the special issue propose datasets or platforms optimization problems et 2D and 3D shapes et Chen and and manufacturing et while not datasets do use platforms such as to data or these papers within the Special is a wide to be by future work that can useful datasets or for machine learning within engineering design. are some of the areas that papers in the Special Issue and have used ML but for good benchmark datasets datasets for human or including how designers or for more complex than optimization do not have the the model of for Engineering problems such as or that link or use multiple of design representations. datasets a CAD and a by the do not datasets have ML approaches in in the have in for and and datasets for approaches to specific in do not have a good of engineering problems and a given is to This is not unique to Engineering Design. How do or the and that it a realistic benchmark for Computer In some to be is the design on of or the of where and design datasets apply and how their results to datasets and understanding of their or be for future of how to such models within Engineering Design. how and of ML models within this on the of or ML models influence the design of a How does this Design are some of the and the papers in this special issue the of research opportunities in this The papers in the special issue many points from future researchers and may set guest the for their in the and feedback to the Special to and for their help with the for the process of of this special issue, and for the
The topic of the talk and its classification are one of the central issues of syntax. This article compares the classification of the Arabic alphabet with the Arabic and Uzbek linguistic norms. In terms of the stylistics of the Uzbek language, it is explained in terms of how the spelling of the Arabic word begins, and the classification of the Muslim in terms of the context. It is emphasized in Maonic science that the most important aspect of non- speaking in other languages, especially in the Arabian minority, is the purpose of the speaker and the state of the listener. In Maonical Science there is information on classification in relation to reality, the goal of the speaker and the status of the listener, and in the so-called interpreter, to be a change in reality, and to choose the types of speech.
In opera theatres, a speech consultant assists singers in pronouncing foreign languages, as well as helping them with the pronunciation of Slovenian. A sung opera text needs to be intelligible and linguistically standardised. The articulation of the sounds of Slovenian literary language in opera singing depends on the linguistic norm and on musical laws. The opera theatre speech consultant is involved in the process of creating a performance from the very beginning. Detailed knowledge of the text, as well as of music and directing concepts, also enables the consultant to proofread the texts in the theatre programme in a quality way and to ensure that the surtitles correspond to the stage action.
This paper presents a case study of the use of the NINJAL Parsed Corpus of Modern Japanese (NPCMJ) for syntactic research. NPCMJ is the first phrase structure-based treebank for Japanese that is specifically designed for application in linguistic (in addition to NLP) research. After discussing some basic methodological issues pertaining to the use of treebanks for theoretical linguistics research, we introduce our case study on the status of the Coordinate Structure Constraint (CSC) in Japanese, showing that NPCMJ enables us to easily retrieve examples that support one of the key claims of Kubota and Lee (2015): that the CSC should be viewed as a pragmatic, rather than a syntactic constraint. The corpus-based study we conducted moreover revealed a previously unnoticed tendency that was highly relevant for further clarifying the principles governing the empirical data in question. We conclude the paper by briefly discussing some further methodological issues brought up by our case study pertaining to the relationship between linguistic research and corpus development.
This chapter aims at presenting different strategies that have been designed to incorporate multiword expression (MWE) identification in the process of syntactic parsing using statistical approaches. We discuss MWE representation in treebanks, pipeline and joint orchestrations, the integration of external lexicons and the evaluation of MWE-aware parsers, concluding with our suggestions for future research.
The goal of this study is two-fold. First, it will reveal to what extent differences in amount of experience with a particular register manifest themselves in different familiarity judgments when faced with word sequences that are characteristic of that register. To this end, three groups of participants –recruiters, job-seekers, and people not (yet) looking for a job– performed a metalinguistic judgment task in which they assigned familiarity ratings to two sets of stimuli – word sequences characteristic of either job ads or news reports. As the three groups differ in experience in the domain of job hunting, they are likely to differ in experience with collocations that are typically used in that domain. According to usage-based theories, these differences in experience lead to differences in mental representations of language. This leads to a testable hypothesis: If familiarity judgments give expression to linguistic representations, the ratings should reflect these differences. That is, the Job ad stimuli ought to be most familiar to the Recruiters and least familiar to the Inexperienced participants. Subsequently, we examined the relationship between metalinguistic judgments and other types of experimental data. The stimuli that were presented in the judgment task have also been used in two other experiments conducted among the same participants: a Completion task and a Voice Onset Time experiment (both described in Verhagen et al, 2019). By analyzing the judgment data in relation to the participants’ Completion task responses, their voice onset times, and corpus-based frequencies, we can answer the second research question: To what extent do someone’s own data from psycholinguistic processing tasks have explanatory power in predicting familiarity judgments in addition to corpus frequencies? If the different types of tasks tap into the same mental representations, one’s performance in the processing tasks should be a significant predictor of one’s familiarity ratings. If it does not prove to be a significant predictor, this means that there are substantial differences between the tasks in the information they provide.
Abstract Normative and cognitive-linguistic accounts of linguistic meaning are often portrayed and conceived as mutually exclusive alternatives. This dichotomy stems from an insufficient understanding of what the phenomenological accessibility of meaning and usage-basedness of language entail. Namely, the theoretical premises of Cognitive Linguistics actually presuppose socially grounded, normative linguistic meanings. The question remains, what kind of entities normative meanings are like. The present chapter makes a case for construal, linguistic perspective-taking usually analyzed as a conceptual phenomenon, as a normative facet of meaning. Analysis presented here suggests that construal emerges as an inherent property of linguistic expressions via conventionalization of intentionality. This analysis does not only expand the area of linguistic normativity but also points to the integral relation between linguistic norms and intentionality.
This work is partly supported by the Asian-MT “Network-based ASEAN Languages Translation Public Service” and the ASEAN IVO Project “Open Collaboration for Developing and Using Asian Language Treebank”.
The reader is typeset (see LaTeX / XeLaTeX source), many errors are corrected. The book is submitted to review at the University of Zagreb. Other resources (treebanks, grammatical annotations) are still in development.
Universal Dependencies Treebank is a cross-linguistic project to annotate Parts-of-Speech and dependency relations universally. The same 17 universal Part-of-Speech tags and 37 universal dependency-relation tags are used for the annotation universally across all languages. However, in fact, each UD treebank is developed by each developer of each language, reflecting “not-universal” treatments of the language. In this paper, we reveal the difference among Classical Chinese UD, Modern Japanese UD, and Modern English UD, upon parallel corpora of 大學(The Great Learning).
Space-valence metaphors (e.g., bad is down) are embedded within cognitive and emotional processing (e.g., negative stimuli at a lower space capture visual attention more than those at an upper space). Previous studies have revealed that motor action to vertical direction affects the emotional valence rating of stimuli in a metaphor-congruent manner only when the action was introduced after the stimuli presentation. In the present study, we hypothesized that motor action before the stimuli presentation does not affect valence rating while it may affect visual selective attention. In Experiment 1 (participants: 28 university students; mean age = 19.50 years), we partially replicated the previous result with repeated ANOVA and t-tests; manual action introduced before the stimuli presentation does not affect the valence rating. Then, in Experiment 2 (participants: 28 university students; mean age = 19.57 years), we employed a modified version of the dot-probe task as a measure of visual selective attention to emotional stimuli, where participants’ vertical or horizontal manual action was introduced before the presentation of a pair of emotional words. The results of the t-tests revealed that an upward manual action promoting selective attention to negative words, which was incongruent with the space-valence metaphorical correspondence. These results suggest that even though manual action does not affect the evaluative process of emotional stimuli prospectively, upward manual action introduced before stimuli presentation can promote visual attention to the subsequent negative stimuli in a way that is incongruent with the space-valence metaphor.
Treebank is one of the important and useful resources in natural language processing represented in two different annotated schemas: phrase and dependency structures. There are many works that convert a phrase structure into a dependency structure and vice versa. Most of them are based that exploit the handcrafted head percolation table and argument table in predefined deterministic ways. In this article, we propose a method to convert a dependency structure into a phrase structure by enriching a trainable model of former hybrid strategy approach. By adding a classifier to the algorithm and using postprocessing modification, the quality of conversion is increased. We evaluate our method in two different languages, English and Persian, and then analyze the errors. The results of our experiments show a 46.01% reduction of error rate in English and 76.50% for Persian compared to our baseline. We build a new phrase structure treebank by converting 10,000 sentences of Persian dependency treebank into corresponding phrase structures and correcting them manually.
Most syntactic dependency parsing models may fall into one of two categories: transition- and graph-based models. The former models enjoy high inference efficiency with linear time complexity, but they rely on the stacking or re-ranking of partially-built parse trees to build a complete parse tree and are stuck with slower training for the necessity of dynamic oracle training. The latter, graph-based models, may boast better performance but are unfortunately marred by polynomial time inference. In this paper, we propose a novel parsing order objective, resulting in a novel dependency parsing model capable of both global (in sentence scope) feature extraction as in graph models and linear time inference as in transitional models. The proposed global greedy parser only uses two arc-building actions, left and right arcs, for projective parsing. When equipped with two extra non-projective arc-building actions, the proposed parser may also smoothly support non-projective parsing. Using multiple benchmark treebanks, including the Penn Treebank (PTB), the CoNLL-X treebanks, and the Universal Dependency Treebanks, we evaluate our parser and demonstrate that the proposed novel parser achieves good performance with faster training and decoding.
The practice of assessing brand management in construction in Ukraine is in a passive stage, but due to the entry into the Ukrainian market of foreign companies for which regular evaluation of their brand — the need for survival in a competitive environment, Ukrainian companies are beginning to pay more attention to the creation and formation of their competitive trading because of its high business image / rating. A modern toolkit based on appropriate approaches is used to form organizational and economic foundations. The article analyzes modern scientific approaches to assessing the economic potential of an enterprise, identifies the main trends and factors that affect the assessment of construction enterprises. The article is devoted to the study of theoretical and methodological assessments of branding of construction enterprises. The role and importance of innovation in ensuring the efficient operation of modern enterprises is emphasized. It is found that construction, especially innovative, is of great social importance and has a significant economic effect. The importance of determining the potential of innovative development of construction in general and construction enterprises in particular is substantiated. The purpose of this article is to systematically investigate the interpretation of the potential of innovative development of construction enterprises. The theoretical basis of the research is the scientific works of foreign and domestic scientists on the problems of identifying the essence of innovative development potential. A systematic study of the general characteristics of the construction company brand was conducted and the priority directions for choosing the development of the economic potential of the enterprises were determined. The necessity to understand the potential of the enterprise in the unity of all its elements, which are subject to the achievement of the overall goals of the enterprise, is substantiated. The weight of the component of the brand in the potential of the construction industry enterprises is substantiated. Some aspects of the development of theoretical and methodological approaches to the estimation of the intellectual capital of construction enterprises are formulated. Existing theoretical and methodological approaches to the brand assessment of construction enterprises are analyzed.
The relationship between words in a sentence often tell us more about the underlying semantic content of a document than its actual words individually. Natural language understanding has seen an increasing effort in the formation of techniques that try to produce non-trivial features, in the last few years, especially after robust word embeddings models became prominent, when they proved themselves able to capture and represent semantic relationships from massive amounts of data. These new dense vector representations indeed leverage the baseline in natural language processing, but they still fall short in dealing with intrinsic issues in linguistics, such as polysemy and homonymy. Systems that make use of natural language at its core, can be affected by a weak semantic representation of human language, resulting in inaccurate outcomes based on poor decisions. In this subject, word sense disambiguation and lexical chains have been exploring alternatives to alleviate several problems in linguistics, such as semantic representation, definitions, differentiation, polysemy, and homonymy. However, little effort is seen in combining recent advances in token embeddings (e.g. words, documents) with word sense disambiguation and lexical chains. To collaborate in building a bridge between these areas, this work proposes a collection of algorithms to extract semantic features from large corpora as its main contributions, named MSSA, MSSA-D, MSSA-NR, FLLC II, and FXLC II. The MSSA techniques focus on disambiguating and annotating each word by its specific sense, considering the semantic effects of its context. The lexical chains group derive the semantic relations between consecutive words in a document in a dynamic and pre-defined manner. These original techniques' target is to uncover the implicit semantic links between words using their lexical structure, incorporating multi-sense embeddings, word sense disambiguation, lexical chains, and lexical databases. A few natural language problems are selected to validate the contributions of this work, in which our techniques outperform state-of-the-art systems. All the proposed algorithms can be used separately as independent components or combined in one single system to improve the semantic representation of words, sentences, and documents. Additionally, they can also work in a recurrent form, refining even more their results.
يعد القرآن الكريم من مصادر المعرفة، وقد تولدت منه فروع واسعة؛ إذ نُزل القرآن الكريم باللغة العربية، ولا يوجد خيار آخر لإتقان المعرفة الواردة فيه إلا من خلال تعلم اللغة العربية. تهدف هذه الدراسة إلى بيان مفهوم المدونة العربية القرآنية ومكوناتها، والكشف عن علاقة تعلم اللغة العربية بالقرآن الكريم، وبيان كيفية تعليم وتعلم القواعد العربية الأساسية عبر المدونة العربية القرآنية، وستتبع الدراسة المنهج الوصفي والتحليلي. إن وجود العلاقة بين اللغة العربية والقرآن الكريم، يدفع الطلبة المتخصصين في اللغة العربية أن يربطوا اللغة العربية بالقرآن؛ لذلك نرى أن المدونة العربية القرآنية تساعدهم على فهم القواعد القرآنية بطريقة مثيرة للاهتمام. في نظرة شاملة يمكن أن نستنتج أن المدونة العربية القرآنية هي واحدة من أهم الأدوات الحسابية التي تم إنتاجها في خدمة اللغة العربية؛ حيث توفر للمتعلمين ما يحتاجون إليه في مجال اللغة واللغويات والدراسات الحاسوبية، كما تمهد الطريق للباحثين لدراسة الهياكل المورفولوجية والنحوية من خلال دراسات الحوسبة العميقة للقرآن.
 الكلمات المفتاحية: المصرف القرآني، نموذج حاسوبي، المدونة العربية القرآنية، المعجم القرآني.
 Abstract 
 The Holy Quran is a source of knowledge and it has generated wide branches of knowledge. The Holy Quran was revealed in Arabic. Hence, there is no other option to master its knowledge except by learning the Arabic language. This study aims at explaining the concept of the Arabic Quranic Corpus and its components, revealing the relationship between learning the Arabic language and the Holy Quran, and showing how to teach and learn basic Arabic grammar through the Quranic Arabic Corpus. The study will follow the descriptive and analytical approach. The existence of the relationship between the Arabic language and the Holy Quran prompts Arabic learners to associate Arabic with the Qur'an. Therefore, we see that the Quranic Arabic Corpus helps them to understand Quranic rules in an interesting way. In a comprehensive view, we can conclude that the Arabic Quranic Corpus is one of the most important web-based medium produced to serve the Arabic language. It provides learners with what they need in the field of language, linguistics and computer studies, and paves the way for researchers to study morphological and grammatical structures through technology with detail description of grammars.
 Keywords: Quranic Treebank, Computational Model, Arabic Quranic Corpus, Qur’anic Dictionary.
Paper is dedicated to the testing of the concept of literacy, based on the questionnaire, carried out in the school year of 2018/2019 among the students of two secondary vocational schools in Vrsac, Belgrade and Grammar School in Vrsac (200 respondents). The primary hypothesis of the research was that detection and detailed study of high school students conceptosphere on literacy identify the fields to improve the teaching of Serbian as a mother tongue in secondary schools and the aim of work that, based on the collected and then processed data in analytical, cognitive and descriptive method, is to (a) isolate the dominant concepts of (non)literacy, (b) look at the tendencies of spreading and shaping the notion of literacy induced by the needs of modern life, and also that, in order to improve linguistic culture in all domains and all educational levels -(c) point to the possibility of improving the teaching of the Serbian language as a mother tongue. According to results of the survey secondary school students experience literacy in the 21 st century as a complex concept; from the one who is literate expecting linguistic knowledge, what are the basic, traditionally accepted parameters, and recognize illiteracy as the lack of ability to apply knowledge in the field of language. They also demonstrated that it is necessary to improve the efficiency of teaching approaches designed to improve functional literacy in a variety of communicative situations; increase the number of hours and exercises in the field of spelling, or nurture and acquire more comprehensive and knowledge in use and skills of different forms of literacy needed for managing in 21 st century; more attention should be paid to including relevant language handbooks in teaching; more explicit, on frequent and more familiar examples to students, point to the advantages of knowing and respecting the linguistic norm, paving the way for a better linguistic culture and enrichment of the mother tongue.
Speech processing systems rely on robust feature extraction to handle phonetic and semantic variations found in natural language. While techniques exist for desensitizing features to common noise patterns produced by Speech-to-Text (STT) and Text-to-Speech (TTS) systems, the question remains how to best leverage state-of-the-art language models (which capture rich semantic features, but are trained on only written text) on inputs with ASR errors. In this paper, we present Telephonetic, a data augmentation framework that helps robustify language model features to ASR corrupted inputs. To capture phonetic alterations, we employ a character-level language model trained using probabilistic masking. Phonetic augmentations are generated in two stages: a TTS encoder (Tacotron 2, WaveGlow) and a STT decoder (DeepSpeech). Similarly, semantic perturbations are produced by sampling from nearby words in an embedding space, which is computed using the BERT language model. Words are selected for augmentation according to a hierarchical grammar sampling strategy. Telephonetic is evaluated on the Penn Treebank (PTB) corpus, and demonstrates its effectiveness as a bootstrapping technique for transferring neural language models to the speech domain. Notably, our language model achieves a test perplexity of 37.49 on PTB, which to our knowledge is state-of-the-art among models trained only on PTB.
Lexical Markup Framework (LMF) or ISO 24613 [1] is a de jure standard that\nprovides a framework for modelling and encoding lexical information in\nretrodigitised print dictionaries and NLP lexical databases. An in-depth review\nis currently underway within the standardisation subcommittee,\nISO-TC37/SC4/WG4, to find a more modular, flexible and durable follow up to the\noriginal LMF standard published in 2008. In this paper we will present some of\nthe major improvements which have so far been implemented in the new version of\nLMF.\n
The intelligent information processing of the standard Zhuang language spoken mainly in Southern China is presently in its infancy, and lacks a well-defined language corpus and automatic part-of-speech tagging methods. Therefore, this study proposes an adversarial part-of-speech tagging method based on reinforcement learning, which solves the problems associated with a lack of a language corpus, time-consuming laborious manual marking, and the low performance of machine marking. Firstly, we construct a markup dictionary based on the grammatical characteristics of standard Zhuang and the Penn Chinese Treebank. Secondly, a dependency syntax analysis is applied for constructing the semantic information feature vectors of sentences, and long short-term memory is adopted as the policy network architecture to enhance available information using recurrent memory, and a conditional random field is employed as the discriminant network to perform label inference with global normalization. Finally, we use reinforcement learning as the model framework, target parts of speech as the feedback of the environment, and then obtain the optimal policy through adversarial learning. The results show that the combination of reinforcement learning and adversarial network alleviates the dependence of the model on the training corpus to some extent, and can quickly and effectively expand the scale of the annotation dictionary for the Zhuang language, thereby obtaining better labeling results.
According to World Intellectual Property Organization (2017) report, over 3 million patents exist in the patent database, but only certain numbers have commercial potential. Generally, to assess the commercial potential of patent, it consumes time and requires various expertise. Currently, several models have been developed to address this matter, which to assess using questionnaire tool for portfolios by human. So that occurs bias any limitation exists that our research will address by artificial intelligence. Hence, this research applies a Natural Language Programming to assess for commercial potential of patent, consisting of five steps - (i) Morphological analysis based on the Lexical database, (ii) Syntactic analysis of sentence to check syntax sentence patterns, (iii) Sematic analysis to interpret the meaning of words derived from the previous step, (iv) Discourse integration from context of domain together with the main sentence providing more accurate sentence analysis, and (v) Pragmatic analysis to ensure the correct meaning of interpretation. Then, the obtained data is used to determine criterion factors and formulate the model for assessing commercial potential of patent using Natural Language programming. This finding should deliver an alternative effective patent assessment system, which addresses some current deficiency in patent's assessment for commercial potential.
espanolResumen en castellano: Esta tesis trata de la morfologia verbal del ingles antiguo para identificar y lematizar los verbos debiles de esta lengua en un corpus al que se accede a traves de una base de datos lexica. La lematizacion es una de las tareas mas importantes a la hora de construir un diccionario. Sin embargo, es una de las tareas pendientes en el campo de la linguistica historica debido a que no existen corpora exhaustivos y lematizados de esta lengua. El enfoque de esta tesis doctoral esta en la lematizacion de las tres clases de verbos debiles del ingles antiguo, aunque las areas de la Lexicografia y la Linguistica de Corpus son tambien relevantes para esta investigacion. Las fuentes principales de esta investigacion son las formas flexivas que estan atestiguadas en el Dictionary of Old English Corpus (DOEC) y que estan disponibles en el lematizador Norna, las fuentes lexicograficas que existen publicadas sobre esta lengua, principalmente el Dictionary of Old English (DOE), y otras fuentes textuales como el York-Toronto-Helsinki Parsed Corpus of Old English (YCOE) y una indexacion de fuentes secundarias del ingles antiguo. El objetivo principal supone la identificacion de las flexiones de los verbos debiles y de su lematizacion con uno de los lemas propuestos en las listas de referencia. Conseguir este objetivo implica manejar las fuentes disponibles en ingles antiguo para poder lematizar y validar los resultados del analisis y el diseno de un metodo que combine busquedas automaticas en la base de datos lexica Nerthus y la revision manual de los resultados. La metodologia incluye cuatro pasos sucesivos con diversas tareas en cada paso. El primero de estos pasos tiene como objetivo la lematizacion de las formas canonicas de los verbos debiles lanzando cadenas de busquedas especificas para cada clase de verbos debiles en el lematizador Norna, donde esta disponible un indice de tipos del DOEC, la fuente de informacion mas fiable de la que se dispone en ingles antiguo. Despues, los resultados se validan con el DOE y se anaden las formas no-canonicas de los verbos debiles entre las letras A y H. El tercer paso tiene como objetivo identificar las formas no-canonicas de las terminaciones flexivas y de las vocales de los radicales que aparecen con mas frecuencia en los verbos debiles para generar patrones de lematizacion. La busqueda de estos patrones y de la lista de prefijos no-canonicos que esta disponible en Norna culmina en la lematizacion de las formas flexivas no transparentes de los verbos debiles. La validacion de los resultados de las letras I a la Y supone el ultimo paso de la metodologia, donde se comparan los datos obtenidos con el analisis sintactico del YCOE y con los datos que se obtienen de una base de datos de indexacion de las fuentes secundarias del ingles antiguo. Los problemas que surgen a lo largo del proceso de lematizacion tienen que ver principalmente con las peculiaridades del ingles antiguo y las limitaciones de la lematizacion de tipos que esta investigacion sigue. La discusion de los resultados del analisis concluye esta tesis. Las principales aportaciones de esta tesis son las listas de lemas y sus formas flexivas, especialmente las de los verbos entre las letras I y la Y ya que no estan disponibles todavia, y el metodo que se ha disenado para identificar estas formas, incluyendo los patrones de lematizacion generados para lematizar las formas con terminaciones no comunes y vocales no canonicas en el radical. EnglishThis thesis deals with the verbal morphology of the Old English language in order to identify and lemmatise weak verbs in a corpus accessed through a lexical database. Lemmatisation is a pending task in the field of historical linguistics given the lack of comprehensive and lemmatised corpora in this language. The focus of this doctoral dissertation is on the lemmatisation of the three classes of weak verbs, although the linguistic fields of Lexicography and Corpus Linguistics are also relevant to this research. The main aim involves the identification of the canonical and non-canonical realisations of the Old English weak verbs and their lemmatisation with a lemma from a reference list of weak verbs. Achieving this goal involves, firstly, the use of the available sources of the Old English language in order to lemmatise and validate the results and, secondly, the design of a semi-automatic research methodology that combines automatic searches in the lexical database Nerthus and the manual revision of the results in order to achieve this task. The sources for this investigation are the inflectional forms that are attested in the Dictionary of Old English Corpus (DOEC) which are available in the lemmatiser Norna, the lexicographical sources published on the Old English language, mainly the Dictionary of Old English (DOE), and other textual sources such as the York-Toronto-Helsinki Parsed Corpus of Old English (YCOE) and an index of secondary sources of Old English. The methodology comprises four successive steps and several tasks within each step. The first step aims at the lemmatisation of the transparent forms of weak verbs with the search of specific query strings for each subclass of weak verbs in the lemmatiser Norna, where an index type of the DOEC, the most reliable source of information regarding the Old English language, is available. Then, the second step validates the results with the DOE and adds to the analysis the non-canonical attestations for the weak verbs from the letter A-H. Thirdly, the identification of the most recurrent non-canonical inflectional endings and stem vowels attested in weak verbs gives rise to lemmatisation patterns. The search of these sets of correspondences and the list of non-canonical prefixes that is available in Norna results in the lemmatisation of the non-canonical inflections of weak verbs. The validation of the results from the letter I-Y concludes the research methodology with the syntactic parsing provided by the YCOE and the data retrieved from the index of secondary sources of Old English Freya. The issues that arise throughout the lemmatisation process mainly concern the idiosyncrasy of the Old English language writing system and the limitations of the lemmatisation by type that this investigation follows. The quantitative and qualitative discussion of the results of the analysis concludes this thesis. The main contributions of this thesis are the lists of weak lemmas and their lemmatised inflectional forms, specially those of the verbs I-Y which are not available yet and the designed research methodology to identify these forms, including the sets of lemmatisation patterns of the non-canonical inflectional endings and stem vowels of weak verbs.
Lexical simplification (LS) aims to replace complex words in a given sentence with their simpler alternatives of equivalent meaning. Recently unsupervised lexical simplification approaches only rely on the complex word itself regardless of the given sentence to generate candidate substitutions, which will inevitably produce a large number of spurious candidates. We present a simple LS approach that makes use of the Bidirectional Encoder Representations from Transformers (BERT) which can consider both the given sentence and the complex word during generating candidate substitutions for the complex word. Specifically, we mask the complex word of the original sentence for feeding into the BERT to predict the masked token. The predicted results will be used as candidate substitutions. Despite being entirely unsupervised, experimental results show that our approach obtains obvious improvement compared with these baselines leveraging linguistic databases and parallel corpus, outperforming the state-of-the-art by more than 12 Accuracy points on three well-known benchmarks.
Named Entity Recognition (NER) for Myanmar Language is essential to Myanmar natural language processing research work. In this work, NER for Myanmar language is treated as a sequence tagging problem and the effectiveness of deep neural networks on NER for Myanmar language has been investigated. Experiments are performed by applying deep neural network architectures on syllable level Myanmar contexts. Very first manually annotated NER corpus for Myanmar language is also constructed and proposed. In developing our in-house NER corpus, sentences from online news website and also sentences supported from ALT-Parallel-Corpus are also used. This ALT corpus is one part of the Asian Language Treebank (ALT) project under ASEAN IVO. This paper contributes the first evaluation of neural network models on NER task for Myanmar language. The experimental results show that those neural sequence models can produce promising results compared to the baseline CRF model. Among those neural architectures, bidirectional LSTM network added CRF layer above gives the highest F-score value. This work also aims to discover the effectiveness of neural network approaches to Myanmar textual processing as well as to promote further researches on this understudied language.
It is widely accepted that the human cognitive system organizes perceptual input into complex hierarchical descriptions which can be represented by tree structures. Tree structures have been used to describe linguistic, musical and visual perception. In this paper, we will investigate whether there exists an underlying model that governs perceptual organization in general. Our key idea is that the cognitive system strives for the simplest structure (the “simplicity principle”), but in doing so it is biased by the likelihood of previous experiences (the “likelihood principle”). We will present a model which combines these two principles by balancing the notion of most likely tree with the notion of shortest derivation. Experiments with linguistic and musical benchmarks (Penn Treebank and Essen Folksong Collection) show that such a combination outperforms models that are based on either simplicity or likelihood alone.
This chapter engages with research in lingua franca scenarios, defining this term and arguing for the value of adopting a linguistic ethnographic approach in this area. The development of the field of English as a Lingua Franca (ELF) is outlined. Historical antecedents for lingua franca studies in interactional sociolinguistics are identified, including work on intercultural miscommunications, the study of English as an international language and the identification of pragmatic strategies in ELF scenarios such as the let-it-pass procedure. The chapter considers critical debates, including the need to challenge the norms of standard language pedagogies, and the importance of maintaining a critical perspective on the current global dominance of English. It provides a review of current areas of focus in this area, including studies of English as lingua franca in higher education in an increasingly internationalised university system, and in workplaces. The range of methods drawn on in ethnographic studies of lingua franca scenarios is described, and implications for practice are identified in relation to the teaching of English and language policy and planning. Directions for future research identified include the emergence of social and linguistic norms in interaction, and developing the focus on lingua francas other than English.
Discourse relation classification has proven to be a hard task, with rather low performance on several corpora that notably differ on the relation set they use. We propose to decompose the task into smaller, mostly binary tasks corresponding to various primitive concepts encoded into the discourse relation definitions. More precisely, we translate the discourse relations into a set of values for attributes based on distinctions used in the mappings between discourse frameworks proposed by This arguably allows for a more robust representation of discourse relations, and enables us to address usually ignored aspects of discourse relation prediction, namely multiple labels and underspecified annotations. We study experimentally which of the conceptual primitives are harder to learn from the Penn Discourse Treebank English corpus, and propose a correspondence to predict the original labels, with preliminary empirical comparisons with a direct model.
When using computer-aided translation systems in a typical, professional translation workflow, there are several stages at which there is room for improvement. The SCATE (Smart Computer-Aided Translation Environment) project investigated several of these aspects, both from a human-computer interaction point of view, as well as from a purely technological side. This paper describes the SCATE research with respect to improved fuzzy matching, parallel treebanks, the integration of translation memories with machine translation, quality estimation, terminology extraction from comparable texts, the use of speech recognition in the translation process, and human computer interaction and interface design for the professional translation environment. For each of these topics, we describe the experiments we performed and the conclusions drawn, providing an overview of the highlights of the entire SCATE project.
In the paper the analysis of the perspective of the technology of intelligent monitoring systems is conducted, as intellectualization is the main direction of development of modern technologies, and the property of intellectuality should be inherent in all the latest information management systems. Various strategies for intellectualization of monitoring are aimed at implementing intellectual information support for decision-makers using monitoring tools. Such support can be realized by building fuzzy linguistic databases/knowledge together with fuzzy inference subsystems, and information for decision making can be displayed on the automated workplace of the decision maker.
Implicit causal relation recognition aims to identify the causal relation between a pair of arguments. It is a challenging task due to the lack of conjunctions and the shortage of labeled data. In order to improve the identification performance, we come up with an approach to expand the training dataset. On the basis of the hypothesis that there inherently exists causal relations in WHY-type Question-Answer (QA) pairs, we utilize WHY-type QA pairs for the training set expansion. In practice, we first collect WHY-type QA pairs from the Knowledge Bases (KBs) of the reading comprehension tasks, and then convert them into narrative argument pairs by Question-Statement Conversion (QSC). In order to alleviate redundancy, we use active learning (AL) to select informative samples from the synthetic argument pairs. The sampled synthetic argument pairs are added to the Penn Discourse Treebank (PDTB), and the expanded PDTB is used to retrain the neural network-based classifiers. Experiments show that our method yields a performance gain of 2.42% F 1-score when AL is used, and 1.61% without using.
The aim of the present study was to examine whether offspring at high and low familial risk for depression differ in the immediate and more lasting behavioural and physiological effects of hedonically-based mood repair. Participants (9- to 22-year olds) included never-depressed offspring at high familial depression risk (high-risk, n = 64), offspring with similar familial background and personal depression histories (high-risk/DEP, n = 25), and never-depressed offspring at low familial risk (controls, n = 62). Offspring provided affect ratings at baseline, after sad mood induction, immediately following hedonically-based mood repair, and at subsequent, post-repair epochs. Physiological reactivity, indexed via respiratory sinus arrhythmia (RSA), was assessed during the protocol. Following mood induction and mood repair, high- and low-risk (control) offspring reported comparable changes in levels of sadness and RSA. However, sadness increased among high-risk offspring following the post-repair epoch, whereas low-risk offspring maintained mood repair benefits. High-risk/DEP offspring also reported higher levels of sadness following the post-repair epoch than did low-risk offspring. Change in RSA did not differ across the three offspring groups. Self-ratings confirm that one source of difficulty associated with depression risk is diminished ability to maintain hedonically-based mood repair gains, which were not apparent at the physiological level.
We present a strategy to automate the extraction of semantic relations from texts. Both machine learning and rule-based techniques are investigated and the impact of different linguistic knowledge is analyzed for the various approaches. To implement the extraction system RExtractor, several natural language processing tools have been improved: from sentence splitting and tokenization modules to dependency syntax parsers. Furthermore, we created the Czech Legal Text Treebank with several layers of linguistic annotation, which is used to train and test each stage of the proposed system. As a result of the performed work, new Semantic Web resources and tools are available for automatic processing of texts.
A workshop on open resources for the original languages of the Bible in Copenhagen in March 2018 was the start of a new Copenhagen Alliance for Open Biblical Resources. The point of departure for the workshop was the need for programs and applications like Paratext and Bible Online Learner to have access to high-quality and reliable open data in order to assist Bible translators, teachers and students of Biblical Hebrew and New Testament Greek. The publication of contributions presents papers on methods for annotation, resources tracing patristic quotations and data for detached constructions in Biblical Hebrew. Reports cover tasks and data for Bible translation and research, treebanks, and applications like STEPBible and Bible Online Learner.
Abstract This chapter analyzes the interface between prosodic features and situational variables in the Rhapsodie treebank. First, we provide general quantitative information. Second, we present a preliminary set of statistical analyses performed on six specific prosodic variables, that provides two types of information. While Principal Component Analysis is useful to identify major trends in the data, inferential analysis is the first step in modeling the effect of the different modalities of each situational variable on the prosodic patterns, before considering predictive modeling. Finally, by focusing on the IPA unit, we present a third statistical method based on the computation of specificity indexes.
In this paper, we extend recent approaches to Lexicalized Tree Adjoining Grammar (LTAG) parsing that combine supertagging with dependency parsing. In other words, we assign supertags (= unanchored elementary trees) to lexical items and we compute substitution/adjunction arcs between them. Kasai et al. (2017, 2018) jointly predict these structures with a neural graph-based parser. Predicting 1-best supertags and dependency arcs (as in Kasai et al. (2017, 2018)) however leads only to partial parsing due to incompatibilities between elementary trees and derivation trees. We therefore extend the approach described in Kasai et al. (2017, 2018) to n-best supertags and k-best dependency arcs and combine it with a subsequent A*-parsing step that extends the TAG parser from Waszczuk (2017). We show that this architecture allows for efficient full TAG parsing while being sufficiently accurate. We test our architecture on an LTAG extracted from the French Treebank (FTB).
We recast dependency parsing as a sequence labeling problem, exploring several encodings of dependency trees as labels. While dependency parsing by means of sequence labeling had been attempted in existing work, results suggested that the technique was impractical. We show instead that with a conventional BiLSTM-based model it is possible to obtain fast and accurate parsers. These parsers are conceptually simple, not needing traditional parsing algorithms or auxiliary structures. However, experiments on the PTB and a sample of UD treebanks show that they provide a good speed-accuracy tradeoff, with results competitive with more complex approaches.