Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This paper presents the first gold-standard resource for Russian annotated with compositionality information of noun compounds. The compound phrases are collected from the Universal Dependency treebanks according to part of speech patterns, such as ADJ+NOUN or NOUN+NOUN, using the gold-standard annotations. Each compound phrase is annotated by two experts and a moderator according to the following schema: the phrase can be either compositional, non-compositional, or ambiguous (i.e., depending on the context it can be interpreted both as compositional or noncompositional). We conduct an experimental evaluation of models and methods for predicting compositionality of noun compounds in unsupervised and supervised setups. We show that methods from previous work evaluated on the proposed Russian-language resource achieve the performance comparable with results on English corpora.
The article refers to the concept of intelligentsia as a social group which exerts significant influence on Polish standard patterns. Although the term intelligentsia is vague and questionable, it is well-established term in Polish linguistics, especially in sociolinguistics. Author argues that science communicators (young professional researchers, science journalists, PhD students) represent the young intelligentsia, because these well-educated people pursue their intellectual development and they have sense of public duty. The article examines standard of popular science texts in Internet, new tendencies in written Polish and attitude of young intelligentsia toward traditional linguistic norm. The errors (esp. punctuation and syntax) exemplify impact of technological changes and phenomenon of secondary orality. It would be useful for science communicators to edit carefully their texts. Both researchers and journalists need to improve their writing skills permanently. Nevertheless it must be emphasized that school education and competent teachers seem to have important influence on the linguistic patterns.
The communicative role of nonlinear vocal phenomena remains poorly understood since they are difficult to manipulate or even measure with conventional tools. In this study parametric voice synthesis was employed to add pitch jumps, subharmonics/sidebands, and chaos to synthetic human nonverbal vocalizations. In Experiment 1 (86 participants, 144 sounds), chaos was associated with lower valence, and subharmonics with higher dominance. Arousal ratings were not noticeably affected by any nonlinear effects, except for a marginal effect of subharmonics. These findings were extended in Experiment 2 (83 participants, 212 sounds) using ratings on discrete emotions. Listeners associated pitch jumps, subharmonics, and especially chaos with aversive states such as fear and pain. The effects of manipulations in both experiments were particularly strong for ambiguous vocalizations, such as moans and gasps, and could not be explained by a non-specific measure of spectral noise (harmonics-to-noise ratio) – that is, they would be missed by a conventional acoustic analysis. In conclusion, listeners interpret nonlinear vocal phenomena quite flexibly, depending on their type and the kind of vocalization in which they occur. These results showcase the utility of parametric voice synthesis and highlight the need for a more fine-grained analysis of voice quality in acoustic research.
Abstract This article presents results from a study on hybrid linguistic norms in translated articles from New York Times made available on UOL website. According to Faraco (2008) and Bagno (2012), there is a difference between norma padrão (a prescriptive norm, but not based on usage) and norma culta (an alternative, usage-based norm). The first one combines normative rules that determine correct linguistic forms, but generally hard to follow by most users, while the second one brings together a set of linguistic forms frequently employed by users, because they are more intuitively accessible, although not subscribed by the conservative standard norm (norma padrão) commonly taught in grammar books and in writing style manuals. The research was meant to verify if journalistic texts translated from English have been as permeable to linguistic forms not subscribed by the prescriptive standard norm, as the ones originally written in Portuguese have proven to be.
Location: Dewberry Hall With over 34,000 students representing 123 countries, George Mason University is a vastly diverse university with students bringing different learning experiences and skill sets with them into the classroom. The study that our team conducted analyzes the essays of native English speakers, as well as the essays of students for whom English is their second language. Our objective when conducting this research was to observe the essays for signs of syntactic complexity and patterns of language errors. Specifically, we looked for subordinating clauses, transitions, subject/verb agreement, article usage, run-on sentences, and fragments. We found that some errors in L1 and L2 populations were consistent with our expectations, but others reveled a more complex understanding of the linguistic norms of the groups studied. The results of these findings will give professors of all disciplines and modalities insight to the challenges that first-year L1 and L2 students confront when faced with a writing assignment.
Contextualized embeddings, which capture appropriate word meaning depending\non context, have recently been proposed. We evaluate two meth ods for\nprecomputing such embeddings, BERT and Flair, on four Czech text processing\ntasks: part-of-speech (POS) tagging, lemmatization, dependency pars ing and\nnamed entity recognition (NER). The first three tasks, POS tagging,\nlemmatization and dependency parsing, are evaluated on two corpora: the Prague\nDependency Treebank 3.5 and the Universal Dependencies 2.3. The named entity\nrecognition (NER) is evaluated on the Czech Named Entity Corpus 1.1 and 2.0. We\nreport state-of-the-art results for the above mentioned tasks and corpora.\n
This article proposes a character-level neural language model (NLM) that is based on quantum theory. The input of the model is the character-level coding represented by the quantum semantic space model. Our model integrates a convolutional neural network (CNN) that is based on network-in-network (NIN). We assessed the effectiveness of our model through extensive experiments based on the English-language Penn Treebank dataset. The experiments results confirm that the quantum semantic inputs work well for the language models. For example, the PPL of our model is 10%–30% less than the states of the arts, while it keeps the relatively smaller number of parameters (i.e., 6 m).
The article considers the modernstate of computer (electronic) lexicography, which is traditionally divided into corpus and electronic ones. At first glance theincreasing global computerization has greatly facilitated the work of lexicographers and linguists, however, a number ofproblems have arisen at once. Among the first ones are the principles of corpora compilation, the requirements of which haveconsiderably expanded since the appearance of the first ones, and at present stage of corpus technologies development corporashould be able to answer a wide range of possible inquiries and meet the needs of different users. Most of the principlesdeveloped up till now are standardized and unified for the convenience of both scientific and non-scientific research. The nextkey issue that exists nowadays is not only the inclusion of the main lexicographic postulates, but also consideringtechnological and operational features of the new media (smartphones and tablets). Provided practical lexicography isprimarily intended to meet the needs of the end user, the procedures for determining the needs of users are becomingincreasingly relevant today by involving them in surveys and analysis of logs. It is noted that more and more non-specialists ofthe field are involved in the actual development of lexicographic projects. They perform simple tasks, mainly as volunteers.Such practices also require well-balanced and clear rules in order to avoid creating a low-quality final product. There exist afew online platforms where you can “hire” a number of volunteers to perform such tasks. Another aspect of the developmentof new lexicography, mainly in the English-speaking world, is a special type of lexicography - lexicography for fun. The mainobject of research here is the vocabulary associated with certain hobbies, books, films or computer games. The articleemphasizes the importance and relevance of this subject with anthropocentric cognitive paradigm of modern linguistics. Thestate of research of this phenomenon abroad and in Ukraine is clarified and disclosed. The novelty of the work is to analyzethe current state of research, development and methods of organization of lexicographic work, as well as to indicate theperspectives of its development in Ukraine. References Burkhanov, Igor. 1999. Linguistic Foundations of Ideography: Semantic Analysis and IdeographicDictionaries. Rzeszow: WSP. Cibej, Jaka, Darja Fise, Kosem Iztok. “The role of crowdsourcing in lexicography” (paper presented at eLEX 2015 conference in Herstmonceux Castle (UK) from 11 to 13 August 2015). Down, Ellie. 2005. The Unofficial Guide to Harry Potter. Chichester: Summersdale Publishers Ltd. Gao, Yongwei “The Appification of Dictionaries: From a Chinese Perspective”. (paper presented at eLex 2013 conference in Tallinn (Estonia) from 17 to 19 October 2013). Garside Roger, Geoffrey Leech, and Tony McEnery. 1997. Corpus Annotation: Linguistic Information from Computer Text Corpora. London: Longman. Gouws, Rufus Hjalmar, Ulrich Heid, Wolfgang Schweickard, et al. 2013. An International Encyclopedia ofLexicography. Supplementary. Volume: Recent Developments with Focus on Electronic and Computational Lexicography. Berlin; Boston: De Gruyter Mouton. Granger, Sylviane. 2012. “Introduction: Electronic lexicography-from challenge to opportunity”. Electronic Lexicography, edited by Sylviane Granger, and Magali Paquot. Oxford: Oxford University Press. Hidalgo, Pablo. 2017. Star Wars: The Last Jedi. The Visual Dictionary. N. Y.: DK Publishing. Holmer Louise, Martens von, Monica, and Skoldberg Emma. Making a dictionary app from a lexical database: the case of the Contemporary Dictionary of the Swedish Academy. (paper presented at eLEX 2015 conference in Herstmonceux Castle (UK) from 11 to 13 August 2015) Horot, Yevheniya, Lesya Malimon. 2015. “Suchasna leksykohrafiya: problemy y perspektyvy”. Aktualnipytannya inozemnoyi filolohiyi 3: 42–48. Danchevska, Yuliya. 2014. “Korpusy tekstiv u linhvodydaktytsi: zdobutky ta perspektyvy”. Novapedahohichna dumka 1: 58–60. Dubichynskyy, Volodymyr. 2013. “Ukraynskaya leksykohrafyya: istoriya i sovremennost”. Slavyanskayaleksykohrafyya. Moskva: Azbukovnik. Kulchytska, Tetyana. 1999. Ukrayinska leksykohrafiya XIX – XX st.: bibliohraf. pokazhchyk. NAN Ukrayiny: Lviv. nauk. b-ka im. V. Stefanyka. Kulchytskyy Ihor, Yulia Danchevska, and Ihor Likhnyakevych. 2013. “Deyaki aspekty stvorennya tavykorystannya paralelnykh korpusiv”. Naukovyy visnyk VNU im. Lesi Ukrayinky 20: 48–52. Kulchytskyy Ihor. 2017. “Informatsiyna tekhnolohiya poperedn’oho opratsyuvannya pryrodomovnykh tekstiv. Informatsiyni tekhnolohiyi ta vzayemodiyi”. Kulchytskyy, Ihor. 2002. Kompyuterno-tekhnolohichni aspekty stvorennya suchasnykh leksykohrafichnykh system. – K.: Nats. b-ka Ukrayiny im. V. I. Vernadskoho NAN Ukrayiny Kulchytskyy, Ihor. 2015. “Tekhnolohichni aspekty ukladannya korpusiv tekstiv”. Dani tekstovykh korpusiv u linhvistychnykh doslidzhennyakh. Lviv: Vydavnytstvo Lvivskoi politekhniky. Kupriyanov, Yevhen. 2008. “Kompyuterna leksykohrafiya yak problema suchasnoho movoznavstva(istorychnyy aspekt)”. Visnyk Kharkivskoho natsionalnoho universytetu im. V.N. Karazina 53: 12–16. Levchenko Olena, Ihor Kulchytskyy. 2013. “Tekhnolohiya peretvorennya p’atymovnoho slovnyka porivnyan u elektronnu formu”. Visnyk Natsionalnoho universytetu Lvivska politekhnika 770: 129–138. Lew, Robert. Space restrictions in paper and electronic dictionaries and their implications for the design of production dictionaries. Accessed February 25, 2019. https://repozytorium.amu.edu.pl/bitstream/10593/799/1/Lew_space_restrictions_in_paper_and_electronic_dictionaries.pdf McArthur, Tom. 1986. Worlds of Reference. Cambridge: Cambridge University Press. Meyer, Christian, Andrea Abel. 2017. “User participation in the Internet era”. The Routledge Handbook of Lexicography. Abingdon; New York: Routledge. Perebyynis, Valentyna, Viktor Sorokin. 2009. Tradytsiyna ta kompyuterna leksykohrafiya. Kyiv: Vyd. tsentr KNLU. Polyuha, Lev. 2006. “Ukrayinske slovnytstvo na perelomi tysyacholit”. Ukrayinoznavchi studiyi 6-7:17–25. Reynolds, David. 2000. Star Wars: the Visual Dictionary. The Ultimate Guide to Star Wars Characters and Creatures. N.Y.: Dorling Kindersley. Rundell, Michael. Redefining the dictionary: From print to digital. Accessed February 20, 2019.https://www.kdictionaries.com/kdn/kdn21_2013.pdf Rusanivskyy, Vitaliy, Volodymyr Shyrokov. 2002. “Informatsiyno-linhvistychni osnovy suchasnoyitlumachnoyi leksykohrafiyi”. Movoznavstvo 6: 7–48. Shyrokov, Volodymyr. 2011. Kompyuterna leksykohrafiya. Kyiv: Nauk. dumka. Syvokozova, Viktoriya. 2013. “Do pytannya periodyzatsiyi ukrayinskoyi tlumachnoyi leksykohrafiyi”. Movni i kontseptualni kartyny svitu 43: 61–62. Snizhko, Nataliya. 2017. “Nova dzherelna baza ukrayinskoyi leksykohrafiyi v systemi intehralnykhlinhvistychnykh doslidzhen”. Lyudyna. Kompyuter. Komunikatsiya. Lviv: Vydavnytstvo Lvivskoyipolitekhniky 47–51. Starko, Vasyl. 2017. “Kompyuterni linhvistychni proekty hurtu r2u: stan ta zastosuvannya”. Ukrayinka mova. 3: 86–97. Svensen, Bo. 1993. Practical Lexicography: Principles and Methods of Dictionary-Making. Translated by John Sykes and Kerstin Schofield. Oxford: Oxford University Press. Tarp, Sven. 2012. “Online dictionaries: today and tomorrow”. Lexicographica. De Gruyter. 28: 253–268. Trap-Jensen, Lars. Lexicography between NLP and Linguistics: Aspects of Theory and Practice (paperpresented at eLEX 2017 conference in Leiden (the Netherlands) from 19 to 21 September 2017.
Abstract This chapter is devoted to the presentation of the tools and methods used for the different steps of the semi-automatic syntactic annotation: automatic preprocessing; microsyntactic parsing with the FRMG tool, correction of the parsing with the Arborator tool, agreement analysis, post-validation correction, and development of the final format of the Rhapsodie syntactic treebank. As FRMG is a parser for written French that was not configured to analyze disfluencies and reformulation, we used our manual pile marking to unfold the piles and produce a series of simplified “sentences” with only government relations. Despite having two annotators plus a validator for the corrections, we found a substantial number of errors in the post-validation procedure by using a set of rules to determine the well-formedness of the trees.
This talk will explore the role of individual social actors and their communities and networks in the formation, maintenance and dissolution of linguistic norms. It will consider several standardisation episodes in the history of Old and Early Middle English and connect them to other unification processes in the political and cultural history of England at the time. It will be suggested that the suppression of variability on the linguistic level often accompanies, or is a symptom of, a similar suppression on the ideological level. Political, religious, legal, and linguistic processes mingle in various ways in this period (as they do today) to replicate and enhance the social order, to support a reform movement, or to refute dissent. \nThree case studies will offer insights into these processes: 1) shire courts and the ‘standardised’ lexis of the Anglo-Saxon Chronicle in the reign of King Alfred and Edward the Elder; 2) chancery norms and charters of the eleventh century; and 3) religious reform and linguistic focusing in the thirteenth century.
This is a work-in-progress report, which aims to share preliminary results of a novel sequence-to-sequence schema for dependency parsing that relies on a combination of a BiLSTM and two Pointer Networks (Vinyals et al., 2015), in which the final softmax function has been replaced with the logistic regression. The two pointer networks co-operate to develop a latent syntactic knowledge, by learning the lexical properties of "selection" and the lexical properties of "selectability", respectively. At the moment and without fine-tuning, the parser implementation gets a UAS of 93.14% on the English Penn-treebank (Marcus et al., 1993) annotated with Stanford Dependencies: 2-3% under the SOTA but yet attractive as a baseline of the approach.
The battle against HIV is one of the objectives in our century. Along these lines, among the HIV-contaminated patients, one of the most perilous and remarkable with its entanglements is those with lung pathologies. Concurring clinical arranging of the illness, such patients may introduce Tuberculosis, Pneumocystis jirovecii, Cytomegaloviruses, Candidiasis, Toxoplasmosis and so on. The exploration by logical examination foundation of lung illness was done among the inpatient people in measure of 48.37 (77%) of them were given tuberculosis and 11 (23%) with Interstitial Lung Disease (ILD). Studies were introduced on HIV-positive patients who were isolated by the randomization methods. Among 37 patients with tuberculosis, 29 (78%) had AFB (corrosive quick bacillius) with Gexpert, HAIN strategies, 6 (22%) were analyzed by imaging techniques (HRCT, chest X-beam) and serum ADA level. As per past investigations, there were no relationships between's serum ADA level rises at HIV-positive patients (p esteem 0.05). Among 11 patients gave ILD Pneumocystis jirovecii were distinguished at 5 (45%), 3 (27.5%) were given every day mortality, 3 took a Co-Trimaxozole treatment analyzed by imaging strategies. Clinical viability was endorsed by the nearness of pneumocystis starting point. At the second phase of the examination was discovered a relationship between's various Cd4 cell check and imaging rating. In this way, among absolute number of 119 HIV-positive patients, 38 (32%) had penetration zones, 53 (44%) had an obliteration, 20 (17%) dispersal, 8 (7%) mediastinal lymphadenopathy. Measurement results p esteem 0.000424, hence there is immediate relationship. There are heap aspiratory conditions related with HIV, extending from intense contaminations to constant noncommunicable infections. The study of disease transmission of these illnesses has changed essentially in the time of far reaching antiretroviral treatment. Assessment of the HIV-tainted patient includes evaluation of the seriousness of ailment and an intensive yet effective quest for authoritative conclusion, which may include various etiologies at the same time. Significant pieces of information to an analysis incorporate clinical and social history, segment subtleties, for example, travel and geology of habitation, substance use, sexual practices, and domiciliary and detainment status. CD4 cell tally is a colossally valuable proportion of resistant capacity and hazard for HIV-related infections, and assists slender with bringing down the differential. Cautious history of current side effects and physical assessment with specific thoughtfulness regarding extrapulmonary signs are vital early advances. Numerous adjunctive research facility studies can recommend or preclude specific findings. Pneumonic capacity testing (PFT) may help in portrayal of a few incessant noninfectious sicknesses quickened by HIV. Chest radiograph and registered tomography (CT) check take into account arrangement of sicknesses by pathognomonic imaging designs, albeit numerous irresistible conditions present atypically, especially with lower CD4 tallies. At last, conclusive finding with sputum, bronchoscopy with bronchoalveolar lavage, or lung tissue is regularly required. It is of most extreme significance to keep up a serious extent of doubt for HIV in any case undiscovered patients, as the principal introduction of HIV might be through an intense pneumonic disease. The assessment of respiratory side effects in HIV-contaminated patients can be trying for various reasons. Respiratory manifestations are a continuous objection among HIV-contaminated people and might be brought about by a wide range of ailments. The range of pneumonic sicknesses in HIV-tainted patients incorporates both HIV-related and non-HIV-related conditions. The HIV-related aspiratory conditions incorporate both entrepreneurial contaminations (OIs) and neoplasms. The OIs include bacterial, mycobacterial, contagious, viral, and parasitic pathogens. Every one of these OIs and neoplasms has a trademark clinical and radiographic introduction. Nonetheless, there can be impressive variety and cover in these introductions. Thusly, no group of stars of manifestations, physical assessment discoveries, research facility anomalies, and chest radiographic discoveries is pathognomonic or explicit for a specific malady. Thus, a complete microbiologic or pathologic determination is desirable over empiric treatment at whatever point conceivable. Symptomatic tests incorporate societies from sputum and blood and from respiratory examples acquired by obtrusive systems, for example, bronchoscopy, thoracentesis, figured tomography (CT)- guided transthoracic needle goal, thoracoscopy, mediastinoscopy, and open-lung biopsy. This section portrays the recurrence of respiratory side effects, the range of pneumonic diseases that can influence HIV-tainted patients, and an indicative way to deal with the assessment of respiratory manifestations in HIV-contaminated patients, featuring certain parts of the clinical introduction that might be valuable in separating the most widely recognized OIs and neoplasms. The trademark chest radiographic introductions of the most well-known OIs and neoplasms are portrayed and blueprints of 3 case situations are introduced to outline differential judgments and significant analytic and restorative choices for an assortment of clinical and radiographic introductions. For subtleties on explicit indicative tests and treatment regimens for every one of the OIs and neoplasms, see explicit parts inside the HIV InSite Knowledge Base.
Do individual sounds carry meaning? The relationship between sound and meaning in human languages is typicallyassumed to be arbitrary, though recent research provides evidence for the existence of both iconicity and systematicitybetween word forms and their meaning. However, this research has not asked whether individual sounds in a languagecovary in systematic ways with aspects of meaning. In two analyses, we find evidence for more systematicity betweenthe initial phones of words and those words concreteness ratings than one would expect in a truly arbitrary lexicon. Thissuggests that initial phones may act as cues to aspects of word meaning, and raises questions about whether languagelearners detect and exploit these cues.
На иранских языках говорили многочисленные племена и народности, сыгравшие важную роль в мировой истории. К основным иранским языкам относятся персидский, таджикский, дари, афганский (пушту), осетинский, курдский, белуджский и др. Наиболее распространенным и статусным иранским языком в настоящее время является персидский. Предок современного персидского языка древнеперсидский сформировался еще в середине I тыс. до н.э. на территории западной части Иранского нагорья в области Фарс. После подчинения Александром Македонским царства Ахеменидов в нем официальным языком стал греческий, функционировавший долгие столетия, и лишь в III в. н.э. с установлением гегемонии Сасанидов официальным языком в государстве стал персидский. В результате завоевания Ирана арабами в 637-652 гг. н.э. официальное функционирование среднеперсидского языка надолго прекратилось. Официальным языком арабского Халифата стал арабский. Это продолжалось до IX в. С начала X в. началось бурное развитие новоперсидского языка и персидской литературы. В настоящее время персидский является государственным языком большого и многонационального государства Иран. Персы доминирующий этнос в государстве. На персидском происходит обучение и в школах, начиная с 1-го класса. Делопроизводство также осуществляется исключительно на персидском. Другие языки в официальной сфере не используются. Исторически персидский язык оказал огромное влияние не только на иранские, но и на многие тюркские и индийские языки. На базе классического персидского языка сформировались современный персидский, таджикский и дари. Высоким является и статус таджикского языка. Он полностью используется во всех сферах деятельности. Объем научных исследований таджикского языка не уступает персидскому. Дариязычное население проживает в Афганистане. Оно составляет около 40 жителей страны. Афганцы (пуштуны) являются одним из крупнейших ираноязычных этносов. Они живут в Афганистане и Пакистане. В Афганистане язык пушту является официальным наряду с дари. Другой иранский народ осетины живет в центральной части Кавказа по обеим сторонам Главного Кавказского хребта. В результате ассимиляционных процессов численность осетиноговорящего населения имеет тенденцию к сокращению. Крупными иранскими языками являются также курдский и белуджский. Несмотря на многочисленность курдоговорящего населения, этот язык не имеет высокого официального статуса. Лишь в Ираке в курдских районах курдский был объявлен официальным наряду с арабским. Другой крупный иранский язык белуджский ни в одном государстве не имеет официального статуса. Однако белуджи хорошо чувствуют и охраняют языковую норму. The Iranian languages were spoken by numerous tribes and nationalities, which played an important role in world history. The main Iranian languages include Persian, Tajik, Dari, Afghan (Pashto), Ossetian, Kurdish, Balochi, etc. The most common and status Iranian language is currently Persian. The ancestor of the modern Persian language ancient Persian was formed in the middle of the first millennium BC in the western part of the Iranian highlands in the region of Fars. After Alexander the Great had subjugated the Achaemenid kingdom, Greek became an official language there, functioning for centuries, and only in the 3rd c. AD with the establishment of hegemony of the Sassanids, Persian became the official language in the state. As a result of the conquest of Iran by the Arabs in 637-652 AD the official functioning of the Middle Persian language ceased for a long time. The official language of the Arabic Caliphate is Arabic. This continued until the 9th century. The rapid development of the New Persian language and Persian literature began in the early 10th century. It is currently the state language of the large and multinational state of Iran. Persians are the dominant nation in the state. It takes place in schools, starting from the 1st grade. Office work is also carried out exclusively in Persian. Other languages are not used in the official sphere. Historically, the Persian language has had a huge impact not only on Iranian, but also on many Turkic and Indian. On the basis of the classical Persian language, modern Persian, Tajik and Dari were formed. The status of the Tajik language is also high. It is fully used in all fields of activity. The volume of scientific studies of the Tajik language is not inferior to the Persian. Daria-speaking population lives in Afghanistan. It makes up about 40 of the countrys population. Afghans (Pashtuns) are one of the largest Iranian-speaking ethnic groups. They live in Afghanistan and Pakistan. In Afghanistan, the Pashto language is an official language alongside with Dari. Another Iranian people, the Ossetians, live in the central part of the Caucasus, on both sides of the Main Caucasian Range. As a result of assimilation processes, the Ossetian-speaking population tends to decline. The major Iranian languages are also Kurdish and Balochi. Despite the large Kurdish-speaking population, this language does not have a high official status. Only in Iraq in Kurdish areas Kurdish was declared official alongside with Arabic. The other major Iranian language, Balochi, has no official status in any state. However, the Baluchis feel and properly preserve the linguistic norm.
This paper is concerned with whether deep syntactic information can help surface parsing, with a particular focus on empty categories. We consider data-driven dependency parsing with both linear and neural disambiguation models. We find that the information about empty categories is helpful to reduce the approximation error in a structured prediction based parsing model, but increases the search space for inference and accordingly the estimation error. To deal with structure-based overfitting, we propose to integrate disambiguation models with and without empty elements. Experiments on English and Chinese TreeBanks indicate that incorporating empty elements consistently improves surface parsing.
Neural models have been investigated for sentiment classification over constituent trees. They learn phrase composition automatically by encoding tree structures but do not explicitly model sentiment composition, which requires to encode sentiment class labels. To this end, we investigate two formalisms with deep sentiment representations that capture sentiment subtype expressions by latent variables and Gaussian mixture vectors, respectively. Experiments on Stanford Sentiment Treebank (SST) show the effectiveness of sentiment grammar over vanilla neural encoders. Using ELMo embeddings, our method gives the best results on this benchmark.
Previous research has shown that that evaluative verbal information (praise and criticism) conveys different affective values: criticism is perceived as unpleasant while praise is generally considered pleasant. Here, using praise and criticism in Chinese, we investigated how affective value is modulated in men and women, depending on the particular attribute (personality vs. appearance) targeted by social comments. Results showed that whereas praise was rated as pleasant and criticism as unpleasant overall, criticizing personality reduced pleasantness more than criticizing appearance. In men, moreover, criticism of personality was deemed more unpleasant than criticism of appearance while personality-targeted praise was rated more pleasant than appearance-targeted praise. This effect was absent in women and consistent with men's higher arousal ratings for personality- relative to appearance-targeted comments. Our findings suggest that men are more concerned about external perception of their personality than that of their appearance whereas women's affective judgment is more balanced. These gender-specific results may have implications for topic selection in evaluative social communication.
Throughout many studies which focus on brain laterality as a key component to the outcome of an experiment, or to a participants’ reaction to a stimulus, it can be noted that the different areas of the brain involved in a task response must work together to produce a viable outcome (i.e. lateralized brain processes). While there are laterality components in relation to the cognitive processes of handed and footed responses, it is still largely unknown how the different areas of the brain interpret the emotional stimuli to then affect these outcomes. The purpose of our study is to determine how emotional context affects a simple cognitive task that includes handed and footed responses, and if any observed differences can be traced back to the different systems at work within the brain. Subjects will be tested over two days for handed and footed responses in a cognitive Simon Task. Subjects will be tested with and without emotional context (i.e. a background image of a specific valence and arousal rating), and any resulting differences between non-emotional context and emotional context reaction times will be compared.
In sequence learning tasks such as language modelling, Recurrent Neural\nNetworks must learn relationships between input features separated by time.\nState of the art models such as LSTM and Transformer are trained by\nbackpropagation of losses into prior hidden states and inputs held in memory.\nThis allows gradients to flow from present to past and effectively learn with\nperfect hindsight, but at a significant memory cost. In this paper we show that\nit is possible to train high performance recurrent networks using information\nthat is local in time, and thereby achieve a significantly reduced memory\nfootprint. We describe a predictive autoencoder called bRSM featuring recurrent\nconnections, sparse activations, and a boosting rule for improved cell\nutilization. The architecture demonstrates near optimal performance on a\nnon-deterministic (stochastic) partially-observable sequence learning task\nconsisting of high-Markov-order sequences of MNIST digits. We find that this\nmodel learns these sequences faster and more completely than an LSTM, and offer\nseveral possible explanations why the LSTM architecture might struggle with the\npartially observable sequence structure in this task. We also apply our model\nto a next word prediction task on the Penn Treebank (PTB) dataset. We show that\na 'flattened' RSM network, when paired with a modern semantic word embedding\nand the addition of boosting, achieves 103.5 PPL (a 20-point improvement over\nthe best N-gram models), beating ordinary RNNs trained with BPTT and\napproaching the scores of early LSTM implementations. This work provides\nencouraging evidence that strong results on challenging tasks such as language\nmodelling may be possible using less memory intensive, biologically-plausible\ntraining regimes.\n
The emergence of deep learning as a commanding technique for learning heterogeneous layers of feature representations have consequently substituted traditional machine learning algorithms which are generally poor in analyzing compound sentences. Additionally, convolutional and recurrent neural networks have auspiciously yielded state-of-the-art results in sentiment classification and Natural Language Processing (NLP). In this paper, a deep sentiment representation model through the combination of multiple Convolutional Neural Networks (CNN) kernels with Long Short-Term Memory (LSTM) is proposed for sentiment classification. Our model gains word vector representation using pre-trained Global Vectors for Word Representation (GloVe) embeddings, thereafter used as input to the CNN layer which extracts higher local text representations. Finally, Bidirectional LSTM (biLSTM) generates sentiment classification of sentence representation based on context dependent features. Our combined approach of CNN and biLSTM was experimented using the Stanford Large Movie Review Dataset (IMDB) and Stanford Sentiment Treebank Dataset (SSTB) for binary classification. The evaluation achieves outstanding results in outperforming several existing approaches with 90.4% accuracy on the Stanford Sentiment Treebank dataset and 94.8% accuracy on the Stanford Large Movie Review dataset. These results are achieved with a drastic reduction of model parameters and without a pooling layer in the CNN architecture, helping to retain local and structural information in comparison to other existing deep neural network frameworks.
The paper explores the »power« names have over historians: they influence historical interpretations and shape academic premises. This article uses the historio-linguistic database »Nomen et gens« (NeG) to investigate the interpretation of early medieval names as historical sources. Its objective is twofold: the inquiry aims to revise the interpretation of names as ethnic markers via the case study of the so called »Anglo-Saxon mission« and thereby to present NeG and its possibilities to the medievalist community. The database NeG is a result of the eponymous research project (www.neg.uni-tuebingen.de), reaching back to the first dawns of digital humanities in the 1970s. Its origins are set in the fierce debate on ethnic identity of that time, addinganonomastic dimension to this discussion: Do specific names indicate specific ethnicities? While originally intended to provide a sound data stock for ethnic interpretations of names, NeG quickly became part of the academic deconstruction of ethnicity as an analytical tool. Today it provides a unique – and growing – data of more than 17.250 single references to names in early medieval sources between 400– 800, assigned to 6.500 historical persons, representing 3.100 early medieval names. The paper starts with the original question of NeG concerning names and ethnicity. A first result is predictable in light of the outcome of the debate on ethnicity: ethnic labels like e. g. »goth« or »franc« are attributed only to a very small number of persons, in total 273 out of 6.528. The impact of this at first glance quite unsurprising result is shown by a reconsideration of the so called »Anglo-Saxon mission«: while seen as one of the driving forces of Frankish history in the 8th century, this »mission« is in urgent need of a reevaluation. First of all Anglo-Saxon identity is a construct of modern scholarship, as has been stressed in the last decades. Moreover, the extensive data provided by NeG comprises only 20 persons that were denominated with one of those labels usually perceived as Anglo-Saxon like »Angles« or – even more specific – »Mercians« or »Bernicians«. No more than 8 of these persons were clerics acting in the eastern borderlands of Frankish power in the 8th century, and only 3 of them did use names that would philologically be described as anglic. Only an insignificant small number of »Anglo-Saxons« in the Frankish world can thus be identified. Apparently such an attribution was of very little interest to contemporary authors when writing about single individuals. Ethnic labels like »the Franks« did matter ona higher level of group identities. But historians should cease assigning ethnic labels to persons based on their names and should stop drawing conclusion from these ascriptions. Just as the »Anglo-Saxon« mission was no mission in its most essential meaning, it was not »Anglo-Saxon«. A re-evaluation could shed new light on the phenomenons of cultural entanglement in the 8th century.
OBJECTIVE: Biased attention for disorder-relevant information plays a crucial role in the maintenance of different mental disorders including eating disorders and might be of use to define recovery beyond symptom-related criteria. METHOD: We assessed attention deployment using eye tracking in a cued choice viewing paradigm to two different categories of disorder-relevant stimuli in 24 individuals with acute anorexia nervosa (AN), 20 weight-recovered individuals with a history of AN (WRAN) and 23 healthy control participants (CG). Picture pairs consisted of a food stimulus or a picture depicting physical activity and a matched control stimulus (household item/physical inactivity). Participants rated the valence of stimuli afterwards. RESULTS: The groups did not differ in initial attention deployment. In later processing stages, AN patients showed a generalized attentional avoidance of food and control pictures as compared to CG, while WRAN individuals were in between. AN patients showed an attentional bias toward physical activity pictures as compared to WRAN individuals, but not the CG. AN individuals rated the food pictures and the pictures showing physical inactivity as less pleasant than the CG, while WRAN individuals were in between. DISCUSSION: Attention deployment is partly changed in WRAN as compared to the acute AN group, especially with regard to a shift away from illness-compatible stimuli (physical activity), and this might be a useful recovery criterion. Valence rating of food stimuli might be an additional useful tool to distinguish between acutely ill and weight-recovered individuals. Attentional biases for illness-compatible stimuli might qualify as a valuable approach to defining recovery in AN.
This thesis studies the connections between parsing friendly representations and interlingua grammars developed for multilingual language generation. Parsing friendly representations refer to dependency tree representations that can be used for robust, accurate and scalable analysis of natural language text. Shared multilingual abstractions are central to both these representations. Universal Dependencies (UD) is a framework to develop cross-lingual representations, using dependency trees for multlingual representations. Similarly, Grammatical Framework (GF) is a framework for interlingual grammars, used to derive abstract syntax trees (ASTs) corresponding to sentences. The first half of this thesis explores the connections between the representations behind these two multilingual abstractions. The first study presents a conversion method from abstract syntax trees (ASTs) to dependency trees and present the mapping between the two abstractions – GF and UD – by applying the conversion from ASTs to UD. Experiments show that there is a lot of similarity behind these two abstractions and our method is used to bootstrap parallel UD treebanks for 31 languages. In the second study, we study the inverse problem i.e. converting UD trees to ASTs. This is motivated with the goal of helping GF-based interlingual translation by using dependency parsers as a robust front end instead of the parser used in GF. \n\nThe second half of this thesis focuses on the topic of data augmentation for parsing – specifically using grammar-based backends for aiding in dependency parsing. We propose a generic method to generate synthetic UD treebanks using interlingua grammars and the methods developed in the first half. Results show that these synthetic treebanks are an alternative to develop parsing models, especially for under-resourced languages without much resources. This study is followed up by another study on out-of-vocabulary words (OOVs) – a more focused problem in parsing. OOVs pose an interesting problem in parser development and the method we present in this paper is a generic simplification that can act as a drop-in replacement for any symbolic parser. Our idea of replacing unknown words with known, similar words results in small but significant improvements in experiments using two parsers and for a range of 7 languages.
Discourse Relations, also known as coherence or rhetorical relations, characterize the semantic or pragmatic relationships between clauses or sentences in discourse.Such relations are established in order to facilitate effective communication.In addition to the inventory of relations, previous research has also investigated how discourse relations are established or signaled.Discourse markers (DMs) are considered to be the most typical signals in discourse; however, focusing merely on DMs is inadequate as they can only account for a small number of relations in discourse.Thus, researchers have been exploring textual signals beyond DMs such as the Penn Discourse Treebank 2.0 (PDTB, Prasad et al. [22]) and the Rhetorical Structure Theory Signalling Corpus (RST-SC, Das and Taboada [5]).Despite their different theoretical groundings and approaches to relation signaling, both corpora annotated the Wall Street Journal (WSJ) section of the Penn Treebank (PTB, Marcus et al. [19]), i.e. the news articles.Nevertheless, previous work has suggested that signaling information is indicative of genres (e.g.Taboada and Lavid [28]; Zeldes [34]).Therefore, this project aims to anchor signaling devices on a more diverse corpus to demonstrate the inadequacy of signaling by DMs only, the abundance of open-class signals, and more importantly, the distribution of signaling devices across genres.
This paper suggests annotation guidelines to build a Universal Dependencies (UD) treebank for Korean. We discuss the part-of-speech annotation of Korean specific-categories such as prenouns, numeral classifiers, and (pre)final endings, and propose how to implement UD scheme in Korean regarding selecting a head and assigning dependency relations to dependents. UD prioritizes content words over functional words since the former exhibits less cross-linguistic variations. In a noun phrase, for instance, a core noun is always a head of the entire noun phrase independently of a language. The rest are treated as a dependent: not only a modifier such as an adjective but also a functional category such as an article, numeral quantifier, demonstrative, and so on. However, when it comes to head-less constructions such as coordination or predicate ellipsis, UD firmly advocates the head-initial strategy. The present application of UD to Korean tries to follow UD’s principles as much as possible. Korean is a head-final language, so that headed constructions are analyzed head-finally. In contrast, head-less ones are tagged head-initially. This might disregard language-specific characteristics from a linguistic perspective, but the strategy allows us to build up a set of treebanks in a cross-linguistically consistent way (i.e., the fundamental purpose of UD).
This repository contains code for reproducing experiments done in Marasovic and Frank (2018).<p> <p> <b>Paper abstract:</b><p> For over a decade, machine learning has been used to extract opinion-holder-target structures from text to answer the question "Who expressed what kind of sentiment towards what?". Recent neural approaches do not outperform the state-of-the-art feature-based models for Opinion Role Labeling (ORL). We suspect this is due to the scarcity of labeled training data and address this issue using different multi-task learning (MTL) techniques with a related task which has substantially more data, i.e. Semantic Role Labeling (SRL). We show that two MTL models improve significantly over the single-task model for labeling of both holders and targets, on the development and the test sets. We found that the vanilla MTL model, which makes predictions using only shared ORL and SRL features, performs the best. With deeper analysis, we determine what works and what might be done to make further improvements for ORL.<p> <p> <b>Data for ORL</b></p> <ul> <li>Download MPQA 2.0 corpus. <li>Check mpqa2-pytools for example usage. <li>Splits can be found in the datasplit folder. </ul> <p> <b>Data for SRL</b><p> The data is provided by: CoNLL-2005 Shared Task, but the original words are from the Penn Treebank dataset, which is not publicly available.<p> <b>How to train models?</b><p> <p> python main.py --adv_coef 0.0 --model fs --exp_setup_id new --n_layers_orl 0 --begin_fold 0 --end_fold 4<p> <p> python main.py --adv_coef 0.0 --model html --exp_setup_id new --n_layers_orl 1 --n_layers_shared 2 --begin_fold 0 --end_fold 4<p> <p> python main.py --adv_coef 0.0 --model sp --exp_setup_id new --n_layers_orl 3 --begin_fold 0 --end_fold 4<p> <p> python main.py --adv_coef 0.1 --model asp --exp_setup_id prior --n_layers_orl 3 --begin_fold 0 --end_fold 10<p>
Modern machine learning (ML) techniques are transforming many disciplines ranging from transportation to healthcare by uncovering patterns in data, developing autonomous systems that mimic human abilities, and supporting human decision-making. Modern ML techniques, such as deep neural networks, are fueling the rapid developments in artificial intelligence. Engineering design researchers have increasingly used and developed ML techniques to support a wide range of activities from preference modeling to uncertainty quantification in high-dimensional design optimization problems. This special issue brings together fundamental scientific contributions across these areas.The special issue consists of 24 papers spread over two issues of the Journal of Mechanical Design. The papers use various ML techniques, including artificial neural networks, Gaussian processes, reinforcement learning, clustering techniques, and natural language processing. Based on their research objective, the papers can be broadly classified into four groups: (i) ML to support surrogate modeling, design exploration, and optimization, (ii) ML for design synthesis, (iii) ML for extracting human preferences and design strategies, and (iv) comparative studies of ML techniques and research platforms to help design researchers. The papers are summarized in Secs. 1–4. An analysis of the themes covered in the special issue and the potential opportunities for future research in ML for Engineering Design are presented in Sec. 5.In the paper titled Multifidelity Physics-Constrained Neural Network and Its Application in Materials Modeling, Liu and Yang address how to incorporate multifidelity, physics-based constraints into neural network predictions. The paper contributes two key insights. First, the paper extends existing Physics-Constraints Neural Network architectures by imposing a multifidelity constraint scheme wherein an auxiliary network minimizes discrepancies between low and high fidelity models—essentially learning how to correct the low-fidelity one. Second, it proposes an adaptive weighting scheme to control the convergence of individual losses among the different fidelities. They demonstrate the impact of these improvements on several fundamental multiscale material modeling challenges including two-dimensional heat transfer, phase transition, and dendritic growth problems. On these problems, the proposed multifidelity, physics-based constraints decrease the prediction error up to order of magnitude compared with networks without such constraints. This achieves comparable accuracy to that of direct numerical solutions of the underlying equations.Sarkar et al. present a multifidelity modeling and information-theoretic sequential sampling strategy for optimization in their paper titled Multifidelity and Multiscale Bayesian Framework for High-Dimensional Engineering Design and Calibration. The approach is based on modeling of the varied fidelity information sources via Gaussian processes, augmented with efficient active learning strategies that involve sequential selection of optimal points in a multiscale architecture. The strategy is demonstrated using the design optimization of a compressor rotor and calibration of a microstructure prediction model.In the paper titled A Case Study of Deep Reinforcement Learning for Engineering Design: Application to Microfluidic Devices for Flow Sculpting, Lee et al. address how to design micro-fluidic flow sculpting devices by overcoming some of the key weaknesses of evolutionary optimization-based methods, namely, poor sample efficiency and slow optimization convergence. The paper adapts deep reinforcement learning (DRL) techniques to the flow sculpting task and also studies the effectiveness of transfer learning on accelerating the design of target flow shapes. The paper demonstrates that DRL is able to match 90% of the target flow shapes using significantly fewer sculpting pillars than comparable GA models as well as provides a means to interpret the learned model (using Principal Components) that existing approaches to fluidic sculpting do not provide.Lynch et al., in their paper Machine Learning to Aid Tuning of Numerical Parameters in Topology Optimization, present an ML-based meta-learning framework to determine tuning parameters in topology optimization. The parameters are learned from similar optimization problems carried out in the past and adjusted for the problem at hand. This helps in avoiding costly trial-and-error involved in manual parameter tuning.In the paper Data-Driven Design Space Exploration and Exploitation for Design for Additive Manufacturing, Xiong et al. present a data-driven approach for design search and optimization at successive stages in the design process. They use Bayesian network classifier in the embodiment design stage and Gaussian process regression in the detailed design phase. The approach is illustrated in the paper through the design of a customized ankle brace design.Odonkor and Lewis apply data-driven design to the design of operational strategies of complex systems, specifically distributed energy resources. The paper is titled Data-Driven Design of Control Strategies for Distributed Energy Systems. The problem of maximizing arbitrage value is formulated as an optimization problem and solved using reinforcement learning. The approach is demonstrated for shared distributed energy resources in multi-building residential clusters.In Globally Approximate Gaussian Processes for Big Data With Application to Data-Driven Metamaterials Design by Bostanabad et al., a globally approximate Gaussian process (GAGP) is introduced for the purpose of handling large datasets. A GAGP is constructed by pooling several Gaussian processes using identical hyperparameters but built from different subsets of the training data. The predictive capability of GAGPs is shown to be at least as good as state-of-the-art supervised learning methods. It is demonstrated on the unit-cell design of metamaterials through inverse optimization.Liu et al. present a method for the design for crashworthiness involving categorical multimaterial structures in their paper titled Design for Crashworthiness of Categorical Multimaterial Structures Using Cluster Analysis and Bayesian Optimization. Following a topology optimization, the dimensionality of the problem is reduced through clustering followed by a Bayesian optimization to assign a given material to a specific cluster. The approach is applied to the maximization of absorbed energy of an S-rail.Garriga et al. propose a framework to assist the optimization of aircraft systems at the early design stages. The approach in their paper titled A Machine Learning Enabled Multifidelity Platform for the Integrated Design of Aircraft Systems is based on the screening of designs using clustering followed by an identification of the best candidate on a Pareto front. The framework enables the use of models of various fidelities and is demonstrated on a primary flight control system and a landing gear.In Synthesizing Designs With Interpart Dependencies Using Hierarchical Generative Adversarial Networks, Chen and Fuge present a method for synthesizing hierarchical designs with inter-part dependencies using generative models learned from examples. The method constructs multiple generative models using generative adversarial networks (GANs) while satisfying the dependencies through part dependency graphs. The paper lays the foundation for extending the use of generative models from creative individual parts to more realistic engineering systems.The objective in Evolving a Psycho-Physical Distance Metric for Generative Design Exploration of Diverse Shapes by Khan et al. is to incorporate humans’ psychological perceptions about design into the design exploration process. A psycho-physical distance metric is proposed that enables the augmentation of CAD designs based on feedback from users. Results reveal that the proposed method generates more distinct variations of CAD designs compared with a baseline Euclidean distance method.Oh et al. in their paper titled Deep Generative Design: Integration of Topology Optimization and Generative Models present a design framework for creating diverse aesthetic designs that are optimized for engineering performance. The framework integrates topology optimization and generative adversarial networks (GANs) to generate large numbers of design options from limited previous design data. The approach is validated using a 2D wheel design problem.Deshpande and Purwar, in their paper Computational Creativity Via Assisted Variational Synthesis of Mechanisms Using Deep Generative Models, present an approach for variational synthesis of mechanisms and an End-to-End synthesis pipeline that accepts raw, high-level input from users and provides them with distinct concept solutions. The approach is based on learning the probability distribution of linkage parameters and their interdependence to perform tasks such as input conditioning, imputation, and variational synthesis. The approach is a step in the direction of enhancing users’ computational creativity for engineering design.Stump et al. in their paper, Spatial Grammar-Based Recurrent Neural Network for Design Form and Behavior Optimization, present a method for simultaneous optimization of form and behavior through a combination of physics-based models and ML techniques. Specifically, they use character-Recurrent Neural Networks to embody spatial grammars and reinforcement learning to optimize the behavior. The design of a modular multi-hull sailing craft is used as a demonstration problem.Suryadi and Kim utilize machine-learning algorithms for customer choice modeling in their paper titled A Data-Driven Methodology to Construct Customer Choice Sets Using Online Data and Customer Reviews. They present an approach that utilizes publicly available online data and customer reviews from e-commerce websites to construct customer choice sets in the absence of both an actual choice set and customer sociodemographic data. The approach consists of clustering (i) products based on their attributes and (ii) customers based on their reviews, and constructing the choice-sets based on a sampling probability scenario that relies on product and customer clusters. The approach generates choice models with higher predictive ability than randomly constructed choice sets.In their paper Extracting Customer Perceptions of Product Sustainability From Online Reviews, El Dehaibi et al. seek to extract perceived sustainable design features from online reviews. Annotators from Amazon’s Mechanical Turk are used to annotate product reviews and develop a natural language processing model that predicts the positive/negative sentiment of sustainable phrases. The results reveal that the model is more efficient at predicting positive sentiment pertaining to sustainable product features compared with negative sentiments.Raina et al. take a step toward transfer learning from human designers to computational agents in their paper titled Transferring Design Strategies From Human To Computer and Across Design Problems. They present an approach where design strategies are represented using a probabilistic model that provides a general mechanism to transfer strategies from human designers to computational design agents and to generate new designs. The approach is illustrated using a configuration design problem.The goal in Learning to Design From Humans: Imitating Human Designers Through Deep Learning by Raina et al. is to teach computational agents to generate designs without the need for explicit information about objective or performance metrics. A deep learning model is proposed that learns from historical human data and identifies the important regions of a design space. The results reveal that the machine learning agent learns to create designs that are comparable to human-generated ones, despite not having the same explicit feedback that humans do to guide them through the design exploration process.He et al. address the challenge of mining large numbers of design ideas generated from the crowd in their paper titled Mining and Representing the Concept Space of Existing Ideas for Directed Ideation. The authors use natural language processing to extract keywords as elementary concepts and represent the concepts in a way that they can be recombined to generate new ideas.In the paper titled A Data-Driven Approach to Product Usage Context Identification From Online Customer Reviews, Suryadi and Kim use machine learning and natural language processing to identify and cluster usage contexts from a large volume of customer reviews. The methodology also captures sentiments toward a particular usage context in a sentence. The methodology enables designers to effectively use online product reviews by focusing on several specific reviews regarding particular usage contexts and potentially to identify market opportunities for new products that excel in specific usage contexts.Sharpe et al. illuminate differences between Supervised Learning algorithms in terms of how and where different algorithms may apply to different Engineering Design applications. Their paper titled A Comparative Evaluation of Supervised Machine Learning Classification Techniques for Engineering Design Applications does this by comparing four common supervised learning approaches—Support Vector Machines, Random Forests, Gaussian Näive Bayes, and shallow depth Neural Networks—across six example problems that demonstrate different facets or challenges classifiers may face within the engineering design. The results from the work are multifaceted with different algorithms performing better or worse under different conditions and performance measures. However, this leads to the general notion of strong problem dependence for the classifier choice and highlights the importance of understanding appropriate benchmark problems within the engineering design that can shed light on such issues in the future.The availability of data enables not just designers but also design researchers. Rahman et al., in their paper A Computer-Aided Design Based Research Platform for Design Thinking Studies, present a research platform to support data-driven design-thinking and decision-making research. Through the use of fine-grained design action data and unsupervised clustering methods in conjunction with design process models, the authors show how the platform enables data-driven research studies on designers’ sequential decision-making behaviors.In Design Repository Effectiveness for 3D Convolutional Neural Networks: Application to Additive Manufacturing, Williams et al. address the question of whether or not a data repository is useful for training effective ML tools. The authors experimentally test the effects of changes in CAD datasets on the precision and generalizability of trained convolutional neural networks (CNNs) for additive manufacturing applications. The study sheds light on how standardization of design repositories can influence the performance of ML tools.Cunningham et al. study the construction of a performance surrogate based on 3D point cloud representations. In their paper titled An Investigation of Surrogate Models for Efficient Performance-Based Decoding of 3D Point Clouds, a radial basis function (RBF) surrogate is used to link performance and cloud representation mapped onto a latent vector. The proposed RBF-based approach was found to be more efficient and accurate than traditional neural network-based approaches.In the call for proposals for this special issue, the guest editors posed three primary questions: How to effectively use ML for new design applications that are not well-supported by existing ML practice or tools?How to leverage the unique aspects of engineering design in creating new ML approaches?How to share benchmark problems or datasets that can measure ML progress in design?Looking back at the papers collectively within this Special Issue helps shed light on the areas that are receiving significant within the design research and the areas where are opportunities for future design research a strong of using machine learning. techniques have used for data-driven techniques such as Gaussian process regression and neural networks have applications in learning complex between design and performance natural language processing used for mining customer and reinforcement learning is an part of control systems design. of these techniques are in the papers for the special issue Bostanabad et et Suryadi and Xiong et of the deep learning techniques such as convolutional neural networks (CNNs) and generative adversarial networks (GANs) are their way into engineering design research and and and et papers in this special issue address diverse applications including modeling, additive distributed energy systems, and topology optimization, synthesis, mechanism preference modeling, and learning from human of the areas where is potential for ML techniques, but are not well represented in this special issue, modeling human design of market systems, of products with humans or the design for use of data from product usage to using data from of the product or data, or across different aspects of the are opportunities for design challenges in engineering such as and and that from the products and ML can also be used to support engineering design for supporting studies and the generalizability of research terms of understanding and of machine learning methods, the papers in the special issue on the the of multifidelity or multiple data or with the common surrogate modeling and physics-based constraints into ML and ML-based models of human preferences and of the papers in the special issue present surrogate modeling this an active of research for the past two the approaches are using ML techniques. A of papers issues that ML models in multifidelity and et The availability of multifidelity models is in engineering design. In some designers’ are to construct multiple and fidelities of models or of a system as to with it at the appropriate This is not that ML systems are to In this of do multiple or important to to this special issue and is an active of of the approaches on existing with changes that constraints. of the is that many ML systems are not used to many of the problems that need for engineering design. in et al., how the between form and behavior the generated or in and Lewis the between Control and Design these of between and within Engineering and existing ML approaches do not need to for papers the of design or constraints in a constraints by or and hierarchical or et Chen and and constraints and This is important the of many ML systems relies on their to but specific of into the This is where engineering design researchers are well to are opportunities for research in engineering with ML models, such as physics-based models from or via more system models, and studies a problem and as a In the can ML be used with a between design and A is learning from multiple of a or learning among design with multiple data structures and of the papers to this special issue used supervised or reinforcement learning where data or available via used Learning clustering and dimensionality in natural and have from unsupervised learning of using deep learning or and this be a for future work in design. are opportunities for approaches to or for in uncertainty with or data. and calibration of ML models an papers on the of ML for optimization and specifically how to about generative models of two of approach to on the of the the approach Optimization as as by approaches that used or latent methods et al. and et or Optimization as such as that formulated the problem using reinforcement learning or inverse problem et are active and areas of in to for optimization that not in this special issue, such as direct inverse design papers how to best or behavior into an ML This is a are many opportunities for human information and strategies into ML models, and models of and computational or for behavior. This to approaches have used to models into the same be for of models, such as involving human or to ML models of human behavior designers or the design research on in both and and this is within the special However, are constraints that engineering design on of a notion that be to function and to that out as a of approaches to do not such aspects and this a that the Engineering Design can This also for are the fundamental of designers as conditions and does this datasets have ML research and a common for performance. collectively the the among have collectively of of They have in and This special issue set out to to new datasets for engineering design papers in the special issue propose datasets or platforms optimization problems et 2D and 3D shapes et Chen and and manufacturing et while not datasets do use platforms such as to data or these papers within the Special is a wide to be by future work that can useful datasets or for machine learning within engineering design. are some of the areas that papers in the Special Issue and have used ML but for good benchmark datasets datasets for human or including how designers or for more complex than optimization do not have the the model of for Engineering problems such as or that link or use multiple of design representations. datasets a CAD and a by the do not datasets have ML approaches in in the have in for and and datasets for approaches to specific in do not have a good of engineering problems and a given is to This is not unique to Engineering Design. How do or the and that it a realistic benchmark for Computer In some to be is the design on of or the of where and design datasets apply and how their results to datasets and understanding of their or be for future of how to such models within Engineering Design. how and of ML models within this on the of or ML models influence the design of a How does this Design are some of the and the papers in this special issue the of research opportunities in this The papers in the special issue many points from future researchers and may set guest the for their in the and feedback to the Special to and for their help with the for the process of of this special issue, and for the
The topic of the talk and its classification are one of the central issues of syntax. This article compares the classification of the Arabic alphabet with the Arabic and Uzbek linguistic norms. In terms of the stylistics of the Uzbek language, it is explained in terms of how the spelling of the Arabic word begins, and the classification of the Muslim in terms of the context. It is emphasized in Maonic science that the most important aspect of non- speaking in other languages, especially in the Arabian minority, is the purpose of the speaker and the state of the listener. In Maonical Science there is information on classification in relation to reality, the goal of the speaker and the status of the listener, and in the so-called interpreter, to be a change in reality, and to choose the types of speech.
In opera theatres, a speech consultant assists singers in pronouncing foreign languages, as well as helping them with the pronunciation of Slovenian. A sung opera text needs to be intelligible and linguistically standardised. The articulation of the sounds of Slovenian literary language in opera singing depends on the linguistic norm and on musical laws. The opera theatre speech consultant is involved in the process of creating a performance from the very beginning. Detailed knowledge of the text, as well as of music and directing concepts, also enables the consultant to proofread the texts in the theatre programme in a quality way and to ensure that the surtitles correspond to the stage action.
This paper presents a case study of the use of the NINJAL Parsed Corpus of Modern Japanese (NPCMJ) for syntactic research. NPCMJ is the first phrase structure-based treebank for Japanese that is specifically designed for application in linguistic (in addition to NLP) research. After discussing some basic methodological issues pertaining to the use of treebanks for theoretical linguistics research, we introduce our case study on the status of the Coordinate Structure Constraint (CSC) in Japanese, showing that NPCMJ enables us to easily retrieve examples that support one of the key claims of Kubota and Lee (2015): that the CSC should be viewed as a pragmatic, rather than a syntactic constraint. The corpus-based study we conducted moreover revealed a previously unnoticed tendency that was highly relevant for further clarifying the principles governing the empirical data in question. We conclude the paper by briefly discussing some further methodological issues brought up by our case study pertaining to the relationship between linguistic research and corpus development.
This chapter aims at presenting different strategies that have been designed to incorporate multiword expression (MWE) identification in the process of syntactic parsing using statistical approaches. We discuss MWE representation in treebanks, pipeline and joint orchestrations, the integration of external lexicons and the evaluation of MWE-aware parsers, concluding with our suggestions for future research.
The goal of this study is two-fold. First, it will reveal to what extent differences in amount of experience with a particular register manifest themselves in different familiarity judgments when faced with word sequences that are characteristic of that register. To this end, three groups of participants –recruiters, job-seekers, and people not (yet) looking for a job– performed a metalinguistic judgment task in which they assigned familiarity ratings to two sets of stimuli – word sequences characteristic of either job ads or news reports. As the three groups differ in experience in the domain of job hunting, they are likely to differ in experience with collocations that are typically used in that domain. According to usage-based theories, these differences in experience lead to differences in mental representations of language. This leads to a testable hypothesis: If familiarity judgments give expression to linguistic representations, the ratings should reflect these differences. That is, the Job ad stimuli ought to be most familiar to the Recruiters and least familiar to the Inexperienced participants. Subsequently, we examined the relationship between metalinguistic judgments and other types of experimental data. The stimuli that were presented in the judgment task have also been used in two other experiments conducted among the same participants: a Completion task and a Voice Onset Time experiment (both described in Verhagen et al, 2019). By analyzing the judgment data in relation to the participants’ Completion task responses, their voice onset times, and corpus-based frequencies, we can answer the second research question: To what extent do someone’s own data from psycholinguistic processing tasks have explanatory power in predicting familiarity judgments in addition to corpus frequencies? If the different types of tasks tap into the same mental representations, one’s performance in the processing tasks should be a significant predictor of one’s familiarity ratings. If it does not prove to be a significant predictor, this means that there are substantial differences between the tasks in the information they provide.
Abstract Normative and cognitive-linguistic accounts of linguistic meaning are often portrayed and conceived as mutually exclusive alternatives. This dichotomy stems from an insufficient understanding of what the phenomenological accessibility of meaning and usage-basedness of language entail. Namely, the theoretical premises of Cognitive Linguistics actually presuppose socially grounded, normative linguistic meanings. The question remains, what kind of entities normative meanings are like. The present chapter makes a case for construal, linguistic perspective-taking usually analyzed as a conceptual phenomenon, as a normative facet of meaning. Analysis presented here suggests that construal emerges as an inherent property of linguistic expressions via conventionalization of intentionality. This analysis does not only expand the area of linguistic normativity but also points to the integral relation between linguistic norms and intentionality.
This work is partly supported by the Asian-MT “Network-based ASEAN Languages Translation Public Service” and the ASEAN IVO Project “Open Collaboration for Developing and Using Asian Language Treebank”.
The reader is typeset (see LaTeX / XeLaTeX source), many errors are corrected. The book is submitted to review at the University of Zagreb. Other resources (treebanks, grammatical annotations) are still in development.
Universal Dependencies Treebank is a cross-linguistic project to annotate Parts-of-Speech and dependency relations universally. The same 17 universal Part-of-Speech tags and 37 universal dependency-relation tags are used for the annotation universally across all languages. However, in fact, each UD treebank is developed by each developer of each language, reflecting “not-universal” treatments of the language. In this paper, we reveal the difference among Classical Chinese UD, Modern Japanese UD, and Modern English UD, upon parallel corpora of 大學(The Great Learning).
Space-valence metaphors (e.g., bad is down) are embedded within cognitive and emotional processing (e.g., negative stimuli at a lower space capture visual attention more than those at an upper space). Previous studies have revealed that motor action to vertical direction affects the emotional valence rating of stimuli in a metaphor-congruent manner only when the action was introduced after the stimuli presentation. In the present study, we hypothesized that motor action before the stimuli presentation does not affect valence rating while it may affect visual selective attention. In Experiment 1 (participants: 28 university students; mean age = 19.50 years), we partially replicated the previous result with repeated ANOVA and t-tests; manual action introduced before the stimuli presentation does not affect the valence rating. Then, in Experiment 2 (participants: 28 university students; mean age = 19.57 years), we employed a modified version of the dot-probe task as a measure of visual selective attention to emotional stimuli, where participants’ vertical or horizontal manual action was introduced before the presentation of a pair of emotional words. The results of the t-tests revealed that an upward manual action promoting selective attention to negative words, which was incongruent with the space-valence metaphorical correspondence. These results suggest that even though manual action does not affect the evaluative process of emotional stimuli prospectively, upward manual action introduced before stimuli presentation can promote visual attention to the subsequent negative stimuli in a way that is incongruent with the space-valence metaphor.
Treebank is one of the important and useful resources in natural language processing represented in two different annotated schemas: phrase and dependency structures. There are many works that convert a phrase structure into a dependency structure and vice versa. Most of them are based that exploit the handcrafted head percolation table and argument table in predefined deterministic ways. In this article, we propose a method to convert a dependency structure into a phrase structure by enriching a trainable model of former hybrid strategy approach. By adding a classifier to the algorithm and using postprocessing modification, the quality of conversion is increased. We evaluate our method in two different languages, English and Persian, and then analyze the errors. The results of our experiments show a 46.01% reduction of error rate in English and 76.50% for Persian compared to our baseline. We build a new phrase structure treebank by converting 10,000 sentences of Persian dependency treebank into corresponding phrase structures and correcting them manually.
Most syntactic dependency parsing models may fall into one of two categories: transition- and graph-based models. The former models enjoy high inference efficiency with linear time complexity, but they rely on the stacking or re-ranking of partially-built parse trees to build a complete parse tree and are stuck with slower training for the necessity of dynamic oracle training. The latter, graph-based models, may boast better performance but are unfortunately marred by polynomial time inference. In this paper, we propose a novel parsing order objective, resulting in a novel dependency parsing model capable of both global (in sentence scope) feature extraction as in graph models and linear time inference as in transitional models. The proposed global greedy parser only uses two arc-building actions, left and right arcs, for projective parsing. When equipped with two extra non-projective arc-building actions, the proposed parser may also smoothly support non-projective parsing. Using multiple benchmark treebanks, including the Penn Treebank (PTB), the CoNLL-X treebanks, and the Universal Dependency Treebanks, we evaluate our parser and demonstrate that the proposed novel parser achieves good performance with faster training and decoding.
This chapter highlights the depth of the relationship between language and emotions. It aims to clarify some important distinctions and terms with respect to the linguistic encoding of emotions. Anthropological research has long confirmed that people in different human groups deal with emotional experience in different ways and different languages also offer very different means to talk about it. The chapter shows that emotional experience involves both some physiological and neurological mechanisms, which – for some of them – may be universal, as well as the cognitive organization of these mechanisms into experience, governed by culturally specific social and linguistic norms. A basic criterion that differentiates between descriptive and expressive linguistic resources is the semiotic status of the linguistic devices in question. Descriptive resources consist mostly of lexical resources, that is words, and some constructions.
<b>Purpose: </b>Verbs with low concreteness are frequent in discourse samples but rarely targeted in aphasia treatments for verbs. These verbs are an important part of functional communication, and recent studies have called for more research regarding aphasia and treatment stimuli with low concreteness. The aim of this study was to pilot the use of verbs with low concreteness in a novel sentence production intervention with persons with aphasia.<b>Method: </b>The study took the form of a single-case experimental design with multiple baselines across behaviors and across participants. Three persons with chronic nonfluent aphasia and apraxia of speech participated in the study. Each participant received treatment designed to increase the semantic networks of verbs with high frequency and low concreteness. Sentence production was closely examined over the course of treatment for treated and untreated verbs of varying concreteness levels. Additional measures of language and cognitive functioning were also taken before and after treatment.<b>Results: </b>Results indicated improved sentence production with target verbs attributable to the treatment in the 1st phase of 2 phases for 2 of the 3 participants. The increases corresponded with the application of treatment, despite the difference in number of baseline sessions for the participants. Where there were treatment effects, there was also considerable generalization to untreated sets of items during the 1st treatment phase. Word retrieval also improved for 2 participants.<b>Conclusions:</b> The results suggest that the novel treatment may improve sentence production and word retrieval in persons with aphasia, even when using target verbs with low concreteness ratings. Future research is warranted into the use of low concreteness verbs.<br><b>Supplemental Figure S1. </b>Sentence frame for treatment. <b>Supplemental Figure S2.</b> Participants 1–3 sentence repetition probe performance. <br><b>Supplemental Figure S3.</b> Participants 1–3 discourse generalization probe performance.<br><b>Supplemental Table S1.</b> Verb stimuli lists for sentence production probes.<br><b>Supplemental Table S2.</b> Discourse probe stimuli.<br><b>Supplemental Table S3.</b> Sentence production scoring rubric.<br>Bailey, D. J., Nessler, C., Berggren, K. N., & Wambaugh, J. L. (2019). An aphasia treatment for verbs with low concreteness: A pilot study. <i>American Journal of Speech-Language Pathology.</i> Advance online publication. https://doi.org/10.1044/2019_AJSLP-18-0257<br>
Deep Universal Dependencies is a collection of treebanks derived semi-automatically from Universal Dependencies (http://hdl.handle.net/11234/1-2988). It contains additional deep-syntactic and semantic annotations. Version of Deep UD corresponds to the version of UD it is based on. Note however that some UD treebanks have been omitted from Deep UD.
People who experience trauma can develop enduring trauma-related symptoms. In daily life, post-trauma symptoms (e.g., elevated physiological arousal) can be triggered by affectively salient cues in the environment, especially by cues that act as trauma reminders. Trauma exposure is associated with enduring changes in two biological stress systems: the sympathetic nervous system (SNS) and the hypothalamic-pituitary-adrenal (HPA) axis. In women, activity in both systems is additionally modulated by fluctuations in levels of sex hormones (e.g., estradiol), which could influence physiological responses to trauma reminders. Additionally, previous work has linked the sex hormone estradiol with affect, suggesting that menstrual cycle might influence trauma-related symptoms or daily affect more broadly within the context of trauma exposure. However, we do not yet have a clear understanding of how estradiol influences affective experiences post-trauma. We used a multi-method approach to examine the influence of estradiol on daily affective experiences in a non-clinical, trauma-exposed sample of 40 naturally cycling premenopausal women. The first specific goal of this study was to test the hypothesis that low estradiol would be related to trauma symptoms, including an asymmetrical profile of SNS and HPA axis stress reactivity to a naturalistic trauma reminder. Lower estradiol was related to greater number and severity of PTSD symptoms, and participants in low versus high estradiol menstrual cycle phases showed higher SNS and reduced HPA axis reactivity to a trauma reminder. These results suggest that lower estradiol is associated with a less adaptive profile of stress system reactivity and increased PTSD symptom expression. The second specific goal of this study was to test the influence of menstrual cycle phase on daily affect in a subset of 30 participants. We assessed affective experience over the course of a 10-day ecological momentary assessment (EMA) period, which included the early follicular (low estradiol) and late follicular (high estradiol) phases. We selected these menstrual cycle phases to capture a portion of the cycle where estradiol increased, whereas progesterone remained low, allowing us to test the effects of estradiol without the confound of progesterone. Participants reported more frequent aversive affective experiences, defined as negatively valenced, high arousal states, including PTSD symptoms, during the early versus late follicular phase. During the early versus late follicular phase, participants also reported greater negative and positive affect and showed greater variability in affective ratings. These results suggest that lower estradiol menstrual cycle phases are characterized by more frequent aversive affective experiences, greater affective lability and increased PTSD symptom severity. Together, these results have potential implications for clinical assessment, as menstrual cycle phase at the time of assessment could influence diagnosis of PTSD or symptom severity. Additionally, clinicians working with women with PTSD might anticipate greater affective lability and increased symptom severity during low estradiol phases of the menstrual cycle.
This dissertation investigates the frequency, semantic, and functional characteristics of recurring discontinuous formulaic language in a learner corpus of argumentative and literary essays. Discontinuous sequences of words, or ‘frames’, are recurrent sequences of words that have one or more variable slots. For example, in the * of and it is * to, where the asterisks represent variable slots in the sequences of words. A corpus of English argumentative essays authored by native speakers of English, Japanese, and Spanish is analyzed using modern methods in corpus linguistics to determine which frames are used frequently. Frequent frames are compared between the L1 groups to investigate trends in structure, frequency, and variability.\nThe focus of analysis then narrows to a group of 30 recurring frames comprised of only function words, that is, function word frames such as in the * of, the * of the, and to the * that. The 30 frames are grouped based on structural characteristics, such as noun and preposition-based frames, for further analysis to better understand their semantic and functional characteristics. A lexical database is used to explore the semantic characteristics of fillers of the frames. To determine discourse functions of all instances of each of the 30 target frames, a well-known taxonomy previously applied to continuous sequences of words is adapted and applied to the present context. Discourse functions of the frames are then used as dependent variables in a multinomial logistic regression conducted with four distinct predictor variables: (1) L1, (2) proficiency level, (3) topic of essay, and (4) specific frame. The purpose of the regression is to see which predictor best accounts for discourse function fulfilled by frames from the structural groups.\nFindings from the various analyses first indicate that Japanese learners of English use function word frames at far lower rates than the L1 English and Spanish speakers. Secondly, fillers of the structural groups of function word frames tend to be abstract nouns and the frames largely serve the discourse function of intangible framing of a following noun phrase. In terms of predictor variables, the frames themselves, particularly prepositions, best predict discourse function. The results lend support to the idea that function words, despite carrying little meaning in isolation, are semantically motivated and systematically contribute meaning to larger sequences of words. Pedagogical implications of this study include the teaching of function words from a more phraseological perspective as well as highlighting the connection between frames, fillers, and discourse functions.
Over the last few decades, corpora with comprehensive syntactic annotation, known as treebanks or parsed corpora, have been created in various formats for major languages of the world (e.g., As modes of accessing annotation have become more linguistically sophisticated, so these corpus resources have become more relevant for linguistics in general by providing sources of insight into factors that only become visible through analysis generalized over structures: phenomena in co-occurrence, frequency, constituency, embeddability, scope, agreement, dependency, etc. These insights are spurring new research and refinements in both corpus techniques and theoretical understanding. While much research has concentrated on challenges inherent in the creation as well as correction of annotated corpora (e.g., Examples include linking corpora to external resources like lexical databases, abstracting the contents sufficiently to be of use to non-experts, exploration of crosslinguistic patterns, etc. This special issue consists of five articles focused on applying parsed corpora research in three areas: (I) enrichment and
Modern health worries (MHW) represent individual differences in the perceived threat posed to health and well-being by aspects of modern life. Current evidence suggests that MHW are positively associated with trait negative emotionality, and given that trait negative emotionality is associated with state emotional reactivity to environmental stressors, it is reasonable to expect that persons with elevated MHW would show increased state emotional reactivity to MHW-related stimuli. Consequently, this study aimed to investigate the association of MHW with state emotional reactivity (i.e., valence and arousal) to MHW-related stimuli (i.e., images of air pollution). Combining these stimuli with other stimuli varying in valence and arousal allowed us to examine whether MHW are specifically associated with emotional reactivity induced by MHW-related stimuli. A total of 73 college students viewed 48 images encompassing eight different content areas, including a subset of MHW-related images (air pollution); each image was rated for valence and arousal. Participants also completed measures of MHW, trait negative emotionality (i.e., neuroticism), and demographics. After controlling for neuroticism and gender, results suggest that MHW only predicted valence rating for images of air pollution; conversely, MHW appears to be associated with arousal ratings in response to a variety of stimuli. Implications and limitations are discussed.