Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Towards an Effective Extension of Activity Streams

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Activity Streams is a data format designed to describe activities. Although its specification is written in natural language, its core vocabulary is formally defined in an OWL (Web Ontology Language) ontology. In this work, we propose a set of additional OWL axioms to extend the existing ontology, with the goal of enabling more detailed, precise, and machine-interpretable descriptions of Activity Streams. This enhancement aims to support real-world applications, such as those related to automatic test generation from OWL specifications.

Similar Papers
  • Research Article
  • 10.1525/collabra.160142
Core Vocabulary Reveals Differences Between Human Word Prediction and Large Language Models
  • Apr 27, 2026
  • Collabra: Psychology
  • Andrew Wang + 3 more

The question of which words are the most central or important to a language has been explored in various ways. In this study, we propose definitions of core vocabulary that are based on how language is learned, represented, and processed from psychological perspectives, and test these on a word prediction task. We aim to (1) compare core vocabulary based on word frequency in natural language, word association network centrality, and age-of-acquisition in terms of how well they are guessed in word prediction contexts, and (2) investigate the extent to which word prediction in large language models aligns with humans, and if there are systematic differences between them, whether these can be captured by core vocabulary measures. Across two experiments, 867 participants completed a task which involved guessing target words that were missing from sentence contexts. Word frequency-based core words were easier to guess overall, but when the predictability of words was taken into account, word association- and acquisition-based core words were easier to predict, indicating that people have a preference for simpler or more basic words over and above distributional predictability. Additionally, we show that there are systematic deviations between the predictions of a large language model (BERT) and humans, and that these are best captured by word association-based coreness. The findings suggest that distributional relationships between words in text is not all there is to human word prediction, but that people may also rely on factors like communicative usefulness and multimodal or extralinguistic information. Implications for core vocabulary, word prediction, and the relationship between large language models and human cognition are discussed.

  • Research Article
  • Cite Count Icon 17
  • 10.1016/j.scico.2019.01.003
Test case generation, selection and coverage from natural language
  • Jan 11, 2019
  • Science of Computer Programming
  • Sidney Nogueira + 4 more

Test case generation, selection and coverage from natural language

  • Research Article
  • Cite Count Icon 15
  • 10.1111/1460-6984.12218
Korean word frequency and commonality study for augmentative and alternative communication.
  • Mar 27, 2016
  • International Journal of Language & Communication Disorders
  • Sangeun Shin + 1 more

Vocabulary frequency results have been reported to design and support augmentative and alternative communication (AAC) interventions. A few studies exist for adult speakers and for other natural languages. With the increasing demand on AAC treatment for Korean adults, identification of high-frequency or core vocabulary (CV) becomes essential. The overall objective was to identify the frequency and commonality of spoken Korean words that occurred in spontaneous conversations for the development of AAC interventions. The specific aims were: (1) to generate a Korean CV list based on the conversations of Korean adults; (2) to address the characteristics of the identified words; and (3) to determine whether a quantitative data analysis procedure, based on a grouped frequency distribution, would support identifying high- and low-frequency words. Language samples were collected from 12 native Korean-speaking adults during conversation. CV words were identified based on a grouped frequency distribution analysis and a word commonality analysis. Results established a Korean CV list of 219 words with high frequency and commonality accounting for 60.82% of the total sample. Analysis of word types showed a wide range of particles and verb endings in the CV list. Finally, a distinct distribution pattern was identified from a frequency of 0.2‰ to support high-frequency word selection. The CV list and consideration of the linguistic characteristics of Korean are expected to be used to develop Korean AAC interventions. The grouped frequency distribution revealed a robust method to distinguish high-frequency words and to improve AAC vocabulary selection and organization.

  • Research Article
  • Cite Count Icon 2
  • 10.1007/s11225-012-9430-y
Tractability and Intractability of Controlled Languages for Data Access
  • Aug 1, 2012
  • Studia Logica
  • Camilo Thorne + 1 more

In this paper we study the semantic data complexity of several controlled fragments of English designed for natural language front-ends to OWL (Web Ontology Language) and description logic ontology-based systems. Controlled languages are fragments of natural languages, obtained by restricting natural language syntax, vocabulary and semantics with the goal of eliminating ambiguity. Semantic complexity arises from the formal logic modelling of meaning in natural language and fragments thereof. It can be characterized as the computational complexity of the reasoning problems associated to their semantic representations. Data complexity (the complexity of answering a question over an ontology, stated in terms of the data items stored therein), in particular, provides a measure of the scalability of controlled languages to ontologies, since tractable data complexity implies scalability of data access. We present maximal tractable controlled languages and minimal intractable controlled languages.

  • Research Article
  • Cite Count Icon 25
  • 10.1080/07434618.2019.1692902
Determining a Zulu core vocabulary for children who use augmentative and alternative communication
  • Oct 2, 2019
  • Augmentative and Alternative Communication
  • Jocelyn Mngomezulu + 3 more

Vocabulary selection is an important aspect to consider when designing augmentative and alternative communication (AAC) systems for children who have not yet developed conventional literacy skills. AAC team members have used core vocabulary lists (representing words most commonly and frequently used by speakers of a natural language) as a resource to assist in this process. To date, there are no core vocabulary lists for Zulu. Therefore, the aim of this study was to identify the vocabulary most frequently and commonly used by Zulu-speaking preschool children, in order to inform vocabulary selection for peers who use AAC. Communication samples from 6 Zulu-speaking participants without disabilities were collected during regular preschool activities. Analyses were conducted both by orthographic words and by morphological analysis of formatives. Due to the linguistic and orthographic structure of Zulu, an analysis by formatives was found to be more useful to determine a core vocabulary. The number of different formatives used, frequency of use, and commonality of use among the participants were identified. A total of 213 core formatives were identified; core formatives related to language structure were used more frequently than those that related to lexical content. The characteristics of this Zulu core vocabulary were consistent with those of core vocabularies established in other languages. Implications for the design of Zulu AAC systems are discussed.

  • Research Article
  • Cite Count Icon 4
  • 10.15446/dyna.v81n183.36348
IDENTIFYING DEAD FEATURES AND THEIR CAUSES IN PRODUCT LINE MODELS: AN ONTOLOGICAL APPROACH
  • Jan 13, 2014
  • DYNA
  • Gloria Lucia Giraldo + 2 more

Feature Models (FMs) are a notation to represent differences and commonalities between products derived from a product line. However, product line modelers could unintentionally incorporate dead features in FMs. A dead feature is a type of defect, which implies that one or more features are not present in any product of the product line. Some authors have used ontologies in product lines, but they have not exploited ontology reasoning to identify and explain causes for defects in FMs in natural language. In this paper, we propose an ontology that represents FMs in OWL (Web Ontology Language). Then, we use SQWRL (Semantic Query-enhanced Web Rule Language) to identify dead features in a FM and identify and explain certain causes of this defect in natural language. Our preliminary empirical evaluation confirms the benefits of our approach.

  • Conference Article
  • Cite Count Icon 3
  • 10.29007/zbb8
Automatic generation of high quality test sets via CBMC
  • May 15, 2012
  • EPiC series in computing
  • Emanuele Di Rosa + 4 more

Software Testing is the most used technique for software verification in industry. In the case of safety critical software, the test set can be required to cover a high percentage (up to 100%) of the software code according to some metrics. Unfortunately, attaining such high percentages is not easy using standard automatic tools for tests generation, and manual generation by domain experts is often necessary, thereby significantly increasing the associated costs.In previous papers, we have shown how it is possible to automatize the test generation process of C programs via the bounded model checker CBMC. In particular, we have shown how it is possible to productively use CBMC for the automatic generation of test sets covering 100% of branches of 5 modules of ERTMS/ETCS, a safety critical industrial software by Ansaldo STS. Unfortunately, the test set we automatically generated, is of lower "quality" if compared to the test set manually generated by domain experts: Both test sets attained the desired 100% branch coverage, but the sizes of the automatically generated test sets are roughly twice the sizes of the corresponding manually generated ones. Indeed, the automatically generated test sets contain redundant tests, i.e. tests that do not contribute to reach the desired 100% branch coverage. These redundant tests are useless from the perspective of the branch coverage, are not easy to detect and then to eliminate a posteriori, and, if maintained, imply additional costs during the verification process.In this paper we present a new methodology for the automatic generation of "high quality" test sets guaranteeing full branch coverage. Given an initially empty test set T, the basic idea is to extend T with a test covering as many as possible of the branches which are not covered by T. This requires an analysis of the control flow graph of the program in order to first individuate a path p with the desired property, and then the run of a tool (CBMC in our case) able to return either a test causing the execution of p or that such a test does not exist (under the given assumptions). We have experimented the methodology on 31 modules of the Ansaldo STS ERTMS/ETCS software, thus greatly extending the benchmarking set. For 27 of the 31 modules we succeeded in our goal to automatically generate "high quality" test sets attaining full branch coverage: All the feasible branches are executed by at least one test and the sizes of our test sets are significantly smaller than the sizes of the test sets manually generated by domain experts (and thus are also significantly smaller than the test sets automatically generated with our previous methodology). However, for 4 modules, we have been unable to automatically generate test sets attaining full branch coverage: These modules contain complex functions falling out of CBMC capacity.Our analysis on 31 modules greatly extends our previous analysis based on 5 modules, confirming that automatic test generation tools based on CBMC can be productively used in industry for attaining full branch coverage. Further, the methodology presented in this paper leads to a further increase in the productivity by substantially reducing the number of generated tests and thus the costs of the testing phase.

  • Research Article
  • Cite Count Icon 132
  • 10.1109/tcs.1979.1084676
Automatic test generation techniques for analog circuits and systems: A review
  • Jul 1, 1979
  • IEEE Transactions on Circuits and Systems
  • P Duhamel + 1 more

The purpose of this paper is both a review and an assessment of techniques presently available for automatic test generation for analog systems. After recalling the general problems of automatic testing (definitions, faults in analog systems, different types of tests, main operations, and diagnosis procedures), characterization and description modes of analog systems, and the main software ingredients of automatic test equipment, a categorization of known techniques along several criteria can be proposed. Then, several techniques, respectively, proceeding from approaches based on deterministic and probabilistic estimation, taxonomical and topological analyses can be detailed. Techniques specific to linear systems (several of them belonging to the above three categories) are dealt with in a separate section. The main features of the techniques that are described are summed up in five synoptic tables. As a conclusion, several research areas that need further investigation in view of a possible industrial implementation of automatic analog test generation techniques are identified. Two appendixes deal briefly with fault tolerance and fault simulation in analog systems. An extensive bibliography ( <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">\sim</tex> 500 entries) is provided.

  • Conference Article
  • Cite Count Icon 2
  • 10.1109/compsac51774.2021.00014
Domain-Agnostic Context-Aware Framework for Natural Language Interface in a Task-Based Environment
  • Jul 1, 2021
  • Sarthak Tiwari + 1 more

Smart home assistants are becoming a norm due to their ease-of-use. They employ spoken language as an interface, facilitating easy interaction with their users. Even with their obvious advantages, natural-language based interfaces are not prevalent outside the domain of home assistants. It is hard to adopt them for computer-controlled systems due to the numerous complexities involved with their implementation in varying fields. The main challenge is the grounding of natural language base terms into the underlying system's primitives. The existing systems that do use natural language interfaces are specific to one problem domain only.In this paper, a domain-agnostic framework that creates natural language interfaces for computer-controlled systems has been developed by creating a customizable mapping between the language constructs and the system primitives. The framework employs ontologies built using OWL (Web Ontology Language) for knowledge representation and machine learning models for language processing tasks.

  • Research Article
  • Cite Count Icon 113
  • 10.1007/s12599-009-0078-8
Semantic Process Modeling – Design and Implementation of an Ontology-based Representation of Business Processes
  • Nov 1, 2009
  • Business &amp; Information Systems Engineering
  • Oliver Thomas + 1 more

An extension of process modeling languages is designed which allows representing the semantics of model element labels which are formulated in natural language by using concepts of a formal ontology. This combination of semiformal models with formal ontologies will be characterized as semantic process modeling. The approach is exemplarily applied to the languages EPC (Event-driven Process Chain), BPMN (Business Process Modeling Notation) and OWL (Web Ontology Language) and is generalized by means of an information model. The proposed formalization of the semantics of individual model elements in conjunction with the usage of inference engines allows the improvement of query functionalities in modeling tools and enables new possibilities of model validation. The integration of the approach in the IT-based work environments of modelers is demonstrated by a system architecture and a prototypical implementation. Evidently, advantages in the areas of modeling, model management, IT-business alignment, and compliance can be achieved by the application of modeling tools augmented with semantic technologies.

  • Conference Article
  • 10.1109/icsc50631.2021.00021
Domain-Agnostic Context-Aware Assistant Framework for Task-Based Environment
  • Jan 1, 2021
  • Sarthak Tiwari + 1 more

Smart home assistants are becoming a norm due to their ease-of-use. They employ spoken language as an interface, facilitating easy interaction with their users. Even with their obvious advantages, natural-language based interfaces are not prevalent outside the domain of home assistants. It is hard to adopt them for computer-controlled systems due to the numerous complexities involved with their implementation in varying fields. The main challenge is the grounding of natural language base terms into the underlying system's primitives. The existing systems that do use natural language interfaces are specific to one problem domain only. This paper presents a domain-agnostic framework that creates natural language interfaces for computer-controlled systems that have been developed by creating a customizable mapping between the language constructs and the system primitives. The framework employs ontologies built using OWL (Web Ontology Language) for knowledge representation and machine learning models for language processing tasks.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 34
  • 10.3390/fi4030830
Semantic Web Approach to Ease Regulation Compliance Checking in Construction Industry
  • Sep 11, 2012
  • Future Internet
  • Khalil Riad Bouzidi + 4 more

Regulations in the Building Industry are becoming increasingly complex and involve more than one technical area, covering products, components and project implementations. They also play an important role in ensuring the quality of a building, and to minimize its environmental impact. Control or conformance checking are becoming more complex every day, not only for industrials, but also for organizations charged with assessing the conformity of new products or processes. This paper will detail the approach taken by the CSTB (Centre Scientifique et Technique du Bâtiment) in order to simplify this conformance control task. The approach and the proposed solutions are based on semantic web technologies. For this purpose, we first establish a domain-ontology, which defines the main concepts involved and the relationships, including one based on OWL (Web Ontology Language) [1]. We rely on SBVR (Semantics of Business Vocabulary and Business Rules) [2] and SPARQL (SPARQL Protocol and RDF Query Language) [3] to reformulate the regulatory requirements written in natural language, respectively, in a controlled and formal language. We then structure our control process based on expert practices. Each elementary control step is defined as a SPARQL query and assembled into complex control processes “on demand”, according to the component tested and its semantic definition. Finally, we represent in RDF (Resource Description Framework) [4] the association between the SBVR rules and SPARQL queries representing the same regulatory constraints.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 11
  • 10.12948/issn14531305/21.3.2017.01
Using Ontologies in Cybersecurity Field
  • Sep 30, 2017
  • Informatica Economica
  • Tiberiu Marian Georgescu + 1 more

1 IntroductionThe evolution and expansion of the internet facilitated by great developments in fields such as Big Data, Artificial Intelligence and Machine Learning led to a great change in the virtual environment. The internet transitioned from an environment designed for humans to one where both humans and machines exist and interact. The Semantic Web was first introduced by Tim Berners-Lee et al back in 2001, Berners-Lee being no one else, but the creator of World Wide Web (abbreviated WWW). [1] Since then, the Semantic Web technologies became widespread with applications in various fields. The transition from machine readable information to machine understandable information is possible by expressing the information in languages such as RDF and OWL. [2]In this article the authors discuss how Semantic Web technologies can be used in Cybersecurity field. Cybersecurity is arguably a very complex and extensive domain, whose activities can be classified in two: those undertaken to design an optimal system, with as few vulnerabilities as possible and those that are taken as a result of the problems that appear after the system is operational. While the first type of activities have a relatively common approach to improve security for programming developers, the second is a cat and mouse game, where as soon as a black hat hacker manages to find (and exploit) a type of vulnerability, the system experts work to solve it. In contrast to the white hat hackers, the black hat hackers access and perform actions on a computer system illegally, without the owner's permission, in order to gain personal advantages. One of the objectives of this article is to recognize and discuss actions that can be done between the two types of activities described above, with the help of Semantic Web technologies.The authors propose a framework based on Semantic Web technologies which aims to extract and analyse text in natural (human) language available online and provide results that can improve Cybersecurity. As Abbasi et al point out, there is a lack of research that explores automated identification and characterization of expert hackers within online communities [3].Section 2 presents the main semantic web standards which are considered for the model proposed. Section 3 discusses the borders in which Semantic Web technologies can be used to improve Cybersecurity. The authors describe the types of results expected, based on different types of online sources. They also analyse the types of input data, which consists in any online source about Cybersecurity which may link with black hat hacking. Section 4 highlights the main solutions for web data extraction, illustrates the main differences between scrapers and crawlers and compares the main characteristics of crawlers. Section 5 presents a framework which detects potential Cybersecurity threats based on Semantic Web technologies, as well as the data flow of the model. The 6th section display the authors' conclusion and future work.2Semantic Web StandardsSemantic Web is an extension of World Wide Web, where unstructured data is interpreted by machines through ontologies. Borrowed from philosophy, in IT, ontologies are considered explicit, formal definitions of the entities of reality, based on classes, relations and individuals. Essentially, ontologies are tools that provide to the machines the means of understanding natural language. If machines can properly interpret hacker community's discussions then it is likely that cybersecurity field can be improved.For the model described below, the authors expect to develop ontologies by using the following standards: XML (Extensible Markup Language), RDF (Resource Description Framework), RDFS (Resource Description Framework Schema), OWL (Web Ontology Language) and SPARQL (SPARQL Protocol and RDF Query Language). Figure 1 illustrates the main concepts and abstractions as well as the semantic web specifications and solutions. …

  • PDF Download Icon
  • Conference Article
  • Cite Count Icon 3
  • 10.18653/v1/n18-6001
Modelling Natural Language, Programs, and their Intersection
  • Jan 1, 2018
  • Graham Neubig + 1 more

As computers and information grow a more integral part of our world, it is becoming more and more important for humans to be able to interact with their computers in complex ways. One way to do so is by programming, but the ability to understand and generate programming languages is a highly specialized skill. As a result, in the past several years there has been an increasing research interest in methods that focus on the intersection of programming and natural language, allowing users to use natural language to interact with computers in the complex ways that programs allow us to do. In this tutorial, we will focus on machine learning models of programs and natural language focused on making this goal a reality. First, we will discuss the similarities and differences between programming and natural language. Then we will discuss methods that have been designed to cover a variety of tasks in this field, including automatic explanation of programs in natural language (code-to-language), automatic generation of programs from natural language specifications (language-to-code), modeling the natural language elements of source code, and analysis of communication in collaborative programming communities. The tutorial will be aimed at NLP researchers and practitioners, aiming to describe the interesting opportunities that models at the intersection of natural and programming languages provide, and also how their techniques could provide benefit to the practice of software engineering as a whole.

  • Conference Article
  • Cite Count Icon 4
  • 10.3115/977180.977233
Towards a core vocabulary for a natural language system
  • Jan 1, 1991
  • Hubert Lehmann

The desire to construct robust and portable natural language systems has led to research on how a core vocabulary for such systems can be defined. Statistical methods and semantic criteria for doing this are discussed and compared. Currently it does not seem possible to precisely define the notion of core vocabulary, but it is argued that workable criteria can nevertheless be found. Finally it is emphasized that the implementation of a core vocabulary must be seen as a long-range research program rather than as a short-term goal.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant