Core concepts and indirect alternatives: on the anti-duality of quantifiers
Abstract This paper proposes an analysis for the long-standing puzzle observed by Chemla (2007) regarding the anti-duality of the French universal quantifier tous, which arises even though French has no word for ‘both’ to feed a Maximize Presupposition competition. This phenomenon has been cited as an example in language where a dual ‘conceptual alternative’ is at play (Buccola et al., 2018), but no formal account of it has been put forth. Furthermore, a naive implementation of the idea overgenerates anti-duality inferences in other expressions, such as each, which, and one in English and French, which might be expected to be observed due to anti-dual counterparts in some languages like Icelandic and Japanese. We propose an account where French tous has an unpronounceable dual universal alternative built from a dual core concept, competition with which is licensed by the existence of a pronounceable expression equivalent in meaning, which we call an ‘Indirect Alternative’. This proposal accounts for tous’s anti-duality and lack of anti-$n$-ality for $n>2$, as well as the lack of anti-duality in other quantifiers.
- Research Article
15
- 10.1109/mc.1983.1654420
- Jun 1, 1983
- Computer
The finite-element analysis, or FEA, method is used to solve a variety of scientific and engineering problems. Most major civil and mechanical engineering design firms use FEA programs as a design tool, since government-imposed safety regulations and the need for lower cost design require greater in-depth analysis. The dual processor concept was implemented at swanson Analysis Systems Inc., which added the FPS-164 attached processor to its DEC VAX-11/780 superminicomputer. Swanson modified the ANSYS general-purpose, finite-element computer program to run on the dual processor system and evaluated its performance. 7 references.
- Research Article
100
- 10.1039/c2tc00185c
- Jan 1, 2013
- J. Mater. Chem. C
We describe two novel blue emission materials based on a new type of dual core concept. 1-Phenyl-6-(10-phenyl-anthracen-9-yl)-pyrene (Ph-AP-Ph) and 1-[1,1′;3′,1′′]terphenyl-5′-yl-6-(10-[1,1′;3′,1′′]terphenyl-5′-yl-anthracen-9-yl)-pyrene (TP-AP-TP) were synthesized through boronylation and Suzuki coupling reactions. The Tg values of Ph-AP-Ph and TP-AP-TP were 228 °C and 243 °C, respectively, compared to values of 135 °C and 139 °C for the single core materials 9-(3′,5′-diphenylphenyl)-10-(3′′′,5′′′-diphenylbiphenyl-4′′-yl)anthracene (MAM) and 1,6-bis-[1,1′;3′,1′′]terphenyl-5′-yl-pyrene (TP-P-TP). One of the dual core derivatives, TP-AP-TP, exhibited an ELmax value of 456 nm and a high luminance EQE of 7.51% when used in an EL device. The dual core chromophore materials had narrower PL and EL spectra and better thermal properties than the single core chromophore materials. A device based on TP-AP-TP showed twice the lifetime of a device based on the commercialized material, 2-methyl-9,10-bis(naphthalen-2-yl)anthracene (MADN).
- Research Article
- 10.53469/jerp.2025.07(08).05
- Aug 31, 2025
- Journal of Educational Research and Policies
Under the background of the deep integration of knowledge explosion and digital technology, the traditional education paradigm with the core of subject and examination is facing severe challenges. Based on the philosophical thinking of the essence of education and the analysis of the realistic dilemma, this paper puts forward the dual core education concept of “encyclopedic learning” and “new tool Mastery”. The research shows that: providing broad knowledge stimulation in stages is the key path to stimulate individual potential, and the acquisition of “new tools” such as digital literacy and critical thinking is the survival cornerstone to control the information flood. The two endow each other, and point to the ultimate goal of education, which is to train modern citizens with freedom of thought, lifelong development and a better life.
- Research Article
- 10.1166/jnn.2018.14948
- Mar 1, 2018
- Journal of nanoscience and nanotechnology
New blue emitting materials based on dual core concept, TP-AF-TP and TP-HAF-TP were synthesized through boronylation and Suzuki coupling reactions. In the thin film state, TP-AF-TP and TP- HAF-TP exhibited maximum PL values at 445 and 440 nm, respectively. A non-doped OLED device based on TP-AF-TP and TP-HAF-TP showed current efficiency of 3.16 and 2.67 cd/A, respectively. TP-AF-TP exhibited a higher EL efficiency than that of TP-HAF-TP.
- Research Article
28
- 10.1016/j.csda.2015.02.010
- Feb 18, 2015
- Computational Statistics & Data Analysis
SIMD parallel MCMC sampling with applications for big-data Bayesian analytics
- Conference Article
- 10.5540/03.2015.003.01.0108
- Aug 25, 2015
- Proceeding Series of the Brazilian Society of Computational and Applied Mathematics
Sparse matrix-vector multiplication (SpMV) is an important operation in sparse linear algebra problems. According to Bell [1], “they represent the dominant cost in many iterative methods for solving large-scale linear systems”. In this paper we propose a dataflow parallel version of the SpMV kernel using Compressed Sparse Row (CSR) storage format. This implementation is executed in a Dataflow Runtime Environment called Trebuchet [3][4][5]. The dataflow model[2] exposes parallelism by essence. Task execution is triggered as soon as the input data is available. A dataflow program can be described as a dataflow graph where nodes represent instructions (or tasks) and edges represent their input/output dependencies. Two instructions can run concurrently if there is no directed path between them and their pieces of data are available. This differs from Von Neumann model where instructions are guided by control flow and their execution is inherently sequential. Our experimental environment consists of a machine with 4 Quad-Core Intel Core i7-3820 processors running at 3.6GHz and 64 GB RAM. We performed the SpMV kernel in three matrices from University of Florida Sparse Matrix Collection1, listed in Table 1 . Figure 1 presents the results for Dataflow(running on Trebuchet), Intel Math Kernel Library(MKL)2 (which uses OpenMP3 in parallel version) implementations. Results show that with a naive dataflow implementation it is already possible to achieve higher speedups than the ones obtained with Intel MKL. For smalled matrices the dataflow runtime environment overheads become more evident, but for larger ones, Trebuchet provides better results. In order to improve performance, we intend to implement other SpMV kernels using different storage formats and other approaches to get better use of cache locality[6]. Moreover, since dataflow allows more flexibility in describing application dependencies, we may find more interesting parallelization patterns that are easier to implement and provide smaller overheads in dataflow.
- Conference Article
4
- 10.1109/hpcc-smartcity-dss50907.2020.00025
- Dec 1, 2020
Nowadays, Kohn-Sham density functional theory (DFT) calculation has drawn more and more attention in chemistry and material science simulations. However, due to the extreme large Hamiltonian matrix needed to be generated during the calculation, when the studied system increases, the cost of calculation becomes unbearable both in ground and excited state electronic structure simulations with large uniform basis. In this paper, we propose a high-performance multi-GPU approach for linear-response time-dependent density functional theory (LR-TDDFT) calculation to compute the excitation energies in molecules and solids with the plane wave basis set under the periodic boundary condition. We carefully design the parallel implementation, calculation steps and data distribution schemes in the naive CPU implementation to maintain good scalability when the studied system expands, then port the most time-consuming part to multi-GPU platform along with several effective optimization steps. The results show that with dual V100 GPUs, the proposed approach can achieve an average of 6.68x speedup compared with dual 12-core Xeon CPU with bulk silicon systems that comprises thousands of atoms (1,024 atoms).
- Conference Article
1
- 10.1145/1272457.1272459
- Jun 25, 2007
The microprocessor industry is rapidly moving towards chip multi-processors (CMPs), commonly referred to as multi-core processors, where multiple cores can independently execute different threads. This change in computer architecture requires corresponding design modifications in programming paradigms, including grid middleware tools, to harness the opportunities presented by multi-core processors. Naive implementations of grid middleware on multi-core systems can severely impact performance because of limitations of shared bus bandwidth, cache size and coherency, and communication between threads. The goal of developing an optimized multi-threaded grid middleware for emerging multi-core processors will be realized only if researchers and developers have access to an in-depth analysis of the impact of several low level microarchitectural parameters on performance. None of the current grid simulators and emulators provide feedback at the microarchitectural level, which is essential for such an analysis. We describe the initial design and implementation of an emulation framework, Multi-core Grid (McGrid), to analyze and provide insightful feedback on the performance limitations, bottlenecks, and optimization opportunities for grid middleware on multi-core systems. We describe early performance results of emulating the processing of some representative XML based grid documents on multi-core nodes using the McGrid framework.
- Conference Article
3
- 10.1109/ipdps.2008.4536304
- Apr 1, 2008
- Proceedings - IEEE International Parallel and Distributed Processing Symposium
Chip multi-processors (CMPs), commonly referred to as multi-core processors, are being widely adopted for deployment as part of the grid infrastructure. This change in computer architecture requires corresponding design modifications in programming paradigms, including grid middleware tools, to harness the opportunities presented by multi-core processors. Simple and naive implementations of grid middleware on multi-core systems can severely impact performance. This is because programming for CMPs requires special consideration for issues such as limitations of shared bus bandwidth, cache size and coherency, and communication between threads. The goal of developing an optimized multi-threaded grid middleware for emerging multi-core processors will be realized only if researchers and developers have access to an in-depth analysis of the impact of several low level microarchitectural parameters on performance. None of the current grid simulators and emulators provide feedback at the microarchitectural level, which is essential for such an analysis. In earlier work we presented our initial results on the design and implementation of such an emulation framework, Multi- core Grid (McGrid). In this paper we extend that work and present a performance study on the effect of cache coherency, scheduling of processing threads to take advantage of data available in the cache of each core, and read and write access patterns for shared data structures. We present the performance results, analysis, and recommendations based on experiments conducted using the McGrid framework for processing XML-based grid data and documents.
- Conference Article
- 10.1109/escience.2008.79
- Dec 1, 2008
Computer architecture is now at an important juncture as single-core CPU power is expected to be nearly constant. The microprocessor industry is rapidly moving towards chip multi-processors (CMPs), commonly referred to as multi-core processors. The transition of CPUs from single to multi-core implementations requires a corresponding shift in the programming paradigm for grid and e-science libraries. Naive implementations of processing on multi-core systems can severely impact performance because of limitations of shared bus bandwidth, cache size and coherency, and communication between threads. To optimize the performance of e-science services, careful application of thread-level parallelism is needed. We study this problem in the context of processing XML data used in grid and e-science applications. The web services model, which strongly leverages XML, has been adopted as the basic architecture for grid and e-science services. As a result, the optimization of separate Web services applications is critical because Web services that are deployed in a longer chain of service processing events must guarantee minimal response times to ensure overall system performance. Our goal is to analyze and provide insightful feedback on cache behavior of each core and reveal performance limitations, bottlenecks, and multi-threaded optimization opportunities for processing XML data relevant to grid and e-science application data formats. We use a micro-architectural emulation framework, Multi-core Grid (McGrid), to generate performance data at various levels of granularity. We analyze cache behavior to quantify the exact gains and present recommendations for processing XML data in grid and e-science applications that will be deployed on emerging multi-core systems.
- Research Article
43
- 10.1007/s00778-015-0409-y
- Nov 12, 2015
- The VLDB Journal
Implementations of relational operators on GPU processors have resulted in order of magnitude speedups compared to their multicore CPU counterparts. Here we focus on the efficient implementation of string matching operators common in SQL queries. Due to different architectural features the optimal algorithm for CPUs might be suboptimal for GPUs. GPUs achieve high memory bandwidth by running thousands of threads, so it is not feasible to keep the working set of all threads in the cache in a naive implementation. In GPUs the unit of execution is a group of threads and in the presence of loops and branches, threads in a group have to follow the same execution path; if some threads diverge, then different paths are serialized. We study the cache memory efficiency of single- and multi-pattern string matching algorithms for conventional and pivoted string layouts in the GPU memory. We evaluate the memory efficiency in terms of memory access pattern and achieved memory bandwidth for different parallelization methods. To reduce thread divergence, we split string matching into multiple steps. We evaluate the different matching algorithms in terms of average- and worst-case performance and compare them against state-of-the-art CPU and GPU libraries. Our experimental evaluation shows that thread and memory efficiency affect performance significantly and that our proposed methods outperform previous CPU and GPU algorithms in terms of raw performance and power efficiency. The Knuth---Morris---Pratt algorithm is a good choice for GPUs because its regular memory access pattern makes it amenable to several GPU optimizations.
- Research Article
10
- 10.1016/j.future.2020.07.044
- Aug 15, 2020
- Future Generation Computer Systems
An efficient shortest path algorithm for content-based routing on 2-D mesh accelerator networks
- Conference Article
7
- 10.1145/3149457.3149478
- Jan 28, 2018
Our aim in this work is to improve the performance of the multi-threaded 3D FDTD solver using time-space tiling techniques that enable tile-level parallelization. The implementation of tile-level parallelization that we have used is based on the so-called diamond tiling technique. In this paper, we present a systematic manner for introducing time-space tiling techniques into the 3D FDTD solver and compare four different approaches. Our performance evaluation on a state-of-the-art multi-core processor demonstrated the effectiveness of the time-space tiling techniques with tile-level parallelism for the 3D FDTD method. For the problem with 2003 grid points, our implementation with two-dimensional tile-level parallelism achieved a speedup of 1.88 times over the naive implementation, while for the problem of 3003 grid points, our implementation with one-dimensional tile-level parallelism showed a speedup of 2.22 times. Both results are better than the speedup obtained from an implementation with intra-tile parallelization presented in a previous work.
- Conference Article
3
- 10.1109/memcod.2015.7340469
- Sep 1, 2015
The continuous Skyline query has recently become the subject of the several researches due to its wide spectrum of applications such as multi-criteria decision making, graph analysis network, wireless sensor network and data exploration. In these applications, the datasets are huge and have various dimensions. Moreover, they constantly change as time passes. Therefore, this query is considered as a computation intensive operation that finding the result in a reasonable time is a challenge. In this paper, we present an efficient parallel continuous Skyline approach. In our suggested method, the dataset points are sorted and pruned based on Manhattan distance. Moreover, we use several optimization methods to optimize memory usage in comparison with naïve implementation. In addition, besides the applied conventional parallelization methods, we partition the time steps based on the number of available cores. The experimental results for a dataset that contains 800k points with 7 dimensions show considerable speedup.
- Conference Article
- 10.1109/icws.2007.84
- Jul 1, 2007
Chip multi-processors (CMPs), commonly referred to as multi-core processors, are being widely adopted for deployment as part of the grid infrastructure. In CMPs, multiple cores can independently execute different threads. This change in computer architecture requires corresponding design modifications in programming paradigms, including grid middleware tools, to harness the opportunities presented by multi- core processors. Simple and naive implementations of grid middleware on multi-core systems can severely impact performance. The goal of developing an optimized multi-threaded grid middleware for emerging multi-core processors will be realized only if researchers and developers have access to an in-depth analysis of the impact of several low level microarchitectural parameters on performance. None of the current grid simulators and emulators provides feedback at the micro-architectural level. We have designed an emulation framework, Multi-core Grid (McGrid), to analyze and provide insightful feedback on the performance limitations, bottlenecks, and optimization opportunities for grid middleware on multi-core systems.