Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

NEUROMORPHIC-INSPIRED HYBRID COGNITIVE MODEL FOR SELF-OPTIMIZING RESOURCE MANAGEMENT IN 6G EDGE NETWORKS

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Introduction: The 6G world requires connected intelligence, but there is a crucial paradox between the standards of Large Language Model (LLM) and edge constraints. The premium devices have up to 6- 12GB of DRAM, whereas the typical 175B models need 350GB of storage, which is 30 times that of the premium version. Literature Survey: It has been proposed that the bandwidth can be reduced by 90 % with Semantic Communication (Scom) and Edge Semantic Cognitive Intelligence (ESCI). Besides, neuromorphic-based Spiking Neural Networks (SNNs)-model quantization (INT4/INT8) are also known to be necessary to achieve order-of-magnitude energy efficiency (J/token) on resource-constrained hardware. Methodology: This paper proposes a Hybrid Cognitive Model utilizing a three-tier CloudEdge-Device hierarchy. The model integrates event-driven neuromorphic principles with self-optimizing resource management, utilizing paged KV-cache and resource-aware agents for dynamic task offloading. Results: Quantitative evidence is used to show that the hybrid strategy helps to address the 30x resource gap by attaining a 10-100x energy-per-token efficiency due to event-driven neuromorphic sparsity. Statistical analysis makes it evident that semantic filtering substantially reduces communication overhead and maintains reasoning faithfulness by 90 %, and, effectively, it keeps the thermal conditions of devices stable in the case of prolonged 6G edge communications. This model can be used to make sustainable and multi-step thinking on the edge. The hybrid solution achieves the 6G vision of pervasive intelligence by bridging the hardware-software gap via cross-layer co-design.

Similar Papers
  • Research Article
  • Cite Count Icon 9
  • 10.3390/app15137227
Large-Language-Model-Enabled Text Semantic Communication Systems
  • Jun 26, 2025
  • Applied Sciences
  • Zhenyi Wang + 6 more

Large language models (LLMs) have recently demonstrated state-of-the-art performance in various natural language processing (NLP) tasks, achieving near-human levels in multiple language understanding challenges and aligning closely with the core principles of semantic communication Inspired by LLMs’ advancements in semantic processing, we propose LLM-SC, an innovative LLM-enabled semantic communication system framework which applies LLMs directly to the physical layer coding and decoding for the first time. By analyzing the relationship between the training process of LLMs and the optimization objectives of semantic communication, we propose training a semantic encoder through LLMs’ tokenizer training and establishing a semantic knowledge base via the LLMs’ unsupervised pre-training process. This knowledge base facilitates the creation of optimal decoder by providing the prior probability of the transmitted language sequence. Based on this, we derive the optimal decoding criteria for the receiver and introduce beam search algorithm to further reduce complexity. Furthermore, we assert that existing LLMs can be employed directly for LLM-SC without extra re-training or fine-tuning. Simulation results reveal that LLM-SC outperforms conventional DeepSC at signal-to-noise ratios (SNRs) exceeding 3 dB, as it enables error-free transmissions of semantic information under high SNRs while DeepSC fails to do so. In addition to semantic-level performance, LLM-SC demonstrates compatibility with technical-level performance, achieving approximately an 8 dB coding gain for a bit error ratio (BER) of 10−3 without any channel coding while maintaining the same joint source–channel coding rate as traditional communication systems.

  • Research Article
  • Cite Count Icon 41
  • 10.1016/j.neucom.2024.128089
BC4LLM: A perspective of trusted artificial intelligence when blockchain meets large language models
  • Jun 22, 2024
  • Neurocomputing
  • Haoxiang Luo + 2 more

BC4LLM: A perspective of trusted artificial intelligence when blockchain meets large language models

  • Research Article
  • Cite Count Icon 14
  • 10.1145/3767742
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
  • Nov 18, 2025
  • ACM Transactions on Internet of Things
  • Erik Johannes Husom + 7 more

Deploying Large Language Models (LLMs) on edge devices presents significant challenges due to computational constraints, memory limitations, inference speed, and energy consumption. Model quantization has emerged as a key technique to enable efficient LLM inference by reducing model size and computational overhead. In this study, we conduct a comprehensive analysis of 28 quantized LLMs from the Ollama library, which applies by default Post-Training Quantization (PTQ) and weight-only quantization techniques, deployed on an edge device (Raspberry Pi 4 with 4 GB RAM). We evaluate energy efficiency, inference performance, and output accuracy across multiple quantization levels and task types. Models are benchmarked on five standardized datasets (CommonsenseQA, BIG-Bench Hard, TruthfulQA, GSM8K, and HumanEval), and we employ a high-resolution, hardware-based energy measurement tool to capture real-world power consumption. Our findings reveal the trade-offs between energy efficiency, inference speed, and accuracy in different quantization settings, highlighting configurations that optimize LLM deployment for resource-constrained environments. By integrating hardware-level energy profiling with LLM benchmarking, this study provides actionable insights for sustainable AI, bridging a critical gap in existing research on energy-aware LLM deployment.

  • Research Article
  • 10.1108/el-03-2025-0099
Automated learning support literature classification using large language models via different strategies: a study of the LIS literature
  • Nov 27, 2025
  • The Electronic Library
  • Jie Zhang + 5 more

Purpose The purpose of this study is to propose a literature classification scheme based on its knowledge type adapted to search as learning scenario, and explore the feasibility of using generative artificial intelligence tools to automatically complete this kind of literature classification. Design/methodology/approach This study mainly includes two parts: (1) this study investigates knowledge classification from the cognitive perspective, then models the knowledge learning process during academic search based on constructivism, and finally proposes the learning support literature classification (LSLC). (2) Based on three open source large language models (DeepSeek-R1-Distill-Owen-7B, LLaMA3.1-8B and Qwen2.5-7B), this study designs a two-stage experiment of single strategies and hybrid strategies. The classification task performance of three large language models under six different strategies is compared and analysed. Findings This study proposes the LSLC. The first-level classification includes four categories of declarative, procedural, deepened and related content. The second-level classification includes 14 categories of literature review, overview research and so on. Then, six strategies are designed to improve large language models’ performance to auto-complete this kind of literature classification. LLaMA-3.1-8B performs best after optimization. For Chinese literature, the F1 values of first-level and second-level classification of fine-tuned LLaMA-3.1-8B are 88.05% and 71.43%, respectively. For English literature, the F1 values of first-level and second-level classification of fine-tuned and simple thinking prompted LLaMA-3.1-8B are 75.26% and 65%, respectively. Research limitations/implications This study proposes a theoretical achievement of LSLC, and verifies that it is feasible to automatically complete literature classification from a cognitive perspective using large language model, which supports the conclusion that generative artificial intelligence can effectively assist social science research. Originality/value This study proposes a theoretical achievement of LSLC and verifies that it is feasible to automatically complete literature classification from a cognitive perspective using a large language model, which supports the conclusion that generative artificial intelligence can effectively assist social science research.

  • Research Article
  • 10.1109/tmc.2026.3676689
E $^{2}$ LLM: Structure-Guided Efficient Inference for LLMs in Distributed Edge-IoT Environments
  • Jan 1, 2026
  • IEEE Transactions on Mobile Computing
  • Xingyu Feng + 7 more

Large language models (LLMs) are increasingly deployed in edge computing environments to reduce latency and preserve privacy. However, their inference process presents fundamental challenges for resource-constrained IoT devices. LLM inference involves computationally asymmetric stages: parallelizable prompt processing and sequential token decoding. This asymmetry creates deployment bottlenecks where IoT devices lack capacity for prompt processing while edge nodes suffer from inefficient sequential decoding. This paper presents <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">E <inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math></inline-formula> LLM</i>, an efficient distributed inference framework for large language models in heterogeneous edge-IoT environments. <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">E <inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math></inline-formula> LLM</i> leverages high-capacity edge devices for structural planning and introduces auxiliary lightweight models to generate segment-specific key-value (KV) caches. These minimal inference artifacts enable collaborative parallel decoding across IoT devices without requiring full model instantiation. The framework employs static-dynamic KV cache separation to minimize communication overhead while maintaining semantic coherence through structure-guided coordination. Extensive evaluation on realistic edge testbeds demonstrates significant performance improvements. Under diverse deployment settings, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">E <inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math></inline-formula> LLM</i> achieves 74%–87.7% end-to-end latency reduction compared with several state-of-the-art baselines, while maintaining comparable generation quality; meanwhile, it also delivers a 34.6%–72.2% reduction in communication overhead, improves 9-12 × in energy efficiency. The framework exhibits strong scalability under bandwidth-limited conditions, enabling efficient LLM deployment across heterogeneous edge-IoT environments.

  • Research Article
  • Cite Count Icon 13
  • 10.1109/jetcas.2024.3427421
HALO: Communication-Aware Heterogeneous 2.5-D System for Energy-Efficient LLM Execution at Edge
  • Sep 1, 2024
  • IEEE Journal on Emerging and Selected Topics in Circuits and Systems
  • Abhi Jaiswal + 6 more

Large Language Models (LLMs) are used to perform various tasks, especially in the domain of natural language processing (NLP). State-of-the-art LLMs consist of a large number of parameters that necessitate a high volume of computations. Currently, GPUs are the preferred choice of hardware platform to execute LLM inference. However, monolithic GPU-based systems executing large LLMs pose significant drawbacks in terms of fabrication cost and energy efficiency. In this work, we propose a heterogeneous 2.5D chiplet-based architecture for accelerating LLM inference. The proposed 2.5D system consists of heterogeneous chiplets connected via a network-on-package (NoP). In the proposed 2.5D system, we leverage the energy efficiency of in-memory computing (IMC) and the general-purpose computing capability of CMOS-based floating point units (FPUs). The 2.5D technology helps to integrate two different technologies (IMC and CMOS) on the same system. Due to a large number of parameters, communication between chiplets becomes a significant performance bottleneck if not optimized while executing LLMs. To this end, we propose a communication-aware scalable technique to map different pieces of computations of an LLM onto different chiplets. The proposed mapping technique minimizes the communication energy and latency over the NoP, and is significantly faster than existing optimization techniques. Thorough experimental evaluations with a wide variety of LLMs show that the proposed 2.5D system provides up to <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$972\times $ </tex-math></inline-formula> improvement in latency and <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$1600\times $ </tex-math></inline-formula> improvement in energy consumption with respect to state-of-the-art edge devices equipped with GPU.

  • Research Article
  • 10.25163/energy.3110401
Exploring Climate Change Prediction and Mitigation Strategies with Large Language Models (LLM)
  • Jan 1, 2025
  • Energy, Environment, and Economy

Concerning climate change, there is a growing demand for accessible tools that can provide reliable future climate information to support planning, finance, and other decision-making processes. Large language models (LLMs), such as GPT-4 and BERT, present a promising approach to bridging the gap between complex climate data, with the potential to revolutionize data analysis and decision-making across various sectors. However, a significant challenge remains in assessing LLMs' accuracy and reliability in predicting future climate trends. In this study, we introduce a hybrid LLM framework that combines Large Language Models (LLMs) with General Circulation Models (GCMs) to improve climate modeling by increasing prediction accuracy, uncovering hidden patterns in historical data, simulating policy outcomes, and encouraging public engagement. A series of experiments was designed to evaluate the performance of GPT-4 as a hybrid LLM and BERT as a traditional LLM under different conditions. Our results indicate that GPT-4, as a hybrid LLM, reduces prediction errors by up to 19% and produces policy analyses aligned with expert assessments. Although challenges remain in energy efficiency and data biases, our findings demonstrate the transformative potential of LLMs in integrating qualitative human insights with quantitative climate science, emphasizing their crucial role in advancing global climate goals.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 15
  • 10.1038/s41598-024-71678-8
Learning long sequences in spiking neural networks
  • Sep 20, 2024
  • Scientific Reports
  • Matei-Ioan Stan + 1 more

Spiking neural networks (SNNs) take inspiration from the brain to enable energy-efficient computations. Since the advent of Transformers, SNNs have struggled to compete with artificial networks on modern sequential tasks, as they inherit limitations from recurrent neural networks (RNNs), with the added challenge of training with non-differentiable binary spiking activations. However, a recent renewed interest in efficient alternatives to Transformers has given rise to state-of-the-art recurrent architectures named state space models (SSMs). This work systematically investigates, for the first time, the intersection of state-of-the-art SSMs with SNNs for long-range sequence modelling. Results suggest that SSM-based SNNs can outperform the Transformer on all tasks of a well-established long-range sequence modelling benchmark. It is also shown that SSM-based SNNs can outperform current state-of-the-art SNNs with fewer parameters on sequential image classification. Finally, a novel feature mixing layer is introduced, improving SNN accuracy while challenging assumptions about the role of binary activations in SNNs. This work paves the way for deploying powerful SSM-based architectures, such as large language models, to neuromorphic hardware for energy-efficient long-range sequence modelling.

  • Research Article
  • 10.1145/3750727
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
  • Jul 24, 2025
  • ACM Transactions on Embedded Computing Systems
  • Weihong Xu + 4 more

Large language models (LLMs), composed of Transformer decoders, have demonstrated unparalleled proficiency in understanding and generating human language. However, efficient LLM inference on resource-constraint embedded devices remains a challenge because of the sheer model size and memory-intensive operations that arise from feedforward network (FFN) and multi-head attention (MHA) layers. Existing accelerations offload LLM inference to heterogeneous computing systems comprising expensive memory and processing units. However, recent studies show that most hardware resources are not used because LLM exhibits significant sparsity during inference. The sparsity of LLMs provides a good opportunity to perform memory-efficient inference. In this work, we propose SLIM, an algorithm and hardware co-design optimized for sparse LLM serving on the edge. SLIM exploits LLM’s sparsity by only fetching activated neurons to significantly reduce data movement. To this end, the efficient inference algorithm based on adaptive thresholding is proposed to support runtime configurable sparsity at the cost of negligible accuracy loss. Then, we present the SLIM heterogeneous hardware architecture that combines the best of both near-storage processing (NSP) and processing-in-memory (PIM). SLIM stores FFN weights in high-density 3D NAND and computes FFN layers in NSP units, alleviating high memory requirements caused by FFN weights. The memory-intensive MHA with low arithmetic density is processed in the PIM module. By leveraging the inherent sparsity observed in LLM operations and integrating NSP with PIM techniques within SSDs, SLIM significantly reduces memory footprint, data movement, and energy consumption. Meanwhile, we present the software support for integrating design into existing SSD system. Our comprehensive analysis and system-level optimization demonstrate the effectiveness of our sparsity-tailored accelerator, offering 13-18 × throughput improvements over SSD-GPU system and 9-10 × better energy efficiency over DRAM-GPU system while maintaining low latency.

  • Research Article
  • 10.70389/pjp.100002
Using Large Language Models in Psychological Research: A New Frontier for Hypothesis Generation
  • Jan 1, 2025
  • Premier Journal of Psychology
  • Vincent Adeyemi

Large language models (LLMs) are revolutionizing psychology research, particularly in the area of hypothesis generation. Psychologists have traditionally depended on established methods like comprehensive literature reviews, empirical research, and theoretical frameworks. Despite their effectiveness, these methods frequently fall short of the increasing expectations of the data-driven world of today. Since the introduction of LLMs, academics have had access to instruments that can analyze large volumes of text, spot complex patterns, and produce original ideas that may inspire new, creative research questions. This article examines how LLMs can improve hypothesis formation in the social, developmental, and therapeutic domains, among the other subfields of psychology. It looks at their main benefits, namely their capacity to automatically scan vast amounts of psychological literature, creatively combine disparate materials, and identify patterns at scale. It also recognizes their drawbacks, including their dependence on perhaps skewed training data, their inability to comprehend complex situations, and their inability to match artificial intelligence (AI)-driven theories with accepted research practices. Their responsible implementation also emphasizes the importance of ethical factors, such as encouraging transparency, guaranteeing accountability, and safeguarding privacy. Future developments could include the development of specialized LLMs for psychology research, hybrid strategies that incorporate AI and human expertise, and interdisciplinary collaboration to optimize the use of these tools. By carefully incorporating LLMs into research procedures, hypothesis development could be transformed, opening the door to a more profound understanding and a more thorough comprehension of human behavior.

  • Conference Article
  • 10.1109/icatc68823.2025.11407841
Mitigating Bias and Hallucinations in LLMs through Prompt Engineering and Knowledge-Grounded Approaches in Healthcare Domain: A Systematic Literature Review
  • Dec 16, 2025
  • A.P.D Ishadya + 2 more

Large Language Models (LLMs) are increasingly applied in healthcare for clinical decision support, patient communication, medical documentation, and research synthesis. However, their reliability is undermined by persistent risks of bias and hallucinations, which can result in misinformation, workflow disruption, and patient harm. This systematic literature review examines peer-reviewed studies published between 2021 and 2025 that focus on mitigating bias and hallucinations in healthcare LLMs. Thirty-four eligible studies were synthesized using thematic analysis across prompt engineering, knowledge-grounded approaches such as Retrieval-Augmented Generation (RAG) and the Model Context Protocol (MCP), evaluation metrics, and clinical applications. Results show that prompt engineering provides lightweight improvements but is highly dependent on prompt design, while RAG consistently enhances factual grounding though at increased computational cost. MCP emerges as a promising but underexplored tool for dynamic knowledge integration. The review highlights key trends including the shift toward hybrid strategies and domain-specific LLMs, while identifying major research gaps: the absence of unified mitigation frameworks, limited evidence on healthcare-specific prompt design, and a lack of practical deployment guidelines for clinical settings. By consolidating current evidence and outlining open challenges, this review contributes to the development of safe, fair, and trustworthy LLM adoption in healthcare.

  • Research Article
  • Cite Count Icon 28
  • 10.1109/jiot.2024.3366906
CASIT: Collective Intelligent Agent System for Internet of Things
  • Jun 1, 2024
  • IEEE Internet of Things Journal
  • Ningze Zhong + 7 more

In the last few years, the bottleneck of bandwidth in Internet of Thing (IoT) has driven expectations to figure out new ways to preprocess the information needed to be transmitted. The ways which were used before are not smart enough and they cannot align to the users’ need. Large language model (LLM)-based intelligent agent is a very hot concept in AI community, which aims to save various problems via adapting LLM to different industries. In this article, we present a collective intelligent agent system for the IoT (CASIT) that is a pioneering LLM-agent-based IoT system. We put forward a IoT framework that can be used to lots of scenarios. CASIT refers to a system based on multiple intelligent LLM agents, which realizes complex tasks through cooperation and makes full use of collective intelligence. In order to solve the problems, we designed the Memory Mechanism and Summary Mechanism that enable LLMs to efficiently process the data by comparing historical data with Local Knowledge and Chat History in the prompt. After experimental verification, we have found that our framework could accurately conclude the abnormal information, and it outperforms the single LLM system when we input 200 sets of temperature and humidity data from five different places. The system provides a new solution and method for information processing in all IoT systems. Our framework may also provide refreshing ideas for edge computing and semantic communication.

  • Conference Article
  • Cite Count Icon 1
  • 10.2118/229540-ms
Prompting Autonomous Agents and LLMs in Energy Operations, Efficiency Gains or Hidden Liabilities?
  • Nov 3, 2025
  • Alessio Faccia

Objectives/Scope This paper evaluates the operational deployment of large language models (LLMs) and agentic AI systems within the energy industry. The primary focus is on their role in automating reporting, compliance monitoring, decision support, and cyber-defence tasks. It explores both the measurable gains in process efficiency and the poorly understood liabilities introduced by autonomous decision-making and unverifiable outputs. Methods, Procedures, Process The study draws on technical analysis of LLMs and agent chains such as ReAct, AutoGPT, and enterprise-tuned proprietary models. A structured evaluation framework is used to classify outputs based on traceability, reproducibility, and risk of hallucination. Benchmark comparisons are performed between human-led, prompt-based LLM interactions and autonomous agent loops with embedded task prioritisation. Use cases include automated compliance drafting, anomaly detection summaries, and initial incident response generation. Legal and audit implications are assessed with respect to digital evidence, explainability, and attribution of action. Results, Observations, Conclusions Findings reveal that while LLMs and autonomous agents significantly reduce drafting time and offer versatile adaptive outputs, they introduce governance risks rarely captured by current digital assurance frameworks. High-value operations, particularly those involving compliance-sensitive tasks, show susceptibility to misdirection when outputs are not independently verified or when agentic loops execute without escalation triggers. Observations indicate an absence of formalised validation layers, leading to potential acceptance of hallucinated or fabricated content under the false presumption of system precision. In energy sector trials, users lacked clarity on when LLMs switched from generative assistance to autonomous recommendation, thereby diffusing responsibility and undermining post-decision traceability. The study calls for strict operational boundaries and layered validation protocols before integration of LLMs into decision-critical workflows. Novel/Additive Information This paper presents a risk-graded framework for evaluating agentic LLM outputs in operational settings, providing the petroleum sector with a method to align automation benefits with responsible governance practices. It is one of the first to link LLM traceability with legal auditability in energy operations.

  • Research Article
  • 10.18063/csa.v3i1.911
Overview of Large Language Models
  • Mar 26, 2025
  • Computer Simulation in Application
  • Jiaxin Li + 5 more

This article focuses on typical large language models and conducts an in-depth analysis of their definitions, typical models and the development status of their technologies. As an advanced artificial intelligence technology, large language models are trained based on huge parameters and massive data, and achieve natural language processing with the converter structure as the core. This article elaborates on the developing history and features of various large language models. Meanwhile, it is pointed out that the development of large language model technology faces problems as well as challenges such as non-authentic output and security risks, and in the future, it will develop in the directions of lightweight, multimodal and vertical specialization. The research aims to provide references for the further study and application of large language models and contribute to promoting the healthy development of this technology in various fields.

  • Research Article
  • 10.1109/tc.2025.3628193
Bit-serial Acceleration of LLM Inference with Mixture-of-Datatype Quantization
  • Jan 1, 2025
  • IEEE Transactions on Computers
  • Yuzong Chen + 6 more

Large language models (LLMs) have achieved significant breakthroughs on machine learning tasks. Yet the substantial memory footprint of LLMs significantly hinders their wide deployment. In this paper, we propose BitMoD, an algorithm-hardware co-design solution for efficient LLM deployment. On the algorithm side, BitMoD introduces “fine-grained data type adaptation”, which uses a different data type to quantize a group (e.g., 128) of weights and key-value-cache (KV-cache). Through the careful design of these data types, BitMoD is able to quantize LLM weights and KV-cache to sub-4-bit precision while maintaining high accuracy. On the hardware side, BitMoD employs the bit-serial computing paradigm to easily support multiple numerical precisions and data types, thus providing a flexible trade-off between model accuracy and hardware efficiency. Furthermore, we design low-cost hardware components to effectively handle online KV-cache quantization and per-group partial sum dequantization. Our evaluation on a diverse set of LLMs demonstrates that BitMoD significantly outperforms state-of-the-art LLM quantization methods on both discriminative and generative tasks. Combining the superior model performance with an efficient accelerator design, BitMoD surpasses the state-of-the-art LLM accelerator in terms of both hardware performance and energy efficiency.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant