Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

DAttnVis: Attention-guided visual diagnostics for stable diffusion inference in image generation

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Stable Diffusion is a widely used text-to-image generation model. However, its outputs are highly sensitive to hyperparameters settings and often suffer issues such as semantic drift, subject misalignment and detail loss. Traditional methods rely on manually adjusting hyperparameters to alter the attention distribution and thus improve the quality of generated images, which is time-consuming and lacks precision. Therefore, we propose an attention-guided visual diagnostic system named DAttnVis, which is designed to assist users in understanding the complex inference process of the Stable Diffusion model and optimizing its parameters. The core idea is to transform high-dimensional attention signals into comparable diagnostic representations across layers using a quantifiable metric—the Attention Concentration Index (ACI). Additionally, an anomaly detection method based on Median Absolute Deviation (MAD) is proposed to accurately identify abnormal attention layers. By linking multiple views, including UNet attention flow, diagnosis and guidance, cross-attention, and historical comparison, DAttnVis constructs a comprehensive diagnostic workflow that covers global screening, structural drilling-down, semantic tracing, and result verification. Quantitative evaluation experiments, case studies and user studies demonstrate that DAttnVis can effectively reduce trial-and-error costs and debugging burdens in the model tuning process, while improving the accuracy of anomalous structure localization and key prompt attribution.

Similar Papers
  • Research Article
  • 10.3724/sp.j.1089.2024-00281
AIGC-Based Image and Video Generation Method: A Review
  • Mar 1, 2025
  • Journal of Computer-Aided Design & Computer Graphics
  • Luyao Zhang + 4 more

Visual generation plays an increasingly important role across diverse fields, from creative domains such as art and entertainment to critical areas such as medical imaging and digital publishing. The development of AIGC in visual generation will potentially revolutionize our interactions with visual data. First, this paper introduces the classical generative models in the deep learning era. Then, based on different input conditions, several important image generation models developed in recent years including unconditional image generation, class-to-image generation, text-to-image generation and image-to-image translation are highlighted, along with their applications in image editing. Next, a detailed summary of video generation and editing models, especially video diffusion models, is provided. And their advantages and disadvantages based on the requirements of training data are outlined. Additionally, this paper reviews the classic datasets for image and video generation and the commonly used evaluation metrics. Finally, the paper summarizes the challenges faced in visual generation in terms of data collection, inference efficiency, long video generation, controllable video generation and security, and discusses potential future research directions.

  • Conference Article
  • Cite Count Icon 4
  • 10.1109/wacv45572.2020.9093308
Jointly Trained Image and Video Generation using Residual Vectors
  • Mar 1, 2020
  • Yatin Dandi + 4 more

In this work, we propose a modeling technique for jointly training image and video generation models by simultaneously learning to map latent variables with a fixed prior onto real images and interpolate over images to generate videos. The proposed approach models the variations in representations using residual vectors encoding the change at each time step over a summary vector for the entire video. We utilize the technique to jointly train an image generation model with a fixed prior along with a video generation model lacking constraints such as disentanglement. The joint training enables the image generator to exploit temporal information while the video generation model learns to flexibly share information across frames. Moreover, experimental results verify our approach’s compatibility with pre-training on videos or images and training on datasets containing a mixture of both. A comprehensive set of quantitative and qualitative evaluations reveal the improvements in sample quality and diversity over both video generation and image generation baselines. We further demonstrate the technique’s capabilities of exploiting similarity in features across frames by applying it to a model based on decomposing the video into motion and content. The proposed model allows minor variations in content across frames while maintaining the temporal dependence through latent vectors encoding the pose or motion features.

  • Preprint Article
  • Cite Count Icon 1
  • 10.2196/preprints.51099
Comparative Analysis of Pretrained Text to Image Models for Accurate Radiological Image Generation for a Single Text Prompt (Preprint)
  • Jul 20, 2023
  • Shashwat Mookherjee + 4 more

BACKGROUND Generative AI is a rapidly advancing field within Artificial Intelligence with wide-ranging applications. In the medical science domain, machine learning and deep learning methods have already found extensive use. This study aims to conduct a comparative analysis of seven freely available pretrained text-to-image models to generate radiological images based on a single text prompt. The primary objective is to determine which among these models produces the most accurate radiological images. The research investigates the effectiveness of generative AI in the medical domain, particularly for generating radiological images. Several text-to-image models are tested, building on previous research that explored DALL-E 2's capabilities in understanding radiological images. By comparing the performance of different models on a single text prompt, the study provides valuable insights into their potential use in medical image generation. Through this investigation, the study seeks to benefit medical professionals, researchers, and the wider AI community. Identifying the most accurate text-to-image model could enhance medical imaging applications, leading to improved diagnostics and treatment planning. OBJECTIVE The objective of this article is to conduct a comparative study of seven existing pretrained text-to-image models available freely on the internet, with the specific aim of generating radiological images based on a single text prompt. The primary goal is to determine which of these text-to-image models is capable of generating the most accurate radiological images for medical applications. By evaluating the performance of various text-to-image models, the research aims to provide insights into the effectiveness of generative AI techniques in the medical domain. Specifically, the study seeks to identify the model that demonstrates superior capabilities in accurately translating textual descriptions into radiological images. The article seeks to contribute to the field of Generative AI, particularly in the context of medical science and radiological imaging. By comparing different models' outcomes on a single prompt, the research aims to offer valuable information for medical professionals and researchers, highlighting the potential applications and limitations of these text-to-image models in medical image generation. Ultimately, the objective of the article is to facilitate advancements in medical imaging technologies, leading to improved diagnostics and medical decision-making processes. Through its comparative analysis, the article endeavors to aid the AI community in selecting the most suitable text-to-image model for generating accurate and reliable radiological images. METHODS Model Selection: Seven existing pretrained text-to-image generative models available freely on the internet were chosen for the study. These models were specifically designed for generating images from textual descriptions. Model Descriptions: A brief description of each text-to-image model used in the study, along with their respective results, was provided to contextualize the findings. Image Generation: The prompt used for testing all models was: "Photorealistic MRI scan of human lungs suffering from pneumonia." Comparative Analysis: The results obtained from each model were compared and analyzed to determine which model produced the most accurate radiological image in response to the given prompt. Keeping in mind the actual imaging of the medical condition, we consulted a physiology expert to compare with the images generated by the seven different models. RESULTS The following comparison results were obtained after the consultation and are presented as follows: Dall-E 2 created the most realistic image of the lungs of a person suffering from pneumonia. It also shows the thoracic cavity with the heart in it which gives more accuracy to the image that is created. Dall-E 2 could also successfully show the difference between the left and right lung. It showed the septum which most of the models could not. Midjourney did a good job at showing the infection even though it failed to create the image as realistically as Dall-E 2. Midjourney did provide a clear image though. It might accurately show the spread of the infection as well. Min-dalle highlighted the infectious parts well, but it failed to give a more realistic image. Carefree Creator did well with the image of the thoracic cavity but it is not very reliable for the detection of infections. Big Sleep is a model which we are unsure of. If the white parts in between show the mucus congestion, then it did a nice job at showing the congestion of the lungs but did a poor job at showing the thoracic cavity. Aphantasia used bright colours which might help the detection of infections even though it failed to show the lungs and the infection accurately. Deep Daze produced a very complicated image which makes identifying parts of the body and the infection very difficult. CONCLUSIONS From the results above, we can conclude that the existing text-to-image generation models are not capable of generating radiological images with 100% accuracy. However, it must be mentioned that some of the models performed better than the others in specific cases. For eg, the image generated by DALL-E 2 was able to show the difference between the left and right lungs properly and also was able to show the thoracic cavity as compared to the image generated by Aphantasia which was able to showcase the detection of infections better than DALL-E 2 even though it failed to show the lungs accurately. This study also indicates the importance for the need of better visualisation of medical conditions in existing radiological methods. For example the use of colours to better showcase the detection of infections as shown by the image generated by Aphantasia. Of course there are many other factors which must be considered while designing a visualisation method and we aren’t suggesting any particular method which needs to be implemented immediately. Proper consultation with an expert is always the first step. These results surely are a starting step in the domain of image generation for radiological images.It is true that in this study we used only one prompt. Further steps would include giving a better text prompt of the medial condition and giving prompts of more varied medical conditions. There are a variety of applications and benefits of generating radiological images. Many Machine Learning tasks like classification and segmentation require a large dataset for training respective models appropriately and an accurate radiological image generated using AI would help in making the dataset of the required size. Further developments in this field can lead to generating radiological images of specific conditions based on the particular prompt of the user.

  • Research Article
  • Cite Count Icon 1
  • 10.54097/hset.v39i.6561
Researches Advanced in Image Generation based on Deep Learning
  • Apr 1, 2023
  • Highlights in Science Engineering and Technology
  • Jianing Duan

Image generation has always been a study hotspot in machine learning, which aims to build models to learn specific semantic distributions from massive image data to generate realistic simulated images. Thanks to the deep learning technology’s quick development, generative models are constantly being developed and huge success has been achieved in image generation tasks. According to difference between generative models, the existing image generation methods based on deep learning can mainly be separated into three models: image generation based on Variational Autoencoder (VAE), image generation based on Generative Adversarial Network (GAN) and image generation combined the VAE and GAN. Focusing on the three frameworks, in this paper, the development process and related principles of each type of generation model are described respectively. After that, the different generation results of different generation models for the agreed training set are compared intuitively, the advantages and problems of various models are proposed, and reasonable improvement measures are proposed for some problems. Finally, the development prospects of various models are prospected.

  • Conference Article
  • Cite Count Icon 9
  • 10.1117/12.188682
Validation of contrast and phenomenology in the Digital Imaging and Remote Sensing (DIRS) lab's image generation (DIRSIG) model
  • Oct 17, 1994
  • Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE
  • John E Mason + 3 more

Comparison of the components and the overall fidelity of infrared synthetic image generation models with truth data and imagery is a crucial part of determining model validity and identifying areas in which improvements can be made. The Rochester Institute of Technology's Digital Imaging and Remote Sensing Image Generation Model, DIRSIG, was validated in the midwave infrared (MWIR) and longwave infrared (LWIR) regions using measured meteorological, material, and radiometric data. Error propagation techniques clearly defined areas where improvements to the model could be made (e.g., inclusion of clouds). An overall comparison of truth and synthetic images yields rms errors of as low as 1.8 degree(s)C for actual temperature, and 5 degree(s)C (LWIR) and 6 degree(s)C (MWIR) for apparent temperatures. Analysis of rank order correlation statistic shows a very high correlation between brightness rank for object in the truth and DIRSIG images for most times of day.

  • Research Article
  • Cite Count Icon 5
  • 10.1002/cav.2232
Diversified realistic face image generationGANfor human subjects in multimedia content creation
  • Mar 1, 2024
  • Computer Animation and Virtual Worlds
  • Lalit Kumar + 1 more

Face image generation plays an important role in generating innovative and unique multimedia content using the GAN model. With these qualities of the GAN model, they have numerous challenges in the human face image generation. The problems encountered in the generation of facial images are like blurriness in images, incomplete details in the generated facial images, high computational power requirements, and so forth. In this manuscript, we proposed a GAN model that utilizes the composite strength of VGG‐16 and ResNet‐50's models to overcome those difficulties. It uses VGG‐16 to build a discriminator model to discriminate between real and fake images. The generator model utilizes a combination of components from the ResNet‐50 and VGG‐16 models to enhance the image generation process at each iteration, resulting in the creation of realistic face images. The proposed DRFI GAN (Diversified and Realistic Face Image Generation GAN) model's generator achieves an impressive low FID score of 20.50, which is less than existing state‐of‐the‐art approaches. Furthermore, our findings indicate that the images generated by the DRFI GAN model exhibit 10%–15% greater efficiency and realism with reduced training time compared to existing state‐of‐the‐art methods with lower FID scores.

  • PDF Download Icon
  • Research Article
  • 10.14569/ijacsa.2024.01505122
Image Generation of Animation Drawing Robot Based on Knowledge Distillation and Semantic Constraints
  • Jan 1, 2024
  • International Journal of Advanced Computer Science and Applications
  • Dujuan Wang

With the development of robot technology, animation drawing robots have gradually appeared in the public eye. Animation drawing robots can generate many types of images, but there are also problems such as poor quality of generated images and long image drawing time. In order to improve the quality of images generated by animation drawing robots, an animation face line drawing generation algorithm based on knowledge distillation was designed to reduce computational complexity through knowledge distillation. To further raise the quality of images generated by robots, the research also designed an unsupervised facial caricature image generation algorithm based on semantic constraints, which uses facial semantic labels to constrain the facial structure of the generated images. The outcomes denote that the max values of the peak signal-to-noise ratio and feature similarity index measurements of the line drawing generation model are 39.45 and 0.7660 respectively, and the mini values are 37.51 and 0.7483 respectively. The average values of the gradient magnitude similarity bias and structural similarity of the loss function used in this model are 0.2041 and 0.8669 respectively. The max and mini values of Fréchet Inception Distance of the face caricature image generation model are 81.60 and 71.32 respectively, and the max and mini time-consuming values are 15.21s and 13.24s respectively. Both the line drawing generation model and the face caricature image generation model have good performance and can provide technical support for the image generation of animation drawing robots.

  • Research Article
  • Cite Count Icon 12
  • 10.1016/j.gexplo.2010.04.011
Defining trace-element alteration halos to skarn deposits hosted in heterogeneous carbonate rocks: Case study from the Cu–Zn Antamina skarn deposit, Peru
  • Apr 30, 2010
  • Journal of Geochemical Exploration
  • A Escalante + 3 more

Defining trace-element alteration halos to skarn deposits hosted in heterogeneous carbonate rocks: Case study from the Cu–Zn Antamina skarn deposit, Peru

  • Research Article
  • Cite Count Icon 3
  • 10.3390/min15060626
Optimized Hydrothermal Alteration Mapping in Porphyry Copper Systems Using a Hybrid DWT-2D/MAD Algorithm on ASTER Satellite Remote Sensing Imagery
  • Jun 9, 2025
  • Minerals
  • Samane Esmaelzade Kalkhoran + 2 more

Copper is typically acknowledged as a critical mineral and one of the vital components of various of today’s fast-growing green technologies. Porphyry copper systems, which are an important source of copper and molybdenum, typically consist of large volumes of hydrothermally altered rocks, mainly around porphyry copper intrusions. Mapping hydrothermal alteration zones associated with porphyry copper systems is one of the most important indicators for copper exploration, especially using advanced satellite remote sensing technology. This paper presents a sophisticated remote sensing-based method that uses ASTER satellite imagery (SWIR bands 4 to 9) to identify hydrothermal alteration zones by combining the discrete wavelet transform (DWT) and the median absolute deviation (MAD) algorithms. All six SWIR bands (bands 4–9) were analyzed independently, and band 9, which showed the most consistent spatial patterns and highest validation accuracy, was selected for final visualization and interpretation. The MAD algorithm is effective in identifying spectral anomalies, and the DWT enables the extraction of features at different scales. The Urmia–Dokhtar magmatic arc in central Iran, which hosts the Zafarghand porphyry copper deposit, was selected as a case study. It is a hydrothermal porphyry copper system with complex alteration patterns that make it a challenging target for copper exploration. After applying atmospheric corrections and normalizing the data, a hybrid algorithm was implemented to classify the alteration zones. The developed classification framework achieved an accuracy of 94.96% for phyllic alteration and 89.65% for propylitic alteration. The combination of MAD and DWT reduced the number of false positives while maintaining high sensitivity. This study demonstrates the high potential of the proposed method as an accurate and generalizable tool for copper exploration, especially in complex and inaccessible geological areas. The proposed framework is also transferable to other porphyry systems worldwide.

  • Research Article
  • Cite Count Icon 41
  • 10.1016/j.gexplo.2014.07.005
Exploratory data analysis and C–A fractal model applied in mapping multi-element soil anomalies for drilling: A case study from the Sari Gunay epithermal gold deposit, NW Iran
  • Jul 15, 2014
  • Journal of Geochemical Exploration
  • Hooshang H Asadi + 3 more

Exploratory data analysis and C–A fractal model applied in mapping multi-element soil anomalies for drilling: A case study from the Sari Gunay epithermal gold deposit, NW Iran

  • Research Article
  • Cite Count Icon 6
  • 10.12815/kits.2013.12.6.010
Combined Filtering Model Using Voting Rule and Median Absolute Deviation for Travel Time Estimation
  • Dec 30, 2013
  • The Journal of The Korea Institute of Intelligent Transport Systems
  • Youngje Jeong + 3 more

This study suggested combined filtering model to eliminate outlier travel time data in transportation information system, and it was based on Median Absolute Deviation and Voting Rule. This model applied Median Absolute Deviation (MAD) method to follow normal distribution as first filtering process. After that, Voting rule is applied to eliminate remaining outlier travel time data after Median Absolute Deviation. In Voting Rule, travel time samples are judged as outliers according to travel-time difference between sample data and mean data. Elimination or not of outliers are determined using a majority rule. In case study of national highway No. 3, combined filtering model selectively eliminated outliers only and could improve accuracy of estimated travel time.Keywords : Travel Time, Outlier Elimination, Filtering Model, Median Absolute Deviation, Voting Rule * 주저자 : 서울시립대학교 교통공학과 연구교수 ** 공저자 : 한국건설기술연구원 연구원 *** 공저자 : 한국건설기술연구원 연구위원**** 공저자 및 교신저자 : 서울시립대학교 교통공학과 교수†논문접수일 : 2013년 10월 21일†논문접수일 : 2013년 11월 11일†논문접수일 : 2013년 11월 22일

  • Research Article
  • Cite Count Icon 4
  • 10.1080/09544828.2024.2411487
Generating tyre tread designs using a sensory evaluation regression model and a generative model
  • Oct 5, 2024
  • Journal of Engineering Design
  • Ren Hayakawa + 2 more

This study introduces a method for generating tyre tread images that are highly rated in terms of ‘slip resistance’ sensory evaluations. It involves constructing a regression model to predict ‘slip resistance’ sensory evaluations in tyre tread images and subsequently developing an image generation model informed by this regression model. The regression model was developed using a machine learning framework based on ensemble learning and SHapley Additive exPlanations. The model was then used to filter training data for the image generation process, leading to the creation of a tyre tread image generation model aimed at producing images with high ‘slip resistance’ sensory evaluations, using StyleGAN. The proposed method was applied in a case study using real data, and its effectiveness was validated by comparison with a multiple regression analysis approach. The results revealed that more accurate results were achieved with the random forest (RF) regression model than with the multiple regression model due to the RF regression model’s emphasis on non-linear features. Furthermore, a significant difference was observed between the ‘slip resistance’ sensory evaluation values of the images generated by StyleGAN trained with training data screened by the RF regression model and those from the training data.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 17
  • 10.1016/j.jksuci.2023.03.021
Image generation models from scene graphs and layouts: A comparative analysis
  • Apr 6, 2023
  • Journal of King Saud University - Computer and Information Sciences
  • Muhammad Umair Hassan + 2 more

An image is the abstraction of a thousand words. The meaning and essence of complex topics, ideas, and concepts can be easily and effectively conveyed visually by a single image rather than a lengthy verbal description. It is not only essential to teach computers how to recognize and classify images but also how to generate them. Controlled image generation depicting complex and multiple objects is a challenging task in computer vision despite the significant advancements in generative modeling. Among the core challenges, scene graph-based and scene layout-based image generation is a significant problem in computer vision and requires generative models to reason about object relationships and compositionality. Due to its ease of use, less time cost, and labor needs, image generation/synthesizing models from scene graphs and layouts are proliferating. In the case of a more significant number of scene graphs and layout to image generation models, a unique experimental evaluation methodology is required to evaluate the controlled image generation. To this extent, we, in this work, present a standard methodology to evaluate the performance of scene graph and scene layout-based image generation models. We perform a comparative analysis of image generation models to evaluate image generation models’ complexity from scene graphs and scene layouts. We analyze the different components of these models on Visual Genome and COCO-Stuff datasets. The experimental results show that the scene layout-based image generation outperforms its graph-based counterpart in most quantitative and qualitative evaluations.

  • Conference Article
  • Cite Count Icon 7
  • 10.1117/12.509858
Performance analysis of improved methodology for incorporation of spatial/spectral variability in synthetic hyperspectral imagery
  • Jan 7, 2004
  • Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE
  • Neil W Scanlan + 2 more

Synthetic imagery has traditionally been used to support sensor design by enabling design engineers to pre-evaluate image products during the design and development stages. Increasingly exploitation analysts are looking to synthetic imagery as a way to develop and test exploitation algorithms before image data are available from new sensors. Even when sensors are available, synthetic imagery can significantly aid in algorithm development by providing a wide range of “ground truthed” images with varying illumination, atmospheric, viewing and scene conditions. One limitation of synthetic data is that the background variability is often too bland. It does not exhibit the spatial and spectral variability present in real data. In this work, four fundamentally different texture modeling algorithms will first be implemented as necessary into the Digital Imaging and Remote Sensing Image Generation (DIRSIG) model environment. Two of the models to be tested are variants of a statistical Z-Score selection model, while the remaining two involve a texture synthesis and a spectral end-member fractional abundance map approach, respectively. A detailed comparative performance analysis of each model will then be carried out on several texturally significant regions of the resultant synthetic hyperspectral imagery. The quantitative assessment of each model will utilize a set of three peformance metrics that have been derived from spatial Gray Level Co-Occurrence Matrix (GLCM) analysis, hyperspectral Signal-to-Clutter Ratio (SCR) measures, and a new concept termed the Spectral Co-Occurrence Matrix (SCM) metric which permits the simultaneous measurement of spatial and spectral texture. Previous research efforts on the validation and performance analysis of texture characterization models have been largely qualitative in nature based on conducting visual inspections of synthetic textures in order to judge the degree of similarity to the original sample texture imagery. The quantitative measures used in this study will in combination attempt to determine which texture characterization models best capture the correct statistical and radiometric attributes of the corresponding real image textures in both the spatial and spectral domains. The motivation for this work is to refine our understanding of the complexities of texture phenomena so that an optimal texture characterization model that can accurately account for these complexities can be eventually implemented into a synthetic image generation (SIG) model. Further, conclusions will be drawn regarding which of the candidate texture models are able to achieve realistic levels of spatial and spectral clutter, thereby permitting more effective and robust testing of hyperspectral algorithms in synthetic imagery.

  • Conference Article
  • 10.1145/800171.809632
Models (fractal and otherwise) for perception and generation of images
  • Jan 1, 1984
  • Alex P Pentland

Perception is best understood as the interpretation of sensory data in terms of models of how the world is structured and how it behaves; these models are exactly those that are most useful for generation of computer images. By recognizing and exploiting this commonality we have been able to make surprising progress in both fields.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant