Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Paper] TPSMG: Text-Controllable Polyphonic Symbolic Music Generation

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

We propose a novel text-controllable polyphonic symbolic music generation method based on diffusion models. Symbolic music generation has garnered significant attention due to its flexibility and seamless integration with Digital Audio Workstations (DAWs), as it enables the generation of MIDI files, facilitating easier modification compared to waveform music. Although existing techniques enable control through chords or other metadata, few methods allow intuitive control via text prompts, which better align with user preferences. To address this limitation, we introduce Text-Controllable Polyphonic Symbolic Music Generation (TPSMG), a diffusion model specifically designed for text-conditioned symbolic music generation. Our approach incorporates a text condition module into a U-Net backbone within a Denoising Diffusion Probabilistic Model. This module translates text prompts into embeddings that steer the denoising process, thereby enabling precise, text-based control over music generation. Experimental results demonstrate that our method generates high-quality polyphonic symbolic music outputs that closely reflect the intended textual input.

Similar Papers
  • Research Article
  • Cite Count Icon 1
  • 10.1371/journal.pone.0283103.r004
An automatic music generation and evaluation method based on transfer learning
  • May 10, 2023
  • PLOS ONE
  • Yi Guo + 5 more

In recent years, deep learning has seen remarkable progress in many fields, especially with many excellent pre-training models emerged in Natural Language Processing(NLP). However, these pre-training models can not be used directly in music generation tasks due to the different representations between music symbols and text. Compared with the traditional presentation method of music melody that only includes the pitch relationship between single notes, the text-like representation method proposed in this paper contains more melody information, including pitch, rhythm and pauses, which expresses the melody in a form similar to text and makes it possible to use existing pre-training models in symbolic melody generation. In this paper, based on the generative pre-training-2(GPT-2) text generation model and transfer learning we propose MT-GPT-2(music textual GPT-2) model that is used in music melody generation. Then, a symbolic music evaluation method(MEM) is proposed through the combination of mathematical statistics, music theory knowledge and signal processing methods, which is more objective than the manual evaluation method. Based on this evaluation method and music theories, the music generation model in this paper are compared with other models (such as long short-term memory (LSTM) model,Leak-GAN model and Music SketchNet). The results show that the melody generated by the proposed model is closer to real music.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 16
  • 10.1371/journal.pone.0283103
An automatic music generation and evaluation method based on transfer learning.
  • May 10, 2023
  • PloS one
  • Yi Guo + 4 more

In recent years, deep learning has seen remarkable progress in many fields, especially with many excellent pre-training models emerged in Natural Language Processing(NLP). However, these pre-training models can not be used directly in music generation tasks due to the different representations between music symbols and text. Compared with the traditional presentation method of music melody that only includes the pitch relationship between single notes, the text-like representation method proposed in this paper contains more melody information, including pitch, rhythm and pauses, which expresses the melody in a form similar to text and makes it possible to use existing pre-training models in symbolic melody generation. In this paper, based on the generative pre-training-2(GPT-2) text generation model and transfer learning we propose MT-GPT-2(music textual GPT-2) model that is used in music melody generation. Then, a symbolic music evaluation method(MEM) is proposed through the combination of mathematical statistics, music theory knowledge and signal processing methods, which is more objective than the manual evaluation method. Based on this evaluation method and music theories, the music generation model in this paper are compared with other models (such as long short-term memory (LSTM) model,Leak-GAN model and Music SketchNet). The results show that the melody generated by the proposed model is closer to real music.

  • Research Article
  • 10.21428/f1f23564.b8af72cf
Creating Sonic Immersiveness: Sound Generators and Digital Audio Workstations for Generating Sound and Music for Digital Stories
  • Sep 13, 2024
  • IDEAH
  • Katrin Fritsche

Humanities Pedagogy, Training, and Mentorship.Aimed primarily at digital humanities educators, these resources explore digital storytelling and have a particular focus on the integration of music and sound within it.Although they focus on GarageBand, which can be used primarily when creating the associated music and sounds, other applications in this area work similarly and according to the same principles (see Supplement 3).At the heart of this article is a video recording of the conference presentation (Figure 1) that explains the concept of digital storytelling and walks the viewer through the possible creation process, with a particular focus on the musical dimension, because music and sound are extremely important for creating immersiveness.The video serves as a source of ideas and examples for digital storytelling with a connection to cultural topics.It also shows examples curated by cultural institutions to illustrate the application of this concept.In addition, the video outlines possible scenarios for implementing digital storytelling in educational environments.A significant part of the video consists of explaining how a Digital Audio Workstation (DAW) works (using the DAW GarageBand as an example).Although DAWs are usually associated with the professional music industry, free or inexpensive applications make it possible to produce music outside of these boundaries.Demonstrations in the video illustrate how even people who do not have in-depth theoretical and methodological knowledge of music can use automated tools for music generation and as input for DAWs.A specific application for the automated generation of tone sequences is described in more detail, while additional tools for music and sound generation are also referenced in the supplementary materials (Supplement 1, 2, and 3).Beyond the realm of music, the video illustrates the integration of images and video elements to increase the overall appeal of digital stories.The video concludes by exploring how teachers can incorporate this technology into their classrooms.In order to make the steps viewed in the video more accessible, a corresponding transcript of the video is provided.With didactic considerations in mind, the materials continue to provide examples and use cases of digital storytelling with a music-centred focus.Supplement 2 presents an example created in a DAW, including a corresponding MP3 sound file.Teachers who also want to gain a deeper insight into the possible didactic conception of teaching settings that incorporate this topic can look at a sample course plan (Supplement 2) for a course on the topic.This is divided into different phases with different levels of involvement on the part of the learners and provides links to further resources.In addition, Supplement 3 lists additional and alternative applications that can be used (various generation methods and music analysis tools, also for different operating systems).

  • Research Article
  • 10.38048/jcp.v5i2.5425
PREFERENSI DIGITAL AUDIO WORKSTATION (DAW) DAN PENGARUHNYA TERHADAP GAYA ESTETIKA MAHASISWA MINAT STUDI MUSIK TEKNOLOGI
  • Apr 30, 2025
  • Jurnal Citra Pendidikan
  • Enry Johan Jaohari + 3 more

This study explores the preference for Digital Audio Workstation (DAW) and its influence on the musical aesthetics of students majoring in Music Technology. Employing a qualitative descriptive approach with a practice-based research framework, the research involved active students at the Music Study Program, School of Art and Design Education, Universitas Pendidikan Indonesia. Data were collected through open-ended questionnaires, semi-structured interviews, and documentation of students’ digital music works. The findings reveal that DAW selection is influenced by a combination of technical, pedagogical, and creative factors, with Studio One being the most preferred. The features most utilized by students include mixing effects, MIDI sequencing, and preset plugins. Approximately one-third of the students acknowledged that DAW significantly impacted the structure and aesthetic style of their works, while others emphasized personal identity over technological influence. These results indicate that DAWs act not only as production tools but also as mediators of musical thought and expression. The study suggests the necessity of pedagogical frameworks that encourage both technical proficiency and critical reflection in digital music education. The findings also highlight the need for curriculum development that provides students with diverse DAW platforms to support artistic exploration and identity formation. Key Words DAW, music technology, musical aesthetics, digital creativity, music education AbstrakPenelitian ini bertujuan untuk mengeksplorasi preferensi penggunaan Digital Audio Workstation (DAW) dan pengaruhnya terhadap estetika musikal mahasiswa aktif Program Studi Musik yang mengambil Minat Studi Musik Teknologi. Pendekatan yang digunakan adalah kualitatif deskriptif dengan kerangka practice-based research. Data dikumpulkan melalui angket terbuka, wawancara semi-terstruktur, dan dokumentasi karya musik digital mahasiswa. Hasil penelitian menunjukkan bahwa pemilihan DAW dipengaruhi oleh kombinasi faktor teknis, pedagogis, dan kreatif, dengan Studio One menjadi platform yang paling dominan. Fitur yang paling sering digunakan mahasiswa meliputi efek mixing, editor MIDI, dan preset plugin. Sekitar sepertiga mahasiswa mengakui bahwa DAW berkontribusi signifikan dalam membentuk struktur dan gaya estetika karya mereka, sedangkan sebagian besar lainnya menilai bahwa identitas musikal pribadi lebih berpengaruh. Temuan ini menunjukkan bahwa DAW berperan tidak hanya sebagai alat produksi, tetapi juga sebagai mediator dalam berpikir dan mengekspresikan ide musikal. Penelitian ini merekomendasikan pengembangan kurikulum musik digital yang menyeimbangkan penguasaan teknis dengan refleksi kritis terhadap penggunaan teknologi. Kata Kunci DAW, teknologi musik, estetika musikal, kreativitas digital, Pendidikan musik.

  • PDF Download Icon
  • Research Article
  • 10.14569/ijacsa.2024.01505101
Exploring Music Style Transfer and Innovative Composition using Deep Learning Algorithms
  • Jan 1, 2024
  • International Journal of Advanced Computer Science and Applications
  • Sujie He

Automatic music generation represents a challenging task within the field of artificial intelligence, aiming to harness machine learning techniques to compose music that is appreciable by humans. In this context, we introduce a text-based music data representation method that bridges the gap for the application of large text-generation models in music creation. Addressing the characteristics of music such as smaller note dimensionality and longer length, we employed a deep generative adversarial network model based on music measures (MT-CHSE-GAN). This model integrates paragraph text generation methods, improves the quality and efficiency of music melody generation through measure-wise processing and channel attention mechanisms. The MT-CHSE-GAN model provides a novel framework for music data processing and generation, offering an effective solution to the problem of long-sequence music generation. To comprehensively evaluate the quality of the generated music, we used accuracy, loss rate, and music theory knowledge as evaluation metrics and compared our model with other music generation models. Experimental results demonstrate our method's significant advantages in music generation quality. Despite progress in the field of automatic music generation, its application still faces challenges, particularly in terms of quantitative evaluation metrics and the breadth of model applications. Future research will continue to explore expanding the model's application scope, enriching evaluation methods, and further improving the quality and expressiveness of the generated music. This study not only advances the development of music generation technology but also provides valuable experience and insights for research in related fields.

  • Research Article
  • Cite Count Icon 91
  • 10.1145/3597493
A Survey on Deep Learning for Symbolic Music Generation: Representations, Algorithms, Evaluations, and Challenges
  • Aug 25, 2023
  • ACM Computing Surveys
  • Shulei Ji + 2 more

Significant progress has been made in symbolic music generation with the help of deep learning techniques. However, the tasks covered by symbolic music generation have not been well summarized, and the evolution of generative models for the specific music generation task has not been illustrated systematically. This paper attempts to provide a task-oriented survey of symbolic music generation based on deep learning techniques, covering most of the currently popular music generation tasks. The distinct models under the same task are set forth briefly and strung according to their motivations, basically in chronological order. Moreover, we summarize the common datasets suitable for various tasks, discuss the music representations and the evaluation methods, highlight current challenges in symbolic music generation, and finally point out potential future research directions.

  • Research Article
  • 10.1109/taslpro.2026.3667433
MelodyGLM: Multi-task Pre-training for Structured Symbolic Melody Generation
  • Jan 1, 2026
  • IEEE Transactions on Audio, Speech and Language Processing
  • Xinda Wu + 8 more

Pre-trained language models in natural language processing (NLP) have substantially advanced music understanding and generation. However, traditional pre-training methods for symbolic melody generation are limited in their ability to capture multi-scale, multi-dimensional musical structures, primarily due to fundamental differences between textual and musical domains. In addition, the scarcity of large-scale symbolic melody datasets constrains further progress. In this paper, we introduce MelodyGLM, a novel multi-task pre-training framework tailored for structured symbolic melody generation. MelodyGLM enhances autoregressive blank infilling pre-training through melodic <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$n$</tex-math></inline-formula>-gram and long-span sampling strategies, which target local and global structural modeling, respectively. Specifically, we define three types of melodic <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$n$</tex-math></inline-formula>-grams (pitch-based, rhythm-based, and pitch–rhythm hybrid) to model multi-dimensional melodic structures and their intra- and inter-dimensional dependencies. Furthermore, we construct MelodyNet, one of the largest symbolic melody datasets for pre-training, which comprises 440,102 melody pieces extracted from nearly 1.6 million MIDI files. Both subjective and objective evaluations demonstrate that MelodyGLM outperforms previous masked pre-training methods and other symbolic music generation models in modeling melodic structure. In particular, subjective evaluations show that MelodyGLM achieves average gains of 0.75 and 0.69 points in melody continuation and inpainting tasks, respectively, across consistency, rhythmicity, structure, and overall quality.

  • Research Article
  • 10.55041/isjem05180
AI System for Music Generation Based on User Preferences
  • Nov 21, 2025
  • International Scientific Journal of Engineering and Management
  • Mr Lanke Ravi Kumar + 6 more

- The intersection of Computational Creativity and Music Information Retrieval (MIR) presents unique challenges in automating music generation while maintaining emotional coherence. While Deep Learning models like Generative Adversarial Networks (GANs) and Transformers have achieved state-of-the-art results in symbolic music generation, they often suffer from high computational costs, "black-box" un-interpretability, and a lack of closed-loop feedback. This paper proposes a lightweight, transparent, rule-based framework for affective melody generation coupled with a deterministic validation engine. The system utilizes a constrained stochastic process (Random Walk) to generate MIDI sequences based on Western music theory, which are immediately synthesized into audio waveforms. Simultaneously, a Digital Signal Processing (DSP) module extracts spectral features—specifically Spectral Centroid, Bandwidth, and RMS energy—to classify the generated audio into "Energetic" or "Calm" affective states. Experimental validation demonstrates that this architecture successfully enforces harmonic consonance while providing objective, quantifiable feedback on the emotional timbre of the generated composition, achieving a 92% classification accuracy against target moods. Keywords— Music Information Retrieval (MIR), Spectral Feature Extraction, Affective Computing, Digital Signal Processing.

  • Research Article
  • 10.1162/comj_r_00601
Products of Interest
  • Jun 1, 2021
  • Computer Music Journal

Products of Interest

  • Research Article
  • Cite Count Icon 23
  • 10.1109/tmm.2023.3276177
EmoMusicTV: Emotion-Conditioned Symbolic Music Generation With Hierarchical Transformer VAE
  • Jan 1, 2024
  • IEEE Transactions on Multimedia
  • Shulei Ji + 1 more

Emotion is one of the most crucial attributes of music. However, due to the scarcity of emotional music datasets, emotion-conditioned symbolic music generation using deep learning techniques has not been investigated in depth. In particular, no study explores conditional music generation with the guidance of emotion, and few studies adopt time-varying emotional conditions. To address these issues, first, we endow three public lead sheet datasets with fine-grained emotions by automatically computing the valence labels from the chord progressions. Second, we propose a novel and effective encoder-decoder architecture named EmoMusicTV to explore the impact of emotional conditions on multiple music generation tasks and to capture the rich variability of musical sequences. EmoMusicTV is a transformer-based variational autoencoder (VAE) that contains a hierarchical latent variable structure to model holistic properties of the music segments and short-term variations within bars. The piece-level and bar-level emotional labels are embedded in their corresponding latent spaces to guide music generation. Third, we pretrain EmoMusicTV with the lead sheet continuation task to further improve its performance on conditional melody or harmony generation. Experimental results demonstrate that EmoMusicTV outperforms previous methods on three tasks, i.e., melody harmonization, melody generation given harmony, and lead sheet generation. Ablation studies verify the significant roles of emotional conditions and hierarchical latent variable structure on conditional music generation. Human listening shows that the lead sheets generated by EmoMusicTV are closer to the ground truth (GT) and perform slightly worse than the GT in conveying emotional polarity.

  • Research Article
  • Cite Count Icon 4
  • 10.1109/tvcg.2024.3506992
Isolated Diffusion: Optimizing Multi-Concept Text-to-Image Generation Training-Freely With Isolated Diffusion Guidance.
  • Sep 1, 2025
  • IEEE transactions on visualization and computer graphics
  • Jingyuan Zhu + 3 more

Large-scale text-to-image diffusion models have achieved great success in synthesizing high-quality and diverse images given target text prompts. Despite the revolutionary image generation ability, current state-of-the-art models still struggle to deal with multi-concept generation accurately in many cases. This phenomenon is known as "concept bleeding" and displays as the unexpected overlapping or merging of various concepts. This paper presents a general approach for text-to-image diffusion models to address the mutual interference between different subjects and their attachments in complex scenes, pursuing better text-image consistency. The core idea is to isolate the synthesizing processes of different concepts. We propose to bind each attachment to corresponding subjects separately with split text prompts. Besides, we introduce a revision method to fix the concept bleeding problem in multi-subject synthesis. We first depend on pre-trained object detection and segmentation models to obtain the layouts of subjects. Then we isolate and resynthesize each subject individually with corresponding text prompts to avoid mutual interference. Overall, we achieve a training-free strategy, named Isolated Diffusion, to optimize multi-concept text-to-image synthesis. It is compatible with the latest Stable Diffusion XL (SDXL) and prior Stable Diffusion (SD) models. We compare our approach with alternative methods using a variety of multi-concept text prompts and demonstrate its effectiveness with clear advantages in text-image consistency and user study.

  • Research Article
  • Cite Count Icon 1
  • 10.52783/pst.1623
Algorithmic Orchestration: Deep Learning Techniques in Music Generation
  • Mar 10, 2025
  • Power System Technology
  • Sampada K S

Music generation using Deep Learning presents a comprehensive system for generating music in ABC notation using character based recurrent neural networks (RNNs), accompanied by a conversion pipeline that transforms the output music from ABC notation to a playable audio format. The integration of character-based RNNs allows for the creation of coherent and melodic musical compositions, while the ABC- to-MIDI-to WAV conversion enhances the usability and accessibility of the generated music. The work starts by preprocessing the ABC notation dataset and representing it at the character level. A character-based RNN, such as a long short-term memory (LSTM) network, is then employed to learn the sequential dependencies within the ABC notation and generate music that follows the learned patterns and structures. The RNN model is trained on a substantial corpus of ABC- encoded music, enabling it to capture the statistical regularities and nuances of the dataset. To make the generated music readily playable, the project incorporates a conversion pipeline that translates the output from ABC notation to MIDI format. MIDI files serve as a widely supported industry-standard representation of music, making them compatible with a variety of digital audio workstations (DAWs) and synthesizers. The MIDI files are subsequently converted to WAV format, a universally recognized audio format suitable for playback on diverse platforms and devices. This project contributes to the field of AI-generated music by providing an integrated system for music generation in ABC notation, accompanied by a seamless conversion pipeline for playback in the universally supported WAV audio format. The combination of character- based RNNs, ABC notation, and the ABC-to-MIDI-to-WAV conversion offers a valuable tool for musicians, composers, and music enthusiasts, facilitating the generation, sharing, and playback of musical compositions. Future research may explore advanced synthesis techniques, incorporate additional musical features, or investigate other music notation systems to further enhance the capabilities and usability of AI generated music. DOI:https://doi.org/10.52783/pst.1623

  • Book Chapter
  • Cite Count Icon 14
  • 10.1007/978-3-031-29956-8_17
GTR-CTRL: Instrument and Genre Conditioning for Guitar-Focused Music Generation with Transformers
  • Jan 1, 2023
  • Pedro Sarmento + 5 more

Recently, symbolic music generation with deep learning techniques has witnessed steady improvements. Most works on this topic focus on MIDI representations, but less attention has been paid to symbolic music generation using guitar tablatures (tabs) which can be used to encode multiple instruments. Tabs include information on expressive techniques and fingerings for fretted string instruments in addition to rhythm and pitch. In this work, we use the DadaGP dataset for guitar tab music generation, a corpus of over 26k songs in GuitarPro and token formats. We introduce methods to condition a Transformer-XL deep learning model to generate guitar tabs (GTR-CTRL) based on desired instrumentation (inst-CTRL) and genre (genre-CTRL). Special control tokens are appended at the beginning of each song in the training corpus. We assess the performance of the model with and without conditioning. We propose instrument presence metrics to assess the inst-CTRL model’s response to a given instrumentation prompt. We trained a BERT model for downstream genre classification and used it to assess the results obtained with the genre-CTRL model. Statistical analyses evidence significant differences between the conditioned and unconditioned models. Overall, results indicate that the GTR-CTRL methods provide more flexibility and control for guitar-focused symbolic music generation than an unconditioned model.

  • Conference Article
  • Cite Count Icon 2
  • 10.1109/ner.2013.6696026
Listening to the music of the brain: Live analysis of ECoG recordings using digital audio workstation software
  • Nov 1, 2013
  • Griffin Milsap + 3 more

A process is presented for analyzing electrocorticographic (ECoG) recordings and prototyping brain computer interfaces in which complex signal processing chains are able to be rapidly developed and iterated in digital audio workstation (DAW) software. DAW software includes many built-in “drag and drop” blocks that perform common, low-level signal processing algorithms such as filtering and envelope extraction. In addition to being optimized for real-time performance, DAW software also produces audio output, allowing for listening to raw and processed signals. Hearing these sonifications can impart new insights that may not be apparent in purely visual representations. A simple functional mapping analysis is performed in a DAW called Pure Data and compared to the results from a more traditional spatiotemporal analysis in MATLAB. Channels exhibiting qualitative activation in the resulting functional maps were further analyzed in another DAW called Renoise, wherein several high frequency (i.e., >400 Hz) features were observed. This study demonstrates an example use of DAW software, which we suggest is an easy-to-use and intuitive environment for real-time exploratory analyses and sophisticated sonification of ECoG recordings.

  • Conference Article
  • 10.1145/3771594.3771644
User-Centered Insights into Analogue and Digital Mixing: Perspectives of Novice Music Producers
  • Jun 30, 2025
  • Vanja Budsberg + 2 more

The layout of the Audio Mixing Interface (AMI) has remained consistent since it was first conceived over 50 years ago. Whilst usability studies have evaluated alternative AMI designs against this established paradigm, they often lack a holistic perspective. Recently, user experience approaches have been used in ethnographic studies to explore the way music producers use technology. This qualitative study explores novice users experience of mixing popular music. Data was collected via a survey, the user experience questionnaire, and semi-structured group interviews after completing a mixing task with a Digital Audio Workstation (DAW) and analogue mixing console. The analysis revealed that the participants favoured the analogue AMI over the DAW but there were numerous issues with both AMIs. Navigating the AMI, particularly with the DAW, was the biggest reported issue. Furthermore, the DAW’s adaptation of the analogue mixing console does not fully translate, partly due to the lack of tactile control. Novice users enjoyed the additional visual feedback provided by the DAW, specifically for EQ. These findings highlight novice AMI users needs, which are only partly met by the DAW and analogue mixing console. Ideally novice users require a combination of the simplified and tactile experience of the analogue AMI combined with the data visualisation enhanced visual feedback of the DAW to provide a more learnable AMI.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant