From Informed Independent Vector Extraction to Hybrid Architectures for Target Source Extraction
This article revises informed independent vector extraction (iIVE) as a framework for connecting model-based blind source extraction (BSE) with deep learning. We introduce the contrast function for iIVE, which is derived by extending IVE with beamforming-based constraints, enabling an interpretable use of reference signals. We also show that structured mixing models implementing physical knowledge can be integrated, which is demonstrated by two far-field models. With the contrast functions, rapidly converging second-order algorithms are developed, whose performance is first verified through simulations. In the experimental part, we refine iIVE by training models containing unrolled iterations of the developed algorithm. The resulting structures achieve performance comparable to state-of-the-art networks while requiring two orders of magnitude fewer trainable parameters and exhibiting strong generalization to unseen conditions.
- Conference Article
- 10.1109/ieeeconf67917.2025.11443881
- Oct 26, 2025
Independent Vector Analysis (IVA) has become a widely used approach for blind source separation in functional Magnetic Resonance Imaging (fMRI) data analysis. However, IVA requires prior specification of the number of sources (model order), which often needs to be overestimated. This is further associated with higher computational complexity and the need to identify target components, as their order is random. To address this problem, we explore an alternative framework based on source extraction, which targets a particular Source of Interest (SOI). Specifically, we employ the informed FastIVE (iFIVE) algorithm, which utilizes a prior spatial template to steer the extraction towards the desired component, enhancing both separation quality and convergence stability. Experiments on simulated fMRI demonstrate that iFIVE achieves superior extraction accuracy and significantly reduced computational requirements compared to state-of-the-art blind and constrained IVA methods.
- Research Article
109
- 10.1109/tasl.2006.872599
- Nov 1, 2006
- IEEE Transactions on Audio, Speech and Language Processing
This paper presents a method for enhancing target sources of interest and suppressing other interference sources. The target sources are assumed to be close to sensors, to have dominant powers at these sensors, and to have non-Gaussianity. The enhancement is performed blindly, i.e., without knowing the position and active time of each source. We consider a general case where the total number of sources is larger than the number of sensors, and neither the number of target sources nor the total number of sources is known. The method is based on a two-stage process where independent component analysis (ICA) is first employed in each frequency bin and then time-frequency masking is used to improve the performance further. We propose a new sophisticated method for deciding the number of target sources and then selecting their frequency components. We also propose a new criterion for specifying time-frequency masks. Experimental results for simulated cocktail party situations in a room, whose reverberation time was 130 ms, are presented to show the effectiveness and characteristics of the proposed method
- Conference Article
14
- 10.1109/iscas.2005.1465977
- May 23, 2005
The paper presents a method for enhancing a target source of interest and suppressing other interference sources. The target source is assumed to be close to sensors, to have dominant power at these sensors, and to have non-Gaussianity. The enhancement is performed blindly, i.e., without knowing the total number of sources or information about each source, such as position and active time. We consider a general case where the number of sources is larger than the number of sensors. We employ a two-stage process where independent component analysis (ICA) is first employed in each frequency bin and time-frequency masking is then used to improve the performance further. We propose a new sophisticated method for selecting the target source frequency components, and also a new criterion for specifying time-frequency masks. Experimental results for simulated cocktail party situations in a room (reverberation time was 130 ms) are presented to show the effectiveness and characteristics of the proposed method.
- Research Article
24
- 10.1093/mnras/stw982
- Apr 26, 2016
- Monthly Notices of the Royal Astronomical Society
Automated source extraction and parameterization represents a crucial challenge for the next-generation radio interferometer surveys, such as those performed with the Square Kilometre Array (SKA) and its precursors. In this paper we present a new algorithm, dubbed CAESAR (Compact And Extended Source Automated Recognition), to detect and parametrize extended sources in radio interferometric maps. It is based on a pre-filtering stage, allowing image denoising, compact source suppression and enhancement of diffuse emission, followed by an adaptive superpixel clustering stage for final source segmentation. A parameterization stage provides source flux information and a wide range of morphology estimators for post-processing analysis. We developed CAESAR in a modular software library, including also different methods for local background estimation and image filtering, along with alternative algorithms for both compact and diffuse source extraction. The method was applied to real radio continuum data collected at the Australian Telescope Compact Array (ATCA) within the SCORPIO project, a pathfinder of the ASKAP-EMU survey. The source reconstruction capabilities were studied over different test fields in the presence of compact sources, imaging artefacts and diffuse emission from the Galactic plane and compared with existing algorithms. When compared to a human-driven analysis, the designed algorithm was found capable of detecting known target sources and regions of diffuse emission, outperforming alternative approaches over the considered fields.
- Research Article
11
- 10.1109/tsp.2022.3216106
- Jan 1, 2022
- IEEE Transactions on Signal Processing
In this article, nonstationary mixing and source models are combined for developing new fast and accurate algorithms for Independent Component or Vector Extraction (ICE/IVE), one of which stands for a new extension of the well-known FastICA. This model allows for a moving source-of-interest (SOI) whose distribution on short intervals can be (non-)circular (non-)Gaussian. A particular Gaussian source model assuming tridiagonal covariance matrix structures is proposed. It is shown to be beneficial in the frequency-domain speaker extraction problem. The algorithms are verified in simulations. In comparison to the state-of-the-art algorithms, they show superior performance in terms of convergence speed and extraction accuracy.
- Research Article
1
- 10.1109/tsp.2025.3620539
- Jan 1, 2025
- IEEE Transactions on Signal Processing
Independent vector analysis (IVA) is an attractive solution to address the problem of joint blind source separation (JBSS), that is, the simultaneous extraction of latent sources from several datasets implicitly sharing some information. Among IVA approaches, we focus here on the celebrated IVA-G model, that describes observed data through the mixing of independent Gaussian source vectors across the datasets. IVA-G algorithms usually seek the values of demixing matrices that maximize the joint likelihood of the datasets, estimating the sources using these demixing matrices. Instead, we write the likelihood of the data with respect to both the demixing matrices and the precision matrices of the source estimates. This allows us to formulate a cost function whose mathematical properties enable the use of a proximal alternating algorithm based on closed form operators with provable convergence to a critical point. After establishing the convergence properties of the new algorithm, we illustrate its desirable performance in separating sources with covariance structures that represent varying degrees of difficulty for JBSS.
- Conference Article
7
- 10.1109/icassp49357.2023.10095128
- Jun 4, 2023
Recent research has shown remarkable performance in leveraging multiple extraneous conditional and non-mutually-exclusive semantic concepts for sound source separation, allowing the flexibility to extract a given target source based on multiple different queries. In this work, we propose a new optimal condition training (OCT) method for single-channel target source separation, based on greedy parameter updates using the highest performing condition among equivalent conditions associated with a given target source. Our experiments show that the complementary information carried by the diverse semantic concepts significantly helps to disentangle and isolate sources of interest much more efficiently compared to single-conditioned models. Moreover, we propose a variation of OCT with condition refinement, in which an initial condition vector is adapted to the given mixture and transformed to a more amenable representation for target source extraction. We showcase the effectiveness of OCT on diverse source separation experiments where it improves upon permutation invariant models with oracle assignment between estimated and target sources and obtains state-of-the-art performance in the more challenging task of text-based source separation, outperforming even dedicated text-only conditioned models.
- Conference Article
3
- 10.1109/icassp.2016.7471662
- Mar 1, 2016
We investigated informative acoustic feature extraction based on dimension reduction for collecting target sources on a noisy sports field. Although a Wiener filter is often used for sound source enhancement, it is difficult to accurately design the Wiener filter by simply using spatial cues because the noise on a sports field (e.g., cheering from spectators) arrives from the same direction as that of the targeted source. A statistical approach is used to estimate the Wiener filter by using pre-trained acoustic feature models. However, an informative acoustic feature, which provides a powerful clue for clear extraction of the target source, is unknown. For this study, we developed a method for optimizing a projection matrix for dimension reduction by maximizing the mutual information between acoustic features and the Wiener filter. Through experiments using two-directional microphones on a mock sports field, we confirmed that the proposed method outperformed previous methods in terms of both the noise reduction and quality of the recovered sound sources.
- Research Article
1
- 10.5281/zenodo.1327682
- Jul 22, 2007
- Zenodo (CERN European Organization for Nuclear Research)
This paper presents the source extraction system which can extract only target signals with constraints on source localization in on-line systems. The proposed system is a kind of methods for enhancing a target signal and suppressing other interference signals. But, the performance of proposed system is superior to any other methods and the extraction of target source is comparatively complete. The method has a beamforming concept and uses an improved time-frequency (TF) mask-based BSS algorithm to separate a target signal from multiple noise sources. The target sources are assumed to be in front and test data was recorded in a reverberant room. The experimental results of the proposed method was evaluated by the PESQ score of real-recording sentences and showed a noticeable speech enhancement. Keywords—Beamforming, Non-stationary noise reduction, Source separation, TF mask.
- Conference Article
15
- 10.1109/icassp.2016.7471712
- Mar 1, 2016
We propose a method for estimating the prior signal-to-noise ratio (SNR), which is used for calculating the Wiener filter for distant sound source extraction, from output signals of beamforming using statistical mapping based on the deep neural network (DNN). Since informative features to estimate the prior SNR are included in multiple beamforming outputs, the SNR can be accurately estimated by this mapping using the DNN. The proposed method was applied to a large microphone array, the design of which was optimized to form effective directivity patterns to extract distant sound sources. Experimental results proved that the target source was clearly extracted with the proposed method.
- Conference Article
5
- 10.1109/icassp39728.2021.9413422
- Jun 6, 2021
This paper is devoted to the recently proposed mixing model with constant separating vector (CSV) for Blind Source Extraction of moving sources using the FastDIVA algorithm, which is an extension of the famous FastICA and FastIVA for static mixtures. The benefits due to the CSV model and FastDIVA are demonstrated in three new applications. First, the extraction of a moving speaker in a noisy reverberant environment using a dense array of 48 MEMS microphones is considered. Second, a case study on the blind extraction of moving brain activity from visually evoked potentials in electroencephalogram is reported. Third, a simulation of block-by-block online extraction of a moving source is demonstrated. In these examples, the CSV and FastDIVA show their new potential and good performance in handling the blind moving source extraction problem.
- Conference Article
3
- 10.1109/iwaenc53105.2022.9914778
- Sep 5, 2022
The manuscript deals with the robust extraction of a speaker of interest (SOI) from a mixture of audio sources. A blind algorithm based on independent vector extraction (IVE) is used, which, by definition, extracts an arbitrary source. To focus the extraction towards the SOI, a prior knowledge identifying the target source is required. To this end, the manuscript exploits speaker-identification based on embedding features computed via a pretrained forward sequential memory network (FSMN). We introduce and experimentally validate three ways how this prior knowledge can be employed in a blind algorithm, namely, speaker-specific initialization, pilot signal, and supervised deflation of the mixture. The experiments show that the proposed techniques complement each other and lead to robust identification/extraction of the SOI in difficult mixtures of three speakers.
- Conference Article
12
- 10.1109/icassp49357.2023.10097106
- Jun 4, 2023
Spatial information can help improve source separation performance. Numerous spatially informed source extraction methods based on the independent vector analysis (IVA) have been developed, which can achieve reasonably good performance in non- or weakly reverberant environments. However, the performance of those methods degrades quickly as the reverberation increases. The underlying reason is that those methods are derived based on the multiplicative transfer function model with a rank-1 assumption, which does not hold true if reverberation is strong. To circumvent this issue, this paper proposes to use the convolutive transfer function (CTF) model to improve the source extraction performance and develop a spatially informed IVA algorithm. Simulations demonstrate the efficacy of the developed method even in highly reverberant environments.
- Research Article
2
- 10.1121/10.0011746
- Jun 1, 2022
- The Journal of the Acoustical Society of America
The complete decomposition performed by blind source separation is computationally demanding and superfluous when only the speech of one specific target speaker is desired. This letter proposes a computationally efficient blind source extraction method based on the fast fixed-point optimization algorithm under the mild assumption that the average power of the source of interest outweighs the interfering sources. Moreover, a one-unit scaling operation is designed to solve the scaling ambiguity for source extraction. Experiments validate the efficacy of the proposed method in extracting the dominant source.
- Book Chapter
1
- 10.1007/978-3-642-10677-4_41
- Jan 1, 2009
This paper introduces a method for selecting a target source of interest. The target source is assumed to be the closest to sensors among all the other sources regardless of the target source not being the dominant power at the sensors. In this paper, we propose a simple method to select the closest source from signals separated by Independent Vector Analysis (IVA). The proposed method is processed in two-stages. Firstly, IVA is used to separate the mixed signals. Secondly, the mixing channel characteristics are used to choose the closest source. Simulated experimental results are presented to show how well the proposed method works.KeywordsBlind Source ExtractionBlind Source Separation(BSS)Closest SourceConvolutive MixtureIndependent Vector Analysis