Publications

Journal articles, conference papers, and theses by the group, with links to code, data, and demos.

146 publications · 55 journal articles · 85 conference papers

2026

  1. Active Acoustic Enhancement Systems: A Review
    Will J. Cassidy, Gian Marco De Bortoli, Karolina Prawda, Philip Coleman, Russell Mason, Tapio Lokki, Sebastian J. Schlecht, Enzo De Sena
    The Journal of the Acoustical Society of America
    Abstract

    Active acoustic enhancement systems (AAESs) use microphones, loudspeakers and electronic processing to modify the reverberation of a space, offering flexible and cost-effective alternatives to passive variable acoustics. These systems can extend the reverberation time of a space and modify perceived characteristics such as wall distance, diffuseness and intimacy. In this article, the current literature is discussed, and common conditions of AAESs are demonstrated using simulations to help researchers to establish a comprehensive understanding of the field. A general model is first defined to approximate any AAES as a linear, time-invariant system of transfer functions. This is used to analyse the general stability condition, which is valuable for system tuning and prediction. The three main topologies of AAESs are presented, namely, in-line, regenerative and hybrid systems, describing the fundamental differences as well as the nuances of commercial implementations with a focus on signal processing techniques. Articles investigating AAESs have been summarised to allow readers to gauge the coverage of experimental research to date. The simulated contribution serves as an exploratory environment to compare AAES conditions, where code and audio examples are available online. Promising future trajectories are identified involving machine learning, artefact perception and expressive performance.

  2. AnyRIR: Robust Non-Intrusive Room Impulse Response Estimation in the Wild
    Kyung Yun Lee, Nils Meyer-Kahlen, Karolina Prawda, Vesa Välimäki, Sebastian J. Schlecht
    ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
    Abstract

    We address the problem of estimating room impulse responses (RIRs) in noisy, uncontrolled environments where non-stationary sounds such as speech or footsteps corrupt conventional deconvolution. We propose AnyRIR, a non-intrusive method that uses music as the excitation signal instead of a dedicated test signal, and formulate RIR estimation as an $$1-norm regression in the time--frequency domain. Solved efficiently with Iterative Reweighted Least Squares (IRLS) and Least-Squares Minimal Residual (LSMR) methods, this approach exploits the sparsity of non-stationary noise to suppress its influence. Experiments on simulated and measured data show that AnyRIR outperforms $$2-based and frequency-domain deconvolution, under in-the-wild noisy scenarios and codec mismatch, enabling robust RIR estimation for AR/VR and related applications.

  3. Best Least Squares Paraunitary Approximation: Analytic Procrustes Problem
    Stephan Weiss, Sebastian J. Schlecht, Marc Moonen
    IEEE Transactions on Signal Processing
    Abstract

    This paper addresses the analytic Procrustes problem, which aims to find the best least-squares paraunitary approximation of a square matrix of analytic transfer functions, or the best paraunitary transformation between two rectangular analytic matrices. This is accomplished by generalising the Procrustes solution from ordinary matrices to the case of matrices of analytic functions via their analytic singular value decomposition (SVD). Different from the ordinary matrix case, the analytic SVD is not restricted to singular values being nonnegative. In the case that singular values do not possess any zero crossings, we can find an analytic paraunitary matrix analogously to the standard Procrustes approach. In the case that singular values exhibit any zero crossings, the solution does not only depend on the left- and right-singular vectors, but also on a discontinuous and hence non-analytic switching function that forces those analytic singular values to become nonnegative real. We show that a close approximation of this switching function can be achieved via a complex-valued allpass filter, for which we suggest a new suitable design to minimise the overall least squares error of the fit. In addition, we propose a DFT domain algorithm to approximate this polynomial Procrustes solution, which avoids ambiguities in the analytic SVD, and possesses proven convergence. Generally, this solution requires a delay for causality, and this delay grows with the approximation order. Examples and simulations demonstrate our proposed method.

  4. Continuation Method for Feedback Delay Network Modal Decomposition
    Jeremy B. Bai, Sebastian J. Schlecht
    ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
    Abstract

    Feedback Delay Networks (FDNs) modal decomposition requires solving large polynomial eigenvalue problems, which is costly when the feedback matrix varies. We propose a continuation approach that tracks poles along matrix homotopies using eigenderivatives and a predictor--corrector scheme. Linear and exponential paths are compared, with phase-only updates preserving lossless cases. Experiments on moderately sized FDNs show smooth trajectories, minor pole loss, and reasonable computational cost compared to the Ehrlich--Aberth method, supporting both modal and gradient-based analysis.

  5. Differentiable Grouped Feedback Delay Networks for Learning Coupled Volume Acoustics
    Orchisama Das, Gloria Dal Santo, Sebastian J. Schlecht, Vesa Välimäki, Zoran Cvetković
    IEEE Transactions on Audio, Speech and Language Processing
    Abstract

    Rendering dynamic reverberation in a complicated acoustic space for moving sources and listeners is challenging but crucial for enhancing user immersion in extended-reality (XR) applications. Capturing spatially varying room impulse responses (RIRs) is costly and often impractical. Moreover, dynamic convolution with measured RIRs is computationally expensive with high memory demands, typically not available on wearable computing devices. Grouped Feedback Delay Networks (GFDNs), on the other hand, allow efficient rendering of coupled room acoustics. However, its parameters need to be tuned to match the reverberation profile of a coupled space. In this work, we propose the concept of Differentiable GFDNs (DiffGFDNs), which have tunable parameters that are optimised to match the late reverberation profile of a set of RIRs captured from a space that exhibits multi-slope decay. Once trained on a finite set of measurements, the DiffGFDN interpolates to unmeasured locations in the space. We propose a parallel processing pipeline that has multiple DiffGFDNs with frequency-independent parameters processing each octave band. The parameters of the DiffGFDN can be updated rapidly during inferencing as sources and listeners move. We evaluate the proposed architecture against the Common Slopes (CS) model on a dataset of RIRs for three coupled rooms. The proposed architecture generates multi-slope late reverberation with low memory and computational requirements, achieving a better energy decay relief (EDR) error and slightly worse octave-band energy decay curve (EDC) errors compared to the CS model. Furthermore, DiffGFDN requires an order of magnitude fewer floating-point operations per sample than the CS renderer.

  6. Differentiable Grouped Feedback Delay Networks for Learning Direction and Position-Dependent Late Reverberation
    Orchisama Das, Sebastian J. Schlecht, Gloria Dal Santo, Zoran Cvetković
    ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
    Abstract

    Late reverberation in coupled spaces shows direction and position-dependent decay. Grouped Feedback Delay Networks (GFDNs) are artificial reverberators capable of efficient rendering of coupled room acoustics by modelling multi-slope decays. We extend the concept of Differentiable GFDN (DiffGFDN) to synthesise position- and direction-dependent late reverberation. The proposed DiffGFDNs operate in octave bands, with spherical harmonic receiver gains predicted from listener positions via a Multi-Layer Perceptron. Training minimises spectral deviation, promotes low sparsity, and matches Directional Energy Decay Curves (DEDCs). At inference, signals are processed in frequency subbands, beamformed, and rendered binaurally or over loudspeakers. The proposed method shows slightly higher DEDC errors (+1.5 dB at 0.6 m, +0.9 dB at 0.9 m) than the best baseline, but is the most computationally efficient, enabling faster 6-DoF rendering than time-space-varying convolution.

  7. Matching Reverberant Speech Through Learned Acoustic Embeddings
    Philipp Götz, Gloria Dal Santo, Sebastian J. Schlecht, Vesa Välimäki, Emanuël A. P. Habets
    ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
    Abstract

    Reverberation conveys critical acoustic cues about the environment, supporting spatial awareness and immersion. For auditory augmented reality (AAR) systems, generating perceptually plausible real-time reverberation remains a key challenge, especially when explicit acoustic measurements are unavailable. We address this by formulating blind estimation of artificial reverberation parameters as a reverberant signal matching task, leveraging a learned room-acoustic prior. Furthermore, we propose a feedback delay network (FDN) structure that reproduces both frequency-dependent decay times and the direct-to-reverberation ratio of a target space. Experimental evaluation against a leading automatic FDN tuning method shows improvements in estimated room-acoustic parameters and in the perceptual plausibility of artificial reverberant speech. These results highlight the potential of our approach for efficient, perceptually consistent reverberation rendering in AAR applications.

  8. On the Role of Speech Similarity in the Detection of Room Acoustic Differences
    Thomas McKenzie, Nils Meyer-Kahlen, Sebastian J. Schlecht
    The Journal of the Acoustical Society of America
    Abstract

    Spatial audio systems are typically evaluated in comparative listening tests using the same source signal for each condition such as ABX: ITU-R BS.1116-3 [(2015a) Methods for the Subjective Assessment of Small Impairments in Audio Systems (International Telecommunication Union, Geneva, Switzerland)] and multiple stimulus with hidden reference and anchor ITU-R BS.1534-3 [(2015b) Methods for the Subjective Assessment of Intermediate Quality Level of Audio Systems (International Telecommunication Union, Geneva, Switzerland)]. However, in augmented reality (AR) scenarios, it is infeasible that the same sound source would exist at the same position in space, both real and virtual; instead, each sound source will emit a different signal. To investigate this discrepancy, a perceptual study is conducted on the effect of source signal similarity when distinguishing different room acoustics conditions. Specifically, these conditions are binaural room impulse responses measured at different distances from the source, modified to all use the same direct sound. Three classes of source signal are investigated in a three-alternative forced choice paradigm: the same speech signal for all conditions, the same speaker but a different sentence for each condition, and a different speaker and a different sentence for each condition. Results show that using different speech recordings significantly reduces the ability to identify differences in room acoustics. This suggests that spatial audio system fidelity requirements could vary depending on the source signals used in the target application; AR audio evaluation should use different signals for comparisons.

  9. Solving Room Impulse Response Inverse Problems Using Flow Matching with Analytic Wiener Denoiser
    Kyung Yun Lee, Nils Meyer-Kahlen, Sebastian J. Schlecht, Vesa Välimäki
    The Journal of the Acoustical Society of America
    Abstract

    Room impulse response (RIR) estimation naturally arises as a class of inverse problems, including denoising and deconvolution. While recent approaches often rely on supervised learning or learned generative priors, such methods require large amounts of training data and may generalize poorly outside the training distribution. In this work, we present RIRFlow, a training-free Bayesian framework for RIR inverse problems using flow matching. We derive a flow-consistent analytic prior from the statistical structure of RIRs, eliminating the need for data-driven priors. Specifically, we model RIR as a Gaussian process with exponentially decaying variance, which yields a closed-form Wiener denoiser. This analytic denoiser is integrated as a prior in an existing flow-based inverse solver, where inverse problems are solved via guided posterior sampling. Furthermore, we extend the solver to nonlinear and non-Gaussian inverse problems via a local Gaussian approximation of the guided posterior, and empirically demonstrate that this approximation remains effective in practice. Experiments on real RIRs across different inverse problems demonstrate robust performance, highlighting the effectiveness of combining a classic RIR model with the recent flow-based generative inference.

2025

  1. Challenges to Subcarrier MIMO Precoding and Equalisation with Smooth Phase Responses
    M.A. Bakhit, F.A. Khattak, S.J. Schlecht, G.W. Rice, S. Weiss
    2025 28th International Workshop on Smart Antennas (WSA)
    Abstract

    Precoding for multiple-input multiple-output orthogonal frequency division multiplexing systems is often based on a per-subcarrier singular value decomposition, where phase smoothing is applied to the singular vectors that form the transmit beamformers. We show that such a smooth solution can ideally be based on an analytic singular value decomposition, but for estimated channel matrices is beset by challenges that deny a smooth or even continuous evolution of singular vectors with frequency. We show how such problems can be bypassed by admitting complex-valued singular values or fractional delays, and by exploiting a method analogous to the analytic eigenvalue decomposition to approximate ground truth analytic singular vectors from estimated channel matrices. We present examples and demonstrate some of the capabilities of a proposed algorithm through simulations.

  2. Cropping Room Impulse Responses Using Unimodal Regression of Their Covariance
    Karolina Prawda, Nils Meyer-Kahlen, Sebastian J. Schlecht
    JASA Express Letters
    Abstract

    The presence of unavoidable background noise limits the signal-to-noise ratio in measured room impulse responses (RIRs). A common solution is to crop the RIR to the time interval where the signal dominates the background noise, but finding the correct onset and truncation points is challenging. It usually requires estimating the sound decay rate and noise floor, which is burdened with uncertainty. In this study, we propose an RIR cropping method based on the covariance between two repeated RIRs and its inherent monotonicity. Evaluation on measured RIRs shows the proposed method is highly robust in different scenarios and outperforms state-of-the-art algorithms.

  3. DataRES and PyRES: A Room Dataset and a Python Library for Reverberation Enhancement System Development, Evaluation, and Simulation
    Gian Marco De Bortoli, Karolina Prawda, Philip Coleman, Sebastian J. Schlecht
    International Conference on Digital Audio Effects (Dafx25)
  4. Deep Room Impulse Response Completion
    Jackie Lin, Georg Götz, Sebastian J. Schlecht
    EURASIP Journal on Audio, Speech, and Music Processing
    Abstract

    Rendering immersive spatial audio in virtual reality (VR) and video games demands a fast and accurate generation of room impulse responses (RIRs) to recreate auditory environments plausibly. However, the conventional methods for simulating or measuring long RIRs are either computationally intensive or challenged by low signal-to-noise ratios. This study is propelled by the insight that direct sound and early reflections encapsulate sufficient information about room geometry and absorption characteristics. Building upon this premise, we propose a novel task termed "RIR completion," aimed at synthesizing the late reverberation given only the early portion (50 ms) of the response. To this end, we introduce DECOR, Deep Exponential Completion Of Room impulse responses, a deep neural network structured as an encoder-decoder designed to predict multi-exponential decay envelopes of filtered noise sequences. The proposed method is compared against a much larger adapted state-of-the-art network, and comparable performance shows promising results supporting the feasibility of the RIR completion task. The RIR completion can be widely adapted to enhance RIR generation tasks where fast late reverberation approximation is required.

  5. Delay Optimization towards Smooth Sparse Noise
    Cristóbal Andrade, Sebastian J. Schlecht
    International Conference on Digital Audio Effects (Dafx25)
  6. Efficient Multichannel Auralization Based on the Modal Decomposition of Acoustic Radiance Transfer (MoD-ART)
    Matteo Scerbo, Sebastian J. Schlecht, Randall Ali, Lauri Savioja, Enzo De Sena
    IEEE Transactions on Audio, Speech and Language Processing
    Abstract

    In complex acoustic environments such as multiple connected rooms, reverberation is highly dependent on the posi tions of sound sources and listeners --- not only in terms of early reflections, but late reverberation as well. Modeling this positional dependency accurately is important for immersive, interactive applications such as virtual reality, augmented reality, and video games, where reverberation needs to be adapted in real time as sound sources and listeners move. The recently proposed modal decomposition of acoustic radiance transfer (MoD-ART) method can evaluate position-dependent late reverberation characteristics in real time, based on physical properties of the modeled environment, and it is specifically designed for complex acoustic environments. The reverberation characteristics' auralization (i.e. their application to audio signals) can be accomplished either with convolution or with delay-based reverberators. In this paper, we propose a method to auralize late reverberation efficiently in the presence of multiple sound sources and listeners, based on the MoD-ART model. The proposed method inherits the favorable complexity scaling of MoD-ART's modeling and extends it to the aspect of auralization, enabling the rendering of late reverberation in scenarios with hundreds of interactive sound sources and listeners. Furthermore, the proposed method can model fully dynamic scenarios (meaning both sources and listeners may move) correctly and with no rendering latency.

  7. Estimation of Multi-Slope Amplitudes in Late Reverberation
    Jeremy B. Bai, Sebastian J. Schlecht
    International Conference on Digital Audio Effects (Dafx25)
  8. Evaluating Reverberation Models for Augmented Reality
    Kyung Yun Lee, Nils Meyer-Kahlen, Sebastian J. Schlecht, Vesa Välimäki
    Journal of the Audio Engineering Society
  9. FLAMO: An Open-Source Library for Frequency-Domain Differentiable Audio Processing
    Gloria Dal Santo, Gian Marco De Bortoli, Karolina Prawda, Sebastian J. Schlecht, Vesa Välimäki
    IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP)
    Abstract

    We present FLAMO, a Frequency-sampling Library for Audio-Module Optimization designed to implement and optimize differentiable linear time-invariant audio systems. The library is open-source and built on the frequency-sampling filter design method, allowing for the creation of differentiable modules that can be used stand-alone or within the computation graph of neural networks, simplifying the development of differentiable audio systems. It includes predefined filtering modules and auxiliary classes for constructing, training, and logging the optimized systems, all accessible through an intuitive interface. Practical application of these modules is demonstrated through two case studies: the optimization of an artificial reverberator and an active acoustics system for improved response coloration.

  10. Joint Spectrogram Separation and TDOA Estimation Using Optimal Transport
    Linda Fabiani, Sebastian J Schlecht, Isabel Haasler, Filip Elvander
    Eusipco 2025
    Abstract

    Separating sources is a common challenge in applications such as speech enhancement and telecommunications, where distinguishing between overlapping sounds helps reduce interference and improve signal quality. Additionally, in multichannel systems, correct calibration and synchronization are essential to separate and locate source signals accurately. This work introduces a method for blind source separation and estimation of the Time Difference of Arrival (TDOA) of signals in the time-frequency domain. Our proposed method effectively separates signal mixtures into their original source spectrograms while simultaneously estimating the relative delays between receivers, using Optimal Transport (OT) theory. By exploiting the structure of the OT problem, we combine the separation and delay estimation processes into a unified framework, optimizing the system through a block coordinate descent algorithm. We analyze the performance of the OT-based estimator under various noise conditions and compare it with conventional TDOA and source separation methods. Numerical simulation results demonstrate that our proposed approach can achieve a significant level of accuracy across diverse noise scenarios for physical speech signals in both TDOA and source separation tasks.

  11. Modeling Nonuniform Energy Decay through the Modal Decomposition of Acoustic Radiance Transfer (MoD-ART)
    Matteo Scerbo, Sebastian J. Schlecht, Randall Ali, Lauri Savioja, Enzo De Sena
    IEEE Transactions on Audio, Speech and Language Processing
    Abstract

    Modeling late reverberation in real-time interactive applications is a challenging task when multiple sound sources and listeners are present in the same environment. This is especially problematic when the environment is geometrically complex and/or features uneven energy absorption (e.g. coupled volumes), because in such cases the late reverberation is dependent on the sound sources' and listeners' positions, and therefore must be adapted to their movements in real time.We present a novel approach to the task, named modal decomposition of acoustic radiance transfer (MoD-ART), which can handle highly complex scenarios with efficiency. The approach is based on the geometrical acousticsmethod of acoustic radiance transfer, fromwhich we extract a set of energy decaymodes and their positional relationships with sources and listeners. In this paper, we describe the physical and mathematical significance of MoD-ART, highlighting its advantages and applicability to different scenarios. Through an analysis of the method's computational complexity, we show that it compares very favorably with ray-tracing.We also present simulation results showing thatMoD-ART can capture multiple decay slopes and flutter echoes.

  12. Multi-Shelf Graphic Equalizer
    Sebastian J. Schlecht, Tantep Sinjanakhom, Vesa Välimäki
    Signal Processing
    Abstract

    A graphic equalizer (GEQ) is a standard tool in audio production and effect design. Adjustable gain control frequencies are fixed along the logarithmic frequency axis, and an automatic design method matches the magnitude response to them whenever target gains are changed. Most commonly, the GEQ comprises a set of peak filters centered an octave apart, possibly with a shelving filter at the bottom and top of the frequency range. While accurate designs were proposed, the dynamic range is typically limited to 24 dB. In this paper, we propose two innovations. First, we introduce a GEQ based on shelving filters only, which can cover an extensive dynamic range of over 60 dB. Secondly, we introduce an order-switching technique that combines shelf filters of different order. We demonstrate the performance and advantages of the proposed filter with design examples. The proposed shelf-filter-based GEQ offers a wider dynamic range and a smoother magnitude response than traditional peak-filter-based GEQ designs.

  13. Neural-Network Based Interpolation of Late Reverberation in Coupled Spaces Using the Common Slopes Model
    Orchisama Das, Gloria Dal Santo, Sebastian J. Schlecht, Zoran Cvetković
    2025 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)
  14. Open Educational Resources for Acoustics and Audio Signal Processing Using Jupyter Notebooks and Pyfar
    Fabian Brinkmann, Marco Berzborn, Anne Heimes, Anton Hoyer, Xenofon Karakonstantis, Simon Kersten, Tim Lübeck, Pascal Palenda, Cristobal Andrades, Johannes M. Arend, Nara Hahn, Tobias Jüterbock, Nils Meyer-Kahlen, Artur Paskiewicz, Sebastian J. Schlecht, Frank Schultz, Stefan Weinzierl
    Forum Acusticum / Euronoise 2025: 11th Convention of the European Acoustics Association
  15. Optimizing Tiny Colorless Feedback Delay Networks
    Gloria Dal Santo, Karolina Prawda, Sebastian J. Schlecht, Vesa Välimäki
    EURASIP Journal on Audio, Speech, and Music Processing
    Abstract

    A common bane of artificial reverberation algorithms is spectral coloration in the synthesized sound, typically manifesting as metallic ringing, leading to a degradation in the perceived sound quality. In delay network methods, coloration is more pronounced when fewer delay lines are used. This paper presents an optimization framework in which a tiny differentiable feedback delay network, with as few as four delay lines, is used to learn a set of parameters to iteratively reduce coloration. The parameters under optimization include the feedback matrix, as well as the input and output gains. The optimization objective is twofold: to maximize spectral flatness through a spectral loss while maintaining temporal density by penalizing sparseness in the parameter values. A favorable narrow distribution of modal excitation is achieved while maintaining the desired impulse response density. In a subjective assessment, the new method proves effective in reducing perceptual coloration of late reverberation. Compared to the author's previous work, which serves as the baseline and utilizes a sparsity loss in the time domain, the proposed method achieves computational savings while maintaining performance. The effectiveness of this work is demonstrated through two application scenarios where smooth-sounding synthetic room impulse responses are obtained via the introduction of attenuation filters and an optimizable scattering feedback matrix.

  16. Perceptual Decorrelator Based on Resonators
    Jon Fagerström, Nils Meyer-Kahlen, Sebastian J. Schlecht, Vesa Välimäki
    International Conference on Digital Audio Effects (Dafx25)
  17. Polynomial Eigenvalue Decomposition for Eigenvalues with Unmajorised Ground Truth -- Reconstructing Analytic Dinosaurs
    Sebastian J. Schlecht, Stephan Weiss
    Science Talks
    Abstract

    When estimated space-time covariance matrices from finite data, any intersections of ground truth eigenvalues will be obscured, and the exact eigenvalues become spectrally majorised with probability one. In this paper, we propose a novel method for accurately extracting the ground truth analytic eigenvalues from such estimated space-time covariance matrices. The approach operates in the discrete Fourier transform (DFT) domain and groups sufficiently eigenvalues over a frequency interval into segments that belong to analytic functions and then solves a permutation problem to align these segments. Utilising an inverse partial DFT and a linear assignment algorithm, the proposed EigenBone method retrieves analytic eigenvalues efficiently and accurately. Experimental results demonstrate the effectiveness of this approach in reconstructing eigenvalues from noisy estimates. Overall, the proposed method offers a robust solution for approximating analytic eigenvalues in scenarios where state-of-the-art methods may fail.

  18. Polynomial Procrustes Solution for Randomly Perturbed Near-Paraunitary Systems
    Stephan Weiss, Sebastian J. Schlecht, Marc Moonen
    2025 IEEE Statistical Signal Processing Workshop (SSP)
    Abstract

    We want to recover paraunitary matrices under small random perturbations. The polynomial Procrustes method, based on the analytic singular value decomposition of the perturbed system, in principle solves this. For small random perturbations, where the analytic singular values are close to unity, we propose a simplified polynomial Procrustes method that exploits this property, but show that the support of the solution is generally increased compared to the perturbed matrix. We therefore embed the simplified Procrustes method into an iterative truncation scheme, which can reduce the support while ensuring that a paraunitary approximation remains within a perimeter that is equivalent to the level of perturbation.

  19. Zero-Phase Sound via Giant FFT
    Vesa Välimäki, Stefan Bilbao, Sebastian J. Schlecht, Roope Salmi, David Zicarelli
    International Conference on Digital Audio Effects (Dafx25)

2024

  1. A Common-Slopes Late Reverberation Model Based on Acoustic Radiance Transfer
    Matteo Scerbo, Sebastian J. Schlecht, Randall Ali, Lauri Savioja, Enzo De Sena
    Proceedings of the 27-Th Int. Conf. on Digital Audio Effects (Dafx24)
  2. A Multi-Room Transition Dataset for Blind Estimation of Energy Decay
    Philipp Götz, Georg Götz, Nils Meyer-Kahlen, Kyung Yun Lee, Karolina Prawda, Emanuël A. P. Habets, Sebastian J. Schlecht
    2024 18th Int. Work. Acoust. Signal Enhanc. (IWAENC)
    Abstract

    We present an acoustic room impulse response dataset focusing on multi-room environments characterized by complex geometries and acoustic conditions. The dataset is accompanied by positional information and visual documentation in the form of 360$^$ photographs. Additionally, we present a method to estimate acoustic energy decay functions from noisy, reverberant speech signals. We demonstrate the effectiveness of the proposed method based on experiments conducted with the presented dataset.

  3. Active Acoustics with a Phase Cancelling Modal Reverberator
    Gian Marco de Bortoli, Karolina Prawda, Sebastian J. Schlecht
    Journal of the Audio Engineering Society
  4. Audiovisual Congruence and Localization Performance in Virtual Reality: 3D Loudspeaker Model vs. Human Avatar
    Anja Hofmann, Nils Meyer-Kahlen, Sebastian J Schlecht, Tapio Lokki
    Journal of the Audio Engineering Society
    Abstract

    This paper investigates audiovisual congruence in virtual reality with both horizontal and vertical offsets between audio and visual rendering. Audiovisual congruence and localization errors are assessed using loudspeaker playback and nonindividualized headphone rendering. To account for the influence of different types of visual information on congruence, presentations of a loudspeaker model and 3D human avatar were compared. Therefore, a new dataset of audiovisual speech was recorded. Results show that human avatar rendering increases perceived congruence, and experienced listeners have an increased tendency to respond with ``incongruent'' when a loudspeaker model is shown but not when the human avatar is presented. Moreover, a correlation is found between localization precision and audiovisual congruence for horizontally offset stimuli and avatar presentation. For vertical offsets, the angular range of congruence is generally large, and localization errors are high, so no correlation can be observed between the two. The paper contributes congruence ranges for audiovisual speech in virtual reality, which also has implications for augmented reality telepresence use.

  5. Binaural Dark-Velvet-Noise Reverberator
    Jon Fagerström, Nils Meyer-Kahlen, Sebastian J. Schlecht, Vesa Välimäki
    Proceedings of the 27-Th Int. Conf. on Digital Audio Effects (Dafx24)
  6. Differentiable Active Acoustics - Optimizing Stability via Gradient Descent
    Gian Marco De Bortoli, Gloria Dal Santo, Karolina Prawda, Tapio Lokki, Vesa Välimäki, Sebastian J. Schlecht
    Proceedings of the 27-Th Int. Conf. on Digital Audio Effects (Dafx24)
  7. Directional Distribution of the Pseudo Intensity Vector in Anisotropic Late Reverberation
    Nils Meyer-Kahlen, Sebastian J. Schlecht
    The Journal of the Acoustical Society of America
    Abstract

    The pseudo intensity vector (PIV) is often used to analyze the directional properties of spatial room impulse responses. In the early part of the response, it is capable of estimating the directions of individual reflections. However, thus far, its behaviour in the late field is unclear. Specifically, it is unknown whether anisotropy, i.e., a direction-dependent energy distribution, is captured by the directional estimates. In this study, a closed-form expression of the directional distribution of the pressure-normalized pseudo intensity vector contingent on a general stochastical model of anisotropic fields was analytically derived. This paper shows that the probability density function of this PIV is a multivariate Cauchy distribution, which does indeed depend on the energy distribution of the field, yet the directional distribution has very limited degrees of freedom. The derived distribution is compared to the results of Monte Carlo simulations and fields captured with a microphone array in a real room. These results facilitate better understanding of the behaviour of parametric spatial room impulse response methods and may enable improved directional estimators for anisotropic fields.

  8. Dynamic Late Reverberation Rendering Using the Common-Slope Model
    Georg Götz, Teodors Kerimovs, Sebastian J. Schlecht, Pulkki, Ville
    6th AES International Conference on Audio for Games
  9. Exploring Sauna Impulse Responses
    Karolina Prawda, Nils Meyer-Kahlen, Sebastian J. Schlecht
    INTER-NOISE NOISE-CON Congr. Conf. Proc.
    Abstract

    Sauna is an important element of Finnish tradition and culture, and its properties and features deserve preservation and archiving. Additionally, extreme temperature and humidity in a sauna comprise a unique environment for researching room acoustics in unusual atmospheric conditions. In the present study, we publish a dataset of room impulse response (RIR) measurements of a sauna during the process of warming up and throwing water on the stove. We explore several features of the dataset, such as the impact of the temperature and humidity distributions on the time-of-arrival of reflections, the atmospheric absorption of sound, and spectral coloration of measured RIRs. The results of this research show the influence of environmental factors on RIRs and allow an insight into the distinct room acoustics of a sauna while heating up.

  10. Fade-in Reverberation in Multi-Room Environments Using the Common-Slope Model
    Kyung Yun Lee, Nils Meyer-Kahlen, Georg Götz, U Peter Svensson, Sebastian J Schlecht, Vesa Välimäki
    AES 5th International Conference on Audio for Virtual and Augmented Reality (AVAR)
  11. KLANN: Linearising Long-Term Dynamics in Nonlinear Audio Effects Using Koopman Networks
    Ville Huhtala, Lauri Juvela, Sebastian J. Schlecht
    IEEE Signal Processing Letters
    Abstract

    In recent years, neural network-based black-box modeling of nonlinear audio effects has improved considerably. Present convolutional and recurrent models can model audio effects with long-term dynamics, but the models require many parameters, thus increasing the processing time. In this letter, we propose KLANN, a Koopman-Linearised Audio Neural Network structure that lifts a one-dimensional signal (mono audio) into a high-dimensional approximately linear state-space representation with nonlinear mapping, and then uses differentiable biquad filters to predict linearly within the lifted state-space. Results show that the proposed models match the high performance of the state-of-the-art neural models while having a more compact architecture, reducing the number of parameters by tenfold, and having interpretable components.

  12. Matching Early Reflections of Simulated and Measured RIRs by Applying Sound-Source Directivity Filters
    Anthony Gallien, Karolina Prawda, Sebastian J. Schlecht
    Audio Engineering Society Conference: AES 2024 International Acoustics & Sound Reinforcement Conference
    Abstract

    Acoustic measurements are susceptible to various sources of measurement uncertainty. One significant factor is loudspeaker directivity, which introduces temporal smearing and spectral coloration into room impulse responses (RIRs), predominantly influencing early reflections. Such an artifact affects parametric processing and perceptual evaluation of RIRs and lowers the measurement reproducibility. This study evaluates the impact of loudspeaker directivity on measured RIRs. We acquire directivity filters via measurements in an anechoic chamber, utilizing a custom-made microphone arc. Subsequently, we both capture a series of RIRs in a typical reverberant room and simulate corresponding RIRs with the image-source method (ISM). By convolving the simulations with the correct directivity filters, we match the early reflections of measured and simulated RIRs. Examining the cross-correlation between the simulated and measured RIRs reveals a pronounced likeness for first-order reflections, indicating a substantial influence of the loudspeaker directivity on recorded RIRs. This study is a step towards accounting for the influence of the sound source type and position on RIRs, resulting in better-informed acoustic measurements and higher fidelity of acoustic simulations.

  13. Method for Audio Peak Reduction Using All-Pass Filter
    Sebastian J. Schlecht, Leonardo Fierro, Vesa Välimäki, Juha Backman
  14. Modal Excitation in Feedback Delay Networks
    Sebastian J. Schlecht, Matteo Scerbo, Enzo De Sena, Vesa Välimäki
    IEEE Signal Processing Letters
    Abstract

    Feedback delay networks (FDNs) are used in audio processing and synthesis. The modal shapes of the system describe the modal excitation by input and output signals. Previously, the Ehrlich-Aberth method was used to find modes in large FDNs. Here, the method is extended to the corresponding eigenvectors indicating the modal shape. In particular, the computational complexity of the proposed analysis method does not depend on the delay-line lengths and is thus suitable for large FDNs, such as artificial reverberators. We show the relation between the compact generalized eigenvectors in the delay state space and the spatially extended modal shapes in the state space. We illustrate this method with an example FDN in which the suggested modal excitation control does not increase the computational cost. The modal shapes can help optimize input and output gains. This letter teaches how selecting the input and output points along the delay lines of an FDN adjusts the spectral shape of the system output.

  15. Non-Exponential Reverberation Modeling Using Dark Velvet Noise
    Jon Fagerström, Sebastian J. Schlecht, Vesa Välimäki
    Journal of the Audio Engineering Society
    Abstract

    Previous research on late-reverberation modeling has mainly focused on exponentially decaying room impulse responses, whereas methods for accurately modeling non-exponential reverberation remain challenging. This paper extends the previously proposed basic darkvelvet- noise reverberation algorithm and proposes a parametrization scheme for modeling late reverberation with arbitrary temporal energy decay. Each pulse in the velvet-noise sequence is routed to a single dictionary filter that is selected from a set of filters based on weighted probabilities. The probabilities control the spectral evolution of the late-reverberation model and are optimized to fit a target impulse response via non-negative least-squares optimization. In this way, the frequency-dependent energy decay of a target late-reverberation impulse response can be fitted with mean and maximum reverberation-time errors of 4% and 8%, respectively, requiring about 50% less coloration filters than a previously proposed filteredvelvet- noise algorithm. Furthermore, the extended dark-velvet-noise reverberation algorithm allows the modeled impulse response to be gated, the frequency-dependent reverberation time to be modified, and the model's spectral evolution and broadband decay to be decoupled. The proposed method is suitable for the parametric late-reverberation synthesis of various acoustic environments, especially spaces that exhibit a non-exponential energy decay, motivating its use in musical audio and virtual reality.

  16. Non-Stationary Noise Removal from Repeated Sweep Measurements
    Karolina Prawda, Sebastian J. Schlecht, Vesa Välimäki
    JASA Express Letters
    Abstract

    Acoustic measurements using sine sweeps are prone to background noise and non-stationary disturbances. Repeated measurements can be averaged to improve the resulting signal-to-noise ratio. However, averaging leads to poor rejection of non-stationary high-energy disturbances and, in the case of a time-variant environment, causes attenuation at high frequencies. This paper proposes a robust method to combine repeated sweep measurements using across-measurement median filtering in the time-frequency domain. The method, called Mosaic, successfully rejects non-stationary noise, suppresses background noise, and is more robust toward time variation than averaging. The proposed method allows high-quality measurement of impulse responses in a noisy environment.

  17. Paraunitary Approximation of Matrices of Analytic Functions - the Polynomial Procrustes Problem
    Stephan Weiss, Sebastian J. Schlecht, Orchisama Das, Enzo De Sena
    Science Talks
    Abstract

    The best least squares approximation of a matrix, typically e.g. characterising gain factors in narrowband problems, by a unitary one is addressed by the Procrustes problem. Here, we extend this idea to the case of matrices of analytic functions, and characterise a broadband equivalent to the narrowband approach which we term the polynomial Procrustes problem. Its solution relies on an analytic singular value decomposition, and for the case of spectrally majorised, distinct singular values, we demonstrate the application of a suitable algorithm to three problems via simulations: (i) time delay estimation, (ii) paraunitary matrix completion, and (iii) general paraunitary approximations.

  18. RIR2FDN: An Improved Room Impulse Response Analysis and Synthesis
    Gloria Dal Santo, Benoit Alary, Karolina Prawda, Sebastian Schlecht, Vesa Välimäki
    Proceedings of the 27-Th Int. Conf. on Digital Audio Effects (Dafx24)
  19. Reconstructing Analytic Dinosaurs: Polynomial Eigenvalue Decomposition for Eigenvalues with Unmajorised Ground Truth
    Sebastian J. Schlecht, Stephan Weiss
    2024 32nd Eur. Signal Process. Conf. (EUSIPCO)
    Abstract

    This paper proposes a novel method for accurately estimating the ground truth analytic eigenvalues from estimated space-time covariance matrices, where the estimation process obscures any intersection of eigenvalues with probability one. The approach involves grouping sufficiently separated, bin-wise eigenvalues into segments that belong to analytic functions and then solves a permutation problem to align these segments. By leveraging an inverse partial discrete Fourier transform and a linear assignment algorithm, the proposed EigenBone method retrieves analytic eigenvalues efficiently and accurately. Experimental results demonstrate the effectiveness of this approach in accurately reconstructing eigenvalues from noisy estimates. Overall, the proposed method offers a robust solution for approximating analytic eigenvalues in scenarios where state-of-the-art methods may fail.

  20. Recording a Dataset of Audiovisual Speech for AR Telepresence Studies
    Anja Hofmann, Nils Meyer-Kahlen, Sebastian J. Schlecht, Tapio Lokki
    DAGA
  21. Short-Time Coherence between Repeated Room Impulse Response Measurements
    Karolina Prawda, Sebastian J. Schlecht, Vesa Välimäki
    The Journal of the Acoustical Society of America
    Abstract

    Room impulse responses (RIRs) vary over time due to fluctuations in atmospheric temperature, humidity, and pressure. This can introduce uncertainties in room transfer-function measurements, which are challenging to account for. Previous methods of identification and compensation of time variance focus on systematic atmospheric changes and do not apply to subtle discrepancies in RIRs. In this work, we address this problem by proposing a model of short-time coherence between repeated RIR measurements as an indicator of time-frequency similarity and as a measure of time-variance-induced changes in RIRs. Atmospheric changes cause fluctuation in sound speed, which, in turn, results in variation in the time-of-arrival of sound reflections following a Generalized Wiener process. We show that the short-time coherence decreases exponentially with the reflection-path length and propose volatility as a single model parameter determining the coherence decay rate. The proposed model is validated on simulations and measurements, showing applicability in indoor scenarios. The method reliably estimates volatility of 10-6 s/s as measured under laboratory conditions. We exemplify the utility of short-time coherence loss by predicting the high-frequency energy loss stemming from RIR averaging. The proposed method is useful in assessing the uncertainty of RIR measurements, especially when repeated measurements are compared or averaged.

  22. Similarity Metrics for Late Reverberation
    Gloria Dal Santo, Karolina Prawda, Sebastian J. Schlecht, Vesa Välimäki
    58th Asilomar Conference on Signals, Systems, and Computers
  23. Testing Auditory Illusions in Augmented Reality: Plausibility, Transfer-Plausibility, and Authenticity
    Nils Meyer-Kahlen, Sebastian J Schlecht, Sebastià Amengual Garí, Tapio Lokki
    Journal of the Audio Engineering Society
    Abstract

    Experiments testing sound for augmented reality can involve real and virtual sound sources. Paradigms are either based on rating various acoustic attributes or testing whether a virtual sound source is believed to be real (i.e., evokes an auditory illusion). This study compares four experimental designs indicating such illusions. The first is an ABX task suitable for evaluation under the authenticity paradigm. The second is a Yes/No task, as proposed to evaluate plausibility. The third is a three-alternative-forced-choice (3AFC) task using different source signals for real and virtual, proposed to evaluate transfer-plausibility. Finally, a 2AFC task was tested. The renderings compared in the tests encompassed mismatches between real and virtual room acoustics. Results confirm that authenticity is hard to achieve under nonideal conditions, and ceiling effects occur because differences are always detected. Thus, the other paradigms are better suited for evaluating practical augmented reality audio systems. Detection analysis further shows that the 3AFC transfer-plausibility test is more sensitive than the 2AFC task. Moreover, participants are more sensitive to differences between real and virtual sources in the Yes/No task than theory predicts. This contribution aims to aid in selecting experimental paradigms in future experiments regarding perceptual and technical requirements for sound in augmented reality.

  24. Time-Frequency Audio Similarity Using Optimal Transport
    Linda Fabiani, Sebastian J. Schlecht, Filip Elvander
    2024 58th Asilomar Conf. Signals, Syst., Comput.
    Abstract

    In audio signal processing, having an effective metric for comparing audio data is essential to ensure an accurate understanding of sound properties and attributes. In this work, we formulate two novel approaches for measuring the similarity between audio signals in the time-frequency domain, taking advantage of principles from classical optimal transport problems and sliced Wasserstein distances. Using optimal transport to construct the metric allows for a more robust signal content comparison, considering not only the signals' individual elements but also the global distribution in the signal space. Additionally, the sliced Wasserstein methods expand the use of the distances to high dimensional problems. By integrating both time and frequency aspects into our metrics, we aim for a more comprehensive comparison that can better handle various types of signal distortions. Results show promising behavior in accurately measuring distances for increasing signal differences and avoiding the presence of local minima in the loss curves.

  25. Two-Stage Attenuation Filter for Artificial Reverberation
    Vesa Välimäki, Karolina Prawda, Sebastian J. Schlecht
    IEEE Signal Processing Letters
    Abstract

    Delay networks are a common parametric method to synthesize the late part of the room reverberation. A delay network consists of several feedback loops, each containing a delay line and an attenuation filter, which approximates the same decay rate by appropriately setting the frequency-dependent loop gain. A remaining challenge is the design of the attenuation filters on a wide frequency range based on a measured room impulse response. This letter proposes a novel two-stage attenuation filter structure, sharpening the design. The first stage is a low-order pre-filter approximating the overall shape and determining the decay at the two ends of the frequency range, namely at the dc and the Nyquist limit. The second filter, an equalizer, fine-tunes the gain at different frequencies, such as on one-third-octave bands. It is shown that the proposed design is more accurate and robust than previous methods. A design example applying the proposed method to an interleaved velvet-noise reverberator is also exhibited. The proposed two-stage attenuation filter is a step toward a realistic parametric simulation of measured room impulse responses.

2023

  1. Auralization of Measured Room Transitions in Virtual Reality
    Thomas McKenzie, Nils Meyer-Kahlen, Christoph Hold, Sebastian J. Schlecht, Ville Pulkki
    Journal of the Audio Engineering Society
  2. Bounded-Magnitude Discrete Fourier Transform [Tips & Tricks]
    Sebastian J Schlecht, Vesa Välimäki, Emanuël A P Habets
    IEEE Signal Processing Magazine
    Abstract

    Analyzing the magnitude response of a finite-length sequence is a ubiquitous task in signal processing. However, the discrete Fourier transform (DFT) provides only discrete sampling points of the response characteristic. This work introduces bounds on the magnitude response, which can be efficiently computed without additional zero padding. The proposed bounds can be used for more informative visualization and inform whether additional frequency resolution or zero padding is required.

  3. Common-Slope Modeling of Late Reverberation
    Georg Götz, Sebastian J. Schlecht, Ville Pulkki
    IEEE/ACM Transactions on Audio, Speech and Language Processing
    Abstract

    The decaying sound field in rooms is typically described by energy decay functions (EDFs). Late reverberation can deviate considerably from the ideal diffuse field, for example, in multiple connected rooms or non-uniform absorption material distributions. This paper proposes the common-slope model of late reverberation. The model describes spatial and directional late reverberation as linear combinations of exponential decays called common slopes. Its fundamental idea is that common slopes have decay times that are invariant across space and direction, while their amplitudes vary across both. We explore different approaches for determining the common slopes for large EDF sets describing different source-receiver configurations of the same environment. Among the presented approaches, the k-means clustering of decay times is the most general. Our evaluation shows that the common-slope model introduces only a small error between the modeled and the true EDF, while being considerably more compact than the traditional multi-exponential model. The amplitude variations of the common slopes yield interpretable room acoustic analyses. The common-slope model has potential applications in all fields relying on late reverberation models, such as source separation, dereverberation, echo cancellation, and parametric spatial audio rendering.

  4. Decorrelation in Feedback Delay Networks
    Sebastian J. Schlecht, Jon Fagerström, Vesa Välimäki
    IEEE/ACM Transactions on Audio, Speech and Language Processing
    Abstract

    The feedback delay network (FDN) is a popular filter structure to generate artificial spatial reverberation. A common requirement for multichannel late reverberation is that the output signals are well decorrelated, as too high a correlation can lead to poor reproduction of source image and uncontrolled coloration. This article presents the analysis of multichannel correlation induced by FDNs. It is shown that the correlation depends primarily on the feedforward paths, while the long reverberation tail produced by the recursive path does not contribute to the inter-channel correlation. The impact of the feedback matrix type, size, and delays on the inter-channel correlation is demonstrated. The results show that small FDNs with a few feedback channels tend to have a high inter-channel correlation, and that the use of a filter feedback matrix significantly improves the decorrelation, often leading to the lowest inter-channel correlation among the tested cases. The learnings of this work support the practical design of multichannel artificial reverberators for immersive audio applications.

  5. Deep Learning for Loudspeaker Digital Twin Creation
    Bryn Louise, Teodors Kerimovs, Sebastian J. Schlecht
    154th Convention of the Audio Engineering Society
  6. Design with Sound: The Relevance of Sound in VR as an Immersive Design Tool for Landscape Architecture
    Loviisa Luoma, Pia Fricker, Sebastian J. Schlecht
    Journal of Digital Landscape Architecture (JoDLA)
  7. Differentiable Feedback Delay Network for Colorless Reverberation
    Gloria Dal Santo, Karolina Prawda, Sebastian J. Schlecht, Vesa Välimäki
    International Conference on Digital Audio Effects (DAF 23)
  8. Distribution of Modal Damping in Absorptive Shoebox Rooms
    Maximilian Schäfer, Karolina Prawda, Rudolf Rabenstein, Sebastian J. Schlecht
    IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)
    Abstract

    The image-source method is widely applied to compute room impulse responses (RIRs) of shoebox rooms with arbitrary damping. However, with increasing RIR lengths, the number of image sources grows rapidly, leading to slow computation. We propose a method to estimate the damping density of a damped shoebox room, which in turn can provide the energy decay necessary to model the stochastic late reverberation. The damping density is derived from a modal decomposition that is compliant with the ISM solution. We show that the proposed method gives a more accurate estimate of the energy decay than previous methods and can be efficiently computed regardless of the RIR lengths. While we focus on the derivation and evaluation, the main practical applications of the proposed model include, e.g., the faster synthesis of late reverb and the analysis of multi-slope decays.

  9. Grouped Feedback Delay Networks with Frequency-Dependent Coupling
    Orchisama Das, Sebastian J. Schlecht, Enzo De Sena
    IEEE/ACM Transactions on Audio, Speech and Language Processing
    Abstract

    Feedback Delay Networks are one of the most popular and efficient means of generating artificial reverberation. Recently, we proposed the Grouped Feedback Delay Network (GFDN), which couples multiple FDNs while maintaining system stability. The GFDN can be used to model reverberation in coupled spaces that exhibit multi-stage decay. The block feedback matrix determines the inter- and intra-group coupling. In this article, we expand on the design of the block feedback matrix to include frequency-dependent coupling among the various FDN groups. We show how paraunitary feedback matrices can be designed to emulate diffraction at the aperture connecting rooms. Several methods for the construction of nearly paraunitary matrices are investigated. The proposed method supports the efficient rendering of virtual acoustics for complex room topologies in games and XR applications.

  10. How Smooth Do You Think I Am: An Analysis on the Frequency Dependent Temporal Roughness of Velvet Noise
    Jade Roberts, Jon Fagerström, Sebastian J. Schlecht, Vesa Välimäki
    Int. Conf. on Digital Audio Effects (Dafx23)
  11. Inside The Quartet - A First-Person Virtual Reality String Quartet Production (Non-Peer-Reviewed)
    Nils Meyer-Kahlen, Petra Piiroinen, Gautam Vishwanath, Petri Juntunen, Eero Tiainen, Sebastian J. Schlecht
    154th Convention of the Audio Engineering Society
  12. Interpolation of Spatial Room Impulse Responses Using Partial Optimal Transport
    Aaron Geldert, Nils Meyer-Kahlen, Sebastian J Schlecht
    IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
  13. Measuring Motion-to-Sound Latency in Virtual Acoustic Rendering Systems
    Nils Meyer-Kahlen, Miranda Kastemaa, Sebastian J. Schlecht, Tapio Lokki
    Journal of the Audio Engineering Society
  14. Modifying Partials for Minimum-Roughness Sound Synthesis
    Simon Schwär, Meinard Müller, Sebastian J. Schlecht
    Proceedings of the 3rd International Conference on Timbre
  15. Polynomial Procrustes Problem: Paraunitary Approximation of Matrices of Analytic Functions
    Stephan Weiss, Sebastian J. Schlecht, Orchisama Das, Enzo de Sena
    European Signal Processing Conference (EUSIPCO)
  16. Short-Term Rule of Two: Localizing Non-Stationary Noise Events in Swept-Sine Measurements
    Karolina Prawda, Sebastian J. Schlecht, Vesa Välimäki
    154th Convention of the Audio Engineering Society
  17. Source Position Interpolation of Spatial Room Impulse Responses
    Thomas McKenzie, Sebastian J. Schlecht
    154th Convention of the Audio Engineering Society
  18. The Role of Source Signal Similarity in Distinguishing between Different Positions in a Room
    Thomas Mckenzie, Nils Meyer-Kahlen, Sebastian J. Schlecht
    AES International Conference on Spatial and Immersive Audio
    Abstract

    Typically, evaluation of spatial audio systems uses the same source signal for each condition in listening comparison tests (such as ABX and MUSHRA). However in an augmented reality scenario, it is unlikely that the exact same source signal would exist at the exact same position in space, both real and virtual: instead, a real source would be in one position in the room and a virtual source in a different position, both with different source signals. A perceptual study is presented on the effect of source signal similarity when distinguishing different positions in a room. Three source signal types (all speech) are investigated in a multiple stimulus paradigm: the same source signal for all conditions, the same speaker but a different sentence for each condition, and a different speaker and different sentence for each condition. Results show that the source signal similarity significantly impacts the similarity rating between different receiver positions in the same room, which suggests that spatial audio system fidelity requirements could vary depending on the source signal types used in the target application.

  19. Time Variance in Measured Room Impulse Responses
    Karolina Prawda, Sebastian J. Schlecht, Vesa Välimäki
    Forum Acusticum

2022

  1. A Variational Y-Autoencoder for Disentangling Gesture and Material of Interaction Sounds
    Simon Schwär, Meinard Müller, Sebastian J. Schlecht
    AES 2022 International Audio for Virtual and Augmented Reality Conference
  2. Analysis of Multi-Exponential and Anisotropic Sound Energy Decay
    Georg Götz, Christoph Hold, Thomas McKenzie, Sebastian J. Schlecht, Ville Pulkki
    DAGA - Jahrestagung Für Akustik
    Abstract

    Reverberation is an important cue for determining the distance and location of sounds, and can be used to infer the size of a space. Although simple models of reverber- ation assume a single exponentially decaying late rever- beration tail that is independent of direction, in practice, many rooms feature multiple exponential decays with di- rectional anisotropy [1, 2, 3]. In this paper, a framework for the directional analysis of spatial room impulse response reverberation decays is presented. The framework uses a recent neural net- work approach for multi-exponential decay analysis in conjunction with a spherical filterbank analysis. Addi- tionally, the paper introduces the common-slope model of directional reverberation, in which directional or spa- tial decay variations are described in terms of exponential amplitudes, while fixing the corresponding decay times.

  3. Audio Peak Reduction Using Ultra-Short Chirps
    Vesa Välimäki, Leonardo Fierro, Sebastian J. Schlecht, Juha Backman
    Journal of the Audio Engineering Society
    Abstract

    Two filtering methods for reducing the peak value of audio signals are studied. Both methods essentially warp the signal phase while leaving its magnitude spectrum unchanged. The first technique, originally proposed by Lynch in 1988, consists of a wideband linear chirp. The listening test presented here shows that the chirp must not be longer than 4 ms so as not to cause any audible change in timbre. The second method, called the phase rotator, put forward in 2001 by Orban and Foti is based on a cascade of second-order allpass filters. This work proposes extensions to improve the performance of the methods, including rules to choose the parameter values. A comparison with previous methods in terms of achieved peak reduction, using a collection of short audio signals, is presented. The computational load of both methods is sufficiently low for real-time application. The extended phase rotator method is found to be superior to the linear chirp method and comparable to the other search methods. The practical peak reduction obtained with the proposed methods spans from 0 to about 3.5 dB. The signal processing methods presented in this work can increase loudness or save power in audio playback.

  4. Audio Peak Reduction Using a Synced Allpass Filter
    Sebastian J Schlecht, Leonardo Fierro, Vesa Valimaki, Juha Backman
    IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
  5. Blind Directional Room Impulse Response Parameterization from Relative Transfer Functions
    Nils Meyer-Kahlen, Sebastian J. Schlecht
    International Workshop on Acoustic Signal Enhancement (IWAENC)
  6. Calibrating the Sabine and Eyring Formulas
    Karolina Prawda, Sebastian J. Schlecht, Vesa Välimäki
    The Journal of the Acoustical Society of America
    Abstract

    Of the many available reverberation time prediction formulas, Sabine's and Eyring's equations are still widely used. The assumptions of homogeneity and isotropy of sound energy during the decay associated with those models are usually recognized as a reason for lack of agreement between predictions and measurements. At the same time, the inaccuracy in the estimation of the sound-absorption coefficient adds to the uncertainty of calculations. This paper shows that the error of incorrectly assumed sound absorption is more detrimental to the prediction precision than the inherent error in the formulas themselves. The proposed absorption calibration procedure reduces the differences between the measured and predicted reverberation time values, showing that an accuracy within 10% from the target reverberation time values can be achieved regardless of the absorption distribution in a room. The paper also discusses the oft neglected air absorption of sound, which may introduce considerable bias to the measurement results. The need for an air-absorption compensation procedure is highlighted, and a method for the estimation of its parameters in octave bands is proposed and compared with other approaches. The results of this study provide justification for the use of the Sabine and Eyring formulas for reverberation time predictions.

  7. Clearly Audible Room Acoustical Differences May Not Reveal Where You Are in a Room
    Nils Meyer-Kahlen, Sebastian J. Schlecht, Tapio Lokki
    The Journal of the Acoustical Society of America
    Abstract

    A common aim in virtual reality room acoustics simulation is accurate listener position dependent rendering. However, it is unclear whether a mismatch between the acoustics and visual representation of a room influences the experience or is even noticeable. Here, we ask if listeners without any special experience in echolocation are able to identify their position in a room based on the acoustics alone. In a first test, direct comparison between acoustic recordings from the different positions in the room revealed clearly audible differences, which subjects described with various acoustic attributes. The design of the subsequent experiment allows participants to move around and explore the sound within different zones in this room while switching between visual renderings of the zones in a head-mounted display. The results show that identification was only possible in some special cases. In about 74% of all trials, listeners were not able to determine where they were in the room. The results imply that audible position dependent room acoustic rendering in virtual reality may not be noticeable under certain conditions, which highlights the importance of evaluation paradigm choice when assessing virtual acoustics.

  8. Colours of Velvet Noise
    Nils Meyer-Kahlen, Sebastian J. Schlecht, Vesa Välimäki
    Electronics Letters
    Abstract

    Velvet noise is a sparse ternary pseudo-random signal containing only a small portion of non-zero values. In this work, the derivation of the spectral properties of velvet noise is presented. In particular, it is shown that the original velvet noise is white, i.e. has a constant power spectrum. For velvet noise variants with altered probability of polarity, the spectral characteristics are analytically derived. Crushed additive velvet noise is shown to have potential in the design of coloured sparse noise sequences, which are useful in acoustic signal processing.

  9. Common-Slope Modeling of Late Reverberation in Coupled Rooms
    Georg Götz, Sebastian J. Schlecht, Ville Pulkki
    Proceedings of the 24th International Congress on Acoustics (ICA)
    Abstract

    Coupled rooms have a distinct sound energy decay behavior, which exhibits more than one decay time under certain conditions. The sound energy decay analysis in such scenarios requires decay models consisting of multiple exponentials with distinct decay rates and amplitudes. While multi-exponential decay analysis is commonly used in room acoustics, the spatial and directional sound energy decay variations in coupled rooms have received little attention. In this work, we introduce the common-slope model of late reverberation for coupled rooms. Common slopes are spatially and directionally invariant decay functions over time, whose amplitudes model all decay variations with respect to the source-receiver configuration. For example, in a scene consisting of two coupled rooms, it is possible to determine two common decay times that approximate the decay for all source-receiver configurations in the scene. Consequently, all spatial and directional decay variations are expressed via decay amplitudes only. We apply the common-slope analysis to measurements of room transitions between coupled rooms. Our analysis shows that the common-slope model approximates the measured sound energy decay with little error. The proposed common-slope model can be used for room acoustic analysis and the efficient synthesis of artificial late reverberation tails.

  10. Dark Velvet Noise
    Jon Fagerström, Nils Meyer-Kahlen, Sebastian J. Schlecht, Vesa Välimäki
    Proceedings of the 25-Th Int. Conf. on Digital Audio Effects (Dafx20in22)
  11. Multichannel Interleaved Velvet Noise
    Karolina Prawda, Sebastian J. Schlecht, Vesa Välimäki
    Proceedings of the 25-Th Int. Conf. on Digital Audio Effects (Dafx20in22)
  12. Neural Network for Multi-Exponential Sound Energy Decay Analysis
    Georg Götz, Ricardo Falcón Pérez, Sebastian J. Schlecht, Ville Pulkki
    The Journal of the Acoustical Society of America
    Abstract

    An established model for sound energy decay functions (EDFs) is the superposition of multiple exponentials and a noise term. This work proposes a neural-network-based approach for estimating the model parameters from EDFs. The network is trained on synthetic EDFs and evaluated on two large datasets of over 20 000 EDF measurements conducted in various acoustic environments. The evaluation shows that the proposed neural network architecture robustly estimates the model parameters from large datasets of measured EDFs while being lightweight and computationally efficient. An implementation of the proposed neural network is publicly available.

  13. Perceptually Informed Interpolation and Rendering of Spatial Room Impulse Responses for Room Transitions
    Thomas McKenzie, Nils Meyer-Kahlen, Rapolas Daugintis, Leo McCormack, Sebastian J. Schlecht, Ville Pulkki
    Proceedings of the 24th International Congress on Acoustics (ICA)
  14. Physical Modeling Using Recurrent Neural Networks with Fast Convolutional Layers
    Julian D. Parker, Sebastian J. Schlecht, Rudolf Rabenstein, Maximilian Schäfer
    Proceedings of the 25-Th Int. Conf. on Digital Audio Effects (Dafx20in22)
  15. Predicting Perceptual Transparency of Head-Worn Devices
    Pedro Lladó, Thomas Mckenzie, Nils Meyer-Kahlen, Sebastian J. Schlecht
    Journal of the Audio Engineering Society
  16. Resynthesis of Spatial Room Impulse Response Tails with Anisotropic Multi-Slope Decays
    Christoph Hold, Thomas Mckenzie, Georg Götz, Sebastian J. Schlecht, Ville Pulkki
    Journal of the Audio Engineering Society
  17. Robust Selection of Clean Swept-Sine Measurements in Non-Stationary Noise
    Karolina Prawda, Sebastian J Schlecht, Vesa Välimäki
    The Journal of the Acoustical Society of America
    Abstract

    The exponential sine sweep is a commonly used excitation signal in acoustic measurements, which, however, is susceptible to non-stationary noise. This paper shows how to detect contaminated sweep signals and select clean ones based on a procedure called the rule of two, which analyzes repeated sweep measurements. A high correlation between a pair of signals indicates that they are devoid of non-stationary noise. The detection threshold for the correlation is determined based on the energy of background noise and time variance. Not being disturbed by non-stationary events, a median-based method is suggested for reliable background noise energy estimation. The proposed method is shown to detect reliably 95% of impulsive noises and 75% of dropouts in the synthesized sweeps. Tested on a large set of measurements and compared with a previous method, the proposed method is shown to be more robust in detecting various non-stationary disturbances, improving the detection rate by 30 percentage points. The rule-of-two procedure increases the robustness of practical acoustic and audio measurements.

  18. The Auditory Perceived Aperture Position of the Transition between Rooms
    Thomas McKenzie, Sebastian J. Schlecht, Ville Pulkki
    The Journal of the Acoustical Society of America
    Abstract

    This exploratory study investigates the phenomenon of the auditory perceived aperture position (APAP): the point at which one feels they are in the boundary between two adjoined spaces, judged only using auditory senses. The APAP is likely the combined perception of multiple simultaneous auditory cue changes, such as energy, reverberation time, envelopment, decay slope shape, and the direction, amplitude, and colouration of direct and reverberant sound arrivals. A framework for a rendering-free listening test is presented and conducted in situ, avoiding possible inaccuracies from acoustic simulations, impulse response measurements, and auralisation to assess how close the APAP is to the physical aperture position under blindfold conditions, for multiple source positions and two room pairs. Results indicate that the APAP is generally within 1 m of the physical aperture position, though reverberation amount, listener orientation, and source position affect precision. Comparison to objective metrics suggests that the APAP generally falls within the period of greatest acoustical change. This study illustrates the non-trivial nature of acoustical room transitions and the detail required for their plausible reproduction in dynamic rendering and game audio engines.

  19. Transfer-Plausibility of Binaural Rendering with Different Real-World References
    Nils Meyer-Kahlen, Thomas McKenzie, Sebastian J. Schlecht, Sebastiá V. Amengual Garí, Tapio Lokki
    DAGA - Jahrestagung Für Akustik
    Abstract

    The evaluation of virtual acoustics in extended realities (XR), where real and virtual sound sources may co-exist, can be performed using three paradigms: authenticity, plausibility and transfer-plausibility, see Fig. 2. In this paper, we first revisit these three paradigms, before describing a transfer-plausibility experiment, in which participants are asked to identify a virtual source amongst different real sources.

2021

  1. A Dataset of Higher-Order Ambisonic Room Impulse Responses and 3D Models Measured in a Room with Varying Furniture
    Georg Götz, Sebastian J. Schlecht, Ville Pulkki
    International Conference on Immersive and 3D Audio (I3DA)
    Abstract

    This paper presents Motus, a new dataset of higher-order Ambisonic room impulse responses. The measurements took place in a single room while varying the amount and placement of furniture. 830 different room configurations were measured with four source-to-receiver configurations, resulting in 3320 room impulse responses in total. The dataset features various furniture object placements, including non-uniform distributions of absorptive material and cases with occluded direct paths between source and receiver. All acoustic measurements are accompanied by matching 3D models and 360$^$-photographs of the room. After describing the dataset, we demonstrate its usage with a reverberation time analysis. The analysis reveals that most of our measurements follow the expected relationship between absorption area and reverberation time. Some exceptional cases feature particular room acoustic phenomena, such as non-uniform absorption area distributions or multi-slope decays. Additionally, we show with a large number of measurements that furniture placement can significantly affect the reverberation time of a room. The dataset can be used to investigate room acoustic topics such as the acoustic effects of absorber placements or the decay behavior of rooms.

  2. Acoustic Analysis and Dataset of Transitions between Coupled Rooms
    Thomas McKenzie, Sebastian J Schlecht, Ville Pulkki
    IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
    Abstract

    The measurement of room acoustics plays a wide role in audio research, from physical acoustics modelling and virtual reality applications to speech enhancement. While vast literature exists on position-dependent room acoustics and coupling of rooms, little has explored the transition from one room to its neighbour. This paper presents the measurement and analysis of a dataset of room impulse responses for the transition between four coupled room pairs. Each transition consists of 101 room impulse responses recorded using a fourth-order spherical microphone array in 5intervals, both with and without a continuous line-of-sight between the source and microphone. A numerical analysis of the room transitions is then presented, including direct-to-reverberant ratio and direction of arrival estimations, along with potential applications and uses of the dataset.

  3. Allpass Feedback Delay Networks
    Sebastian J. Schlecht
    IEEE Transactions on Signal Processing
    Abstract

    In the 1960s, Schroeder and Logan introduced delay line-based allpass filters, which are still popular due to their computational efficiency and versatile applicability in artificial reverberation, decorrelation, and dispersive system design. In this work, we extend the theory of allpass systems to any arbitrary connection of delay lines, namely feedback delay networks (FDNs). We present a characterization of uniallpass FDNs, i.e., FDNs, which are allpass for an arbitrary choice of delays. Further, we develop a solution to the completion problem, i.e., given an FDN feedback matrix to determine the remaining gain parameters such that the FDN is allpass. Particularly useful for the completion problem are feedback matrices, which yield a homogeneous decay of all system modes. Finally, we apply the uniallpass characterization to previous FDN designs, namely, Schroeder's series allpass and Gardner's nested allpass for single-input, single-output systems, and, Poletti's unitary reverberator for multi-input, multi-output systems and demonstrate the significant extension of the design space.

  4. Assessing Room Acoustic Memory Using a Yes/No and a 2-AFC Paradigm
    Madalina Nastasa, Nils Meyer-Kahlen, Sebastian J. Schlecht
    Nordic Sound and Music Conference
    Abstract

    We present a study that tests the ability to remember room acoustics -- a cognitive skill that is one of the guiding mechanisms behind plausible virtual acoustics for extended realities. Room acoustic memory was tested by assessing a person's ability to recognise sound samples, convolved with room impulse responses of everyday rooms presented in a preceding training session. To test a common assumption of detection theory, we conducted two listening tests using both a yes/no and a 2AFC paradigm. Results show that subjects can recognise different rooms above chance level, but even with relatively large differences between the rooms, the accuracy is low in general. Furthermore, the relation between the two test paradigms follows the prediction of detection theory when averaging over all participants, but less so for individual participants.

  5. Assessing Room Acoustic Self-Localization Using a Virtual Blindfold
    Nils Meyer-Kahlen, Sebastian J. Schlecht
    DAGA - Jahrestagung Für Akustik
    Abstract

    ``Can you hear where you are in a room?'', is an important question for determining to which extend the acoustic rendering in extended realities with 6 Degrees-ofFreedom (6DoF) needs to be position-dependent. In our experiment, we assess the ability to understand positiondependent room acoustical differences using a novel ``virtual blindfold'' test design. In this design, subjects are asked to associate the sound they hear when walking around a loudspeaker in a certain part of the room, with different position in the model presented visually using a head mounted display (HMD) to chose from.

  6. Auralisation of the Transition between Coupled Rooms
    Thomas McKenzie, Sebastian J. Schlecht, Ville Pulkki
    International Conference on Immersive and 3D Audio (I3DA)
    Abstract

    The perceptual experience of the transition between coupled rooms remains a little investigated area of research. This paper presents a pipeline for auralising the transition between coupled rooms, utilising a time-varying partitioned convolution for fast position-dependent switching between spatial room impulse responses (SRIRs) and parametric binaural rendering over highly acoustically transparent headphones, with in-situ calibration to the corresponding real-world acoustics. The system is verified by an in-situ listening test with both real and virtual stimuli, conducted in six degrees-of-freedom virtual reality with three-dimensional visuals from measured room models. Results show that the auralisation is rated as highly natural, equalling the naturalness of the corresponding real world auditory stimuli. This pipeline is therefore appropriate for testing of coupled room transition algorithms and SRIR interpolation techniques, as well as non-in-situ testing.

  7. Autonomous Robot Twin System for Room Acoustic Measurements
    Georg Götz, Sebastian J. Schlecht, Abraham Martinez Ornelas, Ville Pulkki
    J. Audio Eng. Soc
    Abstract

    Whilst room acoustic measurements can accurately capture the sound field of real rooms, they are usually time consuming and tedious if many positions need to be measured. Therefore, this contribution presents the Autonomous Robot Twin System for Room Acoustic Measurements (ARTSRAM) to autonomously capture large sets of room impulse responses with variable sound source and receiver positions. The proposed implementation of the system consists of two robots, one of which is equipped with a loudspeaker, while the other one is equipped with a microphone array. Each robot contains collision sensors, thus enabling it to move autonomously within the room. The robots move according to a random walk procedure to ensure a big variability between measured positions. A tracking system provides position data matching the respective measurements. After outlining the robot system, this paper presents a validation, in which anechoic responses of the robots are presented and the movement paths resulting from the random walk procedure are investigated. Additionally, the quality of the obtained room impulse responses is demonstrated with a sound field visualization. In summary, the evaluation of the robot system indicates that large sets of diverse and high-quality room impulse responses can be captured with the system in an automated way. Such large sets of measurements will benefit research in the fields of room acoustics and acoustic virtual reality.

  8. Generating Coherence-Constrained Multisensor Signals Using Balanced Mixing and Spectrally Smooth Filters
    Daniele Mirabilii, Sebastian J Schlecht, Emanuël A P Habets
    The Journal of the Acoustical Society of America
    Abstract

    The spatial properties of a noise field can be described by a spatial coherence function. Synthetic multichannel noise signals exhibiting a specific spatial coherence can be generated by properly mixing a set of uncorrelated, possibly non-stationary, signals. The mixing matrix can be obtained by decomposing the spatial coherence matrix. As proposed in a widely used method, the factorization can be performed by means of a Choleski or eigenvalue decomposition. In this work, the limitations of these two methods are discussed and addressed. In particular, specific properties of the mixing matrix are analyzed, namely, the spectral smoothness and the mix balance. The first quantifies the mixing matrix-filters variation across frequency and the second quantifies the number of input signals that contribute to each output signal. Three methods based on the unitary Procrustes solution are proposed to enhance the spectral smoothness, the mix balance, and both properties jointly. A performance evaluation confirms the improvements of the mixing matrix in terms of objective measures. Furthermore, the evaluation results show that the error between the target and the generated coherence is lowered by increasing the spectral smoothness of the mixing matrix.

  9. Machine Learning Based Auralization of Rigid Sphere Scattering
    Stefan Wirler, Sebastian J. Schlecht, Ville Pulkki
    International Conference on Immersive and 3D Audio (I3DA)
    Abstract

    In this paper, we present a method to auralize acoustic scattering and occlusion of a single rigid sphere with parametric filters and neural networks to provide fast processing and estimation of parameters. The filter parameters are estimated using neural networks based on the geometric parameters of the simulated scene, e.g., relative receiver position and size of the rigid spherical scatterer. The modeling differentiates an unoccluded and an occluded source-receiver path, for which different filter structures were used. In contrast to simulating occlusion and scattering numerically or analytically methods, the proposed approach provides rendering with low computational load making it suitable for real-time auralization in virtual reality. The presented method provides a good fit for modeling the acoustic effects of a rigid sphere. Further, a listening test was conducted, which resulted in plausible reproduction of the scattering and occlusion of a rigid sphere.

  10. One-to-Many Conversion for Percussive Samples
    Jon Fagerström, Sebastian J. Schlecht, Vesa Välimäki
    International Conference on Digital Audio Effects (Dafx)
    Abstract

    A filtering algorithm for generating subtle random variations in sampled sounds is proposed. Using only one recording for impact sound effects or drum machine sounds results in unrealistic repetitiveness during consecutive playback. This paper studies spectral variations in repeated knocking sounds and in three drum sounds: a hihat, a snare, and a tomtom. The proposed method uses a short pseudo-random velvet-noise filter and a low-shelf filter to produce timbral variations targeted at appropriate spectral regions, yielding potentially an endless number of new realistic versions of a single percussive sampled sound. The realism of the resulting processed sounds is studied in a listening test. The results show that the sound quality obtained with the proposed algorithm is at least as good as that of a previous method while using 77% fewer computational operations. The algorithm is widely applicable to computer-generated music and game audio.

  11. Parametric Late Reverberation from Broadband Directional Estimates
    Nils Meyer-Kahlen, Sebastian J. Schlecht, Tapio Lokki
    International Conference on Immersive and 3D Audio (I3DA)
    Abstract

    Several parametric spatial room impulse response rendering methods use broadband directional estimates, whereby based on sample-by-sample direction-of-arrival estimation, a single channel room impulse response is distributed to multiple loudspeakers. To this end, it has been unclear how such simple parametric processing behaves in the late part of the response. To assess this question, we use simulations and a measurement to show that the commonly applied estimation methods based on the pseudo intensity vector and time difference of arrival estimation do preserve the directional information in the late response. Also, we show that estimated directional differences can be audible under best case conditions. As broadband rendering can sound `rough' or `grainy' for transient input signals due to insufficient pulse density in individual reproduction channels, we use a method to synthesize smooth sounding spatial reverberation. For this, the broadband estimates are used to calculate directional energy envelopes, which are applied to filtered noise sequences. The findings presented here help assessing and improving spatial room impulse response processing methods.

  12. Perceptual Analysis of Directional Late Reverberation
    Benoit Alary, Pierre Massé, Sebastian J Schlecht, Markus Noisternig, Vesa Välimäki
    The Journal of the Acoustical Society of America
    Abstract

    The late reverberation characteristics of a sound field are often assumed to be perceptually isotropic, meaning that the decay of energy is perceived as equivalent in every direction. In this paper, we employ Ambisonics reproduction methods to reassess how a decaying sound field is analyzed and characterized and our capacity to hear directional characteristics within late reverberation. We propose the use of objective measures to assess the anisotropy characteristics of a decaying sound field. The energy-decay deviation is defined as the difference of the direction-dependent decay from the average decay. A perceptual study demonstrates a positive link between the range of these energy deviations and their audibility. These results suggest that accurate sound reproduction should account for directional properties throughout the decay.

  13. Perceptual Roughness of Spatially Assigned Sparse Noise for Rendering Reverberation
    Nils Meyer-Kahlen, Sebastian J. Schlecht, Tapio Lokki
    The Journal of the Acoustical Society of America
    Abstract

    Multichannel auralizations based on spatial room impulse responses often employ sample-wise assignment of an omnidirectional response to form loudspeaker responses. This leads to sparse impulse responses in each reproduction loudspeaker and the auralization of transient signals can sound rough. Based on this observation, we conducted a listening test to examine the general phenomenon of roughness due to spatial assignment. First, participants assessed the roughness of both Gaussian noise and velvet noise, assigned sample-wise to up to 36 loudspeakers by two algorithms. The first algorithm assigns channels merely by selecting random indices, while the second one constrains the time between two peaks on each channel. The results show that roughness already occurs when few channels are used and that the assignment algorithm influences it. In a second experiment, virtualizations of the test were used to examine the factors contributing to increased roughness. We systematically show the effect of spatial assignment on noise and conclude that besides time-differences, level-differences caused by head-shadowing are the principal cause for the perceived roughness. The results have significance in spatial room impulse response rendering and spatial reverberator design.

  14. Room Acoustic Parameters Measurements in Variable Acoustic Laboratory Arni
    Karolina Prawda, Sebastian J. Schlecht, Vesa Välimäki
    Akustiikkapäivät
    Abstract

    The possibility to alter a room's acoustic conditions is useful for spaces requiring specific acoustics depending on diverse functions they serve. Such solutions are applied mostly in multipurpose auditoriums and concert halls, as well as in audio research facilities. The variable acoustics laboratory Arni, a facility within the Acoustics Lab of Aalto University in Espoo, is an example of a space where the wall absorption can be altered considerably with the help of specialized wall- and ceiling-mounted panels. The present work shows the results of measurements of over 5000 panel combinations in Arni, showcasing the change in values of the room acoustic parameters, such as reverberation time and clarity, in relation to the variation of the wall absorption.

  15. Space Walk - Visiting the Solar System through an Immersive Sonic Journey in VR
    Andrea Mancianti, Sebastian J. Schlecht, Vesa Välimäki, Riku Järvinen, Esa Kallio
    Nordic Sound and Music Conference
    Abstract

    Space Walk is a navigable virtual planetarium designed for the Oculus Quest VR headset. It provides an educational yet accurate representation of the Solar System, including visualizations of scientific data, such as magnetic field lines and atmospheric phenomena, and accompanying explanatory text. A navigational interface allows the visitor to travel between planets. As a complement to the visual content, an ad hoc modular soundtrack has been composed, meant to characterize sonically each celestial object and to offer an audio counterpart for each of their possible data visualization layers. Each sound layer could work both in isolation and together with all the other layers, still keeping coherence of the musical discourse. It is also meant to pay tribute to a vast network of literature from Sci-Fi film and video game music, remaining appropriate within a rigorous scientific context. Finally, it integrates both stereophonic and immersive sound spatialization techniques. A fixed rendering through the interactive sound journey can be found online 1 . The full VR experience is freely available on the Oculus AppLab and can be played on Oculus Quest 1 and 2 devices 2 .

  16. Spatial Filter Bank in the Spherical Harmonic Domain: Reconstruction and Application
    Christoph Hold, Sebastian J. Schlecht, Archontis Politis, Ville Pulkki
    IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)
    Abstract

    Filter banks are an integral part of modern signal processing. They may also be applied to spatial filtering and the employed spatial filters can be designed with a specific shape for the analysis, e. g. suppressing side-lobes. After extracting spatially constrained signals from spherical harmonic (SH) input, i. e. filter bank analysis, many applications demand for a re-synthesis of the associated sector signals to the SH domain. This paper hence derives the complementary spatial filter bank reconstruction. The criterion for perfect reconstruction, and energy preserving reconstruction are given and implemented into the design. The filter bank is formulated such that for axisymmetric patterns both criteria can be met by only minor modification to the reconstruction stage. Its application is then demonstrated for both scenarios, perfect reconstruction and energy preservation of SH input signals.

  17. The Role of Modal Excitation in Colorless Reverberation
    Janis Heldmann, Sebastian J. Schlecht
    International Conference on Digital Audio Effects (Dafx)
    Abstract

    A method to generate and evaluate colorless artificial reverberation with feedback delay networks (FDNs) is proposed. The coloration of the reverberation tail is quantified by the modal excitation distribution derived from the modal decomposition of the FDN. The colorless FDN is designed to be allpass and homogeneously decaying such that the corresponding narrow modal excitation distribution leads to a high perceived modal density. The modification alters only the standard FDN gains, and no additional processing is introduced. Three listening tests were conducted to demonstrate the correlation between the modal excitation distribution and the perceived degree of coloration. A fourth test shows a significant reduction of coloration by the proposed FDN. The colorless FDN presents a new baseline structure for neutral reverberation without any extra processing and computational cost.

2020

  1. A String in a Room: Mixed-dimensional Transfer Function Models for Sound Synthesis
    Maximilian Schäfer, Rudolf Rabenstein, Sebastian J. Schlecht
    23rd International Conference on Digital Audio Effects (Dafx2020)
    Abstract

    Physical accuracy of virtual acoustics receives increasing attention due to renewed interest in virtual and augmented reality applications. So far, the modeling of vibrating objects as point sources is a common simplification which neglects effects caused by their spatial extent. In this contribution, we propose a technique for the interconnection of a distributed source to a room model, based on a modal representation of source and room. In particular, we derive a connection matrix that describes the coupling between the modes of the source and the room modes in an analytical form. Therefore, we consider the example of a string that is oscillating in a room. Both, room and string rely on well established physical descriptions that are modeled in terms of transfer functions. The derived connection of string and room defines the coupling between the characteristic string and room modes. The proposed structure is analyzed by numerical evaluations and sound examples on the supplementary website.

  2. Apparatus and Method for Reproducing a Spatially Extended Sound Source or Apparatus and Method for Generating a Bitstream from a Spatially Extended Sound Source
    Jürgen Herre, Emanuel Habets, Sebastian J. Schlecht, Alexander Adami
  3. Evaluation of Reverberation Time Models with Variable Acoustics
    Karolina Prawda, Sebastian J. Schlecht, Vesa Välimäki
    Proceedings of the 17th Sound and Music Computing Conference (SMC)
    Abstract

    Reverberation time of a room is the most prominent parameter considered when designing the acoustics of physical spaces. Techniques for predicting reverberation of enclosed spaces started emerging over one hundred years ago. Since then, several formulas to estimate the reverberation time in different room types were proposed. Although validations of those models were conducted in the past, they lack testing in a space with a high granularity of controllable absorptive and reflective conditions. The present study discusses the reverberation time estimation techniques by comparing various formulas. Moreover, the reverberation time measurements in a variable acoustic laboratory for different combinations of reflective and absorptive panels are shown. The values calculated with the presented models are compared with the ones obtained via measurements. The results show that all formulas predict reverberation time values inaccurately, with an average error of 16% or larger. Among the analyzed models, Fitzroy's formula gives the smallest error.

  4. FDNTB: The Feedback Delay Network Toolbox
    Sebastian J. Schlecht
    23rd International Conference on Digital Audio Effects (Dafx2020)
    Abstract

    Feedback delay networks (FDNs) are recursive filters, which are widely used for artificial reverberation and decorrelation. While there exists a vast literature on a wide variety of reverb topologies, this work aims to provide a unifying framework to design and analyze delay-based reverberators. To this end, we present the Feedback Delay Network Toolbox (FDNTB), a collection of the MATLAB functions and example scripts. The FDNTB includes various representations of FDNs and corresponding translation functions. Further, it provides a selection of special feedback matrices, topologies, and attenuation filters. In particular, more advanced algorithms such as modal decomposition, time-varying matrices, and filter feedback matrices are readily accessible. Furthermore, our toolbox contains several additional FDN designs. Providing MATLAB code under a GNU-GPL 3.0 license and including illustrative examples, we aim to foster research and education in the field of audio processing.

  5. Fade-in Control for Feedback Delay Networks
    Nils Meyer-Kahlen, Sebastian J. Schlecht, Tapio Lokki
    Proceedings of the 23rdInternational Conference on Digital Audio Effects (Dafx2020)
    Abstract

    In virtual acoustics, it is common to simulate the early part of a Room Impulse Response using approaches from geometrical acoustics and the late part using Feedback Delay Networks (FDNs). In order to transition from the early to the late part, it is useful to slowly fade-in the FDN response. We propose two methods to control the fade-in, one based on double decays and the other based on modal beating. We use modal analysis to explain the two concepts for incorporating this fade-in behaviour entirely within the IIR structure of a multiple input multiple output FDN. We present design equations, which allow for placing the fade-in time at an arbitrary point within its derived limit.

  6. No Dynamic Visual Capture for Self-Translation Minimum Audible Angle
    Olli S Rummukainen, Sebastian J Schlecht, Emanuël A P Habets
    The Journal of the Acoustical Society of America
    Abstract

    Auditory localization is affected by visual cues. The study at hand focuses on a scenario where dynamic sound localization cues are induced by lateral listener self-translation in relation to a stationary sound source with matching or mismatching dynamic visual cues. The audio-only self-translation minimum audible angle (ST-MAA) is previously shown to be 3.3$^$ in the horizontal plane in front of the listener. The present study found that the addition of visual cues has no significant effect on the ST-MAA.

  7. Scattering in Feedback Delay Networks
    Sebastian J Schlecht, Emanuël A P Habets
    IEEE/ACM Transactions on Audio, Speech, and Language Processing
    Abstract

    Feedback delay networks (FDNs) are recursive filters, which are widely used for artificial reverberation and decorrelation. One central challenge in the design of FDNs is the generation of sufficient echo density in the impulse response without compromising the computational efficiency. In a previous contribution, we have demonstrated that the echo density of an FDN can be increased by introducing so-called delay feedback matrices where each matrix entry is a scalar gain and a delay. In this contribution, we generalize the feedback matrix to arbitrary lossless filter feedback matrices (FFMs). As a special case, we propose the velvet feedback matrix, which can create dense impulse responses at a minimal computational cost. Further, FFMs can be used to emulate the scattering effects of non-specular reflections. We demonstrate the effectiveness of FFMs in terms of echo density and modal distribution.

  8. Towards Transfer-Plausibility for Evaluating Mixed Reality Audio in Complex Scenes
    Stefan A. Wirler, Nils Meyer-Kahlen, Sebastian J. Schlecht
    AES International Conference on Audio for Virtual and Augmented Reality (AVAR)
    Abstract

    The evaluation of mixed reality audio is typically approached under the paradigms of either authenticity or plausibility. While the first refers to the identity of a real and a virtualized sound source, the latter measures the degree of belief in cases where no direct reference is available. We refer to transfer-plausibility as the ability of a virtualized source to stand alongside multiple real sound sources. We present a perceptual experiment where listeners detect and identify a sound source as being virtualized using dynamic non-individualized binaural rendering under varying scene complexity. Scene complexity is controlled by a varying number of loudspeakers. We demonstrate that the presented methodology mitigates ceiling effects, typically encountered in authenticity and plausibility tests.

  9. Velvet-Noise Feedback Delay Network
    Jon Fagerström, Benoit Alary, Sebastian J. Schlecht, Vesa Välimäki
    Proceedings of the 23rd International Conference on Digital Audio Effects (Dafx2020)
    Abstract

    Artificial reverberation is an audio effect used to simulate the acoustics of a space while controlling its aesthetics, particularly on sounds recorded in a dry studio environment. Delay-based methods are a family of artificial reverberators using recirculating delay lines to create this effect. The feedback delay network is a popular delay-based reverberator providing a comprehensive framework for parametric reverberation by formalizing the recirculation of a set of interconnected delay lines. However, one known limitation of this algorithm is the initial slow build-up of echoes, which can sound unrealistic, and overcoming this problem often requires adding more delay lines to the network. In this paper, we study the effect of adding velvet-noise filters, which have random sparse coefficients, at the input and output branches of the reverberator. The goal is to increase the echo density while minimizing the spectral coloration. We compare different variations of velvet-noise filtering and show their benefits. We demonstrate that with velvet noise, the echo density of a conventional feedback delay network can be exceeded using half the number of delay lines and saving over 50% of computing operations in a practical configuration using low-order attenuation filters.

2019

  1. Dense Reverberation with Delay Feedback Matrices
    Sebastian J Schlecht, Emanuël A P Habets
    IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)
    Abstract

    This paper received the [Best Paper Award](http://waspaa.com/). Feedback delay networks (FDNs) belong to a general class of recursive filters which are widely used in artificial reverberation and decorrelation applications. One central challenge in the design of FDNs is the generation of sufficient echo density in the impulse response without compromising the computational efficiency. In a previous contribution, we have demonstrated that the echo density of an FDN grows polynomially over time, and that the growth depends on the number and lengths of the delays. In this work, we introduce so-called delay feedback matrices (DFMs) where each matrix entry is a scalar gain and a delay. While the computational complexity of DFMs is similar to a scalar-only feedback matrix, we show that the echo density grows significantly faster over time, however, at the cost of non-uniform modal decays.

  2. Directional Feedback Delay Network
    Benoit Alary, Archontis Politis, Sebastian J Schlecht, Vesa Välimäki
    Journal of the Audio Engineering Society
    Abstract

    Artificial reverberation algorithms are used to enhance dry audio signals. Delay-based reverberators can produce a realistic effect at a reasonable computational cost. While the recent popularity of spatial audio algorithms is mainly related to the reproduction of the perceived direction of sound sources, there is also a need to spatialize the reverberant sound field. Usually multichannel reverberation algorithms output a series of decorrelated signals yielding an isotropic energy decay. This means that the reverberation time is uniform in all directions. However, the acoustics of physical spaces can exhibit more complex direction-dependent characteristics. This paper proposes a new method to control the directional distribution of energy over time, within a delay-based reverberator, capable of producing a directional impulse response with anisotropic energy decay. We present a method using multichannel delay lines in conjunction with a direction-dependent transform in the spherical harmonic domain to control the direction-dependent decay of the late reverberation. The new reverberator extends the feedback delay network, retaining its time-frequency domain characteristics. The proposed directional feedback delay network reverberator can produce non-uniform direction-dependent decay time, suitable for anisotropic decay reproduction on a loudspeaker array or in binaural playback through the use of ambisonics.

  3. Enhanced Immersion for Binaural Audio Reproduction of Ambisonics in Six-Degrees-of-Freedom: The Effect of Added Distance Information
    Axel Plinge, Sebastian J. Schlecht, Olli Rummukainen, Emanuël A P Habets
    International Conference on Spatial Audio
  4. Feedback Structures for a Transfer Function Model of a Circular Vibrating Membrane
    Maximilian Schäfer, Sebastian J Schlecht, Rudolf Rabenstein
    IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)
    Abstract

    The attachment of feedback loops to physical or musical systems enables a large variety of possibilities for the modification of the system behavior. Feedback loops may enrich the echo density of feedback delay networks (FDN), or enable the realization of complex boundary conditions in physical simulation models for sound synthesis. Inspired by control theory, a general feedback loop is attached to a model of a vibrating membrane. The membrane model is based on the modal expansion of an initial-boundary value problem formulated in a state-space description. The possibilities of the attached feedback loop are shown by three examples, namely by the introduction of additional mode wise damping; modulation and damping inspired by FDN feedback loops; time-varying modification of the system behavior.

  5. Frequency-Dependent Schroeder Allpass Filters
    Sebastian J Schlecht
    Applied Sciences
    Abstract

    Since the introduction of feedforward--feedback comb allpass filters by Schroeder and Logan, its popularity has not diminished due to its computational efficiency and versatile applicability in artificial reverberation, decorrelation, and dispersive system design. In this work, we present an extension to the Schroeder allpass filter by introducing frequency-dependent feedforward and feedback gains while maintaining the allpass characteristic. By this, we directly improve upon the design of Dahl and Jot which exhibits a frequency-dependent absorption but does not preserve the allpass property. At the same time, we also improve upon Gerzon's allpass filter as our design is both less restrictive and computationally more efficient. We provide a complete derivation of the filter structure and its properties. Furthermore, we illustrate the usefulness of the structure by designing an allpass decorrelation filter with frequency-dependent decay characteristics.

  6. Improved Reverberation Time Control for Feedback Delay Networks
    Karolina Prawda, Sebastian J Schlecht, Vesa Välimäki
    22nd International Conference on Digital Audio Effects (Dafx-19)
    Abstract

    Artificial reverberation algorithms generally imitate the frequency-dependent decay of sound in a room quite inaccurately. Previous research suggests that a 5% error in the reverberation time (T60) can be audible. In this work, we propose to use an accurate graphic equalizer as the attenuation filter in a Feedback Delay Network reverberator. We use a modified octave graphic equalizer with a cascade structure and insert a high-shelf filter to control the gain at the high end of the audio range. One such equalizer is placed at the end of each delay line of the Feedback Delay Network. The gains of the equalizer are optimized using a new weighting function that acknowledges nonlinear error propagation from filter magnitude response to reverberation time values. Our experiments show that in real-world cases, the target T60 curve can be reproduced in a perceptually accurate manner at standard octave center frequencies. However, for an extreme test case in which the T60 varies dramatically between neighboring octave bands, the error still exceeds the limit of the just noticeable difference but is smaller than that obtained with previous methods. This work leads to more realistic artificial reverberation.

  7. Modal Decomposition of Feedback Delay Networks
    Sebastian J Schlecht, Emanuël A P Habets
    IEEE Transactions on Signal Processing
    Abstract

    Feedback delay networks (FDNs) belong to a general class of recursive filters which are widely used in sound synthesis and physical modeling applications. We present a numerical technique to compute the modal decomposition of the FDN transfer function. The proposed pole finding algorithm is based on the Ehrlich-Aberth iteration for matrix polynomials and has improved computational performance of up to three orders of magnitude compared to a scalar polynomial root finder. The computational performance is further improved by bounds on the pole location and an approximate iteration step. We demonstrate how explicit knowledge of the FDN's modal behavior facilitates analysis and improvements for artificial reverberation. The statistical distribution of mode frequency and residue magnitudes demonstrate that relatively few modes contribute a large portion of impulse response energy.

  8. Perceptual Study of Near-Field Binaural Audio Rendering in Six-Degrees-of-Freedom Virtual Reality
    Olli S Rummukainen, Sebastian J Schlecht, T Robotham, Axel Plinge, Emanuël A P Habets
    IEEE Virtual Reality (VR)
    Abstract

    Auditory localization cues in the near-field ( 1.0 m) are significantly different than in the far-field. The near-field region is within an arm's length of the listener allowing to integrate proprioceptive cues to determine the location of an object in space. This perceptual study compares three non-individualized methods to apply head-related transfer functions (HRTFs) in six-degrees-of-freedom near-field audio rendering, namely, far-field measured HRTFs, multi-distance measured HRTFs, and spherical-model-based HRTFs with near-field extrapolation. To set our findings in context, we provide a real-world hand-held audio source for comparison along with a distance-invariant condition. Two modes of interaction are compared in an audio-visual virtual reality: one allowing the participant to move the audio object dynamically and the other with a stationary audio object but a freely moving listener.

  9. Towards Measuring Intonation Quality of Choir Recordings: A Case Study on Bruckner's Locus Iste
    Christof Weiß, Sebastian J. Schlecht, Sebastian Rosenzweig, Meinard Müller
    Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)

2018

  1. Audio Quality Evaluation in Virtual Reality: Multiple Stimulus Ranking with Behavior Tracking
    Olli S Rummukainen, Thomas Robotham, Sebastian J Schlecht, Axel Plinge, Jürgen Herre, Emanuël A P Habets
    AES International Conference on Audio for Virtual and Augmented Reality
    Abstract

    Virtual reality systems with multimodal stimulation and up to six degrees-of-freedom movement pose novel challenges to audio quality evaluation. This paper adapts classic multiple stimulus test methodology to virtual reality and adds behavioral tracking functionality. The method is based on ranking by elimination while exploring an audiovisual virtual reality. The proposed evaluation method allows immersion in multimodal virtual scenes while enabling comparative evaluation of multiple binaural renderers. A pilot study is conducted to evaluate feasibility of the proposed method and to identify challenges in virtual reality audio quality evaluation. Finally, the results are compared to a non-immersive off-line evaluation method.

  2. Optimized Velvet-Noise Decorrelator
    Sebastian J Schlecht, Benoit Alary, Vesa Välimäki, Emanuël A P Habets
    Proc. Int. Conf. Digital Audio Effects (Dafx)
    Abstract

    This paper received the [2nd Best Paper Award](http://dafx2018.web.ua.pt/). Decorrelation of audio signals is a critical step for spatial sound reproduction on multichannel configurations. Correlated signals yield a focused phantom source between the reproduction loudspeakers and may produce undesirable comb-filtering artifacts when the signal reaches the listener with small phase differences. Decorrelation techniques reduce such artifacts and extend the spatial auditory image by randomizing the phase of a signal while minimizing the spectral coloration. This paper proposes a method to optimize the decorrelation properties of a sparse noise sequence, called velvet noise, to generate short sparse FIR decorrelation filters. The sparsity allows a highly efficient time-domain convolution. The listening test results demonstrate that the proposed optimization method can yield effective and colorless decorrelation filters. In comparison to a white noise sequence, the filters obtained using the proposed method preserve better the spectrum of a signal and produce good quality broadband decorrelation while using 76% fewer operations for the convolution. Satisfactory results can be achieved with an even lower impulse density which decreases the computational cost by 88%.

  3. Self-Translation Induced Minimum Audible Angle
    Olli S Rummukainen, Sebastian J Schlecht, Emanuël A P Habets
    The Journal of the Acoustical Society of America
    Abstract

    The minimum audible angle has been studied with a stationary listener and a stationary or a moving sound source. The study at hand focuses on a scenario where the angle is induced by listener self-translation in relation to a stationary sound source. First, the classic stationary listener minimum audible angle experiment is replicated using a headphone-based reproduction system. This experiment confirms that the reproduction system is able to produce a localization cue resolution comparable to loudspeaker reproduction. Next, the self-translation minimum audible angle is shown to be 3.3$^$ in the horizontal plane in front of the listener.

  4. Sign-Agnostic Matrix Design for Spatial Artificial Reverberation with Feedback Delay Networks
    Sebastian J Schlecht, Emanuël A P Habets
    AES Conference on Spatial Reproduction
    Abstract

    Feedback delay networks (FDNs) are an efficient tool for creating artificial reverberation. Recently, various designs for spatially extending the FDN were proposed. A central topic in the design of spatial FDNs is the choice of the feedback matrix that governs the interaction between spatially distributed elements and therefore the spatial impression. In the design prototype, the feedback matrix is chosen to be unilossless such that the reverberation time is infinite. However, in physics- and aesthetics-driven design of spatial FDNs, the target feedback matrix is not necessarily unilossless. This contribution proposes an optimization method for finding a close unilossless feedback matrix and improves the accuracy by relaxing the specification of the target matrix phase component and focussing on the sign-agnostic component.

  5. Six-Degrees-of-Freedom Binaural Audio Reproduction of First-Order Ambisonics with Distance Information
    Axel Plinge, Sebastian J Schlecht, Oliver Thiergart, Thomas Robotham, Olli S Rummukainen, Emanuël A P Habets
    AES International Conference on Audio for Virtual and Augmented Reality (AVAR)
    Abstract

    First-order Ambisonics (FOA) recordings can be processed and reproduced over headphones. They can be rotated to account for the listener's head orientation. However, virtual reality (VR) systems allow the listener to move in six-degrees-of-freedom (6DoF), i.e., three rotational plus three transitional degrees of freedom. Here, the apparent angles and distances of the sound sources depend on the listener's position. We propose a technique to facilitate 6DoF. In particular, a FOA recording is described using a parametric model, which is modified based on the listener's position and information about the distances to the sources. We evaluate our method by a listening test, comparing different binaural renderings of a synthetic sound scene in which the listener can move freely.

2017

  1. Accurate Reverberation Time Control in Feedback Delay Networks
    Sebastian J Schlecht, Emanuël A P Habets
    Proc. Int. Conf. Digital Audio Effects (Dafx)
    Abstract

    The reverberation time is one of the most prominent acoustical qualities of a physical room. Therefore, it is crucial that artificial reverberation algorithms match a specified target reverberation time accurately. In feedback delay networks, a popular framework for modeling room acoustics, the reverberation time is determined by combining delay and attenuation filters such that the frequency-dependent attenuation response is proportional to the delay length and by this complying to a global attenuation-per-second. However, only few details are available on the attenuation filter design as the approximation errors of the filter design are often regarded negligible. In this work, we demonstrate that the error of the filter approximation propagates in a non-linear fashion to the resulting reverberation time possibly causing large deviation from the specified target. For the special case of a proportional graphic equalizer, we propose a non-linear least squares solution and demonstrate the improved accuracy with a Monte Carlo simulation.

  2. Evaluating Binaural Reproduction Systems from Behavioral Patterns in a Virtual Reality---A Case Study with Impaired Binaural Cues and Tracking Latency
    Olli S Rummukainen, Sebastian J Schlecht, Axel Plinge, Emanuël A P Habets
    Audio Engineering Society Convention 143
    Abstract

    This paper proposes a method for evaluating real-time binaural reproduction systems by means of a wayfinding task in six degrees of freedom. Participants physically walk to sound objects in a virtual reality created by a head-mounted display and binaural audio. The method allows for comparative evaluation of different rendering and tracking systems. We show how the localization accuracy of spatial audio rendering is reflected by objective measures of the participants' behavior and task performance. As independent variables we add tracking latency or reduce the binaural cues. We provide a reference scenario with loudspeaker reproduction and an anchor scenario with monaural reproduction for comparison.

  3. Evaluation of Binaural Reproduction Systems from Behavioral Patterns in a Six-Degrees-of-Freedom Wayfinding Task
    Olli S Rummukainen, Sebastian J Schlecht, Axel Plinge, Emanuël A P Habets
    Quality of Multimedia Experience (QoMEX)
    Abstract

    This paper proposes a new method for evaluating real-time binaural reproduction systems by means of a wayfinding task in six degrees of freedom. Participants physically walk to sound objects in a virtual reality created by a head-mounted display and binaural audio. We show how the localization accuracy of spatial audio rendering is reflected by objective measures of the participants' behavior. The method allows for comparative evaluation of different rendering systems as well as the subjective assessment of the quality of experience.

  4. Feedback Delay Networks in Artificial Reverberation and Reverberation Enhancement
    Sebastian J Schlecht
    Abstract

    In today's audio production and reproduction as well as in music performance practices it has become common practice to alter reverberation artificially through electronics or electro- acoustics. For music productions, radio plays, and movie soundtracks, the sound is often captured in small studio spaces with little to no reverberation to save real estate and to ensure a controlled environment such that the artistically intended spatial impression can be added during post-production. Spatial sound reproduction systems require flexible adjustment of artificial reverberation to the diffuse sound portion to help the reconstruction of the spatial impression. Many modern performance spaces are multi-purpose, and the reverberation needs to be adjustable to the desired performance style. Employing electro-acoustic feedback, also known as Reverberation Enhancement Systems (RESs), it is possible to extend the physical to the desired reverberation. These examples demonstrate a wide range of applications where reverberation is created and enhanced artificially employing signal processing techniques. A major challenge of designing artificial reverberators is the high complexity of the physical reverberation process. Even small office spaces of 40 m3 exhibit more than 107 acoustic modes, in concert halls the number of acoustic modes can surpass 109 in the audible range. The room geometry, as well as the interaction with the boundary materials, can be as well fairly complex. Whereas these complex considerations are mandatory for simulations of specific spaces, used for example for the acoustic and architectural planning of a concert venue, they are somewhat misleading in the realm of artistic applications. The focus on perceptually convincing artificial reverberation algorithms provides the freedom to make some simplifications to the generation process, leading to the recursive systems, which play a central role in this dissertation.

  5. Feedback Delay Networks: Echo Density and Mixing Time
    Sebastian J Schlecht, Emanuël A P Habets
    IEEE/ACM Transactions on Audio, Speech, and Language Processing
    Abstract

    Feedback delay networks (FDNs) are frequently used to generate artificial reverberation. This paper discusses the temporal features of impulse responses produced by FDNs, i.e., the number of echoes per time unit and its evolution over time. This so-called echo density is related to known measures of mixing time and their psychoacoustic correlates such as auditive perception of the room size. It is shown that the echo density of FDNs follows a polynomial function, whereby the polynomial coefficients can be derived from the lengths of the delays for which an explicit method is given. The mixing time of impulse responses can be predicted from the echo density, and conversely, a desired mixing time can be achieved by a derived mean delay length. A Monte Carlo simulation confirms the accuracy of the derived relation of mixing time and delay lengths.

  6. On Lossless Feedback Delay Networks
    Sebastian J Schlecht, Emanuël A P Habets
    IEEE Transactions on Signal Processing
    Abstract

    Lossless Feedback Delay Networks (FDNs) are commonly used as a design prototype for artificial reverberation algorithms. The lossless property is dependent on the feedback matrix, which connects the output of a set of delays to their inputs, and the lengths of the delays. Both, unitary and triangular feedback matrices are known to constitute lossless FDNs, however, the most general class of lossless feedback matrices has not been identified. In this contribution, it is shown that the FDN is lossless for any set of delays, if all irreducible components of the feedback matrix are diagonally similar to a unitary matrix. The necessity of the generalized class of feedback matrices is demonstrated by examples of FDN designs proposed in literature.

2016

  1. The Stability of Multichannel Sound Systems with Time-Varying Mixing Matrices
    Sebastian J Schlecht, Emanuël A P Habets
    The Journal of the Acoustical Society of America
    Abstract

    Various time-varying algorithms have been applied in multichannel sound systems to improve the system's stability and, among these, frequency shifting has been demonstrated to reach the maximum stability improvement achievable by time-variation in general. However, the modulation artifacts have been found to diminish the gain improvement unusable for a higher number of channels and high-quality applications such as music reproduction. This paper proposes alternatively time-varying mixing matrices, which is an efficient algorithm corresponding to symmetric up and down frequency shifting. It is shown with a statistical approach that time-varying mixing matrices can as well achieve maximum stability improvement for a higher number of channels. A listening test demonstrates the improved quality of time-varying mixing matrices over frequency shifting.

2015

  1. Apparatus and Method for Generating Output Signals Based on an Audio Source Signal, Sound Reproduction System and Loudspeaker Signal
    Sebastian Schlecht, Andreas Silzle, Emanuel Habets, Christian Borss, Bernhard Neugebauer, Hanne Stenzel
    Abstract

    An apparatus for generating a first multitude of output signals based on at least one audio source signal comprising a delay network and a feedback processor. The delay network comprises a second multitude of delay paths, each delay path having a delay line and an attenuation filter. Each delay line is configured for delaying delay line input signals and for combining the at least one audio source signal and a reverberated audio signal to obtain a combined signal, wherein the attenuation filter of a delay path is configured for filtering the combined signal from the delay line of the delay path to obtain an output signal. The first multitude of output signals comprises the output signal. The feedback processor is configured for reverberating the first multitude of output signals to obtain a third multitude of the reverberated audio signals comprising the reverberated audio signal.

  2. Practical Considerations of Time-Varying Feedback Delay Networks
    Sebastian J Schlecht, Emanuël A P Habets
    Proc. Audio Eng. Soc. Conv.
    Abstract

    Feedback delay networks (FDNs) can be efficiently used to generate parametric artificial reverberation. Recently, the authors proposed a novel approach to time-varying FDNs by introducing a time-varying feedback matrix. The formulation of the time-varying feedback matrix was given in the complex eigenvalue domain, whereas this contribution specifies the requirements for real valued time-domain processing. In addition, the computational costs of different time-varying feedback matrices, which depend on the matrix type and modulation function, are discussed. In a performance evaluation, the proposed orthogonal matrix modulation is compared to a direct interpolation of the matrix entries.

  3. Reverberation Enhancement Systems with Time-Varying Mixing Matrices
    Sebastian J Schlecht, Emanuël A P Habets
    Proc. Audio Eng. Soc. Conf.
    Abstract

    This paper presents a technique to increase the stability of a reverberation enhancement system (RES) via time-variation. RESs are installed in rooms to extend the physical reverberation time via electroacoustic feedback. By introducing variations in the feedback path, the probability that energy builds-up is reduced, and therefore the gain before instability is increased. Time-varying mixing of the feedback paths by unitary matrices is guaranteed to be energy-conserving and introduces variations of the added reverberation generated by a feedback delay network. The theory of time-varying matrices is reviewed and a practical demonstration is given.

  4. Time-Varying Feedback Matrices in Feedback Delay Networks and Their Application in Artificial Reverberation
    Sebastian J Schlecht, Emanuël A P Habets
    The Journal of the Acoustical Society of America
    Abstract

    This paper introduces a time-variant reverberation algorithm as an extension of the feedback delay network (FDN). By modulating the feedback matrix nearly continuously over time, a complex pattern of concurrent amplitude modulations of the feedback paths evolves. Due to its complexity, the modulation produces less likely perceivable artifacts and the time-variation helps to increase the liveliness of the reverberation tail. A listening test, which has been conducted, confirms that the perceived quality of the reverberation tail can be enhanced by the feedback matrix modulation. In contrast to the prior art time-varying allpass FDNs, it is shown that unitary feedback matrix modulation is guaranteed to be stable. Analytical constraints on the pole locations of the FDN help to describe the modulation effect in depth. Further, techniques and conditions for continuous feedback matrix modulation are presented.

2012

  1. Connections between Parallel and Serial Combinations of Comb Filters and Feedback Delay Networks
    Sebastian J Schlecht, Emanuël A P Habets
    International Workshop on Acoustic Signal Enhancement (IWAENC)
    Abstract

    Comb filters composed in a parallel or a serial way are a popular part of delay-line based artificial reverberators. Because the analysis of a complex comb filter structure can be tedious, there is a need for transforming such a structure into a compact and general representation. For this a transformation into the feedback delay network (FDN) filter structure is proposed as it is a general and well established framework to investigate the acoustic properties of the filter and therefore allows to compare different approaches.

  2. Reverberation Enhancement from a Feedback Delay Network Perspective
    Sebastian J Schlecht, Emanuël A P Habets
    Convention of Electrical and Electronics Engineers in Israel (IEEEI)
    Abstract

    In reverberation enhancement systems (RESs), sound is constantly fed back from multiple microphones to multiple loudspeakers to enhance reverberation artificially in the target room. This contribution shows that such a system can be understood as an extended feedback delay network (FDN). A tuning process, similar to that of the FDN is presented, allowing arbitrary frequency-dependent reverberation elongation. The cross-talk between the loudspeakers and the microphones leads to comb filtering and isolated ringing modes in the RES, which produce undesired metallic and rough sounds. To mitigate these undesired effects, a cross-talk cancellation system is integrated in the RES. In a simulation example, the benefits of cross-talk cancellation is evaluated.

2011

  1. Source-Filter Separation for Bowed String Instruments and Its Application for Advanced Audio Effects
    Sebastian J Schlecht
    Abstract

    This project explores the possible use of source-filter separation for bowed string instruments in audio post-processing effects. Namely, two scenarios were implemented: a violin to cello and a violin to string section conversion. This is achieved by a combination of body response estimation and pitch-shifting algorithm. Special attention were given to the particular characteristics in construction and sound of bowed string instruments. The source-filter model were performed by a linear body estimation exploiting the harmonic sound nature of the strings. The pitch-shifting were implemented by a phase-vocoder approach. Furthermore, a live capable real-time implementation for Max/MSP were implemented and thoroughly tested. The resulting algorithm is practical on a consumer laptop and gives plausible results.

2010

  1. Options and Limits of Feedback Delay Networks for Artificial Reverberation of Audio Signals
    Sebastian J Schlecht