Investigating Latent Space Representation of HARP HOA-RIR Dataset
* Presenting author
Abstract:
The study examines the potential of a lightweight encoder-decoder model to encode 100,000 Room Impulse Responses (RIRs), represented in Short-Time Fourier Transform (STFT) form, into a compact latent space of 1280 parameters for accurate RIR reconstruction. The proposed model architecture, with only 25,000 trainable parameters, offers a low-complexity alternative to traditional high-parameter RIR estimators. Specifically, we test whether this reduced latent space can maintain fidelity in reconstructing the spectral and temporal characteristics of the original RIRs. Additionally, we explore correlations between the latent representations and established acoustic parameters (e.g., RT60, clarity, direct-to-reverberant ratios) to assess the latent spaces capacity for capturing perceptually relevant acoustic features.A secondary objective is to evaluate the models performance in encoding 7th-order Ambisonic RIRs that are also presented here and investigate if this architecture can extend to spatially distributed acoustic data without significant model reconfiguration. Furthermore, we assess whether this approach can synthesize multi-channel RIRs from a lower number of input channels, potentially streamlining the RIR generation process for complex spatial audio applications. Initial results suggest that the proposed architecture successfully encodes and reconstructs both single and multi-channel RIRs, pointing to applications in on-device acoustic modelling for spatial audio and AR/VR.