Project outline
This project aims to develop AI models inspired by the auditory system for improved auditory scene analysis.
While recent advances in machine learning have produced numerous acoustic AI models for various auditory tasks, these models typically adapt architectures from computer vision or natural language processing. This cross-domain borrowing often results in models that lack interpretability and biological plausibility, despite being tuned for acoustic applications.
The auditory system naturally and efficiently performs complex auditory scene analysis tasks. By incorporating prior knowledge from auditory neuroscience—spanning peripheral, subcortical, and cortical processing stages—we can develop AI models that are not only more parameter-efficient but also more explainable and interpretable. These biologically inspired principles have been refined through evolution and are well-documented in auditory research, providing a solid foundation for computational implementation.
This research will systematically integrate insights from different levels of the auditory pathway into deep learning architectures. The methodology includes: (1) reviewing established computational models of auditory processing, (2) designing neural network components that reflect biological mechanisms such as cochlear filtering, temporal modulation processing, and attention-driven segregation, and (3) validating these models on auditory scene analysis benchmarks.
A key application target is embedded robotics, where conventional AI models are often impractical due to their computational demands. Auditory system-inspired models, with their inherent efficiency, could enable robots to perform real-time auditory scene analysis with limited hardware resources, enhancing their environmental awareness and interaction capabilities.
Project Partners
This project is hosted at the University of Southampton.

Student
Deokki Min