Sound Visualizer is a desktop application that captures real-time system audio and translates its direction, intensity, and meaning into a graphical overlay. Going beyond simple waveform rendering, it integrates an AI sound classifier powered by YAMNet and ONNX Runtime to instantly identify the type of sound playing (ambient, speech, danger) and provide adaptive visual feedback.
- Background & Objectives
- Key Features
- System & AI Architecture
- Installation & Quick Start
- Build Guide (For Developers)
- Team Members
- License
Modern media content, including YouTube, OTT, and gaming, relies heavily on spatial audio to maximize immersion. However, this creates a severe barrier for the deaf and hard of hearing. Traditional subtitle systems only deliver dialogue, completely omitting crucial acoustic information such as sound direction, footsteps, or urgent sound effects (e.g., gunshots). Sound Visualizer was created to break down these barriers by visualizing the intensity, classification, and spatial direction of invisible sounds. Our primary goal is to resolve this information asymmetry and guarantee an equal, fully accessible media consumption environment for everyone.
This overlay technology is also highly beneficial for gamers, providing a tactical visual indicator (situational awareness) in competitive environments. Additionally, it serves as a perfect alternative for users consuming media in public or silent environments where audio output is restricted.
We provide 4 unique rendering modes that intuitively map sound intensity and direction (supporting 2.0, 5.1, and 7.1 channels):
- Wave Mode: Renders dynamic audio waves along the screen edges that fluctuate based on intensity.
- Circle Mode: Radiates circular ripples outward from the center of the screen.
- Pad Mode: Displays glowing pads anchored to specific spatial grid directions.
- Outline Mode: A minimalist variation of Wave mode, showing only the thin borders of the wave to minimize screen occlusion.
Modify settings instantly while in a full-screen application or game without minimizing the window.
- F2 / F3 Hotkeys: Switch between sound modes and visualization modes on the fly. Hotkeys can be customized.
- Editor Mode (Default F4): Drag the guideline boundaries on your screen to physically resize the rendering limits in real-time. Adjust colors, opacity, AI detection sensitivities, and all other settings directly from the pop-up control panel.
- Automatically detects the system's audio configuration (Stereo, 5.1, 7.1 Surround) and accurately maps the sound's origin (Front/Back/Left/Right) to produce a 3D visual effect on a 2D screen.
- 3-Class Detection: Analyzes all incoming audio into
Ambient,Speech, andDangercategories. Each category triggers independent, customizable UI colors and dynamic opacity changes. - Gunshot Booster: An auxiliary, highly-sensitive model designed specifically to ensure critical warning sounds (like gunshots in games/movies) are never missed.
- Fully supports 8 languages: English, Korean, Japanese, Chinese, Spanish, French, German, and Russian.
The project is rigorously engineered for high-performance, real-time background processing with minimal system overhead.
- Core Audio API (WASAPI): Loopback captures system-wide audio with zero latency.
- DSP & FFT Computation: High-speed frequency transformation calculations run entirely on a background audio thread multiple times per second.
- Zero-Allocation Rendering: The WPF/C# rendering loop is designed to minimize Garbage Collector (GC) allocation, entirely preventing frame drops.
- ONNX Inference Pipeline: Audio is converted to 16kHz mono, processed into Log-mel spectrograms, and fed into a custom-trained YAMNet model via the
Microsoft.ML.OnnxRuntimeengine for instantaneous classification.
Sound Visualizer is provided as a lightweight, portable (no-install) open-source package.
- Go to the GitHub Releases page.
- Download and extract the latest
SoundVisualizer.zipfile to any location on your PC. - Run
SoundVisualizer.exe. - Configure your initial settings and language in the launcher, then click Start to activate the overlay.
To build from source to modify the code or contribute to the project:
- Prerequisites: Windows 10/11, Visual Studio 2022 (with .NET Desktop Development workload), and .NET 10.0 SDK.
- Clone the repository:
git clone https://github.com/amophi/SoundVisualizer.git
- Open the
SoundVisualizer.slnxsolution file in Visual Studio 2022. - Select
ReleaseorDebugconfiguration, build (Ctrl + Shift + B), and pressF5to run.
Sound Visualizer is distributed under the AGPL v3 license to encourage a virtuous cycle in the open-source community. Anyone is welcome to modify the code and share custom UI themes or streaming plugins. See the LICENSE file for more information.
For licensing and copyright information regarding third-party libraries (NAudio, ONNX Runtime, etc.) and AI models used in this project, please refer to the THIRD_PARTY.md file.
