Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🎙️ Audio Language Identifier

Project Overview

The Audio Language Identifier is a machine learning application designed to classify spoken audio as either English or Hindi.

Instead of traditional audio processing, this project leverages computer vision to solve an audio problem. Raw audio clips are dynamically converted into Mel-Spectrograms (visual representations of sound frequencies). These spectrogram images are then passed through a Convolutional Neural Network (CNN) capable of distinguishing the unique visual patterns of both languages to give an accurate prediction.

Tech Stack Used

  • Frontend / Deployment: Streamlit
  • Machine Learning: TensorFlow & Keras (CNN Architecture)
  • Audio Processing: Librosa
  • Data Visualization & Image Processing: Matplotlib, Pillow (PIL), NumPy
  • Environment: Python 3.11

Links

About

A CNN-based application that classifies spoken audio as English or Hindi by converting clips into Mel-Spectrograms and analyzing them as images.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages