{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":7634,"databundleVersionId":46676,"sourceType":"competition"}],"dockerImageVersionId":29852,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"* Importing necessary Libraries\n* Data Source \n* 1. Audio Features\n    * Features extraction\n    * Visualization\n* Data preprocessing \n* Datset Investigation\n* 2. VAD\n* 3. Anamoly Detection \n* 4. Frequency components across the words\n\n------------------------------------","metadata":{}},{"cell_type":"markdown","source":"Source for this work \n- Speech representation and data exploration - DAVIDS -  https://www.kaggle.com/davids1992/speech-representation-and-data-exploration?scriptVersionId=1924001\n- voice activity detection example -ANDRE HOLZNER · - https://www.kaggle.com/holzner/voice-activity-detection-example\n- Voice Activity Detection with webrtcVAD|7z archive -ATUL ANAND {JHA} - https://www.kaggle.com/atulanandjha/voice-activity-detection-with-webrtcvad-7z-archive","metadata":{}},{"cell_type":"markdown","source":"### Importing necessary Libraries","metadata":{}},{"cell_type":"code","source":"import os\nfrom os.path import isdir, join\nfrom scipy.io import wavfile\nfrom subprocess import check_output\nfrom pathlib import Path\nimport pandas as pd\n\n\n# Math\nimport numpy as np\nfrom scipy.fftpack import fft\nfrom scipy import signal\nfrom scipy.io import wavfile\nimport librosa\n\nfrom sklearn.decomposition import PCA\n\n# Visualization\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport IPython.display as ipd\nimport librosa.display\n\nimport plotly.offline as py\npy.init_notebook_mode(connected=True)\nimport plotly.graph_objs as go\nimport plotly.tools as tls\nimport pandas as pd\n\n%matplotlib inline\n","metadata":{"execution":{"iopub.status.busy":"2024-08-16T07:58:29.952507Z","iopub.execute_input":"2024-08-16T07:58:29.952866Z","iopub.status.idle":"2024-08-16T07:58:35.205167Z","shell.execute_reply.started":"2024-08-16T07:58:29.952802Z","shell.execute_reply":"2024-08-16T07:58:35.204045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Install Python packages in an internet-enabled notebook","metadata":{}},{"cell_type":"code","source":"!pip install pyunpack\n!pip install patool","metadata":{"execution":{"iopub.status.busy":"2024-08-16T07:58:35.207537Z","iopub.execute_input":"2024-08-16T07:58:35.207865Z","iopub.status.idle":"2024-08-16T07:58:52.809564Z","shell.execute_reply.started":"2024-08-16T07:58:35.207803Z","shell.execute_reply":"2024-08-16T07:58:52.808357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" # Data Source \n \n Unpack .7z file\n","metadata":{}},{"cell_type":"code","source":"from pyunpack import Archive\nimport shutil\nif not os.path.exists('/kaggle/working/train/'):\n    os.makedirs('/kaggle/working/train/')\nArchive('/kaggle/input/train.7z').extractall('/kaggle/working/train/')\n# for dirname, _, filenames in os.walk('/kaggle/working/train/'):\n#     for filename in filenames:\n#         print(os.path.join(dirname, filename))\n","metadata":{"_kg_hide-output":true,"_kg_hide-input":false,"execution":{"iopub.status.busy":"2024-08-16T07:58:52.811728Z","iopub.execute_input":"2024-08-16T07:58:52.812269Z","iopub.status.idle":"2024-08-16T08:00:39.068860Z","shell.execute_reply.started":"2024-08-16T07:58:52.812112Z","shell.execute_reply":"2024-08-16T08:00:39.067445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"after you are finished working with the images you can delete them so that your commit will succeed (max number of files in working directory for a commit = 500)","metadata":{}},{"cell_type":"markdown","source":"이미지 작업을 마친 후에는 커밋이 성공하도록 이미지를 삭제할 수 있습니다(커밋을 위한 작업 디렉토리의 최대 파일 수 = 500)","metadata":{}},{"cell_type":"code","source":"shutil.make_archive('train/', 'zip', 'train')","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:00:39.071262Z","iopub.execute_input":"2024-08-16T08:00:39.071707Z","iopub.status.idle":"2024-08-16T08:02:38.611971Z","shell.execute_reply.started":"2024-08-16T08:00:39.071634Z","shell.execute_reply":"2024-08-16T08:02:38.610741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# deleting unwanted extracted files to avoid memory overflow (maxlimit files = 500) while commiting.\n!rm -rf kaggle/working/train/*","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:38.615530Z","iopub.execute_input":"2024-08-16T08:02:38.615889Z","iopub.status.idle":"2024-08-16T08:02:39.681735Z","shell.execute_reply.started":"2024-08-16T08:02:38.615811Z","shell.execute_reply":"2024-08-16T08:02:39.680388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Loading the trainig Input file.\ntrain_audio_path = \"/kaggle/working/train/train/audio\"","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:39.685422Z","iopub.execute_input":"2024-08-16T08:02:39.685775Z","iopub.status.idle":"2024-08-16T08:02:39.690353Z","shell.execute_reply.started":"2024-08-16T08:02:39.685722Z","shell.execute_reply":"2024-08-16T08:02:39.689297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\n# It is just a checker code to validate the presence of file.\n\nprint(check_output([\"ls\", \"../input/train/audio\"]).decode(\"utf8\"))\nprint(os.listdir(\"../input/train\"))\n\n\"\"\"\n\n\nprint(check_output([\"ls\", \"/kaggle/working/train/train/audio\"]).decode(\"utf8\"))\nprint(os.listdir(\"/kaggle/working/train/train/audio/yes\"))","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-08-16T08:02:39.691931Z","iopub.execute_input":"2024-08-16T08:02:39.692257Z","iopub.status.idle":"2024-08-16T08:02:39.721558Z","shell.execute_reply.started":"2024-08-16T08:02:39.692205Z","shell.execute_reply":"2024-08-16T08:02:39.720435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Example input file to be used here...\nfilename = '/yes/00f0204f_nohash_0.wav'","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:39.723697Z","iopub.execute_input":"2024-08-16T08:02:39.724051Z","iopub.status.idle":"2024-08-16T08:02:39.727843Z","shell.execute_reply.started":"2024-08-16T08:02:39.723990Z","shell.execute_reply":"2024-08-16T08:02:39.726995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dirs = [f for f in os.listdir(train_audio_path) if isdir(join(train_audio_path, f))]\ndirs.sort()\nprint('Number of labels: ' + str(len(dirs)))","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:39.729607Z","iopub.execute_input":"2024-08-16T08:02:39.730067Z","iopub.status.idle":"2024-08-16T08:02:39.741635Z","shell.execute_reply.started":"2024-08-16T08:02:39.729985Z","shell.execute_reply":"2024-08-16T08:02:39.740594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Audio Features\n\n\n## Features extraction\nA generalized feature extraction algorithm for an audio data sample be like that:\n\n1. Resampling\n2. VAD\n3. Maybe padding with 0 to make signals be equal length\n4. Log spectrogram (or MFCC, or PLP)\n5. Features normalization with mean and std\n6. Stacking of a given number of frames to get temporal information\n\n\n sample_rate, samples = wavfile.read(str(train_audio_path) + filename) \n\n The above code line works fine for everything except **Librosa** library MFCC functionality. So, we'll read wave files using librosa only.\n \n Must to read samples in librosa format. Other wise \"librosa\" error:data must be in floating format","metadata":{"trusted":true}},{"cell_type":"markdown","source":"특징 추출\n오디오 데이터 샘플에 대한 일반화된 특징 추출 알고리즘은 다음과 같습니다.\n\n1.리샘플링\n2.음성활동감지\n3.아마도 신호의 길이가 같도록 0으로 패딩하는 것이 좋습니다.\n4.로그 스펙트로그램(또는 MFCC 또는 PLP)\n5.평균 및 표준 편차를 사용한 기능 정규화\n6.시간 정보를 얻기 위해 주어진 수의 프레임을 스택합니다.\n샘플 속도, 샘플 = wavfile.read(str(train_audio_path) + 파일 이름)\n\n위의 코드 줄은 Librosa 라이브러리 MFCC 기능을 제외한 모든 것에 잘 작동합니다. 따라서 librosa만 사용하여 wave 파일을 읽겠습니다.\n\nlibrosa 형식으로 샘플을 읽어야 합니다. 그렇지 않으면 \"librosa\" 오류:데이터는 플로팅 형식이어야 합니다.","metadata":{}},{"cell_type":"code","source":"samples, sample_rate = librosa.load(str(train_audio_path)+filename)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:39.743614Z","iopub.execute_input":"2024-08-16T08:02:39.744270Z","iopub.status.idle":"2024-08-16T08:02:40.598193Z","shell.execute_reply.started":"2024-08-16T08:02:39.743869Z","shell.execute_reply":"2024-08-16T08:02:40.597196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualization\n\nThere are two theories of a [human hearing - place](https://en.wikipedia.org/wiki/Place_theory_(hearing) (frequency-based) and [temporal](https://en.wikipedia.org/wiki/Temporal_theory_(hearing) In speech recognition, I see two main tendencies - to input spectrogram (frequencies), and more sophisticated features MFCC - Mel-Frequency Cepstral Coefficients, PLP. You rarely work with raw, temporal data.\n","metadata":{}},{"cell_type":"markdown","source":"인간의 청력 에는 두 가지 이론이 있습니다. 장소 (주파수 기반)와 시간적. 음성 인식에서 저는 두 가지 주요 경향을 봅니다. 스펙트로그램(주파수)을 입력하고, 더 정교한 기능을 사용합니다. MFCC(Mel-Frequency Cepstral Coefficients, PLP). 원시 시간적 데이터로 작업하는 경우는 드뭅니다.","metadata":{}},{"cell_type":"markdown","source":"### 1.1 Spectogram \n\nDefine a function that calculates spectrogram.\n\nNote, that we are taking logarithm of spectrogram values. It will make our plot much more clear, moreover, it is strictly connected to the way people hear. We need to assure that there are no 0 values as input to logarithm.\n","metadata":{}},{"cell_type":"markdown","source":"스펙트로그램을 계산하는 함수를 정의합니다.\n\n스펙트로그램 값의 로그를 취하고 있다는 점에 유의하세요. 그러면 플롯이 훨씬 더 명확해질 것입니다. 게다가 사람들이 듣는 방식과 엄격하게 연결되어 있습니다. 로그에 입력으로 0 값이 없도록 해야 합니다.","metadata":{}},{"cell_type":"code","source":"def log_specgram(audio, sample_rate, window_size=20,\n                 step_size=10, eps=1e-10):\n    nperseg = int(round(window_size * sample_rate / 1e3))\n    noverlap = int(round(step_size * sample_rate / 1e3))\n    freqs, times, spec = signal.spectrogram(audio,\n                                    fs=sample_rate,\n                                    window='hann',\n                                    nperseg=nperseg,\n                                    noverlap=noverlap,\n                                    detrend=False)\n    return freqs, times, np.log(spec.T.astype(np.float32) + eps)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:40.599671Z","iopub.execute_input":"2024-08-16T08:02:40.600068Z","iopub.status.idle":"2024-08-16T08:02:40.607642Z","shell.execute_reply.started":"2024-08-16T08:02:40.600001Z","shell.execute_reply":"2024-08-16T08:02:40.606263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Frequencies are in range (0, 8000) according to Nyquist theorem.\n\nLet's plot it:","metadata":{}},{"cell_type":"code","source":"freqs, times, spectrogram = log_specgram(samples, sample_rate)\n\nfig = plt.figure(figsize=(14, 8))\nax1 = fig.add_subplot(211)\nax1.set_title('Raw wave of ' + filename)\nax1.set_ylabel('Amplitude')\nax1.plot(np.linspace(0, sample_rate/len(samples), sample_rate), samples)\n\nax2 = fig.add_subplot(212)\nax2.imshow(spectrogram.T, aspect='auto', origin='lower', \n           extent=[times.min(), times.max(), freqs.min(), freqs.max()])\nax2.set_yticks(freqs[::16])\nax2.set_xticks(times[::16])\nax2.set_title('Spectrogram of ' + filename)\nax2.set_ylabel('Freqs in Hz')\nax2.set_xlabel('Seconds')","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:40.609375Z","iopub.execute_input":"2024-08-16T08:02:40.609762Z","iopub.status.idle":"2024-08-16T08:02:41.272439Z","shell.execute_reply.started":"2024-08-16T08:02:40.609692Z","shell.execute_reply":"2024-08-16T08:02:41.271210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"normalizing the audio data. Always a good plan if we gonna feed it into NN.","metadata":{}},{"cell_type":"code","source":"mean = np.mean(spectrogram, axis=0)\nstd = np.std(spectrogram, axis=0)\nspectrogram = (spectrogram - mean) / std","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:41.273947Z","iopub.execute_input":"2024-08-16T08:02:41.274273Z","iopub.status.idle":"2024-08-16T08:02:41.280274Z","shell.execute_reply.started":"2024-08-16T08:02:41.274210Z","shell.execute_reply":"2024-08-16T08:02:41.279440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There is an interesting fact to point out. We have ~160 features for each frame, frequencies are between 0 and 8000. It means, that one feature corresponds to 50 Hz. However, frequency resolution of the ear is 3.6 Hz within the octave of 1000 – 2000 Hz It means, that people are far more precise and can hear much smaller details than those represented by spectrograms like above.","metadata":{}},{"cell_type":"markdown","source":"지적할 만한 흥미로운 사실이 있습니다. 우리는 각 프레임에 대해 160개의 특징을 가지고 있으며, 주파수는 0-8000 사이입니다. 즉, 하나의 특징은 50Hz에 해당합니다. 그러나 귀의 주파수 분해능은 1000-2000Hz의 옥타브 내에서 3.6Hz입니다. 즉, 사람들은 훨씬 더 정확하고 위와 같은 스펙트로그램으로 표현된 것보다 훨씬 더 작은 세부 사항을 들을 수 있다는 것을 의미합니다.","metadata":{}},{"cell_type":"markdown","source":"### MFCC\n\nIf you want to get to know some details about MFCC take a look at this great tutorial. MFCC explained You can see, that it is well prepared to imitate human hearing properties.\n\nYou can calculate Mel power spectrogram and MFCC using for example librosa python package.\n","metadata":{}},{"cell_type":"markdown","source":"MFCC에 대한 자세한 내용을 알고 싶다면 이 훌륭한 튜토리얼을 살펴보세요. MFCC 설명 인간의 청각 특성을 모방하도록 잘 준비된 것을 볼 수 있습니다.\n\n예를 들어 librosa Python 패키지를 사용하면 Mel 전력 스펙트로그램과 MFCC를 계산할 수 있습니다.","metadata":{}},{"cell_type":"code","source":"# From this tutorial\n# https://github.com/librosa/librosa/blob/master/examples/LibROSA%20demo.ipynb\nS = librosa.feature.melspectrogram(samples, sr=sample_rate, n_mels=128)\n\n# Convert to log scale (dB). We'll use the peak power (max) as reference.\nlog_S = librosa.power_to_db(S, ref=np.max)\n\nplt.figure(figsize=(12, 4))\nlibrosa.display.specshow(log_S, sr=sample_rate, x_axis='time', y_axis='mel')\nplt.title('Mel power spectrogram ')\nplt.colorbar(format='%+02.0f dB')\nplt.tight_layout()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-08-16T08:02:41.282020Z","iopub.execute_input":"2024-08-16T08:02:41.282347Z","iopub.status.idle":"2024-08-16T08:02:42.366569Z","shell.execute_reply.started":"2024-08-16T08:02:41.282285Z","shell.execute_reply":"2024-08-16T08:02:42.365272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Now delta- mfcc","metadata":{}},{"cell_type":"code","source":"mfcc = librosa.feature.mfcc(S=log_S, n_mfcc=13)\n\n# Let's pad on the first and second deltas while we're at it\ndelta2_mfcc = librosa.feature.delta(mfcc, order=2)\n\nplt.figure(figsize=(12, 4))\nlibrosa.display.specshow(delta2_mfcc)\nplt.ylabel('MFCC coeffs')\nplt.xlabel('Time')\nplt.title('MFCC')\nplt.colorbar()\nplt.tight_layout()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-08-16T08:02:42.368573Z","iopub.execute_input":"2024-08-16T08:02:42.369223Z","iopub.status.idle":"2024-08-16T08:02:42.682494Z","shell.execute_reply.started":"2024-08-16T08:02:42.369157Z","shell.execute_reply":"2024-08-16T08:02:42.680729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Spectrogram in 3d ","metadata":{}},{"cell_type":"code","source":"# data = [go.Surface(z=spectrogram.T)]\n# layout = go.Layout(\n#     title='Specgtrogram of \"yes\" in 3d',\n#     scene = dict(\n#     yaxis = dict(title='Frequencies', range=freqs),\n#     xaxis = dict(title='Time', range=times),\n#     zaxis = dict(title='Log amplitude'),\n#     ),\n# )\n# fig = go.Figure(data=data, layout=layout)\n# py.iplot(fig)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:42.684822Z","iopub.execute_input":"2024-08-16T08:02:42.685521Z","iopub.status.idle":"2024-08-16T08:02:42.694802Z","shell.execute_reply.started":"2024-08-16T08:02:42.685228Z","shell.execute_reply":"2024-08-16T08:02:42.693310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In classical systems, MFCC or similar features are taken as the input to the system instead of spectrograms.\n\nHowever, in end-to-end (often neural-network based) systems, the most common input features are probably raw spectrograms, or mel power spectrograms. For example MFCC decorrelates features, but NNs deal with correlated features well. ","metadata":{}},{"cell_type":"markdown","source":"기존 시스템에서는 스펙트로그램 대신 MFCC나 이와 유사한 기능이 시스템의 입력으로 사용됩니다.\n\n그러나 엔드투엔드(종종 신경망 기반) 시스템에서 가장 일반적인 입력 기능은 아마도 원시 스펙트로그램 또는 멜 파워 스펙트로그램일 것입니다. 예를 들어 MFCC는 기능을 상관시키지 않지만 NN은 상관된 기능을 잘 처리합니다.","metadata":{}},{"cell_type":"markdown","source":"## 2. Data Preprocessing \n\n### Silence Removal\nAlthough the words are short, there is a lot of silence in them. A decent VAD can reduce training size a lot, accelerating training speed significantly. Let's cut a bit of the file from the beginning and from the end. and listen to it again (based on a plot above, we take from 4000 to 13000):","metadata":{}},{"cell_type":"markdown","source":"침묵 제거\n단어는 짧지만, 그 안에는 침묵이 많이 있습니다. 괜찮은 VAD는 훈련 크기를 크게 줄여 훈련 속도를 상당히 가속화할 수 있습니다. 파일의 시작과 끝 부분을 조금 잘라내 보겠습니다. 그리고 다시 들어보세요(위의 플롯을 기준으로 4000에서 13000까지 가져갑니다):","metadata":{}},{"cell_type":"code","source":"## without silence removal\nipd.Audio(samples, rate=sample_rate)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:42.697069Z","iopub.execute_input":"2024-08-16T08:02:42.697945Z","iopub.status.idle":"2024-08-16T08:02:42.729845Z","shell.execute_reply.started":"2024-08-16T08:02:42.697852Z","shell.execute_reply":"2024-08-16T08:02:42.728604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# With manual silence removal\nsamples_cut = samples[4000:13000]\nipd.Audio(samples_cut, rate=sample_rate)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:42.731995Z","iopub.execute_input":"2024-08-16T08:02:42.732807Z","iopub.status.idle":"2024-08-16T08:02:42.748443Z","shell.execute_reply.started":"2024-08-16T08:02:42.732730Z","shell.execute_reply":"2024-08-16T08:02:42.747216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can agree that the entire word can be heard. It is impossible to cut all the files manually and do this basing on the simple plot. But you can use for example webrtcvad package to have a good VAD.\n\nLet's plot it again, together with guessed alignment of 'y' 'e' 's' graphems","metadata":{}},{"cell_type":"markdown","source":"우리는 단어 전체를 들을 수 있다는 데 동의할 수 있습니다. 모든 파일을 수동으로 잘라내고 간단한 플롯을 기반으로 이를 수행하는 것은 불가능합니다. 하지만 예를 들어 webrtcvad 패키지를 사용하여 좋은 VAD를 가질 수 있습니다.\n\n'y' 'e' 's' 자소의 추측된 정렬과 함께 다시 그래프로 표시해 보겠습니다.","metadata":{}},{"cell_type":"code","source":"freqs, times, spectrogram_cut = log_specgram(samples_cut, sample_rate)\n\nfig = plt.figure(figsize=(14, 8))\nax1 = fig.add_subplot(211)\nax1.set_title('Raw wave of ' + filename)\nax1.set_ylabel('Amplitude')\nax1.plot(samples_cut)\n\nax2 = fig.add_subplot(212)\nax2.set_title('Spectrogram of ' + filename)\nax2.set_ylabel('Frequencies * 0.1')\nax2.set_xlabel('Samples')\nax2.imshow(spectrogram_cut.T, aspect='auto', origin='lower', \n           extent=[times.min(), times.max(), freqs.min(), freqs.max()])\nax2.set_yticks(freqs[::16])\nax2.set_xticks(times[::16])\nax2.text(0.06, 1000, 'Y', fontsize=18)\nax2.text(0.17, 1000, 'E', fontsize=18)\nax2.text(0.36, 1000, 'S', fontsize=18)\n\nxcoords = [0.025, 0.11, 0.23, 0.49]\nfor xc in xcoords:\n    ax1.axvline(x=xc*16000, c='r')\n    ax2.axvline(x=xc, c='r')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-08-16T08:02:42.750736Z","iopub.execute_input":"2024-08-16T08:02:42.751563Z","iopub.status.idle":"2024-08-16T08:02:43.432601Z","shell.execute_reply.started":"2024-08-16T08:02:42.751489Z","shell.execute_reply":"2024-08-16T08:02:43.431499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" ### Resampling - dimensionality reduction\n \n - reduce the dimensionality of our data is to resample recordings.\n - smaller training size.\n \nYou can hear that the recording don't sound very natural, because they are sampled with 16k frequency, and we usually hear much more. \nHowever, the most speech related frequencies are presented in smaller band. That's why you can still understand another person talking to the telephone, where GSM signal is sampled to 8000 Hz.\n\nSummarizing, we could resample our dataset to 8k. We will discard some information that shouldn't be important, and we'll reduce size of the data.\n\n**FFT (Fast Fourier Transform)** ","metadata":{}},{"cell_type":"markdown","source":"리샘플링 - 차원 감소\n데이터의 차원을 줄이는 방법은 녹음을 재샘플링하는 것입니다.\n더 작은 훈련 규모.\n녹음이 자연스럽지 않게 들린다는 것을 알 수 있는데, 16k 주파수로 샘플링했기 때문이며, 우리는 보통 훨씬 더 많은 것을 듣습니다. 그러나 가장 음성과 관련된 주파수는 더 작은 대역으로 표현됩니다. 그래서 GSM 신호가 8000Hz로 샘플링되는 전화에서 다른 사람이 말하는 것을 여전히 이해할 수 있습니다.\n\n요약하자면, 우리는 데이터 세트를 8k로 리샘플링할 수 있습니다. 중요하지 않은 정보를 버리고 데이터 크기를 줄일 것입니다.","metadata":{}},{"cell_type":"code","source":"def custom_fft(y, fs):\n    T = 1.0 / fs\n    N = y.shape[0]\n    yf = fft(y)\n    xf = np.linspace(0.0, 1.0/(2.0*T), N//2)\n    vals = 2.0/N * np.abs(yf[0:N//2])  # FFT is simmetrical, so we take just the first half\n    # FFT is also complex, to we take just the real part (abs)\n    return xf, vals","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:43.434508Z","iopub.execute_input":"2024-08-16T08:02:43.435046Z","iopub.status.idle":"2024-08-16T08:02:43.440731Z","shell.execute_reply.started":"2024-08-16T08:02:43.434981Z","shell.execute_reply":"2024-08-16T08:02:43.439951Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's read some recording, resample it, and listen. We can also compare FFT, Notice, that there is almost no information above 4000 Hz in original signal.","metadata":{}},{"cell_type":"markdown","source":"녹음을 읽고, 리샘플링하고, 들어보죠. FFT도 비교할 수 있습니다. 원래 신호에는 4000Hz 이상에 대한 정보가 거의 없다는 점에 유의하세요.","metadata":{}},{"cell_type":"code","source":"# filename = '/happy/0b09edd3_nohash_0.wav'\nfilename ='/yes/00f0204f_nohash_0.wav'\nnew_sample_rate = 8000\n\nsample_rate, samples = wavfile.read(str(train_audio_path) + filename)\nresampled = signal.resample(samples, int(new_sample_rate/sample_rate * samples.shape[0]))","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:43.442243Z","iopub.execute_input":"2024-08-16T08:02:43.442604Z","iopub.status.idle":"2024-08-16T08:02:43.459309Z","shell.execute_reply.started":"2024-08-16T08:02:43.442540Z","shell.execute_reply":"2024-08-16T08:02:43.458166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# without resampling \nipd.Audio(samples, rate=sample_rate)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:43.460820Z","iopub.execute_input":"2024-08-16T08:02:43.461434Z","iopub.status.idle":"2024-08-16T08:02:43.479629Z","shell.execute_reply.started":"2024-08-16T08:02:43.461341Z","shell.execute_reply":"2024-08-16T08:02:43.478768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# with resampling \nipd.Audio(resampled, rate=new_sample_rate)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:43.481285Z","iopub.execute_input":"2024-08-16T08:02:43.481569Z","iopub.status.idle":"2024-08-16T08:02:43.502320Z","shell.execute_reply.started":"2024-08-16T08:02:43.481515Z","shell.execute_reply":"2024-08-16T08:02:43.501245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# At original Sampling \n\nxf, vals = custom_fft(samples, sample_rate)\nplt.figure(figsize=(12, 4))\nplt.title('FFT of recording sampled with ' + str(sample_rate) + ' Hz')\nplt.plot(xf, vals)\nplt.xlabel('Frequency')\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:43.504059Z","iopub.execute_input":"2024-08-16T08:02:43.504418Z","iopub.status.idle":"2024-08-16T08:02:43.811464Z","shell.execute_reply.started":"2024-08-16T08:02:43.504364Z","shell.execute_reply":"2024-08-16T08:02:43.809925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# After resampling to reduce traning dat size \n\nxf, vals = custom_fft(resampled, new_sample_rate)\nplt.figure(figsize=(12, 4))\nplt.title('FFT of recording sampled with ' + str(new_sample_rate) + ' Hz')\nplt.plot(xf, vals)\nplt.xlabel('Frequency')\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:43.813610Z","iopub.execute_input":"2024-08-16T08:02:43.814012Z","iopub.status.idle":"2024-08-16T08:02:44.121314Z","shell.execute_reply.started":"2024-08-16T08:02:43.813953Z","shell.execute_reply":"2024-08-16T08:02:44.120255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Set Investigation\n\nNumver of Records ","metadata":{}},{"cell_type":"code","source":"dirs = [f for f in os.listdir(train_audio_path) if isdir(join(train_audio_path, f))]\ndirs.sort()\nprint('Number of labels: ' + str(len(dirs)))","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:44.123180Z","iopub.execute_input":"2024-08-16T08:02:44.123849Z","iopub.status.idle":"2024-08-16T08:02:44.132591Z","shell.execute_reply.started":"2024-08-16T08:02:44.123783Z","shell.execute_reply":"2024-08-16T08:02:44.131574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate\nnumber_of_recordings = []\nfor direct in dirs:\n    waves = [f for f in os.listdir(join(train_audio_path, direct)) if f.endswith('.wav')]\n    number_of_recordings.append(len(waves))\n\n# Plot\ndata = [go.Histogram(x=dirs, y=number_of_recordings)]\ntrace = go.Bar(\n    x=dirs,\n    y=number_of_recordings,\n    marker=dict(color = number_of_recordings, colorscale='dense', showscale=True\n    ),\n)\nlayout = go.Layout(\n    title='Number of recordings in given label',\n    xaxis = dict(title='Words'),\n    yaxis = dict(title='Number of recordings')\n)\npy.iplot(go.Figure(data=[trace], layout=layout))","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:44.134368Z","iopub.execute_input":"2024-08-16T08:02:44.135045Z","iopub.status.idle":"2024-08-16T08:02:46.468810Z","shell.execute_reply.started":"2024-08-16T08:02:44.134969Z","shell.execute_reply":"2024-08-16T08:02:46.467861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"split the dataset in a way that one speaker doesn't occur in both train and test sets","metadata":{}},{"cell_type":"markdown","source":"한 화자가 훈련 세트와 테스트 세트 모두에 나타나지 않도록 데이터 세트를 분할합니다","metadata":{}},{"cell_type":"code","source":"filenames = ['/yes/00f0204f_nohash_0.wav', '/yes/8830e17f_nohash_2.wav']\nfor filename in filenames:\n    sample_rate, samples = wavfile.read(str(train_audio_path) + filename)\n    xf, vals = custom_fft(samples, sample_rate)\n    plt.figure(figsize=(12, 4))\n    plt.title('FFT of speaker ' + filename[4:11])\n    plt.plot(xf, vals)\n    plt.xlabel('Frequency')\n    plt.grid()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:46.470334Z","iopub.execute_input":"2024-08-16T08:02:46.470672Z","iopub.status.idle":"2024-08-16T08:02:47.099175Z","shell.execute_reply.started":"2024-08-16T08:02:46.470613Z","shell.execute_reply":"2024-08-16T08:02:47.098132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filenames = ['on/004ae714_nohash_0.wav', 'on/0137b3f4_nohash_0.wav']\n\nprint('Speaker ' + filenames[0][4:11])\n# Female Speaker\nipd.Audio( join(train_audio_path, filenames[0]), \n          rate=8000)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:47.100954Z","iopub.execute_input":"2024-08-16T08:02:47.101640Z","iopub.status.idle":"2024-08-16T08:02:47.114437Z","shell.execute_reply.started":"2024-08-16T08:02:47.101575Z","shell.execute_reply":"2024-08-16T08:02:47.113308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Speaker ' + filenames[1][4:11])\n# Male Speaker\nipd.Audio(join(train_audio_path, filenames[1]))","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:47.116519Z","iopub.execute_input":"2024-08-16T08:02:47.117169Z","iopub.status.idle":"2024-08-16T08:02:47.129612Z","shell.execute_reply.started":"2024-08-16T08:02:47.117105Z","shell.execute_reply":"2024-08-16T08:02:47.128057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filename = '/yes/01bb6a2a_nohash_1.wav'\nsample_rate, samples = wavfile.read(str(train_audio_path) + filename)\nfreqs, times, spectrogram = log_specgram(samples, sample_rate)\n\nplt.figure(figsize=(10, 7))\nplt.title('Spectrogram of ' + filename)\nplt.ylabel('Freqs')\nplt.xlabel('Time')\nplt.imshow(spectrogram.T, aspect='auto', origin='lower', \n           extent=[times.min(), times.max(), freqs.min(), freqs.max()])\nplt.yticks(freqs[::16])\nplt.xticks(times[::16])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:47.131353Z","iopub.execute_input":"2024-08-16T08:02:47.131998Z","iopub.status.idle":"2024-08-16T08:02:47.462658Z","shell.execute_reply.started":"2024-08-16T08:02:47.131933Z","shell.execute_reply":"2024-08-16T08:02:47.460324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Recordings length","metadata":{}},{"cell_type":"code","source":"os.listdir(join(train_audio_path, direct))","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:47.470554Z","iopub.execute_input":"2024-08-16T08:02:47.471388Z","iopub.status.idle":"2024-08-16T08:02:47.497577Z","shell.execute_reply.started":"2024-08-16T08:02:47.471314Z","shell.execute_reply":"2024-08-16T08:02:47.496558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls -la","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:47.499553Z","iopub.execute_input":"2024-08-16T08:02:47.500213Z","iopub.status.idle":"2024-08-16T08:02:48.556201Z","shell.execute_reply.started":"2024-08-16T08:02:47.500149Z","shell.execute_reply":"2024-08-16T08:02:48.554981Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_audio_path)\nprint(direct)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:48.558882Z","iopub.execute_input":"2024-08-16T08:02:48.559355Z","iopub.status.idle":"2024-08-16T08:02:48.565456Z","shell.execute_reply.started":"2024-08-16T08:02:48.559272Z","shell.execute_reply":"2024-08-16T08:02:48.564372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"all the files have 1 second duration:","metadata":{}},{"cell_type":"markdown","source":"모든 파일의 지속 시간은 1초입니다.","metadata":{}},{"cell_type":"code","source":"os.listdir(train_audio_path)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:48.567263Z","iopub.execute_input":"2024-08-16T08:02:48.567561Z","iopub.status.idle":"2024-08-16T08:02:48.579650Z","shell.execute_reply.started":"2024-08-16T08:02:48.567496Z","shell.execute_reply":"2024-08-16T08:02:48.578508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_of_shorter = 0\nfor direct in dirs:\n    waves = [f for f in os.listdir(join(train_audio_path, direct)) if f.endswith('.wav')]\n    for wav in waves:\n#         try:\n            sample_rate, samples = wavfile.read(train_audio_path +'/' +direct + '/' + wav)\n            if samples.shape[0] < sample_rate:\n                num_of_shorter += 1\n#         except:\n#             print(\"this gets executed only if there is an error\")\n              \nprint('Number of recordings shorter than 1 second: ' + str(num_of_shorter))\n# example file :'/kaggle/working/train/train/audio_/background_noise_/doing_the_dishes.wav'","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:48.581413Z","iopub.execute_input":"2024-08-16T08:02:48.581729Z","iopub.status.idle":"2024-08-16T08:02:51.492661Z","shell.execute_reply.started":"2024-08-16T08:02:48.581672Z","shell.execute_reply":"2024-08-16T08:02:51.491598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###  Mean spectrograms and FFT","metadata":{}},{"cell_type":"code","source":"to_keep = 'yes no up down left right on off stop go'.split()\ndirs = [d for d in dirs if d in to_keep]\n\nprint(dirs)\n\nfor direct in dirs:\n    vals_all = []\n    spec_all = []\n\n    waves = [f for f in os.listdir(join(train_audio_path, direct)) if f.endswith('.wav')]\n    for wav in waves:\n        sample_rate, samples = wavfile.read(train_audio_path +'/' + direct + '/' + wav)\n        if samples.shape[0] != 16000:\n            continue\n        xf, vals = custom_fft(samples, 16000)\n        vals_all.append(vals)\n        freqs, times, spec = log_specgram(samples, 16000)\n        spec_all.append(spec)\n\n    plt.figure(figsize=(14, 4))\n    plt.subplot(121)\n    plt.title('Mean fft of ' + direct)\n    plt.plot(np.mean(np.array(vals_all), axis=0))\n    plt.grid()\n    plt.subplot(122)\n    plt.title('Mean specgram of ' + direct)\n    plt.imshow(np.mean(np.array(spec_all), axis=0).T, aspect='auto', origin='lower', \n               extent=[times.min(), times.max(), freqs.min(), freqs.max()])\n    plt.yticks(freqs[::16])\n    plt.xticks(times[::16])\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:02:51.494455Z","iopub.execute_input":"2024-08-16T08:02:51.494762Z","iopub.status.idle":"2024-08-16T08:03:33.762247Z","shell.execute_reply.started":"2024-08-16T08:02:51.494718Z","shell.execute_reply":"2024-08-16T08:03:33.760990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" **Guassian Mixtures modelling** \n \n Kaldi library, that can model words (or smaller parts of words) with GMMs and model temporal dependencies with Hidden Markov Models.","metadata":{}},{"cell_type":"markdown","source":"Kaldi 라이브러리는 GMM을 사용하여 단어(또는 단어의 작은 부분)를 모델링하고 은닉 마르코프 모델을 사용하여 시간적 종속성을 모델링할 수 있습니다.","metadata":{}},{"cell_type":"code","source":"def violinplot_frequency(dirs, freq_ind):\n    \"\"\" Plot violinplots for given words (waves in dirs) and frequency freq_ind\n    from all frequencies freqs.\"\"\"\n\n    spec_all = []  # Contain spectrograms\n    ind = 0\n    # taking first 8 words only to keep the plots clean and unclumsy.\n    for direct in dirs[:8]:\n        spec_all.append([])\n\n        waves = [f for f in os.listdir(join(train_audio_path, direct)) if\n                 f.endswith('.wav')]\n        for wav in waves[:100]:\n            sample_rate, samples = wavfile.read(\n                train_audio_path + '/' + direct + '/' + wav)\n            freqs, times, spec = log_specgram(samples, sample_rate)\n            spec_all[ind].extend(spec[:, freq_ind])\n        ind += 1\n\n    # Different lengths = different num of frames. Make number equal\n    minimum = min([len(spec) for spec in spec_all])\n    spec_all = np.array([spec[:minimum] for spec in spec_all])\n\n    plt.figure(figsize=(13,7))\n    plt.title('Frequency ' + str(freqs[freq_ind]) + ' Hz')\n    plt.ylabel('Amount of frequency in a word')\n    plt.xlabel('Words')\n    sns.violinplot(data=pd.DataFrame(spec_all.T, columns=dirs[:8]))\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:03:33.763903Z","iopub.execute_input":"2024-08-16T08:03:33.764345Z","iopub.status.idle":"2024-08-16T08:03:33.781376Z","shell.execute_reply.started":"2024-08-16T08:03:33.764285Z","shell.execute_reply":"2024-08-16T08:03:33.780214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"violinplot_frequency(dirs, 20)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:03:33.783511Z","iopub.execute_input":"2024-08-16T08:03:33.784139Z","iopub.status.idle":"2024-08-16T08:03:35.269042Z","shell.execute_reply.started":"2024-08-16T08:03:33.783901Z","shell.execute_reply":"2024-08-16T08:03:35.268016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. Voice Activity Detection ( VAD )","metadata":{}},{"cell_type":"markdown","source":"### use the webrtcvad library to identify segments as speech or not","metadata":{}},{"cell_type":"markdown","source":"webrtcvad 라이브러리를 사용하여 세그먼트를 음성인지 아닌지 식별합니다.","metadata":{}},{"cell_type":"code","source":"!pip install webrtcvad","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:03:35.270616Z","iopub.execute_input":"2024-08-16T08:03:35.271634Z","iopub.status.idle":"2024-08-16T08:03:46.930506Z","shell.execute_reply.started":"2024-08-16T08:03:35.271560Z","shell.execute_reply":"2024-08-16T08:03:46.929399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import webrtcvad","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:03:46.932620Z","iopub.execute_input":"2024-08-16T08:03:46.933085Z","iopub.status.idle":"2024-08-16T08:03:47.097015Z","shell.execute_reply.started":"2024-08-16T08:03:46.932873Z","shell.execute_reply":"2024-08-16T08:03:47.096078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### reading the samples and sample_rate feature again to make them compatible with the webrtcvad library. ( it reads at sample_rate = 16000, 32000, 48000; but we had sample_rate = 22050 with librosa)","metadata":{}},{"cell_type":"markdown","source":"샘플과 sample_rate 기능을 다시 읽어서 webrtcvad 라이브러리와 호환되도록 합니다. (sample_rate = 16000, 32000, 48000으로 읽습니다. 하지만 librosa에서는 sample_rate = 22050이었습니다)","metadata":{}},{"cell_type":"code","source":"sample_rate, samples = wavfile.read(str(train_audio_path) + filename)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:03:47.098629Z","iopub.execute_input":"2024-08-16T08:03:47.099020Z","iopub.status.idle":"2024-08-16T08:03:47.105088Z","shell.execute_reply.started":"2024-08-16T08:03:47.098945Z","shell.execute_reply":"2024-08-16T08:03:47.103922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"vad = webrtcvad.Vad()\n# set aggressiveness from 0 to 3\nvad.set_mode(3)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:03:47.106938Z","iopub.execute_input":"2024-08-16T08:03:47.107314Z","iopub.status.idle":"2024-08-16T08:03:47.115994Z","shell.execute_reply.started":"2024-08-16T08:03:47.107257Z","shell.execute_reply":"2024-08-16T08:03:47.115002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"convert samples to raw 16 bit per sample stream needed by webrtcvad( there are other options available too , like 32 )","metadata":{}},{"cell_type":"markdown","source":"webrtcvad에 필요한 샘플 스트림당 샘플을 원시 16비트로 변환합니다(32와 같은 다른 옵션도 사용 가능)","metadata":{}},{"cell_type":"code","source":"import struct\nraw_samples = struct.pack(\"%dh\" % len(samples), *samples)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:03:47.117815Z","iopub.execute_input":"2024-08-16T08:03:47.118241Z","iopub.status.idle":"2024-08-16T08:03:47.129115Z","shell.execute_reply.started":"2024-08-16T08:03:47.118169Z","shell.execute_reply":"2024-08-16T08:03:47.128262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"run the detector on windows of 30 ms (from https://github.com/wiseman/py-webrtcvad/blob/master/example.py)","metadata":{}},{"cell_type":"markdown","source":"30ms 창에서 감지기를 실행합니다( https://github.com/wiseman/py-webrtcvad/blob/master/example.py 에서 )","metadata":{}},{"cell_type":"code","source":"window_duration = 0.03 # duration in seconds\nsamples_per_window = int(window_duration * sample_rate + 0.5)\nbytes_per_sample = 2","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:03:47.130861Z","iopub.execute_input":"2024-08-16T08:03:47.131249Z","iopub.status.idle":"2024-08-16T08:03:47.139881Z","shell.execute_reply.started":"2024-08-16T08:03:47.131189Z","shell.execute_reply":"2024-08-16T08:03:47.138930Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Detect Speech instances in an audio","metadata":{}},{"cell_type":"markdown","source":"오디오에서 음성 인스턴스 감지","metadata":{}},{"cell_type":"code","source":"segments = []\n\nfor start in np.arange(0, len(samples), samples_per_window):\n    stop = min(start + samples_per_window, len(samples))\n    \n    is_speech = vad.is_speech(raw_samples[start * bytes_per_sample: stop * bytes_per_sample], \n                              sample_rate = sample_rate)\n\n    segments.append(dict(\n       start = start,\n       stop = stop,\n       is_speech = is_speech))","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:03:47.141637Z","iopub.execute_input":"2024-08-16T08:03:47.142037Z","iopub.status.idle":"2024-08-16T08:03:47.153891Z","shell.execute_reply.started":"2024-08-16T08:03:47.141966Z","shell.execute_reply":"2024-08-16T08:03:47.152886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"plot segment identifed as speech","metadata":{}},{"cell_type":"markdown","source":"세그먼트는 연설로 식별된 것을 플롯","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (10,7))\nplt.plot(samples)\n\nymax = max(samples)\n\n\nfor segment in segments:\n    if segment['is_speech']:\n        plt.plot([ segment['start'], segment['stop'] - 1], [ymax * 1.1, ymax * 1.1], color = 'orange')\n\nplt.xlabel('sample')\nplt.grid()","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:03:47.155462Z","iopub.execute_input":"2024-08-16T08:03:47.155746Z","iopub.status.idle":"2024-08-16T08:03:47.575269Z","shell.execute_reply.started":"2024-08-16T08:03:47.155696Z","shell.execute_reply":"2024-08-16T08:03:47.574193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" Listen to the speech only segments","metadata":{}},{"cell_type":"markdown","source":"연설 부분만 들어보세요","metadata":{}},{"cell_type":"code","source":"speech_samples = np.concatenate([ samples[segment['start']:segment['stop']] for segment in segments if segment['is_speech']])\n\nimport IPython.display as ipd\nipd.Audio(speech_samples, rate=sample_rate)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:03:47.576831Z","iopub.execute_input":"2024-08-16T08:03:47.577165Z","iopub.status.idle":"2024-08-16T08:03:47.588494Z","shell.execute_reply.started":"2024-08-16T08:03:47.577109Z","shell.execute_reply":"2024-08-16T08:03:47.587560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Till now we have processed for a single audio of any one word : <span style=\"color : blue;\">YES</span> here.\n\n#### Now, its time to have an overall view on other words also. So, lets visualize frequency components for other words as well.","metadata":{}},{"cell_type":"markdown","source":"# 3. Anomaly detection\n\n lower the dimensionality of the dataset and interactively check for any anomaly. We'll use PCA for dimensionality reduction:","metadata":{}},{"cell_type":"markdown","source":"이상 탐지\n데이터 세트의 차원을 낮추고 상호작용적으로 이상 여부를 확인합니다. 차원 감소를 위해 PCA를 사용합니다.","metadata":{}},{"cell_type":"code","source":"fft_all = []\nnames = []\nfor direct in dirs:\n    waves = [f for f in os.listdir(join(train_audio_path, direct)) if f.endswith('.wav')]\n    for wav in waves:\n        sample_rate, samples = wavfile.read(train_audio_path+ '/' + direct + '/' + wav)\n        if samples.shape[0] != sample_rate:\n            samples = np.append(samples, np.zeros((sample_rate - samples.shape[0], )))\n        x, val = custom_fft(samples, sample_rate)\n        fft_all.append(val)\n        names.append(direct + '/' + wav)\n\nfft_all = np.array(fft_all)\n\n# Normalization\nfft_all = (fft_all - np.mean(fft_all, axis=0)) / np.std(fft_all, axis=0)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:03:47.590213Z","iopub.execute_input":"2024-08-16T08:03:47.590535Z","iopub.status.idle":"2024-08-16T08:04:06.240930Z","shell.execute_reply.started":"2024-08-16T08:03:47.590480Z","shell.execute_reply":"2024-08-16T08:04:06.239823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Dimemsionality reduction\npca = PCA(n_components=3)\nfft_all = pca.fit_transform(fft_all)\n\ndef interactive_3d_plot(data, names):\n    scatt = go.Scatter3d(x=data[:, 0], y=data[:, 1], z=data[:, 2], mode='markers', text=names)\n    data = go.Data([scatt])\n    layout = go.Layout(title=\"Anomaly detection\")\n    figure = go.Figure(data=data, layout=layout)\n    py.iplot(figure)\n    \ninteractive_3d_plot(fft_all, names)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:04:06.242824Z","iopub.execute_input":"2024-08-16T08:04:06.243184Z","iopub.status.idle":"2024-08-16T08:04:19.927369Z","shell.execute_reply.started":"2024-08-16T08:04:06.243114Z","shell.execute_reply":"2024-08-16T08:04:19.925559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Some anomalied listed below","metadata":{}},{"cell_type":"markdown","source":"아래에 일부 이상 항목이 나열되어 있습니다.","metadata":{}},{"cell_type":"code","source":"print('Recording go/0487ba9b_nohash_0.wav')\nipd.Audio(join(train_audio_path, 'go/0487ba9b_nohash_0.wav'))","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:04:19.930404Z","iopub.execute_input":"2024-08-16T08:04:19.931481Z","iopub.status.idle":"2024-08-16T08:04:19.950550Z","shell.execute_reply.started":"2024-08-16T08:04:19.931348Z","shell.execute_reply":"2024-08-16T08:04:19.949179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Recording yes/e4b02540_nohash_0.wav')\nipd.Audio(join(train_audio_path, 'yes/e4b02540_nohash_0.wav'))","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:04:19.952380Z","iopub.execute_input":"2024-08-16T08:04:19.952731Z","iopub.status.idle":"2024-08-16T08:04:19.967681Z","shell.execute_reply.started":"2024-08-16T08:04:19.952672Z","shell.execute_reply":"2024-08-16T08:04:19.966693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Recording seven/e4b02540_nohash_0.wav')\nipd.Audio(join(train_audio_path, 'seven/b1114e4f_nohash_0.wav'))","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:04:19.969421Z","iopub.execute_input":"2024-08-16T08:04:19.969751Z","iopub.status.idle":"2024-08-16T08:04:19.985732Z","shell.execute_reply.started":"2024-08-16T08:04:19.969687Z","shell.execute_reply":"2024-08-16T08:04:19.984377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Frequency components across the words\n\n### plotting for first 8 words only to avoid clumsy tight plots.","metadata":{}},{"cell_type":"markdown","source":"단어 전체의 주파수 성분\n어색하고 딱딱한 줄거리를 피하기 위해 처음 8개 단어에 대해서만 줄거리를 구성했습니다.","metadata":{}},{"cell_type":"code","source":"violinplot_frequency(dirs, 20)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:04:19.987407Z","iopub.execute_input":"2024-08-16T08:04:19.987871Z","iopub.status.idle":"2024-08-16T08:04:23.510672Z","shell.execute_reply.started":"2024-08-16T08:04:19.987686Z","shell.execute_reply":"2024-08-16T08:04:23.509361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"violinplot_frequency(dirs, 50)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:04:23.512804Z","iopub.execute_input":"2024-08-16T08:04:23.513497Z","iopub.status.idle":"2024-08-16T08:04:25.174272Z","shell.execute_reply.started":"2024-08-16T08:04:23.513427Z","shell.execute_reply":"2024-08-16T08:04:25.172694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"violinplot_frequency(dirs, 120)","metadata":{"execution":{"iopub.status.busy":"2024-08-16T08:04:25.177464Z","iopub.execute_input":"2024-08-16T08:04:25.177822Z","iopub.status.idle":"2024-08-16T08:04:26.719122Z","shell.execute_reply.started":"2024-08-16T08:04:25.177761Z","shell.execute_reply":"2024-08-16T08:04:26.718059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Copyright 2019 The TensorFlow Authors.\n        #@title Licensed under the Apache License, Version 2.0 (the \"License\");\n        # you may not use this file except in compliance with the License.\n        # You may obtain a copy of the License at\n        #\n        # https://www.apache.org/licenses/LICENSE-2.0\n        #\n        # Unless required by applicable law or agreed to in writing, software\n        # distributed under the License is distributed on an \"AS IS\" BASIS,\n        # WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.\n        # See the License for the specific language governing permissions and\n        # limitations under the License.","metadata":{}},{"cell_type":"markdown","source":"References for this work\n\n- EABDSMD- https://www.kaggle.com/samadi10/trigger-word-detection-transformer ","metadata":{}},{"cell_type":"markdown","source":"--------------------------------------------------\n**Reading material**\n\n* Encoder-decoder: https://arxiv.org/abs/1508.01211\n* RNNs with CTC loss: https://arxiv.org/abs/1412.5567\n* For me, 1 and 2 are a sensible choice for this competition, especially if you do not have background in SR field. They try to be end-to-end solutions. Speech recognition is a really big topic and it would be hard to get to know important things in short time.\n* Classic speech recognition : http://www.ece.ucsb.edu/Faculty/Rabiner/ece259/Reprints/tutorial%20on%20hmm%20and%20applications.pdf\n\n* Kaldi Tutorial for dummies, with a problem similar to this competition in some way.\n\n* Very deep CNN - Large Vocabulary Continuous Speech Recognition Systems (LVCSR). \n","metadata":{}}]}