{"cells": [{"metadata": {"_cell_guid": "d524ad8f-d024-4ba3-8bbc-2435ec3d0dba", "_uuid": "7bb30a6f2a18be45a491092a6e2dfcdcce71a2a0"}, "cell_type": "markdown", "source": ["## Preface\n", "This notebooks aims to build a light-weight CNN.\n", "\n", "It uses specgrams of resampled wav files(rate 8000) as inputs.\n", "\n", "Due to Kaggle cloud hardware limitations, this script is a 'crippled' version of the original one.\n", "\n", "In order to get LB 0.74, you need to set epoch to 5, set chop_audio(num=1000) and double all Conv layer parameters.\n", "\n", "Although this script is a slight imrpovement over Alex Ozerin's baseline, I believe by using original wav files(16000 sample rate) one can achieve higher scores.\n", "\n", "\n", "## File Structure\n", "This script assumes data are stored in following strcuture:\n", "\n", "speech\n", "\n", "\u251c\u2500\u2500 test            \n", "\n", "\u2502   \u2514\u2500\u2500 audio #test wavfiles\n", "\n", "\u251c\u2500\u2500 train           \n", "\n", "\u2502   \u251c\u2500\u2500 audio #train wavfiles\n", "\n", "\u2514\u2500\u2500 model #store models\n", "\n", "\u2502\n", "\n", "\u2514\u2500\u2500 out #store sub.csv\n", "\n", "## Improve This Script\n", "Since this is only a light-weight CNN, it's performance is limited.\n", "Here are some ways to improve it's performance.\n", "1. Use original wav files instead resampled ones.\n", "2. Create more 'silence' wav files using chop_audio.\n", "3. Build deeper CNN or use RNN.\n", "4. Train for longer epochs\n", "\n", "## After Words\n", "It's still a long way to reach LB 0.88.\n", "\n", "In fact, I doubt CNN would ever reach that high.\n", "\n", "Feel free to share your ideas in the comment sections about using CNN to label wav files :)\n", "\n", "## Appendix\n", "Thanks __DavidS__ and __Alex Ozerin__ for their great notebooks!"]}, {"metadata": {"_cell_guid": "f7fd8bcb-4451-4d47-bfe8-491c94b3b4eb", "collapsed": true, "_uuid": "712710f20b00f97271136cfeab9937a4c6a2458b"}, "outputs": [], "execution_count": null, "cell_type": "code", "source": ["import os\n", "import numpy as np\n", "from scipy.fftpack import fft\n", "from scipy.io import wavfile\n", "from scipy import signal\n", "from glob import glob\n", "import re\n", "import pandas as pd\n", "import gc\n", "from scipy.io import wavfile\n", "\n", "from keras import optimizers, losses, activations, models\n", "from keras.layers import Convolution2D, Dense, Input, Flatten, Dropout, MaxPooling2D, BatchNormalization\n", "from sklearn.model_selection import train_test_split\n", "import keras"]}, {"metadata": {"_cell_guid": "fb35a2f1-9301-4693-a9ef-9d180b630f05", "_uuid": "4b1ba61998e14e15c822c605dbe5961bfed36014"}, "cell_type": "markdown", "source": ["The original sample rate is 16000, and we will resample it to 8000 to reduce data size."]}, {"metadata": {"_cell_guid": "dc66e1df-f1eb-4df4-ba1a-65b9f1675953", "collapsed": true, "_uuid": "4cc586519523b28d1d595716d8709ace9f27ac9c"}, "outputs": [], "execution_count": null, "cell_type": "code", "source": ["L = 16000\n", "legal_labels = 'yes no up down left right on off stop go silence unknown'.split()\n", "\n", "#src folders\n", "root_path = r'..'\n", "out_path = r'.'\n", "model_path = r'.'\n", "train_data_path = os.path.join(root_path, 'input', 'train', 'audio')\n", "test_data_path = os.path.join(root_path, 'input', 'test', 'audio')"]}, {"metadata": {"_cell_guid": "e53561e4-1c98-44c0-9245-d87f7957faa5", "_uuid": "d9a08781f22e574bb1eb0dc29adeb8dddebc8b51"}, "cell_type": "markdown", "source": ["Here are custom_fft and log_specgram functions written by __DavidS__."]}, {"metadata": {"_cell_guid": "0fd0b579-8b6f-4253-bf3a-7f75115a42d6", "collapsed": true, "_uuid": "e7ea2c277b6459e532721452ec3cd80d585eae1e"}, "outputs": [], "execution_count": null, "cell_type": "code", "source": ["def custom_fft(y, fs):\n", "    T = 1.0 / fs\n", "    N = y.shape[0]\n", "    yf = fft(y)\n", "    xf = np.linspace(0.0, 1.0/(2.0*T), N//2)\n", "    # FFT is simmetrical, so we take just the first half\n", "    # FFT is also complex, to we take just the real part (abs)\n", "    vals = 2.0/N * np.abs(yf[0:N//2])\n", "    return xf, vals\n", "\n", "def log_specgram(audio, sample_rate, window_size=20,\n", "                 step_size=10, eps=1e-10):\n", "    nperseg = int(round(window_size * sample_rate / 1e3))\n", "    noverlap = int(round(step_size * sample_rate / 1e3))\n", "    freqs, times, spec = signal.spectrogram(audio,\n", "                                    fs=sample_rate,\n", "                                    window='hann',\n", "                                    nperseg=nperseg,\n", "                                    noverlap=noverlap,\n", "                                    detrend=False)\n", "    return freqs, times, np.log(spec.T.astype(np.float32) + eps)"]}, {"metadata": {"_cell_guid": "c54cda36-777e-4129-bac1-af2d1ed2706e", "_uuid": "5a04e71fe7e66e1a31835feebdfef4c63920faf8"}, "cell_type": "markdown", "source": ["Following is the utility function to grab all wav files inside train data folder."]}, {"metadata": {"_cell_guid": "956f3150-544d-46ed-b0ec-da1c1fb142b4", "collapsed": true, "_uuid": "964d71a229e9d4560b9118fa1c80804ebf8d6be8"}, "outputs": [], "execution_count": null, "cell_type": "code", "source": ["def list_wavs_fname(dirpath, ext='wav'):\n", "    print(dirpath)\n", "    fpaths = glob(os.path.join(dirpath, r'*/*' + ext))\n", "    pat = r'.+/(\\w+)/\\w+\\.' + ext + '$'\n", "    labels = []\n", "    for fpath in fpaths:\n", "        r = re.match(pat, fpath)\n", "        if r:\n", "            labels.append(r.group(1))\n", "    pat = r'.+/(\\w+\\.' + ext + ')$'\n", "    fnames = []\n", "    for fpath in fpaths:\n", "        r = re.match(pat, fpath)\n", "        if r:\n", "            fnames.append(r.group(1))\n", "    return labels, fnames"]}, {"metadata": {"_cell_guid": "41025a55-8497-43cf-b316-003af7d9d19f", "_uuid": "fc18e87793888952e81a867dd95b1dcc455f9932"}, "cell_type": "markdown", "source": ["__pad_audio__ will pad audios that are less than 16000(1 second) with 0s to make them all have the same length.\n", "\n", "__chop_audio__ will chop audios that are larger than 16000(eg. wav files in background noises folder) to 16000 in length. In addition, it will create several chunks out of one large wav files given the parameter 'num'.\n", "\n", "__label_transform__ transform labels into dummies values. It's used in combination with softmax to predict the label."]}, {"metadata": {"_cell_guid": "200c34a1-851a-4447-9ff7-b4e541f090c6", "collapsed": true, "_uuid": "94e40aef3899acfd3ed85557caa66fee5dd47db2"}, "outputs": [], "execution_count": null, "cell_type": "code", "source": ["def pad_audio(samples):\n", "    if len(samples) >= L: return samples\n", "    else: return np.pad(samples, pad_width=(L - len(samples), 0), mode='constant', constant_values=(0, 0))\n", "\n", "def chop_audio(samples, L=16000, num=20):\n", "    for i in range(num):\n", "        beg = np.random.randint(0, len(samples) - L)\n", "        yield samples[beg: beg + L]\n", "\n", "def label_transform(labels):\n", "    nlabels = []\n", "    for label in labels:\n", "        if label == '_background_noise_':\n", "            nlabels.append('silence')\n", "        elif label not in legal_labels:\n", "            nlabels.append('unknown')\n", "        else:\n", "            nlabels.append(label)\n", "    return pd.get_dummies(pd.Series(nlabels))"]}, {"metadata": {"_cell_guid": "dae2a45a-f7ab-4e84-bc73-688eda6eca8e", "_uuid": "267314ef41c459c8b6ab903d721980fdd62b4106"}, "cell_type": "markdown", "source": ["Next, we use functions declared above to generate x_train and y_train.\n", "label_index is the index used by pandas to create dummy values, we need to save it for later use."]}, {"metadata": {"_cell_guid": "4c8d9fdf-ea3e-45fa-b7ef-52542c70b9db", "collapsed": true, "_uuid": "81bc9722dfb036c73721ae44829d429489662e75"}, "outputs": [], "execution_count": null, "cell_type": "code", "source": ["labels, fnames = list_wavs_fname(train_data_path)\n", "\n", "new_sample_rate = 8000\n", "y_train = []\n", "x_train = []\n", "\n", "for label, fname in zip(labels, fnames):\n", "    sample_rate, samples = wavfile.read(os.path.join(train_data_path, label, fname))\n", "    samples = pad_audio(samples)\n", "    if len(samples) > 16000:\n", "        n_samples = chop_audio(samples)\n", "    else: n_samples = [samples]\n", "    for samples in n_samples:\n", "        resampled = signal.resample(samples, int(new_sample_rate / sample_rate * samples.shape[0]))\n", "        _, _, specgram = log_specgram(resampled, sample_rate=new_sample_rate)\n", "        y_train.append(label)\n", "        x_train.append(specgram)\n", "x_train = np.array(x_train)\n", "x_train = x_train.reshape(tuple(list(x_train.shape) + [1]))\n", "y_train = label_transform(y_train)\n", "label_index = y_train.columns.values\n", "y_train = y_train.values\n", "y_train = np.array(y_train)\n", "del labels, fnames\n", "gc.collect()"]}, {"metadata": {"_cell_guid": "56921cf3-1269-4b29-876d-abdd31eb150a", "_uuid": "a87a77b76c42da61ca0bec395c71bef795a9e928"}, "cell_type": "markdown", "source": ["CNN declared below.\n", "The specgram created will be of shape (99, 81), but in order to fit into Conv2D layer, we need to reshape it."]}, {"metadata": {"_cell_guid": "b97e8887-b593-4d88-95c8-fc8f1dd5ca72", "collapsed": true, "_uuid": "60af394ad8e91fb868ea32dbb6ac6a725b5935c9"}, "outputs": [], "execution_count": null, "cell_type": "code", "source": ["input_shape = (99, 81, 1)\n", "nclass = 12\n", "inp = Input(shape=input_shape)\n", "norm_inp = BatchNormalization()(inp)\n", "img_1 = Convolution2D(8, kernel_size=2, activation=activations.relu)(norm_inp)\n", "img_1 = Convolution2D(8, kernel_size=2, activation=activations.relu)(img_1)\n", "img_1 = MaxPooling2D(pool_size=(2, 2))(img_1)\n", "img_1 = Dropout(rate=0.2)(img_1)\n", "img_1 = Convolution2D(16, kernel_size=3, activation=activations.relu)(img_1)\n", "img_1 = Convolution2D(16, kernel_size=3, activation=activations.relu)(img_1)\n", "img_1 = MaxPooling2D(pool_size=(2, 2))(img_1)\n", "img_1 = Dropout(rate=0.2)(img_1)\n", "img_1 = Convolution2D(32, kernel_size=3, activation=activations.relu)(img_1)\n", "img_1 = MaxPooling2D(pool_size=(2, 2))(img_1)\n", "img_1 = Dropout(rate=0.2)(img_1)\n", "img_1 = Flatten()(img_1)\n", "\n", "dense_1 = BatchNormalization()(Dense(128, activation=activations.relu)(img_1))\n", "dense_1 = BatchNormalization()(Dense(128, activation=activations.relu)(dense_1))\n", "dense_1 = Dense(nclass, activation=activations.softmax)(dense_1)\n", "\n", "model = models.Model(inputs=inp, outputs=dense_1)\n", "opt = optimizers.Adam()\n", "\n", "model.compile(optimizer=opt, loss=losses.binary_crossentropy)\n", "model.summary()\n", "\n", "x_train, x_valid, y_train, y_valid = train_test_split(x_train, y_train, test_size=0.1, random_state=2017)\n", "model.fit(x_train, y_train, batch_size=16, validation_data=(x_valid, y_valid), epochs=3, shuffle=True, verbose=2)\n", "\n", "model.save(os.path.join(model_path, 'cnn.model'))"]}, {"metadata": {"_cell_guid": "060811f5-34ac-4fc2-92ba-c4276606c2a0", "_uuid": "2e3fa8d9706f47e69d0b74afcb68280f8a5de706"}, "cell_type": "markdown", "source": ["Test data is way too large to fit in RAM, we need to process them one by one.\n", "Generator test_data_generator will create batches of test wav files to feed into CNN."]}, {"metadata": {"_cell_guid": "7dfe0801-a636-4123-8367-ab2f19c97800", "collapsed": true, "_uuid": "646b6bcfbde7eae53cd8822b8838c575859e51ce"}, "outputs": [], "execution_count": null, "cell_type": "code", "source": ["def test_data_generator(batch=16):\n", "    fpaths = glob(os.path.join(test_data_path, '*wav'))\n", "    i = 0\n", "    for path in fpaths:\n", "        if i == 0:\n", "            imgs = []\n", "            fnames = []\n", "        i += 1\n", "        rate, samples = wavfile.read(path)\n", "        samples = pad_audio(samples)\n", "        resampled = signal.resample(samples, int(new_sample_rate / rate * samples.shape[0]))\n", "        _, _, specgram = log_specgram(resampled, sample_rate=new_sample_rate)\n", "        imgs.append(specgram)\n", "        fnames.append(path.split('\\\\')[-1])\n", "        if i == batch:\n", "            i = 0\n", "            imgs = np.array(imgs)\n", "            imgs = imgs.reshape(tuple(list(imgs.shape) + [1]))\n", "            yield fnames, imgs\n", "    if i < batch:\n", "        imgs = np.array(imgs)\n", "        imgs = imgs.reshape(tuple(list(imgs.shape) + [1]))\n", "        yield fnames, imgs\n", "    raise StopIteration()"]}, {"metadata": {"_cell_guid": "22992a27-deda-4a35-b34c-4aa87ad173ec", "_uuid": "c6d2516a6d5bd3c6a108d7e28565edaa65830958"}, "cell_type": "markdown", "source": ["We use the trained model to predict the test data's labels.\n", "However, since Kaggle doesn't provide test data, the following sections won't be executed here."]}, {"metadata": {"_cell_guid": "7fa8feb8-236e-46c5-8432-014f7e27484d", "collapsed": true, "_uuid": "56194039cac16f5d86a322e67641cbeafda9857d"}, "outputs": [], "execution_count": null, "cell_type": "code", "source": ["exit() #delete this\n", "del x_train, y_train\n", "gc.collect()\n", "\n", "index = []\n", "results = []\n", "for fnames, imgs in test_data_generator(batch=32):\n", "    predicts = model.predict(imgs)\n", "    predicts = np.argmax(predicts, axis=1)\n", "    predicts = [label_index[p] for p in predicts]\n", "    index.extend(fnames)\n", "    results.extend(predicts)\n", "\n", "df = pd.DataFrame(columns=['fname', 'label'])\n", "df['fname'] = index\n", "df['label'] = results\n", "df.to_csv(os.path.join(out_path, 'sub.csv'), index=False)"]}], "metadata": {"kernelspec": {"name": "python3", "language": "python", "display_name": "Python 3"}, "language_info": {"nbconvert_exporter": "python", "name": "python", "mimetype": "text/x-python", "codemirror_mode": {"name": "ipython", "version": 3}, "pygments_lexer": "ipython3", "file_extension": ".py", "version": "3.6.3"}}, "nbformat_minor": 1, "nbformat": 4}