{"cells":[{"metadata":{},"cell_type":"markdown","source":"# **Audio Tagging**"},{"metadata":{},"cell_type":"markdown","source":"# Table of content\n\n## **<div id=\"I\">I. Define the problem</div>**\n\n### **<div id=\"I0\">0. Sources</div>**\n\nhttps://medium.com/@ageitgey/machine-learning-is-fun-part-6-how-to-do-speech-recognition-with-deep-learning-28293c162f7a\nhttps://medium.com/nanonets/how-to-use-deep-learning-when-you-have-limited-data-part-2-data-augmentation-c26971dc8ced\nhttps://www.kaggle.com/daisukelab/cnn-2d-basic-solution-powered-by-fast-ai\n\n### **<div id=\"I1\">1. Problem description</div>**\n\nOne year ago, Freesound and Google’s Machine Perception hosted an audio tagging competition challenging Kagglers to build a general-purpose auto tagging system. This year they’re back and taking the challenge to the next level with multi-label audio tagging, doubled number of audio categories, and a noisier than ever training set.\n\n![](https://storage.googleapis.com/kaggle-media/competitions/freesound/task2_freesound_audio_tagging.png)\n\nHere's the background: Some sounds are distinct and instantly recognizable, like a baby’s laugh or the strum of a guitar. Other sounds are difficult to pinpoint. If you close your eyes, could you tell the difference between the sound of a chainsaw and the sound of a blender?\n\nBecause of the vastness of sounds we experience, no reliable automatic general-purpose audio tagging systems exist. A significant amount of manual effort goes into tasks like annotating sound collections and providing captions for non-speech events in audiovisual content.\n\nTo tackle this problem, Freesound (an initiative by MTG-UPF that maintains a collaborative database with over 400,000 Creative Commons Licensed sounds) and Google Research’s Machine Perception Team (creators of AudioSet, a large-scale dataset of manually annotated audio events with over 500 classes) have teamed up to develop the dataset for this new competition.\n\nTo win this competition, Kagglers will develop an algorithm to tag audio data automatically using a diverse vocabulary of 80 categories.\n\nIf successful, your systems could be used for several applications, ranging from automatic labelling of sound collections to the development of systems that automatically tag video content or recognize sound events happening in real time."},{"metadata":{},"cell_type":"markdown","source":"### **<div id=\"I2\">2. Tools importing</div>**\n\nHere we are importing every useful tool needed during our research process."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import time\nstart_time = time.time()\n\n# Data analysis and wrangling\nimport numpy as np\nimport pandas as pd\n\n# Visualization\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport IPython\nimport IPython.display\nimport librosa\nimport librosa.display\nimport random\nfrom tqdm import tqdm_notebook\nfrom fastai import *\nfrom fastai.vision import *\nfrom fastai.vision.data import *\nfrom fastai.imports import *\nfrom fastai.callback import *\nfrom fastai.callbacks import *\n\n# Machine learning\nfrom sklearn import preprocessing\nimport sklearn.metrics\nfrom sklearn.metrics import label_ranking_average_precision_score\n\n# File handling\nfrom pathlib import Path\nimport gc\nimport os\nprint(os.listdir(\"../input\"))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### **<div id=\"I3\">3. Label ranking average precision</div>**\n\nThe task consists of predicting the audio labels (tags) for every test clip. Some test clips bear one label while others bear several labels. The predictions are to be done at the clip level, i.e., no start/end timestamps for the sound events are required.\n\nThe primary competition metric will be label-weighted label-ranking average precision (lwlrap, pronounced \"Lol wrap\"). This measures the average precision of retrieving a ranked list of relevant labels for each test clip (i.e., the system ranks all the available labels, then the precisions of the ranked lists down to each true label are averaged). This is a generalization of the mean reciprocal rank measure (used in last year’s edition of the competition) for the case where there can be multiple true labels per test item. The novel \"label-weighted\" part means that the overall score is the average over all the labels in the test set, where each label receives equal weight (by contrast, plain lrap gives each test item equal weight, thereby discounting the contribution of individual labels when they appear on the same item as multiple other labels).\n\nWe use label weighting because it allows per-class values to be calculated, and still have the overall metric be expressed as simple average of the per-class metrics (weighted by each label's prior in the test set). For participant’s convenience, a Python implementation of lwlrap is provided in this public Google Colab."},{"metadata":{"trusted":true},"cell_type":"code","source":"# from official code https://colab.research.google.com/drive/1AgPdhSp7ttY18O3fEoHOQKlt_3HJDLi8#scrollTo=cRCaCIb9oguU\ndef _one_sample_positive_class_precisions(scores, truth):\n    \"\"\"Calculate precisions for each true class for a single sample.\n\n    Args:\n      scores: np.array of (num_classes,) giving the individual classifier scores.\n      truth: np.array of (num_classes,) bools indicating which classes are true.\n\n    Returns:\n      pos_class_indices: np.array of indices of the true classes for this sample.\n      pos_class_precisions: np.array of precisions corresponding to each of those\n        classes.\n    \"\"\"\n    num_classes = scores.shape[0]\n    pos_class_indices = np.flatnonzero(truth > 0)\n    # Only calculate precisions if there are some true classes.\n    if not len(pos_class_indices):\n        return pos_class_indices, np.zeros(0)\n    # Retrieval list of classes for this sample.\n    retrieved_classes = np.argsort(scores)[::-1]\n    # class_rankings[top_scoring_class_index] == 0 etc.\n    class_rankings = np.zeros(num_classes, dtype=np.int)\n    class_rankings[retrieved_classes] = range(num_classes)\n    # Which of these is a true label?\n    retrieved_class_true = np.zeros(num_classes, dtype=np.bool)\n    retrieved_class_true[class_rankings[pos_class_indices]] = True\n    # Num hits for every truncated retrieval list.\n    retrieved_cumulative_hits = np.cumsum(retrieved_class_true)\n    # Precision of retrieval list truncated at each hit, in order of pos_labels.\n    precision_at_hits = (\n            retrieved_cumulative_hits[class_rankings[pos_class_indices]] /\n            (1 + class_rankings[pos_class_indices].astype(np.float)))\n    return pos_class_indices, precision_at_hits\n\n\ndef calculate_per_class_lwlrap(truth, scores):\n    \"\"\"Calculate label-weighted label-ranking average precision.\n\n    Arguments:\n      truth: np.array of (num_samples, num_classes) giving boolean ground-truth\n        of presence of that class in that sample.\n      scores: np.array of (num_samples, num_classes) giving the classifier-under-\n        test's real-valued score for each class for each sample.\n\n    Returns:\n      per_class_lwlrap: np.array of (num_classes,) giving the lwlrap for each\n        class.\n      weight_per_class: np.array of (num_classes,) giving the prior of each\n        class within the truth labels.  Then the overall unbalanced lwlrap is\n        simply np.sum(per_class_lwlrap * weight_per_class)\n    \"\"\"\n    assert truth.shape == scores.shape\n    num_samples, num_classes = scores.shape\n    # Space to store a distinct precision value for each class on each sample.\n    # Only the classes that are true for each sample will be filled in.\n    precisions_for_samples_by_classes = np.zeros((num_samples, num_classes))\n    for sample_num in range(num_samples):\n        pos_class_indices, precision_at_hits = (\n            _one_sample_positive_class_precisions(scores[sample_num, :],\n                                                  truth[sample_num, :]))\n        precisions_for_samples_by_classes[sample_num, pos_class_indices] = (\n            precision_at_hits)\n    labels_per_class = np.sum(truth > 0, axis=0)\n    weight_per_class = labels_per_class / float(np.sum(labels_per_class))\n    # Form average of each column, i.e. all the precisions assigned to labels in\n    # a particular class.\n    per_class_lwlrap = (np.sum(precisions_for_samples_by_classes, axis=0) /\n                        np.maximum(1, labels_per_class))\n    # overall_lwlrap = simple average of all the actual per-class, per-sample precisions\n    #                = np.sum(precisions_for_samples_by_classes) / np.sum(precisions_for_samples_by_classes > 0)\n    #           also = weighted mean of per-class lwlraps, weighted by class label prior across samples\n    #                = np.sum(per_class_lwlrap * weight_per_class)\n    return per_class_lwlrap, weight_per_class\n\n\n# Wrapper for fast.ai library\ndef lwlrap(scores, truth, **kwargs):\n    score, weight = calculate_per_class_lwlrap(to_np(truth), to_np(scores))\n    return torch.Tensor([(score * weight).sum()])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### **<div id=\"I4\">4. Data</div>**\n\n#### **Train set**\n\nThe train set is meant to be for system development. The idea is to limit the supervision provided (i.e., the manually-labeled data), thus promoting approaches to deal with label noise. The train set is composed of two subsets as follows:\n\n**Curated subset**\n\nThe curated subset is a small set of manually-labeled data from FSD.\n* Number of clips/class: 75 except in a few cases (where there are less)\n* Total number of clips: 4970\n* Avge number of labels/clip: 1.2\n* Total duration: 10.5 hours\n\nThe duration of the audio clips ranges from 0.3 to 30s due to the diversity of the sound categories and the preferences of Freesound users when recording/uploading sounds. It can happen that a few of these audio clips present additional acoustic material beyond the provided ground truth label(s).\n\n**Noisy subset**\n\nThe noisy subset is a larger set of noisy web audio data from Flickr videos taken from the YFCC dataset.\n* Number of clips/class: 300\n* Total number of clips: 19815\n* Avge number of labels/clip: 1.2\n* Total duration: ~80 hours\n\nThe duration of the audio clips ranges from 1s to 15s, with the vast majority lasting 15s.\n\nConsidering the numbers above, per-class data distribution available for training is, for most of the classes, 300 clips from the noisy subset and 75 clips from the curated subset, which means 80% noisy - 20% curated at the clip level (not at the audio duration level, considering the variable-length clips).\n\n#### **Test set**\n\nThe test set is used for system evaluation and consists of manually-labeled data from FSD. Since most of the train data come from YFCC, some acoustic domain mismatch between the train and test set can be expected. All the acoustic material present in the test set is labeled, except human error, considering the vocabulary of 80 classes used in the competition.\n\nThe test set is split into two subsets, for the public and private leaderboards. In this competition, the submission is to be made through Kaggle Kernels. Only the test subset corresponding to the public leaderboard is provided (without ground truth).\n\nSubmissions must be made with inference models running in Kaggle Kernels. However, participants can decide to train also in the Kaggle Kernels or offline (see Kernels Requirements for details).\n\nThis is a kernels-only competition with two stages. The first stage comprehends the submission period until the deadline on June 10th. After the deadline, in the second stage, Kaggle will rerun your selected kernels on an unseen test set. The second-stage test set is approximately three times the size of the first. You should plan your kernel's memory, disk, and runtime footprint accordingly.\n\n#### **Files**\n\n* train_curated.csv - ground truth labels for the curated subset of the training audio files (see Data Fields below)\n* train_noisy.csv - ground truth labels for the noisy subset of the training audio files (see Data Fields below)\n* sample_submission.csv - a sample submission file in the correct format, including the correct sorting of the sound categories; it contains the list of audio files found in the test.zip folder (corresponding to the public leaderboard)\n* train_curated.zip - a folder containing the audio (.wav) training files of the curated subset\n* train_noisy.zip - a folder containing the audio (.wav) training files of the noisy subset\n* test.zip - a folder containing the audio (.wav) test files for the public leaderboard\n\n#### **Columns**\n\nEach row of the train_curated.csv and train_noisy.csv files contains the following information:\n\n* fname: the audio file name, eg, 0006ae4e.wav\n* labels: the audio classification label(s) (ground truth). Note that the number of labels per clip can be one, eg, Bark or more, eg, \"Walk_and_footsteps,Slam\".\n"},{"metadata":{},"cell_type":"markdown","source":"## **<div id=\"II\">II. Gather the data</div>**\n\nWe start by acquiring the training and testing datasets into Pandas DataFrames. We also combine these datasets to run certain operations on both datasets together."},{"metadata":{"trusted":true},"cell_type":"code","source":"training_curated_df = pd.read_csv(\"../input/train_curated.csv\")\ntraining_noisy_df = pd.read_csv(\"../input/train_noisy.csv\")\ntraining_df = [training_curated_df, training_noisy_df]\ntesting_df = pd.read_csv('../input/sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"Path('trn_curated').mkdir(exist_ok=True, parents=True)\nPath('trn_noisy').mkdir(exist_ok=True, parents=True)\nPath('test').mkdir(exist_ok=True, parents=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## **<div id=\"III\">III. Wrangle, cleanse and Prepare Data for Consumption</div>**"},{"metadata":{"trusted":true},"cell_type":"code","source":"# preview the data\ntraining_curated_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# preview the data\ntraining_noisy_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"training_curated_df.info()\nprint('_'*40)\ntraining_noisy_df.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels_curated = training_curated_df['labels'].unique()\nprint(labels_curated.shape)\nprint('_'*40)\nprint(labels_curated)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-output":false},"cell_type":"code","source":"labels_noisy = training_noisy_df['labels'].unique()\nprint(labels_noisy.shape)\nprint('_'*40)\nprint(labels_noisy)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"training_curated_df.describe()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## **<div id=\"IV\">IV. Turning sound into bits</div>**\n\n### **<div id=\"IV1\">1. Sampling sound data</div>**\n\nThe first step in speech recognition is obvious — we need to feed sound waves into a computer. But sound is transmitted as waves. How do we turn sound waves into numbers?\n\n![](https://cdn-images-1.medium.com/max/1200/1*6_q1VIVJuavYa-9Uby_L-A.png)\n\nSound waves are one-dimensional. At every moment in time, they have a single value based on the height of the wave. To turn this sound wave into numbers, we just record of the height of the wave at equally-spaced points.\n\nThis is called **sampling**. We are taking a reading thousands of times a second and recording a number representing the height of the sound wave at that point in time. That’s basically all an uncompressed .wav audio file is.\n\n“CD Quality” audio is sampled at 44.1khz (44,100 readings per second). For our problem, here is how we will proceed:\n* Handle sampling rate 44.1kHz as is, no information loss.\n* Size of each file will be 128 x L, L is audio seconds x 128; [128, 256] if sound is 2s long.\n\n### **<div id=\"IV2\">2. Pre-processing our sampled Sound Data</div>**\n\nWe now have an array of numbers with each number representing the sound wave’s amplitude at 1/44100th of a second intervals.\n\nWe could feed these numbers right into a neural network. But trying to recognize speech patterns by processing these samples directly is difficult. Instead, we can make the problem easier by doing some pre-processing on the audio data.\n\nTo make this data easier for a neural network to process, we are going to break apart this complex sound wave into it’s component parts. We’ll break out the low-pitched parts, the next-lowest-pitched-parts, and so on. Then by adding up how much energy is in each of those frequency bands (from low to high), we create a fingerprint of sorts for this audio snippet.\n\nWe do this using a mathematic operation called a Fourier transform. It breaks apart the complex sound wave into the simple sound waves that make it up. Once we have those individual sound waves, we add up how much energy is contained in each one.\n\nThe end result is a score of how important each frequency range is, from low pitch (i.e. bass notes) to high pitch.\n\n![](https://cdn-images-1.medium.com/max/1200/1*A4CxgdyqYd_nrF3e-7ETWA.png)\n\nIf we repeat this process on every 20 millisecond chunk of audio, we end up with a spectrogram. This is what we have done below on one of our audio sound."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"#EasyDict allows to access dict values as attributes (works recursively). A Javascript-like properties dot notation for python dicts.\n#It is mandatory in order to use the library below\n# Special thanks to https://github.com/makinacorpus/easydict/blob/master/easydict/__init__.py\nclass EasyDict(dict):\n\n    def __init__(self, d=None, **kwargs):\n        if d is None:\n            d = {}\n        if kwargs:\n            d.update(**kwargs)\n        for k, v in d.items():\n            setattr(self, k, v)\n        # Class attributes\n        for k in self.__class__.__dict__.keys():\n            if not (k.startswith('__') and k.endswith('__')) and not k in ('update', 'pop'):\n                setattr(self, k, getattr(self, k))\n\n    def __setattr__(self, name, value):\n        if isinstance(value, (list, tuple)):\n            value = [self.__class__(x)\n                     if isinstance(x, dict) else x for x in value]\n        elif isinstance(value, dict) and not isinstance(value, self.__class__):\n            value = self.__class__(value)\n        super(EasyDict, self).__setattr__(name, value)\n        super(EasyDict, self).__setitem__(name, value)\n\n    __setitem__ = __setattr__\n\n    def update(self, e=None, **f):\n        d = e or dict()\n        d.update(f)\n        for k in d:\n            setattr(self, k, d[k])\n\n    def pop(self, k, d=None):\n        delattr(self, k)\n        return super(EasyDict, self).pop(k, d)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Thanks to https://github.com/daisukelab/ml-sound-classifier\ndef read_audio(conf, pathname, trim_long_data):\n    y, sr = librosa.load(pathname, sr=conf.sampling_rate) #Loads an audio file as a floating point time series. This functions samples the sound\n    # trim silence\n    if 0 < len(y): # workaround: 0 length causes error\n        y, _ = librosa.effects.trim(y) # trim, top_db=default(60)\n    # make it unified length to conf.samples\n    if len(y) > conf.samples: # long enough\n        if trim_long_data:\n            y = y[0:0+conf.samples]\n    else: # pad blank\n        padding = conf.samples - len(y)    # add padding at both ends\n        offset = padding // 2\n        y = np.pad(y, (offset, conf.samples - len(y) - offset), 'constant')\n    return y\n\ndef audio_to_melspectrogram(conf, audio):\n    spectrogram = librosa.feature.melspectrogram(audio, \n                                                 sr=conf.sampling_rate,\n                                                 n_mels=conf.n_mels,\n                                                 hop_length=conf.hop_length,\n                                                 n_fft=conf.n_fft,\n                                                 fmin=conf.fmin,\n                                                 fmax=conf.fmax)\n    spectrogram = librosa.power_to_db(spectrogram)\n    spectrogram = spectrogram.astype(np.float32) #Returns an 128 x L array corresponding to the spectrogram of the sound (L = 128*n° of s)\n    return spectrogram\n\ndef melspectrogram_to_delta(mels):\n    return librosa.feature.delta(mels)\n\ndef show_melspectrogram(conf, mels, title='Log-frequency power spectrogram'):\n    librosa.display.specshow(mels, x_axis='time', y_axis='mel', \n                             sr=conf.sampling_rate, hop_length=conf.hop_length,\n                            fmin=conf.fmin, fmax=conf.fmax)\n    plt.colorbar(format='%+2.0f dB')\n    plt.title(title)\n    plt.show()\n\ndef read_as_melspectrogram(conf, pathname, trim_long_data, debug_display=False):\n    x = read_audio(conf, pathname, trim_long_data)\n    mels = audio_to_melspectrogram(conf, x)\n    if debug_display:\n        delta = melspectrogram_to_delta(mels)\n        delta_squared = melspectrogram_to_delta(delta)\n        IPython.display.display(IPython.display.Audio(x, rate=conf.sampling_rate))\n        show_melspectrogram(conf, mels)\n        show_melspectrogram(conf, delta)\n        show_melspectrogram(conf, delta_squared)\n    return mels\n\nconf = EasyDict()\nconf.sampling_rate = 44100\nconf.duration = 2\nconf.hop_length = 347 * conf.duration # to make time steps 128\nconf.fmin = 20\nconf.fmax = conf.sampling_rate // 2\nconf.n_mels = 128\nconf.n_fft = conf.n_mels * 20\nconf.samples = conf.sampling_rate * conf.duration","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# example\npath = '../input/train_curated/0006ae4e.wav'\nx = read_audio(conf, path, trim_long_data=False)\nprint(x)\nprint('_'*40)\nprint(audio_to_melspectrogram(conf, x))\nprint(audio_to_melspectrogram(conf, x).shape)\nx1 = read_as_melspectrogram(conf, path, trim_long_data=False, debug_display=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* The first array corresponds to our **sampled audio file**. It is an array, where every value corresponds to the sound's wave amplitude (Hz) every 1/44100th of a second.\n* The second array corresponds to the **spectrogram of our audio file**. It is an list of array.\n    * First, we cut our sampled audio file into 20ms pieces (each piece contains 44100/50 = 882 values.\n    * Then, we perform a Fourier transform for each 20ms piece. It breaks apart the complex sound wave into the simple sound waves that make it up.\n    * Once we have those individual sound waves, we add up how much energy is contained in each one. This creates the first list of our array\n    * Last, we do the same thing for each 20ms piece, in order to create our spectrogram.\n* The last item is a **representation of our spectrogram**. Each list in our previous array is displayed as a colored vertical bar. The more bright the color is, the more the frequency is represented for this audio snippet.\n"},{"metadata":{},"cell_type":"markdown","source":"### **<div id=\"IV3\">3. Transforming sound into images</div>**\n\n\nNow that we can convert our sounds into spectrograms, we want to be able to utilize them. The next step is to convert our spectrograms into images. There is a very powerful model for image recognition which is called CNN (for Convolutional Neural Network), and this model gives also ver good results for audio recognition; but before using it, we need to actually transform our sounds into images, using these spectrograms."},{"metadata":{"trusted":true},"cell_type":"code","source":"\"\"\"\nThe mono_to_color function takes as an input the spectrogram of our sound (list of array, see above). \nIt stacks it three times, so that it has the same shape as a classic RGB image.\nThen it standardize the array (take a matrix and change it so that its mean is equal to 0 and variance is 1). This improves performance.\nThen it normalizes each value between 0 and 255 (gray scale). \n\"\"\"\n\ndef mels_preprocessing(X1, X2, X3, mean=None, std=None, norm_max=None, norm_min=None, eps=1e-6):\n    # Stack X as [X,X,X]\n    X = np.stack([X1, X2, X3], axis=-1)\n\n    # Standardize\n    mean = mean or X.mean()\n    std = std or X.std()\n    #Standardization. Xstd has 0 mean and 1 variance\n    Xstd = (X - mean) / (std + eps)\n    _min, _max = Xstd.min(), Xstd.max()\n    norm_max = norm_max or _max\n    norm_min = norm_min or _min\n    if (_max - _min) > eps:\n        # Scale to [0, 255]\n        V = Xstd\n        V[V < norm_min] = norm_min\n        V[V > norm_max] = norm_max\n        V = 255 * (V - norm_min) / (norm_max - norm_min)\n        V = V.astype(np.uint8)\n    else:\n        # Just zero\n        V = np.zeros_like(Xstd, dtype=np.uint8)\n    return V\n\ndef convert_wav_to_image(df, source, img_dest):\n    X = []\n    for i, row in tqdm_notebook(df.iterrows()):\n        x1 = read_as_melspectrogram(conf, source/str(row.fname), trim_long_data=False)\n        x2 = melspectrogram_to_delta(x1)\n        x3 = melspectrogram_to_delta(x2)\n        x_preprocessed = mels_preprocessing(x1, x2, x3)\n        X.append(x_preprocessed)\n    return df, X","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"training_curated_df, X_train_curated = convert_wav_to_image(training_curated_df, source=Path('../input/train_curated'), img_dest=Path('trn_curated'))\ntesting_df, X_test = convert_wav_to_image(testing_df, source=Path('../input/test'), img_dest=Path('test'))\n\nprint(f\"Finished data conversion at {(time.time()-start_time)/3600} hours\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for i in range(0,6):\n    a = np.asarray(X_train_curated[i:i+1])\n    a = np.squeeze(a)\n    print(a.shape)\n    plt.imshow(a)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The array above represents our spectrogram as an RGB image.\n* Each 1x3 list represents **one pixel of our image**. Each value in this list represents respectively the Red, Green and Blue value (between 0 and 255) of the pixel (see image below). In our case, as our \"color\" is \"one dimensional\" (one pixel in our spectrogram is juste a real number representing the db value for a particular frequency at a particular time), our pixel color will be a shade of grey. Thus, the three RGB values will always be the same.\n* Each list of 1x3 list represents **one horizontal line of our image**.\n* **The whole array is our sound**, represented as a **gray-scaled image**. \n\n![image.png](attachment:image.png)","attachments":{"image.png":{"image/png":"iVBORw0KGgoAAAANSUhEUgAAAeAAAAHnCAYAAABg0Y4tAAAgAElEQVR4Ae3dX4gr733f8Y/SXxITYjekFyUhTW1HOheHTQg0F0EiOJQfJdJJyraEvSt7Y6Te7Tbl9OpcnpuwELSUUlYthCV3Sx22pkcypSQ25YgGEjBlWeiRsEns2OSiof7T4qaGKc+MpJ2ZfWY0M5pHMxq9BT/v6JmZ7/d5Xo98vvvMSKuW53meeCCAAAIIIIDAXgV+ZK/ZSIYAAggggAACvgAFmBcCAggggAACFQhQgCtAJyUCCCCAAAIUYF4DCCCAAAIIVCBAAa4AnZQIIIAAAghQgHkNIIAAAgggUIEABbgCdFIigAACCCBAAeY1gAACCCCAQAUCFOAK0EmJAAIIIIAABZjXAAIIIIAAAhUIUIArQCclAggggAACFGBeAwgggAACCFQgQAGuAJ2UCCCAAAIIUIB5DSCAAAIIIFCBAAW4AnRSIoAAAgggQAHmNYAAAggggEAFAhTgCtBJiQACCCCAAAWY1wACCCCAAAIVCFCAK0AnJQIIIIAAAhRgXgMIIIAAAghUIEABrgCdlAgggAACCFCAeQ0ggAACCCBQgQAFuAJ0UiKAAAIIIEAB5jWAAAIIIIBABQIU4ArQt6ecadRqqRX6rze61tKcuLxWr9XSaLY9ivWIbeen7V/ts/YrlGw2Mn0fKdrF1ZjSOp6WOxSfTQQQQKAJAhTgGs9id7yQ53nyFmNpcqnz66XUvtB7z9NNv/yOL2fX6nUuNd8Sejj1gn5Nh5pPLnUVrbTqnw4lTXQfbp/dayJpeOqg41v6y24EEECgjgIU4DrOSrxP7Rc6kTR/XERWwMvrnr9K9heVs5G/3TNF2jxmI3+lHKxWexrNVu3x2Ovny2udD7YX3/XhTz+7etl5euZv9U/ll+BQBZ7d++VX6/o7GwV9N/3brO7DYWKr4Wer6rzjC8dmGwEEEKiBAAW4BpOwtQvLD3qQ1I1VuvbFrcZdafK2p95gInXHur1oB0V6MNHJNFhBL6YnmgzOta7NSflOhlMtpqZ0pj8mg9XlcT/nmV6148f3FSyC74PL0MtrvQ2Wv/LXv7ORBg9nWpjVvb+KvtO7Lb8fRDKY4lxgfJEYPEEAAQQqFqAAVzwBaennl53gPnDnUhpOg+IaOaGti9uxuvO55upqfHshUwuX7+78y8iTQXB+xxRKzWUW0ImP9oVubvr++YnHrHZsLkEvTO5LdSz3ddeXod9eL1f96Wr8enX5uX+jxRvparT6xWFb32IdKjS+WAyeIoAAAlULUICrnoGU/Jt7wJ6n90nFcfG4umcbL7BdjRere7VmpenivnH7lc66CQNYXYae313p6m4udZ9WyubSeWdwJ53e6jbDitueYQ/jsyemFQEEEChFgAJcCmNVQWYa+ZeBhxqaS9GD4J3H7Vdn6mquu9W7o4J7xfF3JZfQ5+U7+bU1dmk8iLy6DD2faOLX31eb1fXi0bzN60Sn/bb0wVxcT348fDDXppeRw/Y2vuRusQcBBBDYWYACvDNhdQFmo4Em/qXnG92YS9GaaGAuB7cvgpXlZOBfwjZXsMeLm+D+awnd3dwDNu+YXt93tsQNLkObHV2dhW4U91+v+tpq6fzxRGYRHRTaUJD2hd4Mu/Ivw/eudBfa5Xp84VRsI4AAAq4EWp65NskDAQQQQAABBPYqwAp4r9wkQwABBBBAIBCgAPNKQAABBBBAoAIBCnAF6KREAAEEEECAAsxroHSBX+//WekxCYgAAgg0TYAC3LQZrcF4vjz7B6II12Ai6AICCNRagAJc6+mhcwgggAACTRWgADd1ZiseF6vgiieA9AggUHsBCnDtp4gOIoAAAgg0UYAC3MRZrcmYWAXXZCLoBgII1FKAAlzLaaFTCCCAAAJNF6AAN32GKx4fq+CKJ4D0CCBQWwEKcG2nho4hgAACCDRZgALc5NmtydhYBddkIugGAgjUSoACXKvpoDMIIIAAAsciQAE+lpmueJysgiueANIjgEDtBCjAtZsSOoQAAgggcAwCFOBjmOWajJFVcE0mgm4ggEAtBCjAtZgGOoEAAgggcGwCFOBjm/GKx8squOIJID0CCNRGgAJcm6mgIwgggAACxyRAAT6m2a7JWFkF12Qi6AYCCFQqQAGulP94k1OEj3fuGTkCCAQCFGBeCQgggAACCFQgQAHOhD7TqNVSK/Rfb3StpTl3ea1eq6XRLFOg5wdtOz9l/2xk+jRSNPWqr7EO5Tk20smU/JHjCjxhFVwAjVMQQKAxAhTgHFPZHS/keZ68xViaXOr8eim1L/Te83TTzxEo46HL2bV6nUvNE47vnw4lTXQfrsCze00kDU+jHcpzbEI6mhFAAAEEShSgABfBbL/QiaT54yKyAl5e9/xVsr/4nI387Z4p0uYxG/kr5WAV3dNotmpPyr+81vkgufj6p/VP5ZfgUAWe3fvlV7H6K2U4djYK+m/6uFnhh/sXWw0/W1XnHaMkVsFhYLYRQOCYBCjARWZ7+UEPkrovO5Gz2xe3GnelydueeoOJ1B3r9qIdFOnBRCfTYAW9mJ5oMjjXujZHgoSenAynWkxNiU169BUsgu+Dy9DLa70Nlr+Krn/N+VuOnY00eDjTwqzwp0PNJ3d6t+V3hEivTHEuMMZIDJ4ggAACRyRAAc4x2fPLTnAfuHMpDadBcY2c39bF7Vjd+VxzdTW+vVDb3CZ+d+dfRp4MgvM7pjhrLrOATny0L3Rz0/fPTzzGlNXVZei318tVnq7Gr5+XXxMj9dj+jRZvpKvR6peHbf2LdarQGFcxWAXHMHmKAAJHIUABzjHNm3vAnqf3ScVx8bi6ZxsvsF2NF15wD9msMsu6b7y6tDy/u9LV3VzqnumVqfq2R8qx5vJ5Z3Annd7qNnXVbQu8bnMzxl/v/9k6AT8RQACBxghQgEudyplG/qXnoYbmUvQgeIdy+9WZuprr7ip4t1Rwrzj+7uWiHVldWp5PNPHr76uUVXPysYtH81avE53229IHc4E9+fHwwVybXkYO23WMrIKTvdmDAALNFKAAlzivs9FAE//S841uzKVoTTQw78hqXwSrysnAv4RtrmCPFzeW+7TFOhNcWjbndnWWuPwNYicd23+96m+rpfPHE3UlBYU21Kf2hd4Mu/IvxfeudBfa5XqM4VRsI4AAAk0QaHnmWigPBGoiYC43m9Xw+hF/vm7nJwIIIHDoAqyAD30GJXGPtAGTyBAQQODoBCjADZjyJt0/bdJYGvDSYggIIOBQgALsEJfQCCCAAAIIJAlwDzhJ5gDbD+V+afiSefh+b5h8fUzS/vCxbCOAAAKHKPDRIXaaPh+2QLiorgtteETh/eF2thFAAIEmCbACbtJsrt6QdegFLFyUD30sDXt5MRwEEChRgBVwiZiEKkeAoluOI1EQQKDeArwJq97zk7t3pniFV5C5A3ACAggggMBeBCjAe2EmCQIIIIAAAlEBCnDUoxHPWAU3YhoZBAIINFyAAtzwCWZ4CCCAAAL1FOBd0PWcl1J6Ze4FN+UNTa1WqxQTgiCAAAJlCxT9SgVWwGXPBPEQQAABBBDIIEABzoB0qIdwL/hQZ45+I4DAMQjwOeBjmOWGjTHv5Z4il+L3dU7DpobhIHA0AmXcFmMF3PCXC6vghk8ww0MAgYMVYAV8sFNHx/Wd/63eF7+veYTiRzX85U/p5hd5aUdYeIIAArUTYAVcuykpv0NNXwV3f/nvyPtnf9f/b/ppafLV7+r6O+U7EhEBBBAoU4ACXKZmjWM1vQiv6ft//29J+n96pACvSfiJAAI1FaAA13Ri6FYxgdmf/0DSJ3T688XO5ywEEEBgXwLcKNuXdA3yrFfB5meTHvOv/k+1vvo0ou6nf1ydp6dsIYAAArUUoADXclroVB4Bcw/4/epNV8u/+I46X/mOzn/qo01bnlgciwACCOxLgEvQ+5KuSZ71Krgm3Sm9G+2f/3ENJc3/1w9Lj01ABBBAoEyBmhbgmUatlswHndf/9UbXWpqRL6/Va7U0mhVk2HZ+2v7VvnWfzM+epSOzken3SNEursZkOX4zkrTcm4PYSBNY/sX/1URS96e4uJPmxD4EEKheoKYFOIDpjhcyf/XIW4ylyaXOr5dS+0LvPU83/fLxlrNr9TqXsc+VPs8znHpBv6ZDzSeDZ78M9E/NGmyi+3AFnt37hWF46qDjz7uY2tK0VbB/D/gP/kqtP/grdb7yA3U//bd1y+eAU18D7EQAgeoFDmOZ0H6hE1PSHhfS8p1fJE+mnl5/6KlzOZcpiDcaqTWYyBTt9xdtaTZSbzBZFdOuhtNb3fTbyeLLa50PthffSIDOS3UlPXxYSuHY/VMNNdHkfqabflBwZ/dmXTbUuv7ORj0NJsGfkOgOx7q9uVCkd2Y13LmUGaf5ZcOsqgeToabejfyIeccX6Xj6kyJ/hjE9Yrl7Tf/Wjx+V9Ln1k83Pb+rzm+1gI3xObFfi032dk9gBdiCAQG0FPvcbf6qvfOlXdurfYRTg5Qc9mMuKL817Wz9sBty+uNX4rqPLtz09zOdSd6xbU3xN8RpMdDJd6H2/reVspM7gXC8X73WxOfv5xslwqtvTe3UGplhmeCwe/QI/fBEpnZL6MovgyeRes5u++strvfXr7+mmeA4ezrTw3qs9M7843Ond64vUvkV6kza+eFciJz49Wa+CD/Ed0YfY5yd5thBAoAkC5hbkro9aF+D5ZUety2CIXVMc/eIaHnJbF7dj3fmXjbsa3waryOW7O78wzgcd/7Lv+gyzgE78fEr7Qjc3kmb368MTf04GrVXcrobjhfVyuH8ZejLR2+vX6sj0p6vx69Xl5/6NFrrW1ainh9UqOLVvsZ6kji9jAY6F5CkCCCCAwJ4FDuMesOfp/U0/eol2DbVahUpz+UVs3W4K3mJ1r9bcRy7xvrF/D3gxVldzf2W+SRne8C9DS/O7K13dmdX5mV6tiuPyuqfO4E46vdXt1NwvLvLYfXzrVXA4e90vP4f7yjYCCCBwyAK1LsDbYWcamcvF3aGGXWkyCN553H515hfHu6vgXVCm4D1/V/L26KlHtC90O+5qfnku896w54/gMrTmE5lFbvfs1eYXiMWjufd7olNz3/iDubie/PDvL2sZOWwv40vuEnsQQAABBEoQOOgCPBsNNDEr3dsb3dyaFelEA/MxH1MczcpyMvA/xtS5lMaL1ZuXSkBbh2hfvNFQc12uCv26ff0zeDe0edbV2Xr5a+4Qv171tdXS+ePJ0xu51iean+0LvRmaAt9Rq3elu9i+ssZnWwWHU7GNAAIIIOBGoOXl/XZzN/0gaoUC4cvO4e0Ku/QsdfgND7xkn/HQgAACexYo498kCvCeJ62KdOGP0yS9g3h9TNL+KvodzlnGiz0cj20EEEBgF4Ey/k2q9bugd8Hh3CeBcFFdF9qnvVJ4f7idbQQQQAABdwIUYHe2tYxsK7a2olzLztMpBBBAoEECXIIuOJmmaNmKWcFwnLZFoIzLPVtSsBsBBBDILFDGv0kH/S7ozFIODjTFl5WjA1hCIoAAAkciQAE+kolmmAgggAAC9RKgAO8wH6yCd8DjVAQQQODIBXgT1pG/ANKGH77HkXbcvvfVtV/7diAfAscg0OTP/bMC3vEVzCp4R0BORwABBI5UgAJ8pBPPsBFAAAEEqhXgY0gl+TfxY0nhS715LwPl9ch7vJm2IueUNN2EQQABhwK7/NvjsFuR0GX0kRVwhJQnCCCAAAII7EeAAlyS87HcC/53X/4t/cIXf09/FHH7E/2rL/6WfuHL/0Ffj7TzBAEEEEAgSYACnCRDu1Xg45/9rKQ/1pe+Hdr97ff6gvmG45/9VX0m1MwmAggggECyAAU42Sb3nmNYBX/mZ35NJ5K+8K0/2fh8/Xt/Lumz+s2f+blNGxsIIIAAAukCFOB0H/bGBT75q/rNT0n6y/ery9Df1H/51tekT/2aPv5k/GCeI4AAAggkCVCAk2QKtjd/FfxzilyG/t5/03/6LpefC75cOA0BBI5YgAJ8xJNfdOjry9D/43vf1Ne//V/1wOXnopSchwACRyxAAXYw+Y1fBa8uQz98607/lsvPDl5BhEQAgWMQoAAfwyyXPsbVZejv/rG+wOXn0nUJiAACxyFAAXY0z01fBa8vQ/PuZ0cvIMIigEDjBfg2JIdTvC7C5mfjHp/8bf3Hf/zbjRsWA0IAAQT2JUAB3pf0gecxf3c57yPvOXmPN/0pck7ecXA8Aggg4EKAL2NwoRqLaYrEIa6Cy/hj4zEKniKAAAJbBQ7h354y+sg94K0vBQ5AAAEEEECgfAEKcPmmzyKu7wU/20EDAggggMDRClCAj3bqGTgCCCCAQJUCFOA96bMK3hM0aRBAAIEDEaAA12SieDdvTSaCbiCAAAJ7EqAA7wnapGEVvEdsUiGAAAI1F6AA13yC6B4CCCCAQDMFKMB7nlfbKvhQPye8ZzrSIYAAAo0SoAA3ajoZDAIIIIDAoQhQgCuYKdsquIJukBIBBBBAoEIBCnCF+KRGAAEEEDheAf4WtKO5D3+syKx4bY/1MUn7befss62Mv3W6z/6SCwEEmiFwCP/2lNFHvg3J0es1XFTXhTacKrw/3M42AggggMBxCFCA9zDPtmJrK8p76AopEEAAAQRqIsAl6JpMRB27UcYlljqOiz4hgEC9BQ7h354y+sibsOr9OqR3CCCAAAINFaAAN3RiGRYCCCCAQL0FKMD1nh96hwACCCDQUAEKcEMnlmEhgAACCNRbgAJc7/mhdwgggAACDRXgY0g1mtjwu+pq1C2/K3XuW92s6A8Chyzged4hd/+g+s4K+KCmi84igAACCDRFgALclJlkHAgggAACByXAJeiaTtcul4HMX9my/fWtLEPd5VwTf9fzs/SRYxBAoFwBbjGV65k1GivgrFIchwACCCCAQIkCFOASMQmFAAIIIIBAVgEKcFYpjkMAAQQQQKBEAQpwiZjOQn3t3+u7f6+jv47890/0/ZuvO0tJYAQQQAABtwIUYLe+pUb/6M1/1k9/Y+H/95Nn0t+8/R394GulpiAYAggggMCeBCjAe4IuO82P9V9IetAPKcBl0xIPAQQQ2IsABXgvzOUn+ZvZH0o60UefLT82ERFAAAEE3AvwOWD3xqVl+OHbf6S/frsK90sn+onf/z19ggJcmi+BEEAAgX0KsALep/aOuYJ7wDf6MRPnv7/Qj3z8mR0jcjoCCCCAQFUCFOCq5Avn/Yf6yd//p5L+UN//l39UOAonIoAAAghUK0ABrta/WPaP/7l+4pck3f1r3gVdTJCzEEAAgcoFKMCVT0GRDnxGn/gXZhX8oP/zb1gFFxHkHAQQQKBqAd6EVfUMZMn/2c/rU9/4fPTIj39XP/2N34228QwBBBBA4GAEWAEfzFTRUQQQQACBJgm0vF2+965JEjUYS/grwT73G39agx7RBQQQOAaBr3zpVzbDrENJCP9bWIf+bHBCG2X0kQIcAq16s4wJrXoM5EcAgcMTqNu/PXXrj21Gy+gjl6BtsrQhgAACCCDgWIAC7BiY8AgggAACCNgEKMA2FdoQQAABBBBwLEABdgxMeAQQQAABBGwCFGCbCm0IIIAAAgg4FqAAOwYmPAIIIIAAAjYBCrBNhTYEEEAAAQQcC1CAHQMTHgEEEEAAAZsABdimQhsCCCCAAAKOBSjAjoEJjwACCCCAgE2AAmxToQ0BBBBAAAHHAhRgx8CERwABBBBAwCZAAbap0IYAAggggIBjAQqwY2DCI4AAAgggYBOgANtUaEMAAQQQQMCxAAXYMTDhEUAAAQQQsAlQgG0qtCGAAAIIIOBYgALsGJjwCCCAAAII2AQowDYV2hBAAAEEEHAsQAF2DEx4BBBAAAEEbAIUYJsKbQgggAACCDgWoAA7BiY8AggggAACNgEKsE2FNgQQQAABBBwLUIAdAxMeAQQQQAABmwAF2KZCGwIIIIAAAo4FKMCOgQmPAAIIIICATYACbFOhDQEEEEAAAccCFGDHwIRHAAEEEEDAJkABtqnQhgACCCCAgGMBCrBjYMIjgAACCCBgE6AA21RoQwABBBBAwLEABdgxMOERQAABBBCwCVCAbSq0IYAAAggg4FiAAuwYmPAIIIAAAgjYBCjANhXaEEAAAQQQcCxAAXYMTHgEEEAAAQRsAhRgmwptCCCAAAIIOBagADsGJjwCCCCAAAI2AQqwTYU2BBBAAAEEHAtQgB0DEx4BBBBAAAGbAAXYpkIbAggggAACjgUowI6BCY8AAggggIBNgAJsU6ENAQQQQAABxwIUYMfAhEcAAQQQQMAmQAG2qdCGAAIIIICAYwEKsGNgwiOAAAIIIGAToADbVGhDAAEEEEDAsQAF2DEw4RFAAAEEELAJUIBtKrQhgAACCCDgWIAC7BiY8AgggAACCNgEKMA2FdoQQAABBBBwLEABdgxMeAQQQAABBGwCFGCbCm0IIIAAAgg4FqAAOwYmPAIIIIAAAjYBCrBNhTYEEEAAAQQcC1CAHQMTHgEEEEAAAZsABdimQhsCCCCAAAKOBSjAjoEJjwACCCCAgE2AAmxToQ0BBBBAAAHHAhRgx8CERwABBBBAwCZAAbap0IYAAggggIBjAQqwY2DCI4AAAgggYBOgANtUaEMAAQQQQMCxAAXYMTDhEUAAAQQQsAlQgG0qtCGAAAIIIOBYgALsGJjwCCCAAAII2AQowDYV2hBAAAEEEHAsQAF2DEx4BBBAAAEEbAIUYJsKbQgggAACCDgWoAA7BiY8AggggAACNgEKsE2FNgQQQAABBBwLUIAdAxMeAQQQQAABmwAF2KZCGwIIIIAAAo4FKMCOgQmPAAIIIICATYACbFOhDQEEEEAAAccCFGDHwIRHAAEEEEDAJkABtqnQhgACCCCAgGMBCrBjYMIjgAACCCBgE6AA21RoQwABBBBAwLEABdgxMOERQAABBBCwCVCAbSq0IYAAAggg4FiAAuwYmPAIIIAAAgjYBCjANhXaEEAAAQQQcCxAAXYMTHgEEEAAAQRsAhRgmwptCCCAAAIIOBagADsGJjwCCCCAAAI2AQqwTYU2BBBAAAEEHAtQgB0DEx4BBBBAAAGbAAXYpkIbAggggAACjgUowI6BCY8AAggggIBNgAJsU6ENAQQQQAABxwIUYMfAhEcAAQQQQMAmQAG2qdCGAAIIIICAYwEKsGNgwiOAAAIIIGAToADbVGhDAAEEEEDAsQAF2DEw4RFAAAEEELAJUIBtKrQhgAACCCDgWIAC7BiY8AgggAACCNgEKMA2FdoQQAABBBBwLEABdgxMeAQQQAABBGwCFGCbCm0IIIAAAgg4FqAAOwYmPAIIIIAAAjYBCrBNhTYEEEAAAQQcC1CAHQMTHgEEEEAAAZsABdimQhsCCCCAAAKOBSjAjoEJjwACCCCAgE2AAmxToQ0BBBBAAAHHAhRgx8CERwABBBBAwCZAAbap0IYAAggggIBjAQqwY2DCI4AAAgggYBOgANtUaEMAAQQQQMCxAAXYMTDhEUAAAQQQsAlQgG0qtCGAAAIIIOBYgALsGJjwCCCAAAII2AQowDYV2hBAAAEEEHAsQAF2DEx4BBBAAAEEbAIUYJsKbQgggAACCDgWoAA7BiY8AggggAACNgEKsE2FNgQQQAABBBwLUIAdAxMeAQQQQAABmwAF2KZCGwIIIIAAAo4FKMCOgQmPAAIIIICATYACbFOhDQEEEEAAAccCFGDHwIRHAAEEEEDAJkABtqnQhgACCCCAgGMBCrBjYMIjgAACCCBgE6AA21RoQwABBBBAwLEABdgxMOERQAABBBCwCVCAbSq0IYAAAggg4FiAAuwYmPAIIIAAAgjYBCjANhXaEEAAAQQQcCxAAXYMTHgEEEAAAQRsAhRgmwptCCCAAAIIOBagADsGJjwCCCCAAAI2AQqwTYU2BBBAAAEEHAtQgB0DEx4BBBBAAAGbAAXYpkIbAggggAACjgUowI6BCY8AAggggIBNgAJsU6ENAQQQQAABxwIUYMfAhEcAAQQQQMAmQAG2qdCGAAIIIICAYwEKsGNgwiOAAAIIIGAToADbVGhDAAEEEEDAsQAF2DEw4RFAAAEEELAJUIBtKrQhgAACCCDgWIAC7BiY8AgggAACCNgEKMA2FdoQQAABBBBwLEABdgxMeAQQQAABBGwCFGCbCm0IIIAAAgg4FqAAOwYmPAIIIIAAAjYBCrBNhTYEEEAAAQQcC1CAHQMTHgEEEEAAAZsABdimQhsCCCCAAAKOBSjAjoEJjwACCCCAgE2AAmxToQ0BBBBAAAHHAhRgx8CERwABBBBAwCZAAbap0IYAAggggIBjAQqwY2DCI4AAAgggYBOgANtUaEMAAQQQQMCxAAXYMTDhEUAAAQQQsAlQgG0qtCGAAAIIIOBYgALsGJjwCCCAAAII2AQowDYV2hBAAAEEEHAsQAF2DEx4BBBAAAEEbAIUYJsKbQgggAACCDgWoAA7BiY8AggggAACNgEKsE2FNgQQQAABBBwLUIAdAxMeAQQQQAABmwAF2KZCGwIIIIAAAo4FKMCOgQmPAAIIIICATYACbFOhDQEEEEAAAccCFGDHwIRHAAEEEEDAJkABtqnQhgACCCCAgGMBCrBjYMIjgAACCCBgE6AA21RoQwABBBBAwLEABdgxMOERQAABBBCwCVCAbSq0IYAAAggg4FiAAuwYmPAIIIAAAgjYBCjANhXaEEAAAQQQcCxAAXYMTHgEEEAAAQRsAhRgmwptCCCAAAIIOBagADsGJjwCCCCAAAI2AQqwTYU2BBBAAAEEHAtQgB0DEx4BBBBAAAGbAAXYpkIbAggggAACjgU+chyf8AgggAACCBQWaLVahc+t+4msgOs+Q/QPAQQQQKCRAhTgRk4rg0IAAQQQqLsAl6DrPkP0DwEEEDhiAc/zGjt6VsCNnVoGhgACCCBQZwEKcJ1nh74hgAACCDRWgALc2KllYAgggOvA1fYAABK/SURBVAACdRagANd5dugbAggggEBjBSjAjZ1aBoYAAgggUGcBCnCdZ4e+IYAAAgg0VoAC3NipZWAIIIAAAnUWoADXeXboGwIIIIBAYwUowI2dWgaGAAIIIFBnAQpwnWeHviGAAAIINFaAAtzYqWVgCCCAAAJ1FqAA13l26BsCCCCAQGMFKMCNnVoGhgACCCBQZwEKcJ1nh74hgAACCDRWgALc2KllYAgggAACdRagANd5dugbAggggEBjBSjAjZ1aBoYAAgggUGcBCnCdZ4e+IYAAAgg0VoAC3NipZWAIIIAAAnUWoADXeXboGwIIIIBAYwUowI2dWgaGAAIIIFBnAQpwnWeHviGAAAIINFaAAtzYqWVgCCCAAAJ1FqAA13l26BsCCCCAQGMFKMCNnVoGhgACCCBQZwEKcJ1nh74hgAACCDRW4KPGjoyBIYAAAgjkFmi1WrnP4YRiAqyAi7lxFgIIIIAAAjsJUIB34uNkBBBAAAEEiglwCbqYG2chgAACjRTwPK+R46rjoFgB13FW6BMCCCCAQOMFKMCNn2IGiAACCCBQRwEKcB1nhT4hgAACCDRegALc+ClmgAgggAACdRSgANdxVugTAggggEDjBSjAjZ9iBogAAgggUEcBCnAdZ4U+IYAAAgg0XoAC3PgpZoAIIIAAAnUUoADXcVboEwIIIIBA4wUowI2fYgaIAAIIIFBHAQpwHWeFPiGAAAIINF6AAtz4KWaACCCAAAJ1FKAA13FW6BMCCCCAQOMFKMCNn2IGiAACCCBQRwEKcB1nhT4hgAACCDRegALc+ClmgAgggAACdRSgANdxVugTAggggEDjBSjAjZ9iBogAAgggUEcBCnAdZ4U+IYAAAgg0XoAC3PgpZoAIIIAAAnUUoADXcVboEwIIIIBA4wUowI2fYgaIAAIIIFBHAQpwHWeFPiGAAAIINF6AAtz4KWaACCCAAAJ1FKAA13FW6BMCCCCAQOMFKMCNn2IGiAACCCBQRwEKcB1nhT4hgAACCDRegALc+ClmgAgggAACdRSgANdxVugTAggggEDjBT5q/AgPdICtVutAe063EUAAAQSyCLACzqLEMQgggAACCJQsQAEuGZRwCCCAAAIIZBHgEnQWpT0d43nenjKRBgEEEECgagFWwFXPAPkRQAABBI5SgAJ8lNPOoBFAAAEEqhagAFc9A+RHAAEEEDhKAQrwUU47g0YAAQQQqFqAAlz1DJAfAQQQQOAoBSjARzntDBoBBBBAoGoBCnDVM0B+BBBAAIGjFKAAH+W0M2gEEEAAgaoFKMBVzwD5EUAAAQSOUoACfJTTzqARQAABBKoWoABXPQPkRwABBBA4SgEK8FFOO4NGAAEEEKhagAJc9Qwk5p9p1GrJfC9wq3etZeS4HPuW1+qt46x+9kazTbTldS/I0eppNHvKktS+OTG0MRuZfo70FNXsXPVxlSspXlJ7KPzTpnUsTzZJsWztQZ9t/X5KF2zFrK19eBq5Ldcm4urcEP9mV3hjF08/TsY8so5lu6ctR3bP8EiTtks0T0phad/Z3RIzsWkXe0vQNP+0fZZQNO1TwONRU4GpN5S87nhh6V+OfYux15W84XQVZjo0X7kUPPf3dT0/hd++2k5qt/TEbwrHXB8TbkuKl9S+jhH/6R9f3lgW464nDb01TTxd8DxmXaQPnuctpsE8bOztyVYpQ3O0Pi6LZ948JY8lm+d6QGk/yzFPy2DdFzZeHxBuy/t6Xcew/Sxobwu1bkvzT9u3Pp+f+xdgBbzP33bqkKvzUl1JDx+WWr6701wnetGW5LfP9bhQYnti9/unGkqa3D+tBGf3E0lDnfaT4yXlT8xj3dHVy47rHNbEQWMGT7PSPB9cap4SJrKroGfuPJGk6yfpnuXkWOcq+DOjubnyMxqNNleAetdPV3ismXdwz53L2oHt9uXksSansQIBCnAF6JWmXDz6heDkRVuLx+clwRTmpPbkfvd1GlTg4DL08lpv/fp7qr6UGC9/nqAHk8Hq0vxgInXP9Kpdfo7kscb2ZPA0Z5wMp1pMDVKWRzFPEzlfnqAveTyL5sgy6szHZDQ38SY61a3n+fbzy47SL/8Xd8+fKxhtXvuieYJs/G/dBCjAdZsRR/15+j/6g4bjhW5MZSzx0Q8qsN5er1fWXY1fl5xk1d/h1JPnefIWY3Xnl+qk/6ta4iifQuXybF/o5qYvc6Eh66OQZ4E8pj+5PAvmyDrutONyma8CDU8D9/ZqdWt+wUx7FHIvmMuclst+hzxpY2ZfdQIU4Ors95rZ/z+6KVia62GVufPSXIyOPszKOKk9emTs2eofuPndla7u5puVqTkqKV5Seyxy8tP2K52thpAUK6k9OWi2PXk8s0WMHVXAMxYh/9MMnvmDlneGc3PT1SrcTd6a25c3i0QKC1CAwxpN325f6Hbc1fzyXOZ2WPvVmbqayL9161/WC+5BJbWv3zVrX3CuLt/NJ5r49ffVZsWXFC+pPT1PaJKW7+TX+ped/GMJhfE3s75zOHxeRs/wKZHt1Jz5PSOxw09S84QOzOAZOjr7Zjh/0nbWaDnNJ2+Dd3UvZ/cyd0XML5jpr6/i7vlzhQadw96aJxSKzcMRoAAfzlyV0tP2xRsNNdfl1Uzy/zEb6sHcUx08aDi91YW5TprUvqUHweU7c1BXZ+bG7PqRFC+pfX1ews/NpcjOpebdsW5Np5NiJbUnxM7bnMkzb9DV8bk9C+bJ5VkwR5mn5TEfnjzqvNVSx7y+M956KepeJFcR+yJ5yvQnVokC+3/jNRmzCcQ+ihE5qei+SJBCT6bD1UeVCp2d/aR95Il/NMOeM806+3iSjrTnTDq6ePs+8sQ9d+htykfwMkaNf8wndlqpHvvKtSVPmn/avhgNT/cowAq4xF9mXIQy79x8/oc4gkxF9xXu5/Ja93oTrJILB8lw4h7ymD9O0LkMvQt8S8406wwjsh+yJaf9pAKte8jzzLNAN+OnODE3SfbgsRnLnnKl+aft2/STjUoEWqbYV5KZpAgggAACCByxACvgI558ho4AAgggUJ0ABbg6ezIjgAACCByxAAX4iCefoSOAAAIIVCdAAa7OvmDm2DfFRKLE9q0+c+l/o5Llm5D8U8Ofy1zFSv1Gn0g+ybzBY9s3IZWRZ/3ZzehYnr65J0+OoM+2fscGt/5Gp/U3UhX0TOpbPJt5votnnjxFPW05snvaRhxvK+c1HI+67fnO7tsShPdbX0fbX8vhEOHtNP+0feEYbFcksMd3XJOqFIG0j8XE9sU/thD+Zpekb87xz7F8Q1JS32Mx/cNibdZvAsqbp+SxZPtYRgmeSc5le+bNU8QzJUc2z6RBh9vLMQ9HzLQde83658TarK/jTMFjBxW0j0WJPE3zT9sXCcKTvQuwAq7oF59K0oa+RcasgGzf0JP7G4pWf7ov6ZuQSstjBQv+cpfbHNbEQWMGz6S+JUYt6Jk7j7UD6Z7l5LAmzt6Y0Tz3twbt4J47l3W02+3LyWNNTmNFAhTgiuArSRv6FhmT3/bNOfm/oSj9G2TKyxOIbf5yUOibkMrOEWTK8L8ZPJP6lhy9mGf+PEEP8ngWzZE81gJ7MpqbyGV+E9K2sefLFYw7r33+MQV5+N/6ClCA6zs3pfXs6f/ooT/HV+K32qR+g0yJeQyI9dtjSs6xDT6XZ4G+FfIskCe3Z8Ec2zyz7M9lvgpY6jchbRl73ly57QuOKYstx1QnQAGuzn5vmW3fIpOUvNC3B6V8g0ypecLBQt8eE24ObxcaSzhAwnYez4QQ6c0FPNMDZtibwTNDFGeHODc3Pa/C3eStub2zSSWwKMDH8iLwv5Tg6ZuQkoZd7BuKkr9Bptw8oWihb48JtUY2E8cSOSr404T+/bVZfEfK84yeiREs7z5/Oja/59O5sa3UPKFjM3iGjs6+Gc6ftJ01Wk5z67cGhfvwLG9x9/y5Qslz2FvzhEKxeVgCFODDmq+dehv5FpmkSAW/PSjxG2RKzrO5FBn+JqSScySFi7dn8oyflPF5bs+MceOH5fKMn1zB8zzmRb41qKh7kVxF7IvkqWCaSJlVYO/vuybhjgKxj2lEoqXtixxY6Emp3yCT0oN95Il/NMOeE8+UaYrsintGduZ6UoJ5/GM+sfz2uY4dlPXpvnJtyZPmn7Yv6zA5zo0AK+Csv6nU7Li0b4pJ21d4GHv6VhfzURfX37hk/jgB34RU+JXw7MRnns+OyN/g5DVsurGH19dmtHvKleaftm/TTzYqE+DbkCqjJzECCCCAwDELsAI+5tln7AgggAAClQlQgCujJzECCCCAwDELUICPefYZOwIIIIBAZQIU4MroSYwAAgggcMwCFOBjnn3GjgACCCBQmQAFuDJ6EiOAAAIIHLMABfiYZ5+xI4AAAghUJkABroyexAgggAACxyxAAT7m2WfsCCCAAAKVCVCAK6MnMQIWgfC39SRtW06jCQEEDk+AAnx4c0aPEUAAAQQaIMDfgm7AJDIEBBBAAIHDE2AFfHhzRo8RQAABBBogQAFuwCQyBAQQQACBwxOgAB/enNFjBBBAAIEGCFCAGzCJDAEBBBBA4PAEKMCHN2f0GAEEEECgAQIU4IObxJlGrZZa5r/etZaR/sf2rT5H6h+7Oqc3mkXOSHqyvO4FOVo9jWbRLPFzZiPTn5GikVd9yZAvTy5ZxxR3iPcweG7LE/Td1v94jPJs/cjhz/jGU4We72qbJ9cutrY82W1DA07cLNk/Mc/zHaXMwfOwyS07vMaTgqbNRdF9SblozyHg8Tgwgak3lLzueGHpd2zfYux1JW84XR06HXoKP7dE8Jv887qen8I/Z7WddLwtrq3Ndn7eXA7GtBh3PWnorZls3fS8kmw9z1tMg3nJNBc2R1ubvdP5chW1TRlTNtuEzkeay/OPhM3yxOZta8sSK8sxO8xDWvi0uSi6Ly0f+7YLsALO8cvKwR/aeamupIcPS61XO73RSKNesKJer46X7+4014letCX558z1uEgZff9UQ0mT+6c18Ox+Immo077KzWXtRlcvO/vIY00eNGa0Ne7ng0vNU0JFdu1gmztXJPH6yXbbcvKs8xX8mcO/12ppNBrJ/DRXh3rX6Vd4tOMc5M5nJcg2D+XksnaARgcCFGAHqLUNuXj0/+E/8Str0Mu5TvX6vafFuKv55K3Mv0WLx+flwS/aiQPr6zSowMFl6OW13vr191Sm/q4f5eQKok0GwT+ercFE6p7plfllYfUoM8865tafGW1NnJPhVIupAcvyKG5roufLFfQnr23RPFlGn/mYHP4m5kSnuvU8fx7mlx2l3ynZbQ7y5wtGXWQeiuYKMvK/+xagAO9bvIJ8T/9HftBwvNBNqCp2X3Zkalf7xclOPesHFVhvr5cKVtBdjV+HEkkqK5fp6HDqyfM8eYuxuvNLdUL/gpaZZxtKbtv2hW5u+r75ttjr/YVtC+QqZFswz3p8u/zM7b9KNjwN5qC9Wt2m/4IpFZ6DgvnMaUVf43nHtos/5+4mQAHeze8gzvb/j2wKleZ6yNDjzktzoTr6CK+ao3tWz1b/kM3vrnR1N3+2KrWe41/hLpArHKz9SmfPQ4SP8LcLjelZlOcNeW2fR8jQUtA2Q+T0QzLapgdxu3cv/mYIVc2ByX0A8+B2lpsbnQLc3LmNjqx9oVtzmfny3L/MHN0ZfdZ+daauJvJv6fqX9qL3n0KLzdCJq8t084kmfv19lWmVVyxXKO3ynfx6798EDrXHNhPzxI5b3xu3jzF+8Op5DtuECJv75Pa8xWyL5QqdldE2dEb2zfC7wJO2s0Yr4D95G7xzfjm7l7lb4v+CGe7Hs9y7zUH+fKEO5JwHa65QODbrI0ABrs9cOO9J++KNhprr8urpzVLWpP4/aEM9mPusgwcNp7e6CN1jtZ7jLxLW9zW7OgvflE06wbQXzLW59Ni51Lw71u22DhbMk9b18L7MtuGTcmwHl0DNCTlsc8QPH5rbNnxyRdt5/YcnjzpvtdQxr+/YbZmkIewyB0XyFZ2HIrmSxky7Y4Htb5TmiHoJxD6OEelc2r7IgYWfTIdbPpJUOPLzE/eVK/4RDHtebJ/P0PaWuO32M5KOKMk//hGfWDr73McOyvN0n/m25Eqbi6L78lBw7HMBVsCOf8FxFd68c/P5H+IIsqXt26k/y2vd602m1fBOeczJe8pl/ghB5zL0ru8tebHNPrPPbLOfmnikM/89vuY2g9vyWtscV8JG2lwU3VdCt44+BN8HfPQvAQAQQAABBKoQYAVchTo5EUAAAQSOXoACfPQvAQAQQAABBKoQoABXoU5OBBBAAIGjF6AAH/1LAAAEEEAAgSoEKMBVqJMTAQQQQODoBSjAR/8SAAABBBBAoAoBCnAV6uREAAEEEDh6AQrw0b8EAEAAAQQQqELg/wOqnbOgGF7ckQAAAABJRU5ErkJggg=="}}},{"metadata":{},"cell_type":"markdown","source":"### **<div id=\"IV4\">4. Normalizing images and performing data augmentation</div>**\n\nNow that we have transformed our sound into images, we want them to have the same scale (for example, 128x128), for training purposes.\nWe will also perform **data augmentation**.\n\n![test](https://cdn-images-1.medium.com/max/800/1*C8hNiOqur4OJyEZmC7OnzQ.png)\n\nData augmentation consists in making minor alterations to our existing dataset. Minor changes such as flips or translations or rotations. Our neural network would think these are distinct images.\nA convolutional neural network that can robustly classify objects even if its placed in different orientations is said to have the property called invariance. More specifically, a CNN can be invariant to translation, viewpoint, size or illumination (Or a combination of the above).\n\nThis essentially is the premise of data augmentation. And augmentation can also help even with a large dataset; it can help to increase the amount of relevant data in your dataset. This is related to the way with which neural networks learn.\n\nOf course, each change are not good for each type of data. For example, in our problem, we may not want to flip or rotate our image, since it would alter our sound in a bad way. But we can for exmple increase the brightness of the image (resulting in a louder sound I guess), or translate it.  \n"},{"metadata":{"trusted":true},"cell_type":"code","source":"CUR_X_FILES, CUR_X = list(training_curated_df.fname.values), X_train_curated\n\ndef open_fat2019_image(fn, convert_mode, after_open)->Image:\n    # open\n    idx = CUR_X_FILES.index(fn.split('/')[-1])\n    x = PIL.Image.fromarray(CUR_X[idx])\n    # crop\n    time_dim, base_dim = x.size\n    crop_x = 0\n    #crop_x = random.randint(0, time_dim - base_dim)\n    x = x.crop([crop_x, 0, crop_x+base_dim, base_dim])    \n    # standardize\n    return Image(pil2tensor(x, np.float32).div_(255))\n\nvision.data.open_image = open_fat2019_image","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Batch size --> How many images are trained at one time. Lower it if you run out of memory\nbs = 64\n#Image size. Square images makes the learning process faster. We can increase the size of the images once our model is stable, in order to improve accuracy.\nsize = 128\n\n#Performing data augmentation\ntfms = get_transforms(do_flip=False, max_rotate=0, max_lighting=0.1, max_zoom=0, max_warp=0.)\n\n#We put the transformed image data into /kaggle/working because ../input is a read-only directory.\nsrc = (ImageList.from_csv('/kaggle/working', '../input/train_curated.csv', folder='../input/train_curated')\n       .split_by_rand_pct(0.2).label_from_df(label_delim=','))\n\n#Creates a databunch, because our cnn_learner below needs a databunch.\ndata = src.transform(tfms, size=size).databunch(bs=bs).normalize()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data.show_batch(4)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## **<div id=\"V\">V. Model data</div>**\n\nNow we will start training our model. We will use a **convolutional neural network** backbone and a fully connected head with a single hidden layer as a classifier. Our model takes images as input and will output the predicted probability for each of the categories.\n\nWe need to feed our learner with a databunch, and an architecture model. **resnet34** is a very good architecture to get started with. We can use **resnet50** to try getting better results once we are happy with our model. If we run out of memory while using resnet50, we can try to lower *bs* (batch size, how many images are trained at one time). Computing time will get a little bit longer, but we won't run out of memory anymore.\n\nWe have to set *pretrained* to *False*, as pretrained models are forbidden in this competition. For our metrics, we use *lwlrap* as it is the metric used in the competition.\n\nWe use **lr_find** to find the best learning rate for our model. It seems to start diverging when lr > 0,1, so we choose a value ten times lower, i.e. **lr = 0,01**.\n\nWe will train for 5 epochs (5 cycles through all our data)."},{"metadata":{"trusted":true},"cell_type":"code","source":"arch = models.resnet18\n\nlearn = cnn_learner(data, arch, pretrained=False, metrics=[lwlrap], wd = 0.1, ps = 0.5)\n\nlearn.lr_find()\nlearn.recorder.plot()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.fit_one_cycle(5, 1e-2)\nlearn.save('first-attempt-128')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.unfreeze()\nlearn.fit_one_cycle(1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.lr_find()\nlearn.recorder.plot()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.fit(20, slice(2e-3, 2e-4))\nlearn.save('second-attempt-128')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.lr_find()\nlearn.recorder.plot()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.fit(20, slice(1e-3, 1e-4))\nlearn.save('third-attempt-128')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### **<div id=\"VI\">VI. New learner with increased size of images</div>**\n\nNow that we have a pretty good learner fed with 128x128 images, let's: \n* Create a new databunch full of 256x256 images, \n* Keep the same learner as before ('third-attempt-128'),\n* Replace the data inside the learner with my new 256x256 data,\n* Freeze the model again (in order to train only the last few layers)."},{"metadata":{"trusted":true},"cell_type":"code","source":"size = 256\n\n#Creates a databunch, because our cnn_learner below needs a databunch.\ndata = src.transform(tfms, size=size).databunch(bs=bs).normalize(imagenet_stats)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Replace with 256x256 databunch\nlearn.data = data\n#Freeze the model\nlearn.freeze()\n#Plot lr_find()\nlearn.lr_find()\nlearn.recorder.plot()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.fit(5, 3e-3)\nlearn.save('first-attempt-256')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.unfreeze()\nlearn.lr_find()\nlearn.recorder.plot()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.fit(7, slice(1e-3, 1e-4))\nlearn.save('second-attempt-256')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.recorder.plot_losses()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.fit_one_cycle(20, slice(5e-4, 5e-5), callbacks=[SaveModelCallback(learn, monitor='lwlrap', mode='max')])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.export()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"CUR_X_FILES, CUR_X = list(testing_df.fname.values), X_test\n\ntest = ImageList.from_csv(Path('/kaggle/working'), Path('../input/sample_submission.csv'), folder=Path('../input/test'))\nlearn = load_learner(Path('/kaggle/working'), test=test)\npreds, _ = learn.get_preds(ds_type=DatasetType.Test)\n\ntesting_df[learn.data.classes] = preds\ntesting_df.to_csv('submission.csv', index=False)\ntesting_df.head()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}