{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## TLDR\n\n- This notebook shows how to implement a on-the-fly audio transcoding and augmenting class using TorchAudio.\n- The implemented class can go through the entire MP3 training set (~353 hrs) of audio and apply augmentations in about 55 minutes with 4 CPUs. Faster iteration can be acheived by using pre-transcoded WAV files directly instead of original MP3 files.\n\n## Transcoding Audio\n\n- The provided dataset is in MP3 format, and sample at 32 kHz.\n- Most model architectures work with 16 kHz audio, and in WAV format with raw audio samples.\n- So we need to **transcode** the files from MP3 @ 32 kHz into WAV @ 16 kHz \n\n## Augmenting Training Audio\n\n- Data augmentation has become a main-stay in machine learning and is necessary to make the model robust to out-of-domain, unseen conditions.\n\n- There are several traditional augmentation techniques that are applied to audio for ASR training:\n\n\n| Augmentation Type | Description |\n| :---------------: |:------------|\n| Volume Scaling | Involves randomly perturbing the audio volume; Makes ASR robust against amplitude of speaker's voice, mic gain variation, distance to mic etc. |\n| Speed Perturbation | Involves speeding up (\\~1.1x) and slowing down (\\~0.9x) the audio; Makes ASR robust against different speaking rates, as well introducing more examples of frequencies that the model should expect. |\n| Adding Noise | Involves introducing different type of noise (white, babble, music, applicances etc.) into the audio at different signal-to-noise ratios (SNR) to simulate noisy environments. |\n| Adding Reverberation | Involves adding echoes into the audio to simulate Room Impulse Respones (RIRs); You may notice that in an empty room or a room without curtains, your voice is usually more echo-ey. This can degrade ASR performance if the model is not trained to handle such conditions |\n| Re-encoding | Involves re-encoding audio files with different compression algorithms / bit-rates; Compressed file formats such as MP3s can change the frequency content of the audio and adversely affect ASR if the model is presented with audio compressed at lower bit rates. |\n  \n\n- Most of these augmentation techniques can be implemented in Python very easily using PyTorch and TorchAudio.","metadata":{}},{"cell_type":"markdown","source":"## Importing Dependencies","metadata":{}},{"cell_type":"code","source":"from typing import Dict, List, Tuple, Any, Union\n\nimport os\nimport time\nimport random\n\nimport numpy as np\nimport pandas as pd\n\nimport torch\nimport torchaudio\nimport torchaudio.functional as F\nimport torchaudio.transforms as T\n\nimport librosa\nimport librosa.display\nfrom matplotlib import pyplot as plt\n\nfrom tqdm import tqdm\nfrom IPython.display import display, Audio, HTML\n\nrandom.seed(1313)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-20T06:33:17.992719Z","iopub.execute_input":"2022-07-20T06:33:17.993278Z","iopub.status.idle":"2022-07-20T06:33:22.881917Z","shell.execute_reply.started":"2022-07-20T06:33:17.993151Z","shell.execute_reply":"2022-07-20T06:33:22.880245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## `AudioConverter` : Class for transcoding and augmenting audio on-the-fly","metadata":{}},{"cell_type":"code","source":"class AudioConverter:\n    \"\"\"\n    AudioConverter offers methods to load, transcode and augment\n    audio data in various ways.\n    \"\"\"\n\n    # Configurations for parameters used in torchaudio's resampling kernel.\n    resampleFilterParams = {\n        \"fast\": {  # Fast and less accurate but still MSE = ~2e-5 compared to librosa.\n            \"lowpass_filter_width\": 16,\n            \"rolloff\": 0.85,\n            \"resampling_method\": \"kaiser_window\",\n            \"beta\": 8.555504641634386,\n        },\n        \"best\": { # Twice as slow, and a little bit more accurate.\n            \"lowpass_filter_width\": 64,\n            \"rolloff\": 0.9475937167399596,\n            \"resampling_method\": \"kaiser_window\",\n            \"beta\": 14.769656459379492,       \n        },\n    }\n\n    def __init__(\n        self,\n        sampleRate: int,\n        speedAugProb: float = 0.5,\n        volAugProb: float = 0.5,\n        reverbAugProb: float = 0.25,\n        noiseAugProb: float = 0.25,\n        speedFactors: Tuple[float, float] = None,\n        volScaleMinMax: Tuple[float, float] = None,\n        reverbRoomScaleMinMax: Tuple[float, float] = None,\n        reverbHFDampingMinMax: Tuple[float, float] = None,\n        reverbSustainMinMax: Tuple[float, float] = None,\n        noiseSNRMinMax: Tuple[float, float] = None,\n        noiseFileList: List[str] = None,\n    ):\n        \"\"\"\n        Initializes AudioConverter.\n\n        Parameters\n        ----------\n        sampleRate: int\n            Sampling rate to convert audio to, if required.\n        \n        speedAugProb: float, optional\n            Probability that speed augmentation will be applied.\n            If <= 0, speed augmentation is disabled.\n\n        volAugProb: float, optional\n            Probability that volume augmentation will be applied.\n            If <= 0, volume augmentation is disabled.\n\n        reverbAugProb: float, optional\n            Probability that reverberation augmentation will be applied.\n            If <= 0, reverberation augmentation is disabled.\n\n        noiseAugProb: float, optional\n            Probability that noise augmentation will be applied.\n            If <= 0, noise augmentation is disabled.\n\n        speedFactors: List[float], optional\n            List of factors by which to speed up (>1) or slow down (<1)\n            audio by. One factor is chosen randomly if provided. Otherwise,\n            default speed factors are [0.9, 1.0, 1.0].\n            \n        volScaleMinMax: Tuple[float, float], optional\n            [Min, Max] range for volume scale factors. One factor is\n            chose randomly with uniform probability from this range.\n            Default range is [0.125, 2.0].\n\n        reverbRoomScaleMinMax: Tuple[float, float], optional\n            [Min, Max] range for room size percentage. Values must be\n            between 0 and 100. Larger room size results in more reverb.\n            Default range is [25, 75].\n\n        reverbHFDampingMinMax: Tuple[float, float], optional\n            [Min, Max] range for high frequency damping percentage. Values must\n            be between 0 and 100. More damping results in muffled sound.\n            Default range is [25, 75].\n        \n        reverbSustainMinMax: Tuple[float, float], optional\n            [Min, Max] range for reverberation sustain percentage. Values must\n            be between 0 and 100. More sustain results in longer lasting echoes.\n            Default range is [25, 75].\n            \n        noiseSNRMinMax: Tuple[float, float], optional\n            [Min, Max] range for signal-to-noise ratio when adding noise. One\n            factor is chose randomly with uniform probability from this range.\n            Lower SNR results in louder noise. Default range is [10.0, 30.0].\n\n        noiseFileList: List[str], optional\n            List of paths to audio files to use as noise samples. If None is provided,\n            noise augmentation will be disabled. Otherwise, the audio files will be assumed\n            to be sources of noise, and be mixed in with speech audio on-the-fly.\n        \"\"\"\n        self.sampleRate = sampleRate\n        self.speedAugProb = speedAugProb\n        self.volAugProb = volAugProb\n        self.reverbAugProb = reverbAugProb\n        self.noiseAugProb = noiseAugProb\n        \n        # Factors by which audio speed is perturbed.\n        self.speedFactors = speedFactors\n        if speedFactors is None:\n            self.speedFactors = [0.9, 1.0, 1.1]\n        \n        # [Min, Max] Volume scale range.\n        self.volScaleRange = volScaleMinMax\n        if volScaleMinMax is None:\n            self.volScaleRange = [0.125, 2.0]\n        \n        # [Min, Max] Room size as a percentage, higher = more reverb\n        self.reverbRoomScaleRange = reverbRoomScaleMinMax\n        if reverbRoomScaleMinMax is None:\n            self.reverbRoomScaleRange = [25, 75]\n        \n        # [Min, Max] High frequency damping as a percentage, higher = more damping.\n        self.reverbHFDampingRange = reverbHFDampingMinMax\n        if reverbHFDampingMinMax is None:\n            self.reverbHFDampingRange = [25, 75]\n        \n        # [Min, Max] How long reverb is sustained as a percentage, higher = lasts longer.\n        self.reverbSustainRange = reverbSustainMinMax \n        if reverbSustainMinMax is None:\n            self.reverbSustainRange = [25, 75]       \n\n        # Audio files to use as source of noise.\n        self.noiseFiles = noiseFileList\n        if self.noiseFiles is None or len(self.noiseFiles) == 0:\n            self.noiseAugProb = -1\n\n        # [Min, Max] Signal to noise ratio range for adding noise to audio.\n        # Lower SNR = noise is more prominent, i.e. speech is more noisy.\n        self.noiseSNRRange = noiseSNRMinMax\n        if noiseSNRMinMax is None:\n            self.noiseSNRRange = [10.0, 30.0]\n        \n        self.validateConfig()\n        \n    def validateConfig(self):\n        \"\"\"\n        Checks configured options and raises an error if they\n        are not consistent with what is expected.\n        \"\"\"\n        if len(self.volScaleRange) != 2:\n            raise ValueError(\"volume scale range must be provided as [min, max]\")\n        if len(self.reverbRoomScaleRange) != 2:\n            raise ValueError(\"reverb room scale range must be provided as [min, max]\")\n        if len(self.reverbHFDampingRange) != 2:\n            raise ValueError(\"reverb high frequency dampling range must be provided as [min, max]\")\n        if len(self.reverbSustainRange) != 2:\n            raise ValueError(\"reverb sustain range must be provided as [min, max]\")\n        if len(self.noiseSNRRange) != 2:\n            raise ValueError(\"noise SNR range must be provided as [min, max]\")\n            \n        for v in self.reverbRoomScaleRange:\n            if v > 100 or v < 0:\n                raise ValueError(\"reverb room scale must be between 0 and 100\")\n        for v in self.reverbHFDampingRange:\n            if v > 100 or v < 0:\n                raise ValueError(\"reverb high frequency dampling must be between 0 and 100\")\n        for v in self.reverbSustainRange:\n            if v > 100 or v < 0:\n                raise ValueError(\"reverb sustain range must be between 0 and 100\")\n\n    @classmethod\n    def loadAudio(\n        cls, audioPath: str, sampleRate: int = None, returnTensor: bool = True, resampleType: str = \"fast\",\n    ) -> Union[torch.Tensor, np.ndarray]:\n        \"\"\"\n        Uses torchaudio to load and resample (if necessary) audio files and returns\n        audio samples as either a numpy.float32 array or a torch.Tensor.\n        \n        Parameters\n        ----------\n        audioPath: str\n            Path to audio file file (wav / mp3 / flac).\n        \n        sampleRate: int, optional\n            Sampling rate to convert audio to. If None,\n            audio is not resampled.\n        \n        returnTensor: bool, optional\n            If True, the audio samples are returned as a torch.Tensor.\n            Otherwise, the samples are returned as a numpy.float32 array.\n            \n        resampleType: str, optional\n            Either \"fast\" or \"best\" - sets the quality of resampling.\n            \"best\" is twice as slow as \"fast\" but more accurate. \"fast\"\n            is still comparable to librosa's resampled output though,\n            in terms of MSE.\n\n        Returns\n        -------\n        Union[torch.Tensor, np.ndarray]\n            Audio waveform scaled between +/- 1.0 as either a numpy.float32 array,\n            or torch.Tensor, with shape (channels, numSamples)\n        \"\"\"\n        x, sr = torchaudio.load(audioPath)\n        if sampleRate is not None or sr != sampleRate:\n            x = F.resample(x, sr, sampleRate, **cls.resampleFilterParams[resampleType])\n        \n        if returnTensor:\n            return x\n        \n        return x.numpy()\n\n    def getAudio(self, audioPath: str, returnTensor: bool = False) -> Union[np.ndarray, torch.Tensor]:\n        \"\"\"\n        Loads audio from specified path and applies augmentations randomly\n        on-the-fly. Audio samples scaled between -1.0 and +1.0 are returned\n        as a numpy.float32 array or torch.Tensor with shape (numSamples,).\n\n        Parameters\n        ----------\n        audioPath: str\n            Path to audio file file (wav / mp3 / flac).\n        \n        returnTensor: bool, optional\n            If True, the audio samples are returned as a torch.Tensor.\n            Otherwise, the samples are returned as a numpy.float32 array.\n        \n        Returns\n        ------- \n        Union[torch.Tensor, np.ndarray]\n            Audio waveform scaled between +/- 1.0 as either a numpy.float32 array,\n            or torch.Tensor, with shape (channels, numSamples)\n        \"\"\"\n        wav = self.loadAudio(\n            audioPath, sampleRate=self.sampleRate, returnTensor=True, resampleType=\"fast\",\n        )\n\n        # Applying sox-based effects first.\n        effects = []\n        \n        if random.uniform(0, 1) <= self.speedAugProb:\n            effects.extend([\n                [\"speed\", f\"{random.choice(self.speedFactors)}\"],\n                [\"rate\", f\"{self.sampleRate}\"],\n            ])\n\n        if random.uniform(0, 1) <= self.reverbAugProb:\n            effects.append([\n                \"reverb\",\n                f\"{random.uniform(*self.reverbSustainRange)}\",\n                f\"{random.uniform(*self.reverbHFDampingRange)}\",\n                f\"{random.uniform(*self.reverbRoomScaleRange)}\",\n            ])\n        \n        # If no effects are selected, this is a no-op.\n        wav = self.applySoxEffects(wav, effects)\n\n        if random.uniform(0, 1) <= self.noiseAugProb:\n            noiseFile = random.choice(self.noiseFiles)\n            noiseSNR = random.uniform(*self.noiseSNRRange)\n            wav = self.addNoiseFromFile(wav, noiseFile, noiseSNR)\n\n        if random.uniform(0, 1) <= self.volAugProb:\n            volScale = random.uniform(*self.volScaleRange)\n            wav = self.scaleVolume(wav, volScale)\n        \n        if returnTensor:\n            return wav\n        \n        return wav.numpy()\n\n\n    def scaleVolume(self, wav: Union[np.ndarray, torch.Tensor], scale: float) -> torch.Tensor:\n        \"\"\"\n        Scales the amplitude (with clipping) of the provided audio signal\n        by the given scale factor.\n        \n        Parameters\n        ----------\n        wav: Union[np.ndarray, torch.Tensor]\n             Audio samples scaled between -1.0 and +1.0, with shape\n             (channels, numSamples).\n\n        Returns\n        -------\n        torch.Tensor\n            Audio samples with perturbed volume.\n        \"\"\"\n        if scale == 1.0:\n            return wav\n\n        return torch.clamp(wav * scale, -1.0, 1.0)\n\n    def addNoiseFromFile(\n        self, wav: Union[np.ndarray, torch.Tensor], noiseFile: str, snr: float,\n    ) -> torch.Tensor:\n        \"\"\"\n        Adds noise signal from provided noise audio file at the \n        specified SNR to the speech signal.\n        \n        Parameters\n        ----------\n        wav: Union[np.ndarray, torch.Tensor]\n             Audio samples scaled between -1.0 and +1.0, with shape\n             (channels, numSamples).\n\n        snr: float\n            Signal-to-Noise ratio at which to mix in the noise signal.\n        \n        Returns\n        -------\n        torch.Tensor\n            Audio samples with noise added at specified SNR.\n        \"\"\"\n        # Loading noise signal.\n        noiseSig = self.loadAudio(\n            noiseFile, sampleRate=self.sampleRate, returnTensor=True, resampleType=\"fast\",\n        )\n\n        # Computing noise power.\n        noisePower = torch.mean(torch.pow(noiseSig, 2))\n        \n        # Computing signal power.\n        signalPower = torch.mean(torch.pow(wav, 2))\n\n        # Noise Coefficient for target SNR; amplitude coeff is sqrt of power coeff.\n        noiseScale = torch.sqrt((signalPower / noisePower) / (10 ** (snr / 20.0)))\n        \n        # Add noise at random location in speech signal.\n        nWav, nNoise = wav.shape[-1], noiseSig.shape[-1]\n\n        if nWav < nNoise:\n            a = random.randint(0, nNoise-nWav)\n            b = a + nWav\n            return wav + (noiseSig[..., a:b] * noiseScale)\n        \n        a = random.randint(0, nWav-nNoise)\n        b = a + nNoise          \n        wav[..., a:b] += (noiseSig * noiseScale)\n\n        return wav\n    \n        \n    def applySoxEffects(self, wav: Union[np.ndarray, torch.Tensor], effects: List[List[str]]) -> torch.Tensor:\n        \"\"\"\n        Applies different audio manipulation effects to provided audio, like\n        speed and volume perturbation, reverberation etc. For a full list of\n        supported effects, check torchaudio.sox_effects.\n\n        Parameters\n        ----------\n        wav: Union[np.ndarray, torch.Tensor]\n             Audio samples scaled between -1.0 and +1.0, with shape\n             (channels, numSamples).\n        \n        effects: List[List[str]]\n            List of sox effects and associated arguments, example:\n            '[ [\"speed\", \"1.2\"], [\"vol\", \"0.5\"] ]'\n\n        Returns\n        -------\n        torch.Tensor\n            Audio samples with effects applied. May not be the same\n            number of samples as input sample array, depending on types\n            of effects applied (e.g. speed perturbation may reduce or\n            increase the number of samples).\n        \"\"\"\n        if effects is None or len(effects) == 0:\n            return wav\n\n        wav, _ = torchaudio.sox_effects.apply_effects_tensor(\n            wav, sample_rate=self.sampleRate, effects=effects,\n        )\n\n        return wav\n    \n    def perturbSpeed(self, wav: Union[np.ndarray, torch.Tensor], factor: float) -> torch.Tensor:\n        \"\"\"\n        Perturbs the speed of the provided audio signal by the given factor.\n        \n        Parameters\n        ----------\n        wav: Union[np.ndarray, torch.Tensor]\n             Audio samples scaled between -1.0 and +1.0, with shape\n             (channels, numSamples).\n\n        Returns\n        -------\n        torch.Tensor\n            Audio samples with perturbed speed. Will have more or less\n            samples than input depending on whether slowed down or\n            sped up.\n        \"\"\"\n        effects = [\n            [\"speed\", f\"{factor}\"],\n            [\"rate\", f\"{self.sampleRate}\"],\n        ]\n        \n        return self.applySoxEffects(wav, effects)\n    \n    def addReverb(\n        self, wav: Union[np.ndarray, torch.Tensor], roomSize: float, hfDamping: float, sustain: float,\n    ) -> torch.Tensor:\n        \"\"\"\n        Adds reverberation to the provided audio signal using given parameters.\n        \n        Parameters\n        ----------\n        wav: Union[np.ndarray, torch.Tensor]\n             Audio samples scaled between -1.0 and +1.0, with shape\n             (channels, numSamples).\n        \n        roomSize: float\n            Room size as a percentage between 0 and 100,\n            higher = more reverb\n\n        hfDamping: float\n            High Frequency damping as a percentage between 0 and 100,\n            higher = more damping.\n\n        sustain: float\n            How long reverb is sustained as a percentage between 0 and 100,\n            higher = lasts longer.\n\n        Returns\n        -------\n        torch.Tensor\n            Audio samples with reverberated audio.\n        \"\"\"\n        effects = [[\"reverb\", f\"{roomSize}\", f\"{hfDamping}\", f\"{sustain}\"]]\n        return self.applySoxEffects(wav, effects)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T06:33:22.884539Z","iopub.execute_input":"2022-07-20T06:33:22.885269Z","iopub.status.idle":"2022-07-20T06:33:22.927788Z","shell.execute_reply.started":"2022-07-20T06:33:22.885227Z","shell.execute_reply":"2022-07-20T06:33:22.925871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Generating augmented audio samples","metadata":{}},{"cell_type":"code","source":"# Load training dataset metadata.\ntrainMetaFile = \"../input/dlsprint/train.csv\"\ntrainAudioDir = \"../input/dlsprint/train_files\"\n\ntrainMeta = pd.read_csv(trainMetaFile)\ntrainMeta['path'] = [ os.path.join(trainAudioDir, f) for f in trainMeta['path'] ]","metadata":{"execution":{"iopub.status.busy":"2022-07-20T06:33:22.930819Z","iopub.execute_input":"2022-07-20T06:33:22.931942Z","iopub.status.idle":"2022-07-20T06:33:25.414765Z","shell.execute_reply.started":"2022-07-20T06:33:22.931878Z","shell.execute_reply":"2022-07-20T06:33:25.413585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Instantiating AudioConverter with mostly default arugments.\n# Will use training audio files as noise sources as well.\nac = AudioConverter(\n    sampleRate=16000,\n    noiseFileList=trainMeta['path'].tolist(),\n)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T06:33:25.417812Z","iopub.execute_input":"2022-07-20T06:33:25.418345Z","iopub.status.idle":"2022-07-20T06:33:25.431305Z","shell.execute_reply.started":"2022-07-20T06:33:25.418289Z","shell.execute_reply":"2022-07-20T06:33:25.429814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Checking augmented samples","metadata":{}},{"cell_type":"code","source":"# Getting a random file from the dataset and genrating augmented samples.\nsample = trainMeta.iloc[random.randint(0, len(trainMeta))]\nnoiseSample = trainMeta.iloc[random.randint(0, len(trainMeta))]\n\nsent = sample[\"sentence\"]\ndisplay(HTML(f\"<h3>{sent}</h3>\"))\ndisplay(HTML(\"<br>\"))\n\nx = ac.loadAudio(sample[\"path\"], ac.sampleRate)\ndisplay(HTML(\"<h4>Original Audio</h4>\"))\ndisplay(Audio(x, rate=ac.sampleRate))\n\n# When using getAudio(), augmentations are mostly done in-place, and compounded on top of each other. \n# For demonstrating the output of each augmentation separately, we are reloading the original audio\n# each time before applying augmentation.\n\nfor snr in [30, 20, 10]:\n    x = ac.loadAudio(sample[\"path\"], ac.sampleRate)\n    xNoise = ac.addNoiseFromFile(x, noiseSample[\"path\"], snr=snr)\n    display(HTML(f\"<h4>With Background Noise (SNR = {snr})</h4>\"))\n    display(Audio(xNoise, rate=ac.sampleRate))\n\n\nfor factor in [0.9, 1.1]:\n    x = ac.loadAudio(sample[\"path\"], ac.sampleRate)\n    xSpeed = ac.perturbSpeed(x, factor)\n    display(HTML(f\"<h4>With Speed Perturbation (factor = {factor}x)</h4>\"))\n    display(Audio(xSpeed, rate=ac.sampleRate))\n\nfor reverb in [25, 75]:\n    x = ac.loadAudio(sample[\"path\"], ac.sampleRate)\n    xReverb = ac.addReverb(x, reverb, reverb, reverb)\n    display(HTML(f\"<h4>With Reverberation (room spacing, hf-damping, sustain = {reverb}%)</h4>\"))\n    display(Audio(xReverb, rate=ac.sampleRate))\n\ndisplay(HTML(f\"<h4>With Volume Scaling</h4>\"))\nfig, ax = plt.subplots(nrows=3, sharex=True)\n\nfor i, factor in enumerate([0.125, 0.5, 1.0]):\n    x = ac.loadAudio(sample[\"path\"], ac.sampleRate)\n    xVol = ac.scaleVolume(x, factor)\n    librosa.display.waveshow(xVol.numpy(), sr=ac.sampleRate, ax=ax[i])\n    ax[i].set_title(f\"factor = {factor}x\")\n    ax[i].set_ylim([-1, 1])\n    \nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T06:33:25.432807Z","iopub.execute_input":"2022-07-20T06:33:25.433199Z","iopub.status.idle":"2022-07-20T06:33:27.294933Z","shell.execute_reply.started":"2022-07-20T06:33:25.433158Z","shell.execute_reply":"2022-07-20T06:33:27.293658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Benchmarking iteration over full train set\n\n- To take advantage of multiple CPU cores, we will wrap our `AudioConverter` with the a `torch.Dataset` class, and give to `torch.DataLoader`.\n- `torch.DataLoader` will take care of parallelization for us.\n- With 4 workers, we can through about 40 ~ 60 files a minute with all augmentations having a chance of being applied. If this is not fast enough, we can either get more CPUs or disable some of the augmentations. It also ought to be faster using pre-transcoded WAV files directly instead of using original MP3 files.","metadata":{}},{"cell_type":"code","source":"class DLSprintDataset(torch.utils.data.Dataset):\n        \n    def __init__(self, df: pd.DataFrame, ac: AudioConverter):\n        self.df = df\n        self.ac = ac\n\n        self.paths = df['path'].tolist()\n        self.sentences = df['sentence'].tolist()\n        self.len = len(self.df)\n\n    def __len__(self):\n        return self.len\n    \n    def __getitem__(self, idx): \n        if idx >= self.len:\n            raise IndexError(\"index out of range\")\n        return {\n            \"audio\": self.ac.getAudio(self.paths[idx]),\n            \"sentence\": self.sentences[idx],\n        }\n\n# Getting number of available CPUs.\nnumCPU = !nproc\nnumCPU = int(numCPU[0])\n\n# Setting up dataset and data loader. Ignoring batching and padding for now.\ntrainDS = DLSprintDataset(trainMeta, ac)\ntrainLoader = torch.utils.data.DataLoader(trainDS, batch_size=1, shuffle=False, num_workers=numCPU)\n\ntotalAudioSamples = 0\nfor sample in tqdm(trainLoader):\n    totalAudioSamples += sample[\"audio\"][0].shape[-1]\n\nprint(f\"total audio duration = {totalAudioSamples / (16000 * 3600)} hrs\")","metadata":{"execution":{"iopub.status.busy":"2022-07-20T06:33:27.296854Z","iopub.execute_input":"2022-07-20T06:33:27.297578Z","iopub.status.idle":"2022-07-20T07:33:13.615190Z","shell.execute_reply.started":"2022-07-20T06:33:27.297529Z","shell.execute_reply":"2022-07-20T07:33:13.611890Z"},"trusted":true},"execution_count":null,"outputs":[]}]}