{
  "id": 233269,
  "title": "5 Second Audio Clips Dataset",
  "url": "/competitions/birdclef-2021/discussion/233269",
  "author_name": "",
  "post_date": "2021-04-18T09:35:19.594412900Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>The <code>train_short_audio</code> can be used to train a multi-label classifier. However, the model needs to be trained on 5 second audio clips. I went through the trouble of creating 5 second audio clips using individual audio files in the <code>train_short_audio</code> directory.</p>\n<p>You can either create the clips on the fly or use this dataset. <strong>By using this dataset your input pipeline will be spared of clipping overhead.</strong> </p>\n<p>I have used a 16 CPU core, 60 GB RAM virtual instance on GCP to create this dataset. You can use <a href=\"https://www.kaggle.com/ayuraj/multiprocessing-5-second-audio-clips\" target=\"_blank\">Multiprocessing 5 Second Audio Clips</a> kernel to reproduce the result. It uses multiprocessing to create the dataset. The resulting dataset is 20.5 GB in size. </p>\n<p><strong>Note</strong>: I tried to upload the dataset as a Kaggle dataset but all attempts failed - Kaggle API as well as manually uploading the <code>.zip</code> files. </p>\n<p>I have created a new <a href=\"https://docs.wandb.ai/ref/app/features/teams\" target=\"_blank\">W&amp;B team account</a> to upload the entire dataset as an <a href=\"https://docs.wandb.ai/guides/artifacts\" target=\"_blank\">Artifact</a>. If you want to download this dataset please use the code snippet below:</p>\n<pre><code>import wandb\nrun = wandb.init()\nartifact = run.use_artifact('birdclef21/audio_clips_zip/bird_audio_clips_zip:v0', type='dataset')\nartifact_dir = artifact.download()\n</code></pre>\n<p><strong>Note</strong>: You will require a <a href=\"https://wandb.ai/site\" target=\"_blank\">W&amp;B account</a>. <br>\n<strong>Note</strong>: You will find 398 zipped files where each zipped file is associated to a bird species. I have mistakenly uploaded <code>dataset-metadata.json</code> file. Please ignore it. <br>\n<strong>Note</strong>: The download might take time according to the internet speed available to your system.</p>\n<p>These are raw clips. It's up to you to do pre-processing, filter out unwanted clips, etc. I will share some of my strategies in near future. </p>\n<p>I have also created <code>train_clips.csv</code> file that contains metadata for each clip. To download use the code snippet below:</p>\n<pre><code>import wandb\nrun = wandb.init()\nartifact = run.use_artifact('birdclef21/audio_clips_zip/bird_audio_clip_csv:v0', type='dataset')\nartifact_dir = artifact.download()\n</code></pre>\n<p>Do let me know if you face any difficulty downloading the files. </p>",
  "messages": [
    {
      "id": "1277020",
      "postDate": "04/18/2021 09:35:19",
      "content": "<p>The <code>train_short_audio</code> can be used to train a multi-label classifier. However, the model needs to be trained on 5 second audio clips. I went through the trouble of creating 5 second audio clips using individual audio files in the <code>train_short_audio</code> directory.</p>\n<p>You can either create the clips on the fly or use this dataset. <strong>By using this dataset your input pipeline will be spared of clipping overhead.</strong> </p>\n<p>I have used a 16 CPU core, 60 GB RAM virtual instance on GCP to create this dataset. You can use <a href=\"https://www.kaggle.com/ayuraj/multiprocessing-5-second-audio-clips\" target=\"_blank\">Multiprocessing 5 Second Audio Clips</a> kernel to reproduce the result. It uses multiprocessing to create the dataset. The resulting dataset is 20.5 GB in size. </p>\n<p><strong>Note</strong>: I tried to upload the dataset as a Kaggle dataset but all attempts failed - Kaggle API as well as manually uploading the <code>.zip</code> files. </p>\n<p>I have created a new <a href=\"https://docs.wandb.ai/ref/app/features/teams\" target=\"_blank\">W&amp;B team account</a> to upload the entire dataset as an <a href=\"https://docs.wandb.ai/guides/artifacts\" target=\"_blank\">Artifact</a>. If you want to download this dataset please use the code snippet below:</p>\n<pre><code>import wandb\nrun = wandb.init()\nartifact = run.use_artifact('birdclef21/audio_clips_zip/bird_audio_clips_zip:v0', type='dataset')\nartifact_dir = artifact.download()\n</code></pre>\n<p><strong>Note</strong>: You will require a <a href=\"https://wandb.ai/site\" target=\"_blank\">W&amp;B account</a>. <br>\n<strong>Note</strong>: You will find 398 zipped files where each zipped file is associated to a bird species. I have mistakenly uploaded <code>dataset-metadata.json</code> file. Please ignore it. <br>\n<strong>Note</strong>: The download might take time according to the internet speed available to your system.</p>\n<p>These are raw clips. It's up to you to do pre-processing, filter out unwanted clips, etc. I will share some of my strategies in near future. </p>\n<p>I have also created <code>train_clips.csv</code> file that contains metadata for each clip. To download use the code snippet below:</p>\n<pre><code>import wandb\nrun = wandb.init()\nartifact = run.use_artifact('birdclef21/audio_clips_zip/bird_audio_clip_csv:v0', type='dataset')\nartifact_dir = artifact.download()\n</code></pre>\n<p>Do let me know if you face any difficulty downloading the files. </p>",
      "rawMarkdown": "The `train_short_audio` can be used to train a multi-label classifier. However, the model needs to be trained on 5 second audio clips. I went through the trouble of creating 5 second audio clips using individual audio files in the `train_short_audio` directory.\n\nYou can either create the clips on the fly or use this dataset. **By using this dataset your input pipeline will be spared of clipping overhead.** \n\nI have used a 16 CPU core, 60 GB RAM virtual instance on GCP to create this dataset. You can use [Multiprocessing 5 Second Audio Clips](https://www.kaggle.com/ayuraj/multiprocessing-5-second-audio-clips) kernel to reproduce the result. It uses multiprocessing to create the dataset. The resulting dataset is 20.5 GB in size. \n\n**Note**: I tried to upload the dataset as a Kaggle dataset but all attempts failed - Kaggle API as well as manually uploading the `.zip` files. \n\nI have created a new [W&B team account](https://docs.wandb.ai/ref/app/features/teams) to upload the entire dataset as an [Artifact](https://docs.wandb.ai/guides/artifacts). If you want to download this dataset please use the code snippet below:\n\n```\nimport wandb\nrun = wandb.init()\nartifact = run.use_artifact('birdclef21/audio_clips_zip/bird_audio_clips_zip:v0', type='dataset')\nartifact_dir = artifact.download()\n```\n\n**Note**: You will require a [W&B account](https://wandb.ai/site). \n**Note**: You will find 398 zipped files where each zipped file is associated to a bird species. I have mistakenly uploaded `dataset-metadata.json` file. Please ignore it. \n**Note**: The download might take time according to the internet speed available to your system.\n\nThese are raw clips. It's up to you to do pre-processing, filter out unwanted clips, etc. I will share some of my strategies in near future. \n\nI have also created `train_clips.csv` file that contains metadata for each clip. To download use the code snippet below:\n\n```\nimport wandb\nrun = wandb.init()\nartifact = run.use_artifact('birdclef21/audio_clips_zip/bird_audio_clip_csv:v0', type='dataset')\nartifact_dir = artifact.download()\n```\n\nDo let me know if you face any difficulty downloading the files.",
      "votes": null
    },
    {
      "id": "1277319",
      "postDate": "04/18/2021 16:14:36",
      "content": "<p>I believe the size limit on kaggle datasets is 20GB - likely this is the reason you cannot load the data set.  I used some different logic and my 5 sec set of clips is 27GB.  </p>\n<p>I am planning to run mine again (I do it on local PC and takes almost a day on the machine I used).  I am going to change my logic to extract one less 5 second increment (per file) and see if that reduction of 68K clips will be enough to get me down to less than 20GB.</p>",
      "rawMarkdown": "I believe the size limit on kaggle datasets is 20GB - likely this is the reason you cannot load the data set.  I used some different logic and my 5 sec set of clips is 27GB.  \n\nI am planning to run mine again (I do it on local PC and takes almost a day on the machine I used).  I am going to change my logic to extract one less 5 second increment (per file) and see if that reduction of 68K clips will be enough to get me down to less than 20GB.",
      "votes": null
    },
    {
      "id": "1277733",
      "postDate": "04/19/2021 06:27:17",
      "content": "<p>Good luck with your endeavor. :)</p>\n<p>Try to use multiprocessing to utilize all the available CPU cores if not already. Also if you think that my data creation logic is okay you can simply use the download the dataset and move on.</p>\n<p>Also I don't think it was the size limit due to which I couldn't upload the dataset. I tried uploading individual <code>.zip</code> files (say 40 MB in size) but to no luck.</p>",
      "rawMarkdown": "Good luck with your endeavor. :)\n\nTry to use multiprocessing to utilize all the available CPU cores if not already. Also if you think that my data creation logic is okay you can simply use the download the dataset and move on.\n\nAlso I don't think it was the size limit due to which I couldn't upload the dataset. I tried uploading individual `.zip` files (say 40 MB in size) but to no luck.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1277319,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "04/18/2021 16:14:36",
      "content": "<p>I believe the size limit on kaggle datasets is 20GB - likely this is the reason you cannot load the data set.  I used some different logic and my 5 sec set of clips is 27GB.  </p>\n<p>I am planning to run mine again (I do it on local PC and takes almost a day on the machine I used).  I am going to change my logic to extract one less 5 second increment (per file) and see if that reduction of 68K clips will be enough to get me down to less than 20GB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1277733,
          "author_name": "ayuraj",
          "author_url": "",
          "post_date": "04/19/2021 06:27:17",
          "content": "<p>Good luck with your endeavor. :)</p>\n<p>Try to use multiprocessing to utilize all the available CPU cores if not already. Also if you think that my data creation logic is okay you can simply use the download the dataset and move on.</p>\n<p>Also I don't think it was the size limit due to which I couldn't upload the dataset. I tried uploading individual <code>.zip</code> files (say 40 MB in size) but to no luck.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1277020": "The `train_short_audio` can be used to train a multi-label classifier. However, the model needs to be trained on 5 second audio clips. I went through the trouble of creating 5 second audio clips using individual audio files in the `train_short_audio` directory.\n\nYou can either create the clips on the fly or use this dataset. **By using this dataset your input pipeline will be spared of clipping overhead.** \n\nI have used a 16 CPU core, 60 GB RAM virtual instance on GCP to create this dataset. You can use [Multiprocessing 5 Second Audio Clips](https://www.kaggle.com/ayuraj/multiprocessing-5-second-audio-clips) kernel to reproduce the result. It uses multiprocessing to create the dataset. The resulting dataset is 20.5 GB in size. \n\n**Note**: I tried to upload the dataset as a Kaggle dataset but all attempts failed - Kaggle API as well as manually uploading the `.zip` files. \n\nI have created a new [W&B team account](https://docs.wandb.ai/ref/app/features/teams) to upload the entire dataset as an [Artifact](https://docs.wandb.ai/guides/artifacts). If you want to download this dataset please use the code snippet below:\n\n```\nimport wandb\nrun = wandb.init()\nartifact = run.use_artifact('birdclef21/audio_clips_zip/bird_audio_clips_zip:v0', type='dataset')\nartifact_dir = artifact.download()\n```\n\n**Note**: You will require a [W&B account](https://wandb.ai/site). \n**Note**: You will find 398 zipped files where each zipped file is associated to a bird species. I have mistakenly uploaded `dataset-metadata.json` file. Please ignore it. \n**Note**: The download might take time according to the internet speed available to your system.\n\nThese are raw clips. It's up to you to do pre-processing, filter out unwanted clips, etc. I will share some of my strategies in near future. \n\nI have also created `train_clips.csv` file that contains metadata for each clip. To download use the code snippet below:\n\n```\nimport wandb\nrun = wandb.init()\nartifact = run.use_artifact('birdclef21/audio_clips_zip/bird_audio_clip_csv:v0', type='dataset')\nartifact_dir = artifact.download()\n```\n\nDo let me know if you face any difficulty downloading the files.",
    "1277319": "I believe the size limit on kaggle datasets is 20GB - likely this is the reason you cannot load the data set.  I used some different logic and my 5 sec set of clips is 27GB.  \n\nI am planning to run mine again (I do it on local PC and takes almost a day on the machine I used).  I am going to change my logic to extract one less 5 second increment (per file) and see if that reduction of 68K clips will be enough to get me down to less than 20GB.",
    "1277733": "Good luck with your endeavor. :)\n\nTry to use multiprocessing to utilize all the available CPU cores if not already. Also if you think that my data creation logic is okay you can simply use the download the dataset and move on.\n\nAlso I don't think it was the size limit due to which I couldn't upload the dataset. I tried uploading individual `.zip` files (say 40 MB in size) but to no luck."
  },
  "source": "meta"
}