{
  "id": 89157,
  "title": "Downsampling",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/89157",
  "author_name": "",
  "post_date": "2019-04-11T21:12:44.337331500Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Does anyone have suggestions for preprocessing?</p>\n\n<p>I am downsampling the audio to ~15khz using <a href=\"https://docs.scipy.org/doc/scipy/reference/signal.html\">scipy signal</a> to speed up training.  However, it seems to be taking about two days on my poor i3 just for the curated set.  I was wondering how low you could really go with the sample rate without losing too much important information, and if there might be any GPU accelerated or otherwise optimized libraries for doing it?</p>\n\n<p>I'll respond back if I find anything myself!  :)</p>",
  "messages": [
    {
      "id": "514695",
      "postDate": "04/11/2019 21:12:44",
      "content": "<p>Does anyone have suggestions for preprocessing?</p>\n\n<p>I am downsampling the audio to ~15khz using <a href=\"https://docs.scipy.org/doc/scipy/reference/signal.html\">scipy signal</a> to speed up training.  However, it seems to be taking about two days on my poor i3 just for the curated set.  I was wondering how low you could really go with the sample rate without losing too much important information, and if there might be any GPU accelerated or otherwise optimized libraries for doing it?</p>\n\n<p>I'll respond back if I find anything myself!  :)</p>",
      "rawMarkdown": "Does anyone have suggestions for preprocessing?\n\nI am downsampling the audio to ~15khz using [scipy signal](https://docs.scipy.org/doc/scipy/reference/signal.html) to speed up training.  However, it seems to be taking about two days on my poor i3 just for the curated set.  I was wondering how low you could really go with the sample rate without losing too much important information, and if there might be any GPU accelerated or otherwise optimized libraries for doing it?\n\nI'll respond back if I find anything myself!  :)",
      "votes": null
    },
    {
      "id": "514756",
      "postDate": "04/11/2019 22:20:41",
      "content": "<p>Hi <a href=\"/jamesdonconley\">@jamesdonconley</a> I think the best is try everything on Kaggle kernel or Google colab or any other services for both preprocessing and training. Then you might be free from resource issues, I personally found it is possible thanks to the restriction of this competition rules.\nYou can preprocess data in one kernel, then you can either download processed data to train locally or continue using it on another kernel sessions.\nRegarding sampling rate, even though using higher rate, we could make it lower sized. I’m basically using 44.1kHz converted to 128 freq bins mel-spectrogram, but it can be size of 64 or the less if I want to. I think downsampling also takes time, then I think selecting sampling rate is free of choices for us instead of restricted by computational resource.</p>",
      "rawMarkdown": "Hi @jamesdonconley I think the best is try everything on Kaggle kernel or Google colab or any other services for both preprocessing and training. Then you might be free from resource issues, I personally found it is possible thanks to the restriction of this competition rules.\nYou can preprocess data in one kernel, then you can either download processed data to train locally or continue using it on another kernel sessions.\nRegarding sampling rate, even though using higher rate, we could make it lower sized. I’m basically using 44.1kHz converted to 128 freq bins mel-spectrogram, but it can be size of 64 or the less if I want to. I think downsampling also takes time, then I think selecting sampling rate is free of choices for us instead of restricted by computational resource.",
      "votes": null
    },
    {
      "id": "514799",
      "postDate": "04/11/2019 23:34:36",
      "content": "<p>Hi, You only need preprocessing one time, save in files or memory and use a train generator. With a GPU training is much faster than preprocessing with a cpu.</p>",
      "rawMarkdown": "Hi, You only need preprocessing one time, save in files or memory and use a train generator. With a GPU training is much faster than preprocessing with a cpu.",
      "votes": null
    },
    {
      "id": "514820",
      "postDate": "04/12/2019 00:16:40",
      "content": "<p>I actually found that it took longer to downsample versus just using the original 44.1k and I also got a boost in score when I used 44.1k. Like others mentioned you can just preprocess everything once prior to training, as it would slow down training way too much to do this each time. Convert all the wav files to mel-spectogram and save the arrays to disk to be loaded in batches by a generator during training.</p>",
      "rawMarkdown": "I actually found that it took longer to downsample versus just using the original 44.1k and I also got a boost in score when I used 44.1k. Like others mentioned you can just preprocess everything once prior to training, as it would slow down training way too much to do this each time. Convert all the wav files to mel-spectogram and save the arrays to disk to be loaded in batches by a generator during training.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 514756,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "04/11/2019 22:20:41",
      "content": "<p>Hi <a href=\"/jamesdonconley\">@jamesdonconley</a> I think the best is try everything on Kaggle kernel or Google colab or any other services for both preprocessing and training. Then you might be free from resource issues, I personally found it is possible thanks to the restriction of this competition rules.\nYou can preprocess data in one kernel, then you can either download processed data to train locally or continue using it on another kernel sessions.\nRegarding sampling rate, even though using higher rate, we could make it lower sized. I’m basically using 44.1kHz converted to 128 freq bins mel-spectrogram, but it can be size of 64 or the less if I want to. I think downsampling also takes time, then I think selecting sampling rate is free of choices for us instead of restricted by computational resource.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 514799,
      "author_name": "pierretisseur",
      "author_url": "",
      "post_date": "04/11/2019 23:34:36",
      "content": "<p>Hi, You only need preprocessing one time, save in files or memory and use a train generator. With a GPU training is much faster than preprocessing with a cpu.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 514820,
      "author_name": "jamesrequa",
      "author_url": "",
      "post_date": "04/12/2019 00:16:40",
      "content": "<p>I actually found that it took longer to downsample versus just using the original 44.1k and I also got a boost in score when I used 44.1k. Like others mentioned you can just preprocess everything once prior to training, as it would slow down training way too much to do this each time. Convert all the wav files to mel-spectogram and save the arrays to disk to be loaded in batches by a generator during training.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "514695": "Does anyone have suggestions for preprocessing?\n\nI am downsampling the audio to ~15khz using [scipy signal](https://docs.scipy.org/doc/scipy/reference/signal.html) to speed up training.  However, it seems to be taking about two days on my poor i3 just for the curated set.  I was wondering how low you could really go with the sample rate without losing too much important information, and if there might be any GPU accelerated or otherwise optimized libraries for doing it?\n\nI'll respond back if I find anything myself!  :)",
    "514756": "Hi @jamesdonconley I think the best is try everything on Kaggle kernel or Google colab or any other services for both preprocessing and training. Then you might be free from resource issues, I personally found it is possible thanks to the restriction of this competition rules.\nYou can preprocess data in one kernel, then you can either download processed data to train locally or continue using it on another kernel sessions.\nRegarding sampling rate, even though using higher rate, we could make it lower sized. I’m basically using 44.1kHz converted to 128 freq bins mel-spectrogram, but it can be size of 64 or the less if I want to. I think downsampling also takes time, then I think selecting sampling rate is free of choices for us instead of restricted by computational resource.",
    "514799": "Hi, You only need preprocessing one time, save in files or memory and use a train generator. With a GPU training is much faster than preprocessing with a cpu.",
    "514820": "I actually found that it took longer to downsample versus just using the original 44.1k and I also got a boost in score when I used 44.1k. Like others mentioned you can just preprocess everything once prior to training, as it would slow down training way too much to do this each time. Convert all the wav files to mel-spectogram and save the arrays to disk to be loaded in batches by a generator during training."
  },
  "source": "meta"
}