{
  "id": 159943,
  "title": "resampled and reorganized train data ready for download",
  "url": "/competitions/birdsong-recognition/discussion/159943",
  "author_name": "Radek Osmulski",
  "post_date": "2020-06-19T08:48:09.182000",
  "votes": 50,
  "comment_count": 13,
  "views": 0,
  "content": "<p>The train data is organized in directories with ebird_code as the name. Each directory contains quite a few files at various sample rates.</p>\n\n<p>This is not easy to work with. Also, very importantly, resampling is an expensive operation, meaning it will require quite a few CPU cycles and might slow down the training. It makes sense to offload this expensive operation and and perform it once before training.</p>\n\n<p>I went through all the dirs, resampled the files and concatenated the recordings so that for each ebird_code we get a single wav file. This IMHO should make it easier to work with the data, though we are trading some minimum amount of granularity in creating our datasets for convenience - I feel this is the right way to go though.</p>\n\n<p>These are the sample rates of the original recordings:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F5920c73254bb57caa5a651c9cc978c1d%2Fsrs.png?generation=1592556069054901&amp;alt=media\" alt=\"\"></p>\n\n<p>I opted to go with a sample rate of 48kHz.</p>\n\n<p>Here is the <a href=\"https://storage.googleapis.com/birdcall_competition/train.zip\">link</a> to download the files. The link points to a zip file that is just under 80 GB (reason being I went with <code>wav</code> instead of <code>mp3</code>). After unpacking it should be around 107 GB worth of data.</p>\n\n<p>This is the code I used for the processing:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F77d7ad1a8476d92876736910199d7c51%2Fprocessing.png?generation=1592556211037800&amp;alt=media\" alt=\"\"></p>\n\n<p>It took nearly 2 hrs on a GCP vm with 96 vCPUs. </p>\n\n<p>Currently in the process of building datasets and training - will be sharing the code when I have it ready 😊</p>\n\n<p><strong>UPDATE:</strong> Please check the discussion below, I resampled the files to 32kHz and reuploaded them to the same location.</p>",
  "messages": [
    {
      "id": 892911,
      "postDate": "2020-06-19T08:48:09.183Z",
      "content": "<p>The train data is organized in directories with ebird_code as the name. Each directory contains quite a few files at various sample rates.</p>\n\n<p>This is not easy to work with. Also, very importantly, resampling is an expensive operation, meaning it will require quite a few CPU cycles and might slow down the training. It makes sense to offload this expensive operation and and perform it once before training.</p>\n\n<p>I went through all the dirs, resampled the files and concatenated the recordings so that for each ebird_code we get a single wav file. This IMHO should make it easier to work with the data, though we are trading some minimum amount of granularity in creating our datasets for convenience - I feel this is the right way to go though.</p>\n\n<p>These are the sample rates of the original recordings:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F5920c73254bb57caa5a651c9cc978c1d%2Fsrs.png?generation=1592556069054901&amp;alt=media\" alt=\"\"></p>\n\n<p>I opted to go with a sample rate of 48kHz.</p>\n\n<p>Here is the <a href=\"https://storage.googleapis.com/birdcall_competition/train.zip\">link</a> to download the files. The link points to a zip file that is just under 80 GB (reason being I went with <code>wav</code> instead of <code>mp3</code>). After unpacking it should be around 107 GB worth of data.</p>\n\n<p>This is the code I used for the processing:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F77d7ad1a8476d92876736910199d7c51%2Fprocessing.png?generation=1592556211037800&amp;alt=media\" alt=\"\"></p>\n\n<p>It took nearly 2 hrs on a GCP vm with 96 vCPUs. </p>\n\n<p>Currently in the process of building datasets and training - will be sharing the code when I have it ready 😊</p>\n\n<p><strong>UPDATE:</strong> Please check the discussion below, I resampled the files to 32kHz and reuploaded them to the same location.</p>",
      "rawMarkdown": "The train data is organized in directories with ebird_code as the name. Each directory contains quite a few files at various sample rates.\n\nThis is not easy to work with. Also, very importantly, resampling is an expensive operation, meaning it will require quite a few CPU cycles and might slow down the training. It makes sense to offload this expensive operation and and perform it once before training.\n\nI went through all the dirs, resampled the files and concatenated the recordings so that for each ebird_code we get a single wav file. This IMHO should make it easier to work with the data, though we are trading some minimum amount of granularity in creating our datasets for convenience - I feel this is the right way to go though.\n\nThese are the sample rates of the original recordings:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F5920c73254bb57caa5a651c9cc978c1d%2Fsrs.png?generation=1592556069054901&amp;alt=media)\n\nI opted to go with a sample rate of 48kHz.\n\nHere is the [link](https://storage.googleapis.com/birdcall_competition/train.zip) to download the files. The link points to a zip file that is just under 80 GB (reason being I went with `wav` instead of `mp3`). After unpacking it should be around 107 GB worth of data.\n\nThis is the code I used for the processing:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F77d7ad1a8476d92876736910199d7c51%2Fprocessing.png?generation=1592556211037800&amp;alt=media)\n\nIt took nearly 2 hrs on a GCP vm with 96 vCPUs. \n\nCurrently in the process of building datasets and training - will be sharing the code when I have it ready 😊\n\n**UPDATE:** Please check the discussion below, I resampled the files to 32kHz and reuploaded them to the same location.",
      "votes": 50
    },
    {
      "id": 892974,
      "postDate": "2020-06-19T09:43:16.370Z",
      "content": "<p>Thanks <a href=\"/radek1\">@radek1</a> great work! My only concern is if this process is too slow for test data since we have the kernel time limitation. I'm assuming the test data may also have different sampling rates.</p>",
      "rawMarkdown": "Thanks @radek1 great work! My only concern is if this process is too slow for test data since we have the kernel time limitation. I'm assuming the test data may also have different sampling rates.",
      "votes": 3,
      "replies": [
        {
          "id": 892978,
          "postDate": "2020-06-19T09:46:08.827Z",
          "content": "<p>I also did resampling on test set to 32 kHz and it took around 20 - 30 minutes (srry I didn't measure the exact time)</p>",
          "rawMarkdown": "I also did resampling on test set to 32 kHz and it took around 20 - 30 minutes (srry I didn't measure the exact time)",
          "votes": 3
        },
        {
          "id": 892992,
          "postDate": "2020-06-19T10:05:45.017Z",
          "content": "<p>That is a good point <a href=\"/mnpinto\">@mnpinto</a>! The big question is how many cores does the kernel use when performing inference? My guess is its probably 2 or 4 so not really a whole lot of room for spreading the workload across multiple cores...</p>\n\n<p><a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> did you use librosa for the resampling?</p>\n\n<p>I'm thinking at ~30 mins that is okay, especially if we can figure out how to perform the resampling before running whatever amount of models we will on the data. </p>",
          "rawMarkdown": "That is a good point @mnpinto! The big question is how many cores does the kernel use when performing inference? My guess is its probably 2 or 4 so not really a whole lot of room for spreading the workload across multiple cores...\n\n@hidehisaarai1213 did you use librosa for the resampling?\n\nI'm thinking at ~30 mins that is okay, especially if we can figure out how to perform the resampling before running whatever amount of models we will on the data. ",
          "votes": 1
        },
        {
          "id": 893003,
          "postDate": "2020-06-19T10:14:40.087Z",
          "content": "<blockquote>\n  <p>did you use librosa for the resampling?</p>\n</blockquote>\n\n<p>Yep. I forgot to note that I didn't use multi-processing, so it can be faster.</p>",
          "rawMarkdown": "&gt; did you use librosa for the resampling?\n\nYep. I forgot to note that I didn't use multi-processing, so it can be faster.",
          "votes": 1
        },
        {
          "id": 893049,
          "postDate": "2020-06-19T10:42:30.417Z",
          "content": "<p>The test data should be sampled at 32 kHz. It might make sense to use that for training, too. Librosa supports different resampling modes, I would recommend 'kaiser_fast' instead of the default ('kaiser_best') to speed things up.</p>",
          "rawMarkdown": "The test data should be sampled at 32 kHz. It might make sense to use that for training, too. Librosa supports different resampling modes, I would recommend 'kaiser_fast' instead of the default ('kaiser_best') to speed things up.",
          "votes": 35
        },
        {
          "id": 893550,
          "postDate": "2020-06-19T17:57:09.223Z",
          "content": "<p>Thank you so much for your comment <a href=\"/stefankahl\">@stefankahl</a>, this is a super useful piece of information!</p>\n\n<p>I resampled the data to 32 kHz and reuploaded it, but as I started training a model on this, it turns out that maybe my choice of concatenating all the recordings together wasn't the best... it would make sense to make sure the data in the validation set doesn't come from the same recordings that appear in the train set...</p>\n\n<p>Anyhow, this will be something to address in a subsequent iteration once I get the whole pipeline running from training a model locally all the way to making a submission on kaggle. Getting this end to end pipeline working is what I am focusing on atm.</p>\n\n<p>Thanks again for the piece of information that you shared with us!</p>",
          "rawMarkdown": "Thank you so much for your comment @stefankahl, this is a super useful piece of information!\n\nI resampled the data to 32 kHz and reuploaded it, but as I started training a model on this, it turns out that maybe my choice of concatenating all the recordings together wasn't the best... it would make sense to make sure the data in the validation set doesn't come from the same recordings that appear in the train set...\n\nAnyhow, this will be something to address in a subsequent iteration once I get the whole pipeline running from training a model locally all the way to making a submission on kaggle. Getting this end to end pipeline working is what I am focusing on atm.\n\nThanks again for the piece of information that you shared with us!",
          "votes": 2
        }
      ]
    },
    {
      "id": 968363,
      "postDate": "2020-08-12T23:44:52.910Z",
      "content": "<p>I'm curious as to why you felt concatenating the audio was best practice? </p>\n<p>Also, a little googling revealed that most birds have an upper range of around 8khz. This means a minimum sampling rate of 22050 would probably suffice to minimally resolve bird calls with frequencies up to 11 kHz, which should cover the vast majority of species (according to a quick google)</p>\n<p>Let me know your thoughts!</p>",
      "rawMarkdown": "I'm curious as to why you felt concatenating the audio was best practice? \n\nAlso, a little googling revealed that most birds have an upper range of around 8khz. This means a minimum sampling rate of 22050 would probably suffice to minimally resolve bird calls with frequencies up to 11 kHz, which should cover the vast majority of species (according to a quick google)\n\nLet me know your thoughts!",
      "votes": 2,
      "replies": [
        {
          "id": 968390,
          "postDate": "2020-08-13T01:02:48.753Z",
          "content": "<p>I think he concatenated the audio because that is what the winner of the 2018 and 2019 competitions did, you can read it in his papers. I think that is why Radek decided to concatenate them.</p>\n<p>The competition organizers said that the test was recorded at 32kHz, so we should prepare our train dataset similarly.</p>",
          "rawMarkdown": "I think he concatenated the audio because that is what the winner of the 2018 and 2019 competitions did, you can read it in his papers. I think that is why Radek decided to concatenate them.\n\nThe competition organizers said that the test was recorded at 32kHz, so we should prepare our train dataset similarly.",
          "votes": 2
        }
      ]
    },
    {
      "id": 971802,
      "postDate": "2020-08-15T22:44:35.910Z",
      "content": "<p>As a non-audio expert could you please clarify: when taking a 44100Hz file and resampling it at 32kHz, aren't you introducing some noise to melspectrogram from quantization? Specifically wouldn't it add disonances frequences arising from disonance between 32kHz and 44100Hz?</p>",
      "rawMarkdown": "As a non-audio expert could you please clarify: when taking a 44100Hz file and resampling it at 32kHz, aren't you introducing some noise to melspectrogram from quantization? Specifically wouldn't it add disonances frequences arising from disonance between 32kHz and 44100Hz?"
    },
    {
      "id": 965129,
      "postDate": "2020-08-10T12:03:00.923Z",
      "content": "<p>i tried to upload the resampled data from your GCP  link(<a href=\"https://storage.googleapis.com/birdcall_competition/train_resampled.zip\">https://storage.googleapis.com/birdcall_competition/train_resampled.zip</a>). But the resampled data seems to be greater than 50GB. SO how to handle this?</p>",
      "rawMarkdown": "i tried to upload the resampled data from your GCP  link(https://storage.googleapis.com/birdcall_competition/train_resampled.zip). But the resampled data seems to be greater than 50GB. SO how to handle this?\n"
    },
    {
      "id": 915553,
      "postDate": "2020-07-04T20:23:32.643Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 894678,
      "postDate": "2020-06-20T16:40:05.903Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 895176,
      "postDate": "2020-06-21T06:50:17.240Z",
      "content": "<p>Thanks <a href=\"/radek1\">@radek1</a> great work</p>",
      "rawMarkdown": "Thanks @radek1 great work",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 892974,
      "author_name": "Miguel Pinto",
      "author_url": "",
      "post_date": "2020-06-19T09:43:16.370000",
      "content": "<p>Thanks <a href=\"/radek1\">@radek1</a> great work! My only concern is if this process is too slow for test data since we have the kernel time limitation. I'm assuming the test data may also have different sampling rates.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 892978,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-06-19T09:46:08.827000",
          "content": "<p>I also did resampling on test set to 32 kHz and it took around 20 - 30 minutes (srry I didn't measure the exact time)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 892992,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2020-06-19T10:05:45.017000",
          "content": "<p>That is a good point <a href=\"/mnpinto\">@mnpinto</a>! The big question is how many cores does the kernel use when performing inference? My guess is its probably 2 or 4 so not really a whole lot of room for spreading the workload across multiple cores...</p>\n\n<p><a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> did you use librosa for the resampling?</p>\n\n<p>I'm thinking at ~30 mins that is okay, especially if we can figure out how to perform the resampling before running whatever amount of models we will on the data. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 893003,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-06-19T10:14:40.087000",
          "content": "<blockquote>\n  <p>did you use librosa for the resampling?</p>\n</blockquote>\n\n<p>Yep. I forgot to note that I didn't use multi-processing, so it can be faster.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 893049,
          "author_name": "Stefan Kahl",
          "author_url": "",
          "post_date": "2020-06-19T10:42:30.417000",
          "content": "<p>The test data should be sampled at 32 kHz. It might make sense to use that for training, too. Librosa supports different resampling modes, I would recommend 'kaiser_fast' instead of the default ('kaiser_best') to speed things up.</p>",
          "votes": 35,
          "replies": []
        },
        {
          "id": 893550,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2020-06-19T17:57:09.223000",
          "content": "<p>Thank you so much for your comment <a href=\"/stefankahl\">@stefankahl</a>, this is a super useful piece of information!</p>\n\n<p>I resampled the data to 32 kHz and reuploaded it, but as I started training a model on this, it turns out that maybe my choice of concatenating all the recordings together wasn't the best... it would make sense to make sure the data in the validation set doesn't come from the same recordings that appear in the train set...</p>\n\n<p>Anyhow, this will be something to address in a subsequent iteration once I get the whole pipeline running from training a model locally all the way to making a submission on kaggle. Getting this end to end pipeline working is what I am focusing on atm.</p>\n\n<p>Thanks again for the piece of information that you shared with us!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 968363,
      "author_name": "Wesley Neill",
      "author_url": "",
      "post_date": "2020-08-12T23:44:52.910000",
      "content": "<p>I'm curious as to why you felt concatenating the audio was best practice? </p>\n<p>Also, a little googling revealed that most birds have an upper range of around 8khz. This means a minimum sampling rate of 22050 would probably suffice to minimally resolve bird calls with frequencies up to 11 kHz, which should cover the vast majority of species (according to a quick google)</p>\n<p>Let me know your thoughts!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 968390,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2020-08-13T01:02:48.753000",
          "content": "<p>I think he concatenated the audio because that is what the winner of the 2018 and 2019 competitions did, you can read it in his papers. I think that is why Radek decided to concatenate them.</p>\n<p>The competition organizers said that the test was recorded at 32kHz, so we should prepare our train dataset similarly.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 971802,
      "author_name": "leodav",
      "author_url": "",
      "post_date": "2020-08-15T22:44:35.910000",
      "content": "<p>As a non-audio expert could you please clarify: when taking a 44100Hz file and resampling it at 32kHz, aren't you introducing some noise to melspectrogram from quantization? Specifically wouldn't it add disonances frequences arising from disonance between 32kHz and 44100Hz?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 965129,
      "author_name": "Aishwarya S",
      "author_url": "",
      "post_date": "2020-08-10T12:03:00.923000",
      "content": "<p>i tried to upload the resampled data from your GCP  link(<a href=\"https://storage.googleapis.com/birdcall_competition/train_resampled.zip\">https://storage.googleapis.com/birdcall_competition/train_resampled.zip</a>). But the resampled data seems to be greater than 50GB. SO how to handle this?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 915553,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-04T20:23:32.643000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 894678,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-20T16:40:05.903000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 895176,
      "author_name": "Satya Muralidhar",
      "author_url": "",
      "post_date": "2020-06-21T06:50:17.240000",
      "content": "<p>Thanks <a href=\"/radek1\">@radek1</a> great work</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "892911": "The train data is organized in directories with ebird_code as the name. Each directory contains quite a few files at various sample rates.\n\nThis is not easy to work with. Also, very importantly, resampling is an expensive operation, meaning it will require quite a few CPU cycles and might slow down the training. It makes sense to offload this expensive operation and and perform it once before training.\n\nI went through all the dirs, resampled the files and concatenated the recordings so that for each ebird_code we get a single wav file. This IMHO should make it easier to work with the data, though we are trading some minimum amount of granularity in creating our datasets for convenience - I feel this is the right way to go though.\n\nThese are the sample rates of the original recordings:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F5920c73254bb57caa5a651c9cc978c1d%2Fsrs.png?generation=1592556069054901&amp;alt=media)\n\nI opted to go with a sample rate of 48kHz.\n\nHere is the [link](https://storage.googleapis.com/birdcall_competition/train.zip) to download the files. The link points to a zip file that is just under 80 GB (reason being I went with `wav` instead of `mp3`). After unpacking it should be around 107 GB worth of data.\n\nThis is the code I used for the processing:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F77d7ad1a8476d92876736910199d7c51%2Fprocessing.png?generation=1592556211037800&amp;alt=media)\n\nIt took nearly 2 hrs on a GCP vm with 96 vCPUs. \n\nCurrently in the process of building datasets and training - will be sharing the code when I have it ready 😊\n\n**UPDATE:** Please check the discussion below, I resampled the files to 32kHz and reuploaded them to the same location.",
    "892974": "Thanks @radek1 great work! My only concern is if this process is too slow for test data since we have the kernel time limitation. I'm assuming the test data may also have different sampling rates.",
    "968363": "I'm curious as to why you felt concatenating the audio was best practice? \n\nAlso, a little googling revealed that most birds have an upper range of around 8khz. This means a minimum sampling rate of 22050 would probably suffice to minimally resolve bird calls with frequencies up to 11 kHz, which should cover the vast majority of species (according to a quick google)\n\nLet me know your thoughts!",
    "971802": "As a non-audio expert could you please clarify: when taking a 44100Hz file and resampling it at 32kHz, aren't you introducing some noise to melspectrogram from quantization? Specifically wouldn't it add disonances frequences arising from disonance between 32kHz and 44100Hz?",
    "965129": "i tried to upload the resampled data from your GCP  link(https://storage.googleapis.com/birdcall_competition/train_resampled.zip). But the resampled data seems to be greater than 50GB. SO how to handle this?\n",
    "915553": "",
    "894678": "",
    "895176": "Thanks @radek1 great work"
  }
}