{
  "id": 164197,
  "title": "Resampled Train Audio on Kaggle Dataset",
  "url": "/competitions/birdsong-recognition/discussion/164197",
  "author_name": "Tawara",
  "post_date": "2020-07-05T06:35:42.484000",
  "votes": 115,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hi, Kagglers.</p>\n\n<p>I'm new comer in this competition, and see that resampling recordings is important in several notebooks and discussions.</p>\n\n<p><a href=\"/radek1\">@radek1</a> kindly provide resampled data (see <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159943\">this topic</a>), but I cannot access it now and it is hard to use in kaggle notebooks.</p>\n\n<p>I decided to make resampled data for the above reason and this is now available on Kaggle Dataset.\n(Preprocessing took less than an hour, but uploading each part took about 3 hours, totally 15 hours😓 )</p>\n\n<p>Because of file size (72 GB in total), resampled data is split into five parts:  </p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-00\">birdsong resampled train audio 00 (a ~ b)</a></li>\n<li><a href=\"https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-01\">birdsong resampled train audio 01 (c ~ f)</a></li>\n<li><a href=\"https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-02\">birdsong resampled train audio 02 (g ~ m)</a></li>\n<li><a href=\"https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-03\">birdsong  resampled train audio 03 (n ~ r)</a></li>\n<li><a href=\"https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-04\">birdsong resampled train audio 04 (s ~ y)</a></li>\n</ul>\n\n<p>I'm afraid that I have made some mistakes in that process. Now I'm working on making baseline to make sure the data is okay.</p>\n\n<p>I used preprocessing code shared in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/162304\">this topic</a>. Great thanks to <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> for sharing baseline code.</p>\n\n<h3>Update</h3>\n\n<p>I shared training &amp; inference notebook (see <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/164345\">this toric</a>) .</p>\n\n<p>I think the data is ok from the result (Public 2nd (0.56) now).\n<br>\nI have realized that <a href=\"/radek1\">@radek1</a> shares download link in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/160321\">another topic</a>. Thanks!</p>",
  "messages": [
    {
      "id": 915852,
      "postDate": "2020-07-05T06:35:42.483Z",
      "content": "<p>Hi, Kagglers.</p>\n\n<p>I'm new comer in this competition, and see that resampling recordings is important in several notebooks and discussions.</p>\n\n<p><a href=\"/radek1\">@radek1</a> kindly provide resampled data (see <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159943\">this topic</a>), but I cannot access it now and it is hard to use in kaggle notebooks.</p>\n\n<p>I decided to make resampled data for the above reason and this is now available on Kaggle Dataset.\n(Preprocessing took less than an hour, but uploading each part took about 3 hours, totally 15 hours😓 )</p>\n\n<p>Because of file size (72 GB in total), resampled data is split into five parts:  </p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-00\">birdsong resampled train audio 00 (a ~ b)</a></li>\n<li><a href=\"https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-01\">birdsong resampled train audio 01 (c ~ f)</a></li>\n<li><a href=\"https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-02\">birdsong resampled train audio 02 (g ~ m)</a></li>\n<li><a href=\"https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-03\">birdsong  resampled train audio 03 (n ~ r)</a></li>\n<li><a href=\"https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-04\">birdsong resampled train audio 04 (s ~ y)</a></li>\n</ul>\n\n<p>I'm afraid that I have made some mistakes in that process. Now I'm working on making baseline to make sure the data is okay.</p>\n\n<p>I used preprocessing code shared in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/162304\">this topic</a>. Great thanks to <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> for sharing baseline code.</p>\n\n<h3>Update</h3>\n\n<p>I shared training &amp; inference notebook (see <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/164345\">this toric</a>) .</p>\n\n<p>I think the data is ok from the result (Public 2nd (0.56) now).\n<br>\nI have realized that <a href=\"/radek1\">@radek1</a> shares download link in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/160321\">another topic</a>. Thanks!</p>",
      "rawMarkdown": "Hi, Kagglers.\n\nI'm new comer in this competition, and see that resampling recordings is important in several notebooks and discussions.\n\n@radek1 kindly provide resampled data (see [this topic](https://www.kaggle.com/c/birdsong-recognition/discussion/159943)), but I cannot access it now and it is hard to use in kaggle notebooks.\n\nI decided to make resampled data for the above reason and this is now available on Kaggle Dataset.\n(Preprocessing took less than an hour, but uploading each part took about 3 hours, totally 15 hours😓 )\n\nBecause of file size (72 GB in total), resampled data is split into five parts:  \n \n* [birdsong resampled train audio 00 (a ~ b)](https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-00)\n* [birdsong resampled train audio 01 (c ~ f)](https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-01)\n* [birdsong resampled train audio 02 (g ~ m)](https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-02)\n* [birdsong  resampled train audio 03 (n ~ r)](https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-03)\n* [birdsong resampled train audio 04 (s ~ y)](https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-04)\n\nI'm afraid that I have made some mistakes in that process. Now I'm working on making baseline to make sure the data is okay.\n\nI used preprocessing code shared in [this topic](https://www.kaggle.com/c/birdsong-recognition/discussion/162304). Great thanks to @hidehisaarai1213 for sharing baseline code.\n\n### Update\n\nI shared training &amp; inference notebook (see [this toric](https://www.kaggle.com/c/birdsong-recognition/discussion/164345)) .\n\nI think the data is ok from the result (Public 2nd (0.56) now).\n<br>\nI have realized that @radek1 shares download link in [another topic](https://www.kaggle.com/c/birdsong-recognition/discussion/160321). Thanks!",
      "votes": 115
    },
    {
      "id": 951910,
      "postDate": "2020-07-30T13:48:44.963Z",
      "content": "<p>Thanks for sharing this great data and baseline!\nIs the <code>train_mod.csv</code> file in each part the same?</p>",
      "rawMarkdown": "Thanks for sharing this great data and baseline!\nIs the `train_mod.csv` file in each part the same?",
      "votes": 2,
      "replies": [
        {
          "id": 952270,
          "postDate": "2020-07-30T19:06:27.333Z",
          "content": "<p>Yes, I put the same <code>.csv</code> file in each part.</p>",
          "rawMarkdown": "Yes, I put the same `.csv` file in each part.",
          "votes": 2
        },
        {
          "id": 952583,
          "postDate": "2020-07-31T04:10:55.320Z",
          "content": "<p>Got it. Thanks again.</p>",
          "rawMarkdown": "Got it. Thanks again."
        },
        {
          "id": 960721,
          "postDate": "2020-08-06T16:12:16.063Z",
          "content": "<p>And the <code>train_mod.csv</code> file is, in turn, the same as the original <code>train.csv</code> file, but with additional columns:</p>\n\n<p><code>\n['resampled_sampling_rate', 'resampled_filename', 'resampled_channels']\n</code></p>",
          "rawMarkdown": "And the `train_mod.csv` file is, in turn, the same as the original `train.csv` file, but with additional columns:\n\n```\n['resampled_sampling_rate', 'resampled_filename', 'resampled_channels']\n```"
        }
      ]
    },
    {
      "id": 929342,
      "postDate": "2020-07-14T15:56:38.667Z",
      "content": "<p>Thanks for your work, Tawara! Could you please share the code you used to resample the data or did you use the one provided <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159943\">here</a>? </p>",
      "rawMarkdown": "Thanks for your work, Tawara! Could you please share the code you used to resample the data or did you use the one provided [here](https://www.kaggle.com/c/birdsong-recognition/discussion/159943)? ",
      "votes": 2,
      "replies": [
        {
          "id": 929911,
          "postDate": "2020-07-15T04:17:56.887Z",
          "content": "<p>I used a code based on the one shared in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/162304\">this topic</a>.</p>\n\n<p>An example is as follows (you have to change paths  depending on your environment):</p>\n\n<p>```python\nimport warnings\nfrom pathlib import Path\nfrom joblib import delayed, Parallel</p>\n\n<p>import librosa\nimport audioread\nimport soundfile as sf</p>\n\n<p>import pandas as pd</p>\n\n<p>TRAIN_AUDIO_DIR = \"path/to/train_audio\"\nTRAIN_RESAMPLED_DIR = \"path/to/train_audio_resampled\"</p>\n\n<p>TARGET_SR = 32000\nNUM_THREAD = 8  # for joblib.Parallel</p>\n\n<h1># read train.csv</h1>\n\n<p>train = pd.read_csv(\"path/to/train.csv\")</p>\n\n<h1># extract \"ebird_code\" and  \"filename\"</h1>\n\n<p>train_audio_infos = train[[\"ebird_code\", \"filename\"]].values.tolist()</p>\n\n<h1># make directories for saving resampled audio</h1>\n\n<p>TRAIN_RESAMPLED_DIR.mkdir(parents=True)\nfor ebird_code in train.ebird_code.unique():\n    ebird_dir = TRAIN_RESAMPLED_DIR / ebird_code\n    ebird_dir.mkdir()</p>\n\n<h1># define resampling function</h1>\n\n<p>warnings.simplefilter(\"ignore\")\ndef resample(ebird_code: str, filename: str, target_sr: int): <br>\n    audio_dir = TRAIN_AUDIO_DIR\n    resample_dir = TRAIN_RESAMPLED_DIR\n    ebird_dir = resample_dir / ebird_code</p>\n\n<pre><code>try:\n    y, _ = librosa.load(\n        audio_dir / ebird_code / filename,\n        sr=target_sr, mono=True, res_type=\"kaiser_fast\")\n\n    filename = filename.replace(\".mp3\", \".wav\")\n    sf.write(ebird_dir / filename, y, samplerate=target_sr)\n    return \"OK\"\nexcept Exception as e:\n    with open(resample_dir / \"skipped.txt\", \"a\") as f:\n        file_path = str(audio_dir / ebird_code / filename)\n        f.write(file_path + \"\\n\")\n    return str(e)\n</code></pre>\n\n<h1># resample and save audio using Parallel</h1>\n\n<p>msg_list = Parallel(n_jobs=NUM_THREAD, verbose=1)(\n    delayed(resample)(ebird_code, file_name, TARGET_SR) for ebird_code, file_name in train_audio_infos)</p>\n\n<h1># add information of resampled audios to train.csv</h1>\n\n<p>train[\"resampled_sampling_rate\"] = TARGET_SR\ntrain[\"resampled_filename\"] = train[\"filename\"].map(\n    lambda x: x.replace(\".mp3\", \".wav\"))\ntrain[\"resampled_channels\"] = \"1 (mono)\"</p>\n\n<p>train.to_csv(TRAIN_RESAMPLED_DIR / \"train_mod.csv\", index=False)\n```</p>",
          "rawMarkdown": "I used a code based on the one shared in [this topic](https://www.kaggle.com/c/birdsong-recognition/discussion/162304).\n\nAn example is as follows (you have to change paths  depending on your environment):\n\n```python\nimport warnings\nfrom pathlib import Path\nfrom joblib import delayed, Parallel\n\nimport librosa\nimport audioread\nimport soundfile as sf\n\nimport pandas as pd\n\nTRAIN_AUDIO_DIR = \"path/to/train_audio\"\nTRAIN_RESAMPLED_DIR = \"path/to/train_audio_resampled\"\n\nTARGET_SR = 32000\nNUM_THREAD = 8  # for joblib.Parallel\n\n# # read train.csv\ntrain = pd.read_csv(\"path/to/train.csv\")\n\n# # extract \"ebird_code\" and  \"filename\"\ntrain_audio_infos = train[[\"ebird_code\", \"filename\"]].values.tolist()\n\n# # make directories for saving resampled audio\nTRAIN_RESAMPLED_DIR.mkdir(parents=True)\nfor ebird_code in train.ebird_code.unique():\n    ebird_dir = TRAIN_RESAMPLED_DIR / ebird_code\n    ebird_dir.mkdir()\n\n# # define resampling function\nwarnings.simplefilter(\"ignore\")\ndef resample(ebird_code: str, filename: str, target_sr: int):    \n    audio_dir = TRAIN_AUDIO_DIR\n    resample_dir = TRAIN_RESAMPLED_DIR\n    ebird_dir = resample_dir / ebird_code\n    \n    try:\n        y, _ = librosa.load(\n            audio_dir / ebird_code / filename,\n            sr=target_sr, mono=True, res_type=\"kaiser_fast\")\n\n        filename = filename.replace(\".mp3\", \".wav\")\n        sf.write(ebird_dir / filename, y, samplerate=target_sr)\n        return \"OK\"\n    except Exception as e:\n        with open(resample_dir / \"skipped.txt\", \"a\") as f:\n            file_path = str(audio_dir / ebird_code / filename)\n            f.write(file_path + \"\\n\")\n        return str(e)\n\n# # resample and save audio using Parallel\nmsg_list = Parallel(n_jobs=NUM_THREAD, verbose=1)(\n    delayed(resample)(ebird_code, file_name, TARGET_SR) for ebird_code, file_name in train_audio_infos)\n\n# # add information of resampled audios to train.csv\ntrain[\"resampled_sampling_rate\"] = TARGET_SR\ntrain[\"resampled_filename\"] = train[\"filename\"].map(\n    lambda x: x.replace(\".mp3\", \".wav\"))\ntrain[\"resampled_channels\"] = \"1 (mono)\"\n\ntrain.to_csv(TRAIN_RESAMPLED_DIR / \"train_mod.csv\", index=False)\n```\n",
          "votes": 6
        },
        {
          "id": 930099,
          "postDate": "2020-07-15T07:39:03.617Z",
          "content": "<p>I see, thanks a lot!</p>",
          "rawMarkdown": "I see, thanks a lot!"
        }
      ]
    },
    {
      "id": 916017,
      "postDate": "2020-07-05T09:28:44.257Z",
      "content": "<p>Seems the link to g - m dataset is wrong, but thanks anyway!</p>",
      "rawMarkdown": "Seems the link to g - m dataset is wrong, but thanks anyway!\n",
      "votes": 2,
      "replies": [
        {
          "id": 916030,
          "postDate": "2020-07-05T09:36:57.977Z",
          "content": "<p>Sorry, I correct it and now it is accessible.</p>\n\n<p>Thank you for reply 😃 </p>",
          "rawMarkdown": " Sorry, I correct it and now it is accessible.\n\nThank you for reply 😃 ",
          "votes": 1
        },
        {
          "id": 999748,
          "postDate": "2020-09-06T01:06:34.133Z",
          "content": "<p>Hi! <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> </p>\n<p>Looks like g - m dataset is still showing version 1. Did you put the correct one somewhere else? </p>\n<p><a href=\"https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-02\" target=\"_blank\">https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-02</a><br>\nCan you please confirm that the given link has a valid g -m data</p>\n<p>Thanks</p>",
          "rawMarkdown": "Hi! @ttahara \n\nLooks like g - m dataset is still showing version 1. Did you put the correct one somewhere else? \n\nhttps://www.kaggle.com/ttahara/birdsong-resampled-train-audio-02\nCan you please confirm that the given link has a valid g -m data\n\nThanks",
          "votes": 1
        }
      ]
    },
    {
      "id": 978912,
      "postDate": "2020-08-20T13:48:15.427Z",
      "content": "<p>For people who only want to use the data from all datasets and want to concat the <code>train_mod.csv</code> files. Every <code>train_mod.csv</code> file contains all filenames not only the ones in the dataset. To fix this simply:</p>\n<pre><code>df_1 = df_1.iloc[:4524]\ndf_2 = df_2.iloc[4524:8452]\ndf_3 = df_3.iloc[8452:13164]\ndf_4 = df_4.iloc[13164:16982]\ndf_5 = df_5.iloc[16982:]\n</code></pre>\n<p>where df_1, df_2, … is simply:</p>\n<pre><code>df_1 = pd.read_csv(\"../input/birdsong-resampled-train-audio-00/train_mod.csv\")\ndf_2 = pd.read_csv(\"../input/birdsong-resampled-train-audio-01/train_mod.csv\")\n...\n</code></pre>",
      "rawMarkdown": "For people who only want to use the data from all datasets and want to concat the `train_mod.csv` files. Every `train_mod.csv` file contains all filenames not only the ones in the dataset. To fix this simply:\n\n```\ndf_1 = df_1.iloc[:4524]\ndf_2 = df_2.iloc[4524:8452]\ndf_3 = df_3.iloc[8452:13164]\ndf_4 = df_4.iloc[13164:16982]\ndf_5 = df_5.iloc[16982:]\n```\n\nwhere df_1, df_2, ... is simply:\n```\ndf_1 = pd.read_csv(\"../input/birdsong-resampled-train-audio-00/train_mod.csv\")\ndf_2 = pd.read_csv(\"../input/birdsong-resampled-train-audio-01/train_mod.csv\")\n...\n```"
    },
    {
      "id": 929004,
      "postDate": "2020-07-14T11:51:21.387Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 928993,
      "postDate": "2020-07-14T11:41:57.273Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 916855,
      "postDate": "2020-07-06T04:46:49.883Z",
      "content": "<p>Thanks Tawara!</p>",
      "rawMarkdown": "Thanks Tawara!",
      "votes": 2
    },
    {
      "id": 979051,
      "postDate": "2020-08-20T15:30:40.207Z",
      "content": "<p>Great. Thank for sharing</p>",
      "rawMarkdown": "Great. Thank for sharing"
    }
  ],
  "comments": [
    {
      "id": 951910,
      "author_name": "Yu Kang",
      "author_url": "",
      "post_date": "2020-07-30T13:48:44.963000",
      "content": "<p>Thanks for sharing this great data and baseline!\nIs the <code>train_mod.csv</code> file in each part the same?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 952270,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-07-30T19:06:27.333000",
          "content": "<p>Yes, I put the same <code>.csv</code> file in each part.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 952583,
          "author_name": "Yu Kang",
          "author_url": "",
          "post_date": "2020-07-31T04:10:55.320000",
          "content": "<p>Got it. Thanks again.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 960721,
          "author_name": "Marco Gorelli",
          "author_url": "",
          "post_date": "2020-08-06T16:12:16.063000",
          "content": "<p>And the <code>train_mod.csv</code> file is, in turn, the same as the original <code>train.csv</code> file, but with additional columns:</p>\n\n<p><code>\n['resampled_sampling_rate', 'resampled_filename', 'resampled_channels']\n</code></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 929342,
      "author_name": "Akim Tsvigun",
      "author_url": "",
      "post_date": "2020-07-14T15:56:38.667000",
      "content": "<p>Thanks for your work, Tawara! Could you please share the code you used to resample the data or did you use the one provided <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159943\">here</a>? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 929911,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-07-15T04:17:56.887000",
          "content": "<p>I used a code based on the one shared in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/162304\">this topic</a>.</p>\n\n<p>An example is as follows (you have to change paths  depending on your environment):</p>\n\n<p>```python\nimport warnings\nfrom pathlib import Path\nfrom joblib import delayed, Parallel</p>\n\n<p>import librosa\nimport audioread\nimport soundfile as sf</p>\n\n<p>import pandas as pd</p>\n\n<p>TRAIN_AUDIO_DIR = \"path/to/train_audio\"\nTRAIN_RESAMPLED_DIR = \"path/to/train_audio_resampled\"</p>\n\n<p>TARGET_SR = 32000\nNUM_THREAD = 8  # for joblib.Parallel</p>\n\n<h1># read train.csv</h1>\n\n<p>train = pd.read_csv(\"path/to/train.csv\")</p>\n\n<h1># extract \"ebird_code\" and  \"filename\"</h1>\n\n<p>train_audio_infos = train[[\"ebird_code\", \"filename\"]].values.tolist()</p>\n\n<h1># make directories for saving resampled audio</h1>\n\n<p>TRAIN_RESAMPLED_DIR.mkdir(parents=True)\nfor ebird_code in train.ebird_code.unique():\n    ebird_dir = TRAIN_RESAMPLED_DIR / ebird_code\n    ebird_dir.mkdir()</p>\n\n<h1># define resampling function</h1>\n\n<p>warnings.simplefilter(\"ignore\")\ndef resample(ebird_code: str, filename: str, target_sr: int): <br>\n    audio_dir = TRAIN_AUDIO_DIR\n    resample_dir = TRAIN_RESAMPLED_DIR\n    ebird_dir = resample_dir / ebird_code</p>\n\n<pre><code>try:\n    y, _ = librosa.load(\n        audio_dir / ebird_code / filename,\n        sr=target_sr, mono=True, res_type=\"kaiser_fast\")\n\n    filename = filename.replace(\".mp3\", \".wav\")\n    sf.write(ebird_dir / filename, y, samplerate=target_sr)\n    return \"OK\"\nexcept Exception as e:\n    with open(resample_dir / \"skipped.txt\", \"a\") as f:\n        file_path = str(audio_dir / ebird_code / filename)\n        f.write(file_path + \"\\n\")\n    return str(e)\n</code></pre>\n\n<h1># resample and save audio using Parallel</h1>\n\n<p>msg_list = Parallel(n_jobs=NUM_THREAD, verbose=1)(\n    delayed(resample)(ebird_code, file_name, TARGET_SR) for ebird_code, file_name in train_audio_infos)</p>\n\n<h1># add information of resampled audios to train.csv</h1>\n\n<p>train[\"resampled_sampling_rate\"] = TARGET_SR\ntrain[\"resampled_filename\"] = train[\"filename\"].map(\n    lambda x: x.replace(\".mp3\", \".wav\"))\ntrain[\"resampled_channels\"] = \"1 (mono)\"</p>\n\n<p>train.to_csv(TRAIN_RESAMPLED_DIR / \"train_mod.csv\", index=False)\n```</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 930099,
          "author_name": "Akim Tsvigun",
          "author_url": "",
          "post_date": "2020-07-15T07:39:03.617000",
          "content": "<p>I see, thanks a lot!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 916017,
      "author_name": "Hidehisa Arai",
      "author_url": "",
      "post_date": "2020-07-05T09:28:44.257000",
      "content": "<p>Seems the link to g - m dataset is wrong, but thanks anyway!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 916030,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2020-07-05T09:36:57.977000",
          "content": "<p>Sorry, I correct it and now it is accessible.</p>\n\n<p>Thank you for reply 😃 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 999748,
          "author_name": "Shamrat Akbar",
          "author_url": "",
          "post_date": "2020-09-06T01:06:34.133000",
          "content": "<p>Hi! <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> </p>\n<p>Looks like g - m dataset is still showing version 1. Did you put the correct one somewhere else? </p>\n<p><a href=\"https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-02\" target=\"_blank\">https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-02</a><br>\nCan you please confirm that the given link has a valid g -m data</p>\n<p>Thanks</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 978912,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2020-08-20T13:48:15.427000",
      "content": "<p>For people who only want to use the data from all datasets and want to concat the <code>train_mod.csv</code> files. Every <code>train_mod.csv</code> file contains all filenames not only the ones in the dataset. To fix this simply:</p>\n<pre><code>df_1 = df_1.iloc[:4524]\ndf_2 = df_2.iloc[4524:8452]\ndf_3 = df_3.iloc[8452:13164]\ndf_4 = df_4.iloc[13164:16982]\ndf_5 = df_5.iloc[16982:]\n</code></pre>\n<p>where df_1, df_2, … is simply:</p>\n<pre><code>df_1 = pd.read_csv(\"../input/birdsong-resampled-train-audio-00/train_mod.csv\")\ndf_2 = pd.read_csv(\"../input/birdsong-resampled-train-audio-01/train_mod.csv\")\n...\n</code></pre>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 929004,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-14T11:51:21.387000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 928993,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-14T11:41:57.273000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 916855,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-06T04:46:49.883000",
      "content": "<p>Thanks Tawara!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 979051,
      "author_name": "Lan Dinh",
      "author_url": "",
      "post_date": "2020-08-20T15:30:40.207000",
      "content": "<p>Great. Thank for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "915852": "Hi, Kagglers.\n\nI'm new comer in this competition, and see that resampling recordings is important in several notebooks and discussions.\n\n@radek1 kindly provide resampled data (see [this topic](https://www.kaggle.com/c/birdsong-recognition/discussion/159943)), but I cannot access it now and it is hard to use in kaggle notebooks.\n\nI decided to make resampled data for the above reason and this is now available on Kaggle Dataset.\n(Preprocessing took less than an hour, but uploading each part took about 3 hours, totally 15 hours😓 )\n\nBecause of file size (72 GB in total), resampled data is split into five parts:  \n \n* [birdsong resampled train audio 00 (a ~ b)](https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-00)\n* [birdsong resampled train audio 01 (c ~ f)](https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-01)\n* [birdsong resampled train audio 02 (g ~ m)](https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-02)\n* [birdsong  resampled train audio 03 (n ~ r)](https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-03)\n* [birdsong resampled train audio 04 (s ~ y)](https://www.kaggle.com/ttahara/birdsong-resampled-train-audio-04)\n\nI'm afraid that I have made some mistakes in that process. Now I'm working on making baseline to make sure the data is okay.\n\nI used preprocessing code shared in [this topic](https://www.kaggle.com/c/birdsong-recognition/discussion/162304). Great thanks to @hidehisaarai1213 for sharing baseline code.\n\n### Update\n\nI shared training &amp; inference notebook (see [this toric](https://www.kaggle.com/c/birdsong-recognition/discussion/164345)) .\n\nI think the data is ok from the result (Public 2nd (0.56) now).\n<br>\nI have realized that @radek1 shares download link in [another topic](https://www.kaggle.com/c/birdsong-recognition/discussion/160321). Thanks!",
    "951910": "Thanks for sharing this great data and baseline!\nIs the `train_mod.csv` file in each part the same?",
    "929342": "Thanks for your work, Tawara! Could you please share the code you used to resample the data or did you use the one provided [here](https://www.kaggle.com/c/birdsong-recognition/discussion/159943)? ",
    "916017": "Seems the link to g - m dataset is wrong, but thanks anyway!\n",
    "978912": "For people who only want to use the data from all datasets and want to concat the `train_mod.csv` files. Every `train_mod.csv` file contains all filenames not only the ones in the dataset. To fix this simply:\n\n```\ndf_1 = df_1.iloc[:4524]\ndf_2 = df_2.iloc[4524:8452]\ndf_3 = df_3.iloc[8452:13164]\ndf_4 = df_4.iloc[13164:16982]\ndf_5 = df_5.iloc[16982:]\n```\n\nwhere df_1, df_2, ... is simply:\n```\ndf_1 = pd.read_csv(\"../input/birdsong-resampled-train-audio-00/train_mod.csv\")\ndf_2 = pd.read_csv(\"../input/birdsong-resampled-train-audio-01/train_mod.csv\")\n...\n```",
    "929004": "",
    "928993": "",
    "916855": "Thanks Tawara!",
    "979051": "Great. Thank for sharing"
  }
}