{
  "id": 89336,
  "title": "Preprocessed Mel-spectrogram Dataset is now Ready",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/89336",
  "author_name": "",
  "post_date": "2019-04-13T04:02:29.769996300Z",
  "votes": 14,
  "comment_count": 11,
  "views": 0,
  "content": "<h2>Dataset</h2>\n\n<p><a href=\"https://www.kaggle.com/daisukelab/fat2019_prep_mels1\">https://www.kaggle.com/daisukelab/fat2019_prep_mels1</a></p>\n\n<h2>Kernel to create this dataset</h2>\n\n<p><a href=\"https://www.kaggle.com/daisukelab/creating-fat2019-preprocessed-data\">https://www.kaggle.com/daisukelab/creating-fat2019-preprocessed-data</a></p>\n\n<h2>Sample kernel</h2>\n\n<p><a href=\"https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-2-w-preprocessed-dataset\">https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-2-w-preprocessed-dataset</a></p>\n\n<p>Unfortunately noisy data is too big, then I will split it to multi-parts.</p>\n\n<p>But for now, you can at least save 10 minutes for preprocessing curated and test data.</p>",
  "messages": [
    {
      "id": "515749",
      "postDate": "04/13/2019 04:02:29",
      "content": "<h2>Dataset</h2>\n\n<p><a href=\"https://www.kaggle.com/daisukelab/fat2019_prep_mels1\">https://www.kaggle.com/daisukelab/fat2019_prep_mels1</a></p>\n\n<h2>Kernel to create this dataset</h2>\n\n<p><a href=\"https://www.kaggle.com/daisukelab/creating-fat2019-preprocessed-data\">https://www.kaggle.com/daisukelab/creating-fat2019-preprocessed-data</a></p>\n\n<h2>Sample kernel</h2>\n\n<p><a href=\"https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-2-w-preprocessed-dataset\">https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-2-w-preprocessed-dataset</a></p>\n\n<p>Unfortunately noisy data is too big, then I will split it to multi-parts.</p>\n\n<p>But for now, you can at least save 10 minutes for preprocessing curated and test data.</p>",
      "rawMarkdown": "## Dataset\nhttps://www.kaggle.com/daisukelab/fat2019_prep_mels1\n\n## Kernel to create this dataset\nhttps://www.kaggle.com/daisukelab/creating-fat2019-preprocessed-data\n\n## Sample kernel\nhttps://www.kaggle.com/daisukelab/cnn-2d-basic-solution-2-w-preprocessed-dataset\n\nUnfortunately noisy data is too big, then I will split it to multi-parts.\n\nBut for now, you can at least save 10 minutes for preprocessing curated and test data.",
      "votes": null
    },
    {
      "id": "515818",
      "postDate": "04/13/2019 07:04:36",
      "content": "<p>Hi, <a href=\"/daisukelab\">@daisukelab</a> </p>\n\n<p>Thank you for preparing dataset . <br>\nJust confirm...,  </p>\n\n<p>As in Timeline page, <br>\n<code>\nNote, as this competition is a Kernels-only, two-stage competition, following the final submission deadline for the competition, your kernel code will be re-run on a privately-held test set that is not provided to you. It is your model's score against this private test set that will determine your ranking on the private leaderboard and final standing in the competition. The leaderboard will be updated in the days following the competition's completion, and our team will announce that the re-run has been completed and leaderboard finalized with an announcement made on the competition forums.\n</code>\nmy understanding about stage 2 is that we will not be able to prepare pre-processed spectrograms. <br>\nIs my understanding correct ? <br>\nIf I'm right, just preparing dataset in advance will not save our inference time, because test data that will be used in the evaluation for stage 2 is not available now.  </p>",
      "rawMarkdown": "Hi, @daisukelab \n\nThank you for preparing dataset .  \nJust confirm...,  \n\nAs in Timeline page,  \n```\nNote, as this competition is a Kernels-only, two-stage competition, following the final submission deadline for the competition, your kernel code will be re-run on a privately-held test set that is not provided to you. It is your model's score against this private test set that will determine your ranking on the private leaderboard and final standing in the competition. The leaderboard will be updated in the days following the competition's completion, and our team will announce that the re-run has been completed and leaderboard finalized with an announcement made on the competition forums.\n```\nmy understanding about stage 2 is that we will not be able to prepare pre-processed spectrograms.  \nIs my understanding correct ?  \nIf I'm right, just preparing dataset in advance will not save our inference time, because test data that will be used in the evaluation for stage 2 is not available now.",
      "votes": null
    },
    {
      "id": "515822",
      "postDate": "04/13/2019 07:17:55",
      "content": "<p>I got your point. Competition will replace test set when they run 2nd stage evaluation, then,</p>\n\n<p><strong>Our kernel should be able to prepare for preprocessing at the 2nd test run...</strong></p>\n\n<p>It would be. Then I guess the sample kernel would not be ready for the final submission. I will just clarify on the sample kernel...</p>\n\n<p>Thank you for pointing it out!</p>",
      "rawMarkdown": "I got your point. Competition will replace test set when they run 2nd stage evaluation, then,\n\n__Our kernel should be able to prepare for preprocessing at the 2nd test run...__\n\nIt would be. Then I guess the sample kernel would not be ready for the final submission. I will just clarify on the sample kernel...\n\nThank you for pointing it out!",
      "votes": null
    },
    {
      "id": "515845",
      "postDate": "04/13/2019 08:03:57",
      "content": "<p><a href=\"/daisukelab\">@daisukelab</a> \nThanks, I feel relieved now!</p>",
      "rawMarkdown": "daisukelab \nThanks, I feel relieved now!",
      "votes": null
    },
    {
      "id": "515872",
      "postDate": "04/13/2019 08:35:42",
      "content": "<p><a href=\"/maxwell110\">@maxwell110</a> Excuse me, I re-wrote my comment above. This is just in case you might have read old one. Thanks again for pointing out.</p>",
      "rawMarkdown": "maxwell110 Excuse me, I re-wrote my comment above. This is just in case you might have read old one. Thanks again for pointing out.",
      "votes": null
    },
    {
      "id": "517293",
      "postDate": "04/15/2019 21:11:11",
      "content": "<p>So to confirm, the actual kernel does not need to preprocess the training set and we can use your dataset for that, but the testing dataset preprocessing must be done in the kernel, correct?</p>",
      "rawMarkdown": "So to confirm, the actual kernel does not need to preprocess the training set and we can use your dataset for that, but the testing dataset preprocessing must be done in the kernel, correct?",
      "votes": null
    },
    {
      "id": "517338",
      "postDate": "04/15/2019 22:58:25",
      "content": "<p><a href=\"/tanlikesmath\">@tanlikesmath</a> \nTo my understanding, yes. you are right. <br>\nAnd we must care about run time, <strong>1 GPU HOUR</strong> for inference. <br>\nThe 2nd stage test data will be about 3360 samples ( about 3 times as 1st stage test ). <br>\nSo you can estimate time to transform those into mel-spectrogram or others through a simple experiment using available data.</p>",
      "rawMarkdown": "tanlikesmath \nTo my understanding, yes. you are right.  \nAnd we must care about run time, **1 GPU HOUR** for inference.  \nThe 2nd stage test data will be about 3360 samples ( about 3 times as 1st stage test ).  \nSo you can estimate time to transform those into mel-spectrogram or others through a simple experiment using available data.",
      "votes": null
    },
    {
      "id": "517637",
      "postDate": "04/16/2019 09:47:28",
      "content": "<h2><em>Update Notice: New Version 2</em></h2>\n\n<p>Fixed preprocessing parameter as follows, converted mel-spectrogram has now half length from version 1; Size 128x128 was 1s sound, now 2s in length 128. This change is basically a bug fix, since duration is set to 2s, then 128x128 should have 2s. It would boost performance basically.</p>\n\n<pre><code>conf.sampling_rate = 44100\nconf.duration = 2 # sec\nconf.hop_length = 347*conf.duration # &amp;lt;==== updated from `conf.hop_length = 347`\nconf.fmin = 20\nconf.fmax = conf.sampling_rate // 2\nconf.n_mels = 128\nconf.n_fft = conf.n_mels * 20\n</code></pre>",
      "rawMarkdown": "## _Update Notice: New Version 2_\nFixed preprocessing parameter as follows, converted mel-spectrogram has now half length from version 1; Size 128x128 was 1s sound, now 2s in length 128. This change is basically a bug fix, since duration is set to 2s, then 128x128 should have 2s. It would boost performance basically.\n\n    conf.sampling_rate = 44100\n    conf.duration = 2 # sec\n    conf.hop_length = 347*conf.duration # &lt;==== updated from `conf.hop_length = 347`\n    conf.fmin = 20\n    conf.fmax = conf.sampling_rate // 2\n    conf.n_mels = 128\n    conf.n_fft = conf.n_mels * 20",
      "votes": null
    },
    {
      "id": "523767",
      "postDate": "04/27/2019 01:14:42",
      "content": "<p>I cannot recall who requested me to, here I made a kernel to create dataset.\n<a href=\"https://www.kaggle.com/daisukelab/creating-fat2019-preprocessed-data\">https://www.kaggle.com/daisukelab/creating-fat2019-preprocessed-data</a></p>",
      "rawMarkdown": "I cannot recall who requested me to, here I made a kernel to create dataset.\nhttps://www.kaggle.com/daisukelab/creating-fat2019-preprocessed-data",
      "votes": null
    },
    {
      "id": "526481",
      "postDate": "05/03/2019 06:30:44",
      "content": "<p>Hi, <a href=\"/daisukelab\">@daisukelab</a> \nThank you for preparing the dataset!</p>\n\n<p>As I observed your kernel which created the dataset, you are just taking 50 singled data per class, right?</p>\n\n<p>I am asking this because the dataset contains some duplicated data. I'd like to know whether you made this overwrap in purpose or not.</p>\n\n<p>Thanks in advance!</p>",
      "rawMarkdown": "Hi, @daisukelab \nThank you for preparing the dataset!\n\nAs I observed your kernel which created the dataset, you are just taking 50 singled data per class, right?\n\nI am asking this because the dataset contains some duplicated data. I'd like to know whether you made this overwrap in purpose or not.\n\nThanks in advance!",
      "votes": null
    },
    {
      "id": "526709",
      "postDate": "05/03/2019 15:14:28",
      "content": "<p>Hi, there's no trick for that...\n- It's just making list of single labeled samples for all single classes, then picking 50 randomly.\n- Nothing more than this, it doesn't care about duplicated data you mention.</p>",
      "rawMarkdown": "Hi, there's no trick for that...\n- It's just making list of single labeled samples for all single classes, then picking 50 randomly.\n- Nothing more than this, it doesn't care about duplicated data you mention.",
      "votes": null
    },
    {
      "id": "526713",
      "postDate": "05/03/2019 15:25:41",
      "content": "<p>Thank you for the clarification!</p>",
      "rawMarkdown": "Thank you for the clarification!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 515818,
      "author_name": "maxwell110",
      "author_url": "",
      "post_date": "04/13/2019 07:04:36",
      "content": "<p>Hi, <a href=\"/daisukelab\">@daisukelab</a> </p>\n\n<p>Thank you for preparing dataset . <br>\nJust confirm...,  </p>\n\n<p>As in Timeline page, <br>\n<code>\nNote, as this competition is a Kernels-only, two-stage competition, following the final submission deadline for the competition, your kernel code will be re-run on a privately-held test set that is not provided to you. It is your model's score against this private test set that will determine your ranking on the private leaderboard and final standing in the competition. The leaderboard will be updated in the days following the competition's completion, and our team will announce that the re-run has been completed and leaderboard finalized with an announcement made on the competition forums.\n</code>\nmy understanding about stage 2 is that we will not be able to prepare pre-processed spectrograms. <br>\nIs my understanding correct ? <br>\nIf I'm right, just preparing dataset in advance will not save our inference time, because test data that will be used in the evaluation for stage 2 is not available now.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 515822,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "04/13/2019 07:17:55",
          "content": "<p>I got your point. Competition will replace test set when they run 2nd stage evaluation, then,</p>\n\n<p><strong>Our kernel should be able to prepare for preprocessing at the 2nd test run...</strong></p>\n\n<p>It would be. Then I guess the sample kernel would not be ready for the final submission. I will just clarify on the sample kernel...</p>\n\n<p>Thank you for pointing it out!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 515845,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "04/13/2019 08:03:57",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a> \nThanks, I feel relieved now!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 515872,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "04/13/2019 08:35:42",
          "content": "<p><a href=\"/maxwell110\">@maxwell110</a> Excuse me, I re-wrote my comment above. This is just in case you might have read old one. Thanks again for pointing out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 517293,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "04/15/2019 21:11:11",
      "content": "<p>So to confirm, the actual kernel does not need to preprocess the training set and we can use your dataset for that, but the testing dataset preprocessing must be done in the kernel, correct?</p>",
      "votes": null,
      "replies": [
        {
          "id": 517338,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "04/15/2019 22:58:25",
          "content": "<p><a href=\"/tanlikesmath\">@tanlikesmath</a> \nTo my understanding, yes. you are right. <br>\nAnd we must care about run time, <strong>1 GPU HOUR</strong> for inference. <br>\nThe 2nd stage test data will be about 3360 samples ( about 3 times as 1st stage test ). <br>\nSo you can estimate time to transform those into mel-spectrogram or others through a simple experiment using available data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 517637,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "04/16/2019 09:47:28",
      "content": "<h2><em>Update Notice: New Version 2</em></h2>\n\n<p>Fixed preprocessing parameter as follows, converted mel-spectrogram has now half length from version 1; Size 128x128 was 1s sound, now 2s in length 128. This change is basically a bug fix, since duration is set to 2s, then 128x128 should have 2s. It would boost performance basically.</p>\n\n<pre><code>conf.sampling_rate = 44100\nconf.duration = 2 # sec\nconf.hop_length = 347*conf.duration # &amp;lt;==== updated from `conf.hop_length = 347`\nconf.fmin = 20\nconf.fmax = conf.sampling_rate // 2\nconf.n_mels = 128\nconf.n_fft = conf.n_mels * 20\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 523767,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "04/27/2019 01:14:42",
      "content": "<p>I cannot recall who requested me to, here I made a kernel to create dataset.\n<a href=\"https://www.kaggle.com/daisukelab/creating-fat2019-preprocessed-data\">https://www.kaggle.com/daisukelab/creating-fat2019-preprocessed-data</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 526481,
      "author_name": "hidehisaarai1213",
      "author_url": "",
      "post_date": "05/03/2019 06:30:44",
      "content": "<p>Hi, <a href=\"/daisukelab\">@daisukelab</a> \nThank you for preparing the dataset!</p>\n\n<p>As I observed your kernel which created the dataset, you are just taking 50 singled data per class, right?</p>\n\n<p>I am asking this because the dataset contains some duplicated data. I'd like to know whether you made this overwrap in purpose or not.</p>\n\n<p>Thanks in advance!</p>",
      "votes": null,
      "replies": [
        {
          "id": 526709,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "05/03/2019 15:14:28",
          "content": "<p>Hi, there's no trick for that...\n- It's just making list of single labeled samples for all single classes, then picking 50 randomly.\n- Nothing more than this, it doesn't care about duplicated data you mention.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526713,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "05/03/2019 15:25:41",
          "content": "<p>Thank you for the clarification!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "515749": "## Dataset\nhttps://www.kaggle.com/daisukelab/fat2019_prep_mels1\n\n## Kernel to create this dataset\nhttps://www.kaggle.com/daisukelab/creating-fat2019-preprocessed-data\n\n## Sample kernel\nhttps://www.kaggle.com/daisukelab/cnn-2d-basic-solution-2-w-preprocessed-dataset\n\nUnfortunately noisy data is too big, then I will split it to multi-parts.\n\nBut for now, you can at least save 10 minutes for preprocessing curated and test data.",
    "515818": "Hi, @daisukelab \n\nThank you for preparing dataset .  \nJust confirm...,  \n\nAs in Timeline page,  \n```\nNote, as this competition is a Kernels-only, two-stage competition, following the final submission deadline for the competition, your kernel code will be re-run on a privately-held test set that is not provided to you. It is your model's score against this private test set that will determine your ranking on the private leaderboard and final standing in the competition. The leaderboard will be updated in the days following the competition's completion, and our team will announce that the re-run has been completed and leaderboard finalized with an announcement made on the competition forums.\n```\nmy understanding about stage 2 is that we will not be able to prepare pre-processed spectrograms.  \nIs my understanding correct ?  \nIf I'm right, just preparing dataset in advance will not save our inference time, because test data that will be used in the evaluation for stage 2 is not available now.",
    "515822": "I got your point. Competition will replace test set when they run 2nd stage evaluation, then,\n\n__Our kernel should be able to prepare for preprocessing at the 2nd test run...__\n\nIt would be. Then I guess the sample kernel would not be ready for the final submission. I will just clarify on the sample kernel...\n\nThank you for pointing it out!",
    "515845": "daisukelab \nThanks, I feel relieved now!",
    "515872": "maxwell110 Excuse me, I re-wrote my comment above. This is just in case you might have read old one. Thanks again for pointing out.",
    "517293": "So to confirm, the actual kernel does not need to preprocess the training set and we can use your dataset for that, but the testing dataset preprocessing must be done in the kernel, correct?",
    "517338": "tanlikesmath \nTo my understanding, yes. you are right.  \nAnd we must care about run time, **1 GPU HOUR** for inference.  \nThe 2nd stage test data will be about 3360 samples ( about 3 times as 1st stage test ).  \nSo you can estimate time to transform those into mel-spectrogram or others through a simple experiment using available data.",
    "517637": "## _Update Notice: New Version 2_\nFixed preprocessing parameter as follows, converted mel-spectrogram has now half length from version 1; Size 128x128 was 1s sound, now 2s in length 128. This change is basically a bug fix, since duration is set to 2s, then 128x128 should have 2s. It would boost performance basically.\n\n    conf.sampling_rate = 44100\n    conf.duration = 2 # sec\n    conf.hop_length = 347*conf.duration # &lt;==== updated from `conf.hop_length = 347`\n    conf.fmin = 20\n    conf.fmax = conf.sampling_rate // 2\n    conf.n_mels = 128\n    conf.n_fft = conf.n_mels * 20",
    "523767": "I cannot recall who requested me to, here I made a kernel to create dataset.\nhttps://www.kaggle.com/daisukelab/creating-fat2019-preprocessed-data",
    "526481": "Hi, @daisukelab \nThank you for preparing the dataset!\n\nAs I observed your kernel which created the dataset, you are just taking 50 singled data per class, right?\n\nI am asking this because the dataset contains some duplicated data. I'd like to know whether you made this overwrap in purpose or not.\n\nThanks in advance!",
    "526709": "Hi, there's no trick for that...\n- It's just making list of single labeled samples for all single classes, then picking 50 randomly.\n- Nothing more than this, it doesn't care about duplicated data you mention.",
    "526713": "Thank you for the clarification!"
  },
  "source": "meta"
}