{
  "id": 90536,
  "title": "Use noisy set ",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/90536",
  "author_name": "",
  "post_date": "2019-04-24T16:08:29.329984300Z",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>What is your experience with the noisy set? Do you use it for ensemble?\nI  did not succeed to improve prediction with that set.  I wonder if there is some noisy data (sound)  in the private test set.</p>",
  "messages": [
    {
      "id": "522549",
      "postDate": "04/24/2019 16:08:29",
      "content": "<p>What is your experience with the noisy set? Do you use it for ensemble?\nI  did not succeed to improve prediction with that set.  I wonder if there is some noisy data (sound)  in the private test set.</p>",
      "rawMarkdown": "What is your experience with the noisy set? Do you use it for ensemble?\nI  did not succeed to improve prediction with that set.  I wonder if there is some noisy data (sound)  in the private test set.",
      "votes": null
    },
    {
      "id": "522766",
      "postDate": "04/25/2019 01:24:32",
      "content": "<p>When I load noisy dataset to training, the result becomes worse. </p>",
      "rawMarkdown": "When I load noisy dataset to training, the result becomes worse.",
      "votes": null
    },
    {
      "id": "527116",
      "postDate": "05/04/2019 16:13:11",
      "content": "<p>I'm using noisy set to pre-train model from scratch, it makes final training (with various parameters) much quicker and performance can be extracted stably.</p>",
      "rawMarkdown": "I'm using noisy set to pre-train model from scratch, it makes final training (with various parameters) much quicker and performance can be extracted stably.",
      "votes": null
    },
    {
      "id": "528046",
      "postDate": "05/07/2019 00:46:56",
      "content": "<p>Hi, </p>\n\n<p>as explained in <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/data\">Data section</a>, the test set does not contain noisy data from YFCC (only curated data from FSD). However, this is one of the main challenges of this evaluation: how to exploit a large amount of noisy data (amounting to 80 hours of audio) plus only a small quantity of curated data (~10 hours).</p>\n\n<p>Note that the effort needed to curate the latter 10 hours is much more than to gather the former 80 hours of audio! Coming up with ways to exploit the noisy data can be very relevant.</p>",
      "rawMarkdown": "Hi, \n\nas explained in [Data section](https://www.kaggle.com/c/freesound-audio-tagging-2019/data), the test set does not contain noisy data from YFCC (only curated data from FSD). However, this is one of the main challenges of this evaluation: how to exploit a large amount of noisy data (amounting to 80 hours of audio) plus only a small quantity of curated data (~10 hours).\n\nNote that the effort needed to curate the latter 10 hours is much more than to gather the former 80 hours of audio! Coming up with ways to exploit the noisy data can be very relevant.",
      "votes": null
    },
    {
      "id": "528410",
      "postDate": "05/07/2019 17:56:59",
      "content": "<p>At first, I've found noisy dataset quiet useless since it is not improving my score,but after I've got couple thoughts about its application:\n1. Train on large noisy set and finetune on curated set. That can make results more stable and reduce overfitting.\n2. Train a lot on noisy data and freeze some low-level layers of CNN. So, the model will be less vulnerable to overfitting during treaining  on curated data (since lower weights number to train) and still has the low-level features extracted. \nAlso there should be also other non-standard ways to use this data</p>",
      "rawMarkdown": "At first, I've found noisy dataset quiet useless since it is not improving my score,but after I've got couple thoughts about its application:\n1. Train on large noisy set and finetune on curated set. That can make results more stable and reduce overfitting.\n2. Train a lot on noisy data and freeze some low-level layers of CNN. So, the model will be less vulnerable to overfitting during treaining  on curated data (since lower weights number to train) and still has the low-level features extracted. \nAlso there should be also other non-standard ways to use this data",
      "votes": null
    },
    {
      "id": "528641",
      "postDate": "05/08/2019 08:35:53",
      "content": "<p>May I ask, how do you know it doesn't help without even making a single submit? </p>",
      "rawMarkdown": "May I ask, how do you know it doesn't help without even making a single submit?",
      "votes": null
    },
    {
      "id": "530361",
      "postDate": "05/12/2019 16:17:56",
      "content": "<p>I'm using noisy set to pre-train model from scratch too. However, Lwlrap of validation data does not rise at all at about 0.08. Is there a way to get a good pre-train model from noise data?</p>",
      "rawMarkdown": "I'm using noisy set to pre-train model from scratch too. However, Lwlrap of validation data does not rise at all at about 0.08. Is there a way to get a good pre-train model from noise data?",
      "votes": null
    },
    {
      "id": "530469",
      "postDate": "05/12/2019 23:17:31",
      "content": "<p>I’m doing very simple training steps, and confirming models as described here:\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/89055#latest-515314\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/89055#latest-515314</a></p>\n\n<p>My validation lwlrap is around 0.4 to 0.5 for noisy set only, for me it seems to be good enough if I see it. \nUsing training code which is almost the same with my public kernels.</p>",
      "rawMarkdown": "I’m doing very simple training steps, and confirming models as described here:\nhttps://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/89055#latest-515314\n\nMy validation lwlrap is around 0.4 to 0.5 for noisy set only, for me it seems to be good enough if I see it. \nUsing training code which is almost the same with my public kernels.",
      "votes": null
    },
    {
      "id": "530473",
      "postDate": "05/13/2019 00:04:40",
      "content": "<p>Thank you!!\nDid you use the noisy set all? Or did you use something like trn_noisy_best50?\nI will try to reference</p>",
      "rawMarkdown": "Thank you!!\nDid you use the noisy set all? Or did you use something like trn_noisy_best50?\nI will try to reference",
      "votes": null
    },
    {
      "id": "530481",
      "postDate": "05/13/2019 01:42:18",
      "content": "<p>Using all noisy samples, nothing else. Simply replaced curated samples with noisy ones.</p>",
      "rawMarkdown": "Using all noisy samples, nothing else. Simply replaced curated samples with noisy ones.",
      "votes": null
    },
    {
      "id": "530493",
      "postDate": "05/13/2019 02:46:30",
      "content": "<p>thank you!!\nBoth your kernel and discussion are really helpful.</p>",
      "rawMarkdown": "thank you!!\nBoth your kernel and discussion are really helpful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 522766,
      "author_name": "gmhost",
      "author_url": "",
      "post_date": "04/25/2019 01:24:32",
      "content": "<p>When I load noisy dataset to training, the result becomes worse. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 527116,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "05/04/2019 16:13:11",
      "content": "<p>I'm using noisy set to pre-train model from scratch, it makes final training (with various parameters) much quicker and performance can be extracted stably.</p>",
      "votes": null,
      "replies": [
        {
          "id": 530361,
          "author_name": "bossimuimu",
          "author_url": "",
          "post_date": "05/12/2019 16:17:56",
          "content": "<p>I'm using noisy set to pre-train model from scratch too. However, Lwlrap of validation data does not rise at all at about 0.08. Is there a way to get a good pre-train model from noise data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 530469,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "05/12/2019 23:17:31",
          "content": "<p>I’m doing very simple training steps, and confirming models as described here:\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/89055#latest-515314\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/89055#latest-515314</a></p>\n\n<p>My validation lwlrap is around 0.4 to 0.5 for noisy set only, for me it seems to be good enough if I see it. \nUsing training code which is almost the same with my public kernels.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 530473,
          "author_name": "bossimuimu",
          "author_url": "",
          "post_date": "05/13/2019 00:04:40",
          "content": "<p>Thank you!!\nDid you use the noisy set all? Or did you use something like trn_noisy_best50?\nI will try to reference</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 530481,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "05/13/2019 01:42:18",
          "content": "<p>Using all noisy samples, nothing else. Simply replaced curated samples with noisy ones.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 530493,
          "author_name": "bossimuimu",
          "author_url": "",
          "post_date": "05/13/2019 02:46:30",
          "content": "<p>thank you!!\nBoth your kernel and discussion are really helpful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 528046,
      "author_name": "eduardofonseca",
      "author_url": "",
      "post_date": "05/07/2019 00:46:56",
      "content": "<p>Hi, </p>\n\n<p>as explained in <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/data\">Data section</a>, the test set does not contain noisy data from YFCC (only curated data from FSD). However, this is one of the main challenges of this evaluation: how to exploit a large amount of noisy data (amounting to 80 hours of audio) plus only a small quantity of curated data (~10 hours).</p>\n\n<p>Note that the effort needed to curate the latter 10 hours is much more than to gather the former 80 hours of audio! Coming up with ways to exploit the noisy data can be very relevant.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 528410,
      "author_name": "alexanderkhar",
      "author_url": "",
      "post_date": "05/07/2019 17:56:59",
      "content": "<p>At first, I've found noisy dataset quiet useless since it is not improving my score,but after I've got couple thoughts about its application:\n1. Train on large noisy set and finetune on curated set. That can make results more stable and reduce overfitting.\n2. Train a lot on noisy data and freeze some low-level layers of CNN. So, the model will be less vulnerable to overfitting during treaining  on curated data (since lower weights number to train) and still has the low-level features extracted. \nAlso there should be also other non-standard ways to use this data</p>",
      "votes": null,
      "replies": [
        {
          "id": 528641,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "05/08/2019 08:35:53",
          "content": "<p>May I ask, how do you know it doesn't help without even making a single submit? </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "522549": "What is your experience with the noisy set? Do you use it for ensemble?\nI  did not succeed to improve prediction with that set.  I wonder if there is some noisy data (sound)  in the private test set.",
    "522766": "When I load noisy dataset to training, the result becomes worse.",
    "527116": "I'm using noisy set to pre-train model from scratch, it makes final training (with various parameters) much quicker and performance can be extracted stably.",
    "528046": "Hi, \n\nas explained in [Data section](https://www.kaggle.com/c/freesound-audio-tagging-2019/data), the test set does not contain noisy data from YFCC (only curated data from FSD). However, this is one of the main challenges of this evaluation: how to exploit a large amount of noisy data (amounting to 80 hours of audio) plus only a small quantity of curated data (~10 hours).\n\nNote that the effort needed to curate the latter 10 hours is much more than to gather the former 80 hours of audio! Coming up with ways to exploit the noisy data can be very relevant.",
    "528410": "At first, I've found noisy dataset quiet useless since it is not improving my score,but after I've got couple thoughts about its application:\n1. Train on large noisy set and finetune on curated set. That can make results more stable and reduce overfitting.\n2. Train a lot on noisy data and freeze some low-level layers of CNN. So, the model will be less vulnerable to overfitting during treaining  on curated data (since lower weights number to train) and still has the low-level features extracted. \nAlso there should be also other non-standard ways to use this data",
    "528641": "May I ask, how do you know it doesn't help without even making a single submit?",
    "530361": "I'm using noisy set to pre-train model from scratch too. However, Lwlrap of validation data does not rise at all at about 0.08. Is there a way to get a good pre-train model from noise data?",
    "530469": "I’m doing very simple training steps, and confirming models as described here:\nhttps://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/89055#latest-515314\n\nMy validation lwlrap is around 0.4 to 0.5 for noisy set only, for me it seems to be good enough if I see it. \nUsing training code which is almost the same with my public kernels.",
    "530473": "Thank you!!\nDid you use the noisy set all? Or did you use something like trn_noisy_best50?\nI will try to reference",
    "530481": "Using all noisy samples, nothing else. Simply replaced curated samples with noisy ones.",
    "530493": "thank you!!\nBoth your kernel and discussion are really helpful."
  },
  "source": "meta"
}