{
  "id": 94996,
  "title": "Do yo use  curated data or full(curated + noisy) data to achieve the LB score",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/94996",
  "author_name": "",
  "post_date": "2019-06-09T04:53:21.783044900Z",
  "votes": 2,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I used curated nosiy data </p>",
  "messages": [
    {
      "id": "548264",
      "postDate": "06/09/2019 04:53:21",
      "content": "<p>I used curated nosiy data </p>",
      "rawMarkdown": "I used curated nosiy data",
      "votes": null
    },
    {
      "id": "548296",
      "postDate": "06/09/2019 06:14:10",
      "content": "<p>Nothing different with kernels but ensemble 50 models.  Ensemble is all you need,</p>",
      "rawMarkdown": "Nothing different with kernels but ensemble 50 models.  Ensemble is all you need,",
      "votes": null
    },
    {
      "id": "548302",
      "postDate": "06/09/2019 06:26:02",
      "content": "<p>I ensemble  models with different folds.</p>",
      "rawMarkdown": "I ensemble  models with different folds.",
      "votes": null
    },
    {
      "id": "548307",
      "postDate": "06/09/2019 06:41:19",
      "content": "<p>I tried ensembling but was unsuccessful. I used different model arch. with 5-fold and ensembled the average results from the 5-folds. Was unsuccessful.</p>",
      "rawMarkdown": "I tried ensembling but was unsuccessful. I used different model arch. with 5-fold and ensembled the average results from the 5-folds. Was unsuccessful.",
      "votes": null
    },
    {
      "id": "548319",
      "postDate": "06/09/2019 07:21:38",
      "content": "<p>how much  time your inference time is?  I'm afraid you spend more than 20m.</p>",
      "rawMarkdown": "how much  time your inference time is?  I'm afraid you spend more than 20m.",
      "votes": null
    },
    {
      "id": "548320",
      "postDate": "06/09/2019 07:23:39",
      "content": "<p>We took our best single model and predicted all the noisy data.\nThose sounds that have the correct class in top3 predictions of our model we added for uptraining\nAt the same time, we tried to keep our classes balanced and gave priority to those classes that very rarely appeared in the top predictions for a test.\nWe had several choices for noisy files (1500, 2500 and 7500 files)\nNoisy data in our pipeline gave a significant gain</p>",
      "rawMarkdown": "We took our best single model and predicted all the noisy data.\nThose sounds that have the correct class in top3 predictions of our model we added for uptraining\nAt the same time, we tried to keep our classes balanced and gave priority to those classes that very rarely appeared in the top predictions for a test.\nWe had several choices for noisy files (1500, 2500 and 7500 files)\nNoisy data in our pipeline gave a significant gain",
      "votes": null
    },
    {
      "id": "548334",
      "postDate": "06/09/2019 07:54:01",
      "content": "<p>I was unsuccessful in that it did not improve my results.</p>",
      "rawMarkdown": "I was unsuccessful in that it did not improve my results.",
      "votes": null
    },
    {
      "id": "548781",
      "postDate": "06/09/2019 22:17:32",
      "content": "<p>LB went from 0.692 to 0.720 when I added noisy data. We don't remove any data in noisy set.</p>",
      "rawMarkdown": "LB went from 0.692 to 0.720 when I added noisy data. We don't remove any data in noisy set.",
      "votes": null
    },
    {
      "id": "548811",
      "postDate": "06/10/2019 00:26:14",
      "content": "<p>Do you just add the curated noisy data to your training with no changes?</p>",
      "rawMarkdown": "Do you just add the curated noisy data to your training with no changes?",
      "votes": null
    },
    {
      "id": "548824",
      "postDate": "06/10/2019 01:18:18",
      "content": "<p>I used all the curated data and noisy data. For me, noisy data make the model robuster which could reduce the gap between local lwlrap and LB score.</p>",
      "rawMarkdown": "I used all the curated data and noisy data. For me, noisy data make the model robuster which could reduce the gap between local lwlrap and LB score.",
      "votes": null
    },
    {
      "id": "548828",
      "postDate": "06/10/2019 01:32:14",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "549645",
      "postDate": "06/10/2019 22:38:20",
      "content": "<p>We used both curated data and noisy data. \nIt's interesting to see that the score improvements caused by the usage of noisy data differ a lot between models, although we used noisy data in the same manner.</p>",
      "rawMarkdown": "We used both curated data and noisy data. \nIt's interesting to see that the score improvements caused by the usage of noisy data differ a lot between models, although we used noisy data in the same manner.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 548296,
      "author_name": "baomengjiao",
      "author_url": "",
      "post_date": "06/09/2019 06:14:10",
      "content": "<p>Nothing different with kernels but ensemble 50 models.  Ensemble is all you need,</p>",
      "votes": null,
      "replies": [
        {
          "id": 548302,
          "author_name": "zxyu1995",
          "author_url": "",
          "post_date": "06/09/2019 06:26:02",
          "content": "<p>I ensemble  models with different folds.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 548307,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "06/09/2019 06:41:19",
          "content": "<p>I tried ensembling but was unsuccessful. I used different model arch. with 5-fold and ensembled the average results from the 5-folds. Was unsuccessful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 548319,
          "author_name": "zxyu1995",
          "author_url": "",
          "post_date": "06/09/2019 07:21:38",
          "content": "<p>how much  time your inference time is?  I'm afraid you spend more than 20m.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 548334,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "06/09/2019 07:54:01",
          "content": "<p>I was unsuccessful in that it did not improve my results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 548320,
      "author_name": "sneddy",
      "author_url": "",
      "post_date": "06/09/2019 07:23:39",
      "content": "<p>We took our best single model and predicted all the noisy data.\nThose sounds that have the correct class in top3 predictions of our model we added for uptraining\nAt the same time, we tried to keep our classes balanced and gave priority to those classes that very rarely appeared in the top predictions for a test.\nWe had several choices for noisy files (1500, 2500 and 7500 files)\nNoisy data in our pipeline gave a significant gain</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 548781,
      "author_name": "jihangz",
      "author_url": "",
      "post_date": "06/09/2019 22:17:32",
      "content": "<p>LB went from 0.692 to 0.720 when I added noisy data. We don't remove any data in noisy set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 548828,
          "author_name": "hongxiaofeng",
          "author_url": "",
          "post_date": "06/10/2019 01:32:14",
          "content": "",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 548811,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "06/10/2019 00:26:14",
      "content": "<p>Do you just add the curated noisy data to your training with no changes?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 548824,
      "author_name": "aadan2017",
      "author_url": "",
      "post_date": "06/10/2019 01:18:18",
      "content": "<p>I used all the curated data and noisy data. For me, noisy data make the model robuster which could reduce the gap between local lwlrap and LB score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 549645,
      "author_name": "hidehisaarai1213",
      "author_url": "",
      "post_date": "06/10/2019 22:38:20",
      "content": "<p>We used both curated data and noisy data. \nIt's interesting to see that the score improvements caused by the usage of noisy data differ a lot between models, although we used noisy data in the same manner.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "548264": "I used curated nosiy data",
    "548296": "Nothing different with kernels but ensemble 50 models.  Ensemble is all you need,",
    "548302": "I ensemble  models with different folds.",
    "548307": "I tried ensembling but was unsuccessful. I used different model arch. with 5-fold and ensembled the average results from the 5-folds. Was unsuccessful.",
    "548319": "how much  time your inference time is?  I'm afraid you spend more than 20m.",
    "548320": "We took our best single model and predicted all the noisy data.\nThose sounds that have the correct class in top3 predictions of our model we added for uptraining\nAt the same time, we tried to keep our classes balanced and gave priority to those classes that very rarely appeared in the top predictions for a test.\nWe had several choices for noisy files (1500, 2500 and 7500 files)\nNoisy data in our pipeline gave a significant gain",
    "548334": "I was unsuccessful in that it did not improve my results.",
    "548781": "LB went from 0.692 to 0.720 when I added noisy data. We don't remove any data in noisy set.",
    "548811": "Do you just add the curated noisy data to your training with no changes?",
    "548824": "I used all the curated data and noisy data. For me, noisy data make the model robuster which could reduce the gap between local lwlrap and LB score.",
    "548828": "",
    "549645": "We used both curated data and noisy data. \nIt's interesting to see that the score improvements caused by the usage of noisy data differ a lot between models, although we used noisy data in the same manner."
  },
  "source": "meta"
}