{
  "id": 88497,
  "title": "How to break the baseline",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/88497",
  "author_name": "",
  "post_date": "2019-04-09T03:06:56.986253300Z",
  "votes": 17,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Here's my method so far:\n- I used Melspectrogram features like the Basic Solution kernel. <a href=\"https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-powered-by-fast-ai\">https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-powered-by-fast-ai</a>\n- I used both the noisy and curated set, as reported in <a href=\"https://www.kaggle.com/davids1992/i-ve-listened-to-the-entire-test-set-bonus\">this kernel </a>. Not only are the noisy set, well, noisy, they seem to contain quite different kind of sounds from the curated one. I got about 0.45 LOL score on that set, which only translated to 0.25 on the curated set, so I'm not sure it will be a good idea to combine them together. Instead, I used the noisy set to warm up my CNN, which would be then finetuned on the curated set (who said you aren't allowed to use pretrained models?). This way I think we can also avoid overfitting.\n- My single submission so far scored about 0.5+, blending a few of them will easily break the baseline.</p>\n\n<p>Any way, good luck with the competition :)</p>",
  "messages": [
    {
      "id": "510394",
      "postDate": "04/09/2019 03:06:56",
      "content": "<p>Here's my method so far:\n- I used Melspectrogram features like the Basic Solution kernel. <a href=\"https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-powered-by-fast-ai\">https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-powered-by-fast-ai</a>\n- I used both the noisy and curated set, as reported in <a href=\"https://www.kaggle.com/davids1992/i-ve-listened-to-the-entire-test-set-bonus\">this kernel </a>. Not only are the noisy set, well, noisy, they seem to contain quite different kind of sounds from the curated one. I got about 0.45 LOL score on that set, which only translated to 0.25 on the curated set, so I'm not sure it will be a good idea to combine them together. Instead, I used the noisy set to warm up my CNN, which would be then finetuned on the curated set (who said you aren't allowed to use pretrained models?). This way I think we can also avoid overfitting.\n- My single submission so far scored about 0.5+, blending a few of them will easily break the baseline.</p>\n\n<p>Any way, good luck with the competition :)</p>",
      "rawMarkdown": "Here's my method so far:\n- I used Melspectrogram features like the Basic Solution kernel. https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-powered-by-fast-ai\n- I used both the noisy and curated set, as reported in [this kernel ](https://www.kaggle.com/davids1992/i-ve-listened-to-the-entire-test-set-bonus). Not only are the noisy set, well, noisy, they seem to contain quite different kind of sounds from the curated one. I got about 0.45 LOL score on that set, which only translated to 0.25 on the curated set, so I'm not sure it will be a good idea to combine them together. Instead, I used the noisy set to warm up my CNN, which would be then finetuned on the curated set (who said you aren't allowed to use pretrained models?). This way I think we can also avoid overfitting.\n- My single submission so far scored about 0.5+, blending a few of them will easily break the baseline.\n\nAny way, good luck with the competition :)",
      "votes": null
    },
    {
      "id": "510403",
      "postDate": "04/09/2019 03:22:25",
      "content": "<p>Nice work. The transfer learning approach that you've described is pretty much how the baseline also works and is a good starting point for tackling the domain mismatch and noise.  We hope that all of you can do a lot better than just beat the baseline, though! There are more ways of tackling the mismatch and noise problems and taking advantage of the nearly 80 hours of noisy audio.</p>",
      "rawMarkdown": "Nice work. The transfer learning approach that you've described is pretty much how the baseline also works and is a good starting point for tackling the domain mismatch and noise.  We hope that all of you can do a lot better than just beat the baseline, though! There are more ways of tackling the mismatch and noise problems and taking advantage of the nearly 80 hours of noisy audio.",
      "votes": null
    },
    {
      "id": "510422",
      "postDate": "04/09/2019 04:02:02",
      "content": "<p>Hi, thanks for sharing. I was curious how the other competitors were doing. :)\nI also agree your point.</p>",
      "rawMarkdown": "Hi, thanks for sharing. I was curious how the other competitors were doing. :)\nI also agree your point.",
      "votes": null
    },
    {
      "id": "515477",
      "postDate": "04/12/2019 17:13:34",
      "content": "<p>Hey ! Nice job, I'm also 0.5+ and we have basically the same approach. Have you found an efficient way to convert the sounds into images ? Doing it on the curated set takes me about 10min and the noisy ones just takes forever.</p>\n\n<p>Good luck !</p>",
      "rawMarkdown": "Hey ! Nice job, I'm also 0.5+ and we have basically the same approach. Have you found an efficient way to convert the sounds into images ? Doing it on the curated set takes me about 10min and the noisy ones just takes forever.\n\nGood luck !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 510403,
      "author_name": "plakal",
      "author_url": "",
      "post_date": "04/09/2019 03:22:25",
      "content": "<p>Nice work. The transfer learning approach that you've described is pretty much how the baseline also works and is a good starting point for tackling the domain mismatch and noise.  We hope that all of you can do a lot better than just beat the baseline, though! There are more ways of tackling the mismatch and noise problems and taking advantage of the nearly 80 hours of noisy audio.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 510422,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "04/09/2019 04:02:02",
      "content": "<p>Hi, thanks for sharing. I was curious how the other competitors were doing. :)\nI also agree your point.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 515477,
      "author_name": "nathanh12",
      "author_url": "",
      "post_date": "04/12/2019 17:13:34",
      "content": "<p>Hey ! Nice job, I'm also 0.5+ and we have basically the same approach. Have you found an efficient way to convert the sounds into images ? Doing it on the curated set takes me about 10min and the noisy ones just takes forever.</p>\n\n<p>Good luck !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "510394": "Here's my method so far:\n- I used Melspectrogram features like the Basic Solution kernel. https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-powered-by-fast-ai\n- I used both the noisy and curated set, as reported in [this kernel ](https://www.kaggle.com/davids1992/i-ve-listened-to-the-entire-test-set-bonus). Not only are the noisy set, well, noisy, they seem to contain quite different kind of sounds from the curated one. I got about 0.45 LOL score on that set, which only translated to 0.25 on the curated set, so I'm not sure it will be a good idea to combine them together. Instead, I used the noisy set to warm up my CNN, which would be then finetuned on the curated set (who said you aren't allowed to use pretrained models?). This way I think we can also avoid overfitting.\n- My single submission so far scored about 0.5+, blending a few of them will easily break the baseline.\n\nAny way, good luck with the competition :)",
    "510403": "Nice work. The transfer learning approach that you've described is pretty much how the baseline also works and is a good starting point for tackling the domain mismatch and noise.  We hope that all of you can do a lot better than just beat the baseline, though! There are more ways of tackling the mismatch and noise problems and taking advantage of the nearly 80 hours of noisy audio.",
    "510422": "Hi, thanks for sharing. I was curious how the other competitors were doing. :)\nI also agree your point.",
    "515477": "Hey ! Nice job, I'm also 0.5+ and we have basically the same approach. Have you found an efficient way to convert the sounds into images ? Doing it on the curated set takes me about 10min and the noisy ones just takes forever.\n\nGood luck !"
  },
  "source": "meta"
}