{
  "id": 241736,
  "title": "Handle Domain Shift from the Test Data",
  "url": "/competitions/birdclef-2021/discussion/241736",
  "author_name": "",
  "post_date": "2021-05-25T21:39:19.035334200Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi all, I'm quite new at Kaggle. To make sure that my model works, I create the validation set using <code>train_sounscape</code>. I train the model using <code>train_short_audio</code>. </p>\n<p>My trained model can achieve good performance (more than 90% f1-score) on the training data <code>train_short_audio</code>. But it doesn't get any good performance on the validation set. It never achieves more than 0.7 on <code>train_sounscape</code>. I've performed some data augmentation such as AddGaussianNoise, AddGaussionSNR, AddPinkNoise, and AddBackgroundNoise. </p>\n<p>I would really appreciate everyone for sharing some ideas/solutions here. Many thanks 🙏</p>",
  "messages": [
    {
      "id": "1323008",
      "postDate": "05/25/2021 21:39:19",
      "content": "<p>Hi all, I'm quite new at Kaggle. To make sure that my model works, I create the validation set using <code>train_sounscape</code>. I train the model using <code>train_short_audio</code>. </p>\n<p>My trained model can achieve good performance (more than 90% f1-score) on the training data <code>train_short_audio</code>. But it doesn't get any good performance on the validation set. It never achieves more than 0.7 on <code>train_sounscape</code>. I've performed some data augmentation such as AddGaussianNoise, AddGaussionSNR, AddPinkNoise, and AddBackgroundNoise. </p>\n<p>I would really appreciate everyone for sharing some ideas/solutions here. Many thanks 🙏</p>",
      "rawMarkdown": "Hi all, I'm quite new at Kaggle. To make sure that my model works, I create the validation set using `train_sounscape`. I train the model using `train_short_audio`. \n\nMy trained model can achieve good performance (more than 90% f1-score) on the training data `train_short_audio`. But it doesn't get any good performance on the validation set. It never achieves more than 0.7 on `train_sounscape`. I've performed some data augmentation such as AddGaussianNoise, AddGaussionSNR, AddPinkNoise, and AddBackgroundNoise. \n\nI would really appreciate everyone for sharing some ideas/solutions here. Many thanks 🙏",
      "votes": null
    },
    {
      "id": "1323218",
      "postDate": "05/26/2021 05:01:51",
      "content": "<p>train short audios are very noisy, how do you know your labels for cropped train short audios are correct? And 0.7 on train soundscapes is good</p>",
      "rawMarkdown": "train short audios are very noisy, how do you know your labels for cropped train short audios are correct? And 0.7 on train soundscapes is good",
      "votes": null
    },
    {
      "id": "1323254",
      "postDate": "05/26/2021 05:37:54",
      "content": "<blockquote>\n  <p>train short audios are very noisy, how do you know your labels for cropped train short audios are correct?</p>\n</blockquote>\n<p>I didn't realize that the train short audios are very noisy. Just randomly crop using kkiller notebook</p>\n<blockquote>\n  <p>And 0.7 on train soundscapes is good</p>\n</blockquote>\n<p>btw I create my own dataset. 4 models for each class (SSW, SNE, COR, COL). I got 0.7 for only 1 class from it. the remaining are near 0.6. </p>",
      "rawMarkdown": "> train short audios are very noisy, how do you know your labels for cropped train short audios are correct?\n\nI didn't realize that the train short audios are very noisy. Just randomly crop using kkiller notebook\n\n> And 0.7 on train soundscapes is good\n\nbtw I create my own dataset. 4 models for each class (SSW, SNE, COR, COL). I got 0.7 for only 1 class from it. the remaining are near 0.6.",
      "votes": null
    },
    {
      "id": "1325299",
      "postDate": "05/27/2021 16:28:13",
      "content": "<p>0.7&gt; is good on the train soundscapes well done and a good dataset for validation but it's only 40 birds of 397 in it, to have in mind.</p>",
      "rawMarkdown": "0.7> is good on the train soundscapes well done and a good dataset for validation but it's only 40 birds of 397 in it, to have in mind.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1323218,
      "author_name": "superchenhao",
      "author_url": "",
      "post_date": "05/26/2021 05:01:51",
      "content": "<p>train short audios are very noisy, how do you know your labels for cropped train short audios are correct? And 0.7 on train soundscapes is good</p>",
      "votes": null,
      "replies": [
        {
          "id": 1323254,
          "author_name": "mhilmiasyrofi",
          "author_url": "",
          "post_date": "05/26/2021 05:37:54",
          "content": "<blockquote>\n  <p>train short audios are very noisy, how do you know your labels for cropped train short audios are correct?</p>\n</blockquote>\n<p>I didn't realize that the train short audios are very noisy. Just randomly crop using kkiller notebook</p>\n<blockquote>\n  <p>And 0.7 on train soundscapes is good</p>\n</blockquote>\n<p>btw I create my own dataset. 4 models for each class (SSW, SNE, COR, COL). I got 0.7 for only 1 class from it. the remaining are near 0.6. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1325299,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "05/27/2021 16:28:13",
          "content": "<p>0.7&gt; is good on the train soundscapes well done and a good dataset for validation but it's only 40 birds of 397 in it, to have in mind.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1323008": "Hi all, I'm quite new at Kaggle. To make sure that my model works, I create the validation set using `train_sounscape`. I train the model using `train_short_audio`. \n\nMy trained model can achieve good performance (more than 90% f1-score) on the training data `train_short_audio`. But it doesn't get any good performance on the validation set. It never achieves more than 0.7 on `train_sounscape`. I've performed some data augmentation such as AddGaussianNoise, AddGaussionSNR, AddPinkNoise, and AddBackgroundNoise. \n\nI would really appreciate everyone for sharing some ideas/solutions here. Many thanks 🙏",
    "1323218": "train short audios are very noisy, how do you know your labels for cropped train short audios are correct? And 0.7 on train soundscapes is good",
    "1323254": "> train short audios are very noisy, how do you know your labels for cropped train short audios are correct?\n\nI didn't realize that the train short audios are very noisy. Just randomly crop using kkiller notebook\n\n> And 0.7 on train soundscapes is good\n\nbtw I create my own dataset. 4 models for each class (SSW, SNE, COR, COL). I got 0.7 for only 1 class from it. the remaining are near 0.6.",
    "1325299": "0.7> is good on the train soundscapes well done and a good dataset for validation but it's only 40 birds of 397 in it, to have in mind."
  },
  "source": "meta"
}