{
  "id": 253758,
  "title": "Old dataset vs New dataset",
  "url": "/competitions/seti-breakthrough-listen/discussion/253758",
  "author_name": "shiroe",
  "post_date": "2021-07-18T12:35:06.568000",
  "votes": 11,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I did a quick experiment using old leaky data and new data to find out what works best and also tried to use them together <a href=\"https://www.kaggle.com/mrigendraagrawal/old-data-vs-new-data\" target=\"_blank\">notebook</a><br>\n<strong>Settings:-</strong></p>\n<ul>\n<li>Model:- Efficientnet B0</li>\n<li>CV:- stratified train test split with test size 0.05</li>\n<li>Augmentation:-  no agumentation</li>\n<li>Channels or spatial:- spatial</li>\n</ul>\n<p><strong>Results:-</strong></p>\n<table>\n<thead>\n<tr>\n<th>Dataset</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Old Dataset</td>\n<td>0.9679</td>\n<td>0.692</td>\n</tr>\n<tr>\n<td>New Dataset</td>\n<td>0.7754</td>\n<td>0.680</td>\n</tr>\n<tr>\n<td>Combined</td>\n<td>0.8187</td>\n<td>0.715</td>\n</tr>\n</tbody>\n</table>\n<p>by increasing no. of epochs I found new dataset to give a bit better score compared to old dataset  but they are still very close to each other and combined still give better  result</p>\n<p><strong>Note:-</strong><br>\nI am using both old train and old test labels and to minimize the leak I am normalizing each of the 6 snippets separately as suggested by <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> in <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/248194\" target=\"_blank\">discussion</a></p>",
  "messages": [
    {
      "id": 1392194,
      "postDate": "2021-07-18T12:35:06.570Z",
      "content": "<p>I did a quick experiment using old leaky data and new data to find out what works best and also tried to use them together <a href=\"https://www.kaggle.com/mrigendraagrawal/old-data-vs-new-data\" target=\"_blank\">notebook</a><br>\n<strong>Settings:-</strong></p>\n<ul>\n<li>Model:- Efficientnet B0</li>\n<li>CV:- stratified train test split with test size 0.05</li>\n<li>Augmentation:-  no agumentation</li>\n<li>Channels or spatial:- spatial</li>\n</ul>\n<p><strong>Results:-</strong></p>\n<table>\n<thead>\n<tr>\n<th>Dataset</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Old Dataset</td>\n<td>0.9679</td>\n<td>0.692</td>\n</tr>\n<tr>\n<td>New Dataset</td>\n<td>0.7754</td>\n<td>0.680</td>\n</tr>\n<tr>\n<td>Combined</td>\n<td>0.8187</td>\n<td>0.715</td>\n</tr>\n</tbody>\n</table>\n<p>by increasing no. of epochs I found new dataset to give a bit better score compared to old dataset  but they are still very close to each other and combined still give better  result</p>\n<p><strong>Note:-</strong><br>\nI am using both old train and old test labels and to minimize the leak I am normalizing each of the 6 snippets separately as suggested by <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> in <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/248194\" target=\"_blank\">discussion</a></p>",
      "rawMarkdown": "I did a quick experiment using old leaky data and new data to find out what works best and also tried to use them together [notebook](https://www.kaggle.com/mrigendraagrawal/old-data-vs-new-data)\n**Settings:-**\n- Model:- Efficientnet B0\n- CV:- stratified train test split with test size 0.05\n- Augmentation:-  no agumentation\n- Channels or spatial:- spatial\n\n**Results:-**\n\n| Dataset | CV |LB\n|---  | --- |\n| Old Dataset|0.9679  | 0.692\n| New Dataset|0.7754|0.680\n|Combined|0.8187|0.715\n\nby increasing no. of epochs I found new dataset to give a bit better score compared to old dataset  but they are still very close to each other and combined still give better  result\n\n**Note:-**\nI am using both old train and old test labels and to minimize the leak I am normalizing each of the 6 snippets separately as suggested by @cpmpml in [discussion](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/248194)",
      "votes": 11
    },
    {
      "id": 1394233,
      "postDate": "2021-07-20T08:43:53.320Z",
      "content": "<p>After Combined  cv=0.9205 lb=0.725.<br>\nIt can be seen that there is a score gap of &gt; 0.2 between CV and lb. generally, there is a score gap of 0.1 between CV and lb.<br>\nSo I think it is necessary to give up the old data.</p>",
      "rawMarkdown": "After Combined  cv=0.9205 lb=0.725.\nIt can be seen that there is a score gap of > 0.2 between CV and lb. generally, there is a score gap of 0.1 between CV and lb.\nSo I think it is necessary to give up the old data.",
      "votes": 2,
      "replies": [
        {
          "id": 1394268,
          "postDate": "2021-07-20T09:02:33.657Z",
          "content": "<p>I think new data is much more difficult to classify compared to old data that's why using old data should help. also, the gap between lb and cv could be due to cv containing old data as well</p>",
          "rawMarkdown": "I think new data is much more difficult to classify compared to old data that's why using old data should help. also, the gap between lb and cv could be due to cv containing old data as well",
          "votes": 2
        },
        {
          "id": 1421034,
          "postDate": "2021-08-03T03:45:38.763Z",
          "content": "<p><a href=\"https://www.kaggle.com/zhangeng\" target=\"_blank\">@zhangeng</a> It seems that you have achieved 0.78+ on the leaderboard, which is quite a boost than 0.725. So are you using the combined dataset or new dataset only now?</p>",
          "rawMarkdown": "@zhangeng It seems that you have achieved 0.78+ on the leaderboard, which is quite a boost than 0.725. So are you using the combined dataset or new dataset only now?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1394233,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "2021-07-20T08:43:53.320000",
      "content": "<p>After Combined  cv=0.9205 lb=0.725.<br>\nIt can be seen that there is a score gap of &gt; 0.2 between CV and lb. generally, there is a score gap of 0.1 between CV and lb.<br>\nSo I think it is necessary to give up the old data.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1394268,
          "author_name": "shiroe",
          "author_url": "",
          "post_date": "2021-07-20T09:02:33.657000",
          "content": "<p>I think new data is much more difficult to classify compared to old data that's why using old data should help. also, the gap between lb and cv could be due to cv containing old data as well</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1421034,
          "author_name": "YTEP (Jiazhi Yang)",
          "author_url": "",
          "post_date": "2021-08-03T03:45:38.763000",
          "content": "<p><a href=\"https://www.kaggle.com/zhangeng\" target=\"_blank\">@zhangeng</a> It seems that you have achieved 0.78+ on the leaderboard, which is quite a boost than 0.725. So are you using the combined dataset or new dataset only now?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1392194": "I did a quick experiment using old leaky data and new data to find out what works best and also tried to use them together [notebook](https://www.kaggle.com/mrigendraagrawal/old-data-vs-new-data)\n**Settings:-**\n- Model:- Efficientnet B0\n- CV:- stratified train test split with test size 0.05\n- Augmentation:-  no agumentation\n- Channels or spatial:- spatial\n\n**Results:-**\n\n| Dataset | CV |LB\n|---  | --- |\n| Old Dataset|0.9679  | 0.692\n| New Dataset|0.7754|0.680\n|Combined|0.8187|0.715\n\nby increasing no. of epochs I found new dataset to give a bit better score compared to old dataset  but they are still very close to each other and combined still give better  result\n\n**Note:-**\nI am using both old train and old test labels and to minimize the leak I am normalizing each of the 6 snippets separately as suggested by @cpmpml in [discussion](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/248194)",
    "1394233": "After Combined  cv=0.9205 lb=0.725.\nIt can be seen that there is a score gap of > 0.2 between CV and lb. generally, there is a score gap of 0.1 between CV and lb.\nSo I think it is necessary to give up the old data."
  }
}