{
  "id": 55193,
  "title": "Blends and overfitting",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55193",
  "author_name": "",
  "post_date": "2018-04-23T10:29:25.911562400Z",
  "votes": 1,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Are we overfitting to 18% of test data?</p>",
  "messages": [
    {
      "id": "318177",
      "postDate": "04/23/2018 10:29:25",
      "content": "<p>Are we overfitting to 18% of test data?</p>",
      "rawMarkdown": "Are we overfitting to 18% of test data?",
      "votes": null
    },
    {
      "id": "318194",
      "postDate": "04/23/2018 11:22:01",
      "content": "<p>It depend on how you select your models.  If you use the public Lb score to select your models then you probably are overfitting.  </p>",
      "rawMarkdown": "It depend on how you select your models.  If you use the public Lb score to select your models then you probably are overfitting.",
      "votes": null
    },
    {
      "id": "318198",
      "postDate": "04/23/2018 11:35:40",
      "content": "<p>@CPMP many blended kernels i see are using LB score as the guiding principle.</p>",
      "rawMarkdown": "CPMP many blended kernels i see are using LB score as the guiding principle.",
      "votes": null
    },
    {
      "id": "318205",
      "postDate": "04/23/2018 11:48:15",
      "content": "<p>I know...  </p>\n\n<p>This happens in every competition. </p>",
      "rawMarkdown": "I know...  \n\nThis happens in every competition.",
      "votes": null
    },
    {
      "id": "318216",
      "postDate": "04/23/2018 12:18:10",
      "content": "<p>I learnt this the hard way in Mercedes-Benz competition. From 35th rank on Public to 1400 on private. Even though the size of the public set is substantial and much much greater than the Mercedes one, still it is only for an hour. Please don't rely totally on the Public LB. <a href=\"https://www.kaggle.com/c/mercedes-benz-greener-manufacturing/discussion/36136\">in CV you must trust</a></p>",
      "rawMarkdown": "I learnt this the hard way in Mercedes-Benz competition. From 35th rank on Public to 1400 on private. Even though the size of the public set is substantial and much much greater than the Mercedes one, still it is only for an hour. Please don't rely totally on the Public LB. [in CV you must trust][1]\n\n\n  [1]: https://www.kaggle.com/c/mercedes-benz-greener-manufacturing/discussion/36136",
      "votes": null
    },
    {
      "id": "318239",
      "postDate": "04/23/2018 13:08:52",
      "content": "<p>I suffered a similar fate in Mercari contest :(</p>",
      "rawMarkdown": "I suffered a similar fate in Mercari contest :(",
      "votes": null
    },
    {
      "id": "318274",
      "postDate": "04/23/2018 14:14:22",
      "content": "<p>You can refer this kernel: <a href=\"https://www.kaggle.com/alexvonrass/how-to-get-0-5-public-auc/output\">https://www.kaggle.com/alexvonrass/how-to-get-0-5-public-auc/output</a>.</p>\n\n<p>That way I think we could get simulated public / private lb score in validation.</p>",
      "rawMarkdown": "You can refer this kernel: https://www.kaggle.com/alexvonrass/how-to-get-0-5-public-auc/output.\n\nThat way I think we could get simulated public / private lb score in validation.",
      "votes": null
    },
    {
      "id": "318624",
      "postDate": "04/24/2018 06:41:39",
      "content": "<p>thanks</p>",
      "rawMarkdown": "thanks",
      "votes": null
    },
    {
      "id": "318937",
      "postDate": "04/24/2018 19:44:59",
      "content": "<blockquote>\n  <p>Are we overfitting to 18% of test data?</p>\n</blockquote>\n\n<p>I think it depends on who you ask.</p>\n\n<p>From the comments I have read, it seems that a majority of users are training on an incomplete train dataset and doing simple validation rather than N-fold cross-validation. Normally, that would be a recipe for disaster, but it may work out OK here because the dataset is so big. It will depend on whether the (semi-)random portions of data used for training/validation capture well enough the properties of the whole dataset.</p>",
      "rawMarkdown": "&gt; Are we overfitting to 18% of test data?\n\nI think it depends on who you ask.\n\nFrom the comments I have read, it seems that a majority of users are training on an incomplete train dataset and doing simple validation rather than N-fold cross-validation. Normally, that would be a recipe for disaster, but it may work out OK here because the dataset is so big. It will depend on whether the (semi-)random portions of data used for training/validation capture well enough the properties of the whole dataset.",
      "votes": null
    },
    {
      "id": "318941",
      "postDate": "04/24/2018 19:51:33",
      "content": "<p>I am quite curious how that one plays out - my validation is pretty consistent with the public lb, but the performance on the first hour of the holdout and the remaining period is vastly different (as in: 0.979 on the first hour, 0.9817 on the remaining ones). If this is the real deal, there might be some turbulence.</p>",
      "rawMarkdown": "I am quite curious how that one plays out - my validation is pretty consistent with the public lb, but the performance on the first hour of the holdout and the remaining period is vastly different (as in: 0.979 on the first hour, 0.9817 on the remaining ones). If this is the real deal, there might be some turbulence.",
      "votes": null
    },
    {
      "id": "318945",
      "postDate": "04/24/2018 19:58:51",
      "content": "<blockquote>\n  <p>the performance on the first hour of the holdout and the remaining period is vastly different (as in: 0.979 on the first hour, 0.9817 on the remaining ones).</p>\n</blockquote>\n\n<p>Doesn't surprise me, though I haven't tried it directly. I'd rather train on full train data and get a lower score that I can understand than train on a part of data and get a higher score.</p>",
      "rawMarkdown": "&gt; the performance on the first hour of the holdout and the remaining period is vastly different (as in: 0.979 on the first hour, 0.9817 on the remaining ones).\n\nDoesn't surprise me, though I haven't tried it directly. I'd rather train on full train data and get a lower score that I can understand than train on a part of data and get a higher score.",
      "votes": null
    },
    {
      "id": "318958",
      "postDate": "04/24/2018 21:17:13",
      "content": "<p>What if once we become confident on our validation setting and it is consistent with LB, then do a N-Fold average of your model or better train our model for each day and average the results. I may bet on that.</p>",
      "rawMarkdown": "What if once we become confident on our validation setting and it is consistent with LB, then do a N-Fold average of your model or better train our model for each day and average the results. I may bet on that.",
      "votes": null
    },
    {
      "id": "318991",
      "postDate": "04/25/2018 01:27:04",
      "content": "<p>@Konrad, my local validation on full day 9 is also significantly higher than on hour 4 of that day.  I think others also reported it.</p>",
      "rawMarkdown": "Konrad, my local validation on full day 9 is also significantly higher than on hour 4 of that day.  I think others also reported it.",
      "votes": null
    },
    {
      "id": "319001",
      "postDate": "04/25/2018 02:10:29",
      "content": "<p>Completely agree with you.It's very important that the sample we are using for training is representattive of the whole training set. I took a simple correlation matrix to check whether the sample and test data were similar. But, it seems otherwise.</p>",
      "rawMarkdown": "Completely agree with you.It's very important that the sample we are using for training is representattive of the whole training set. I took a simple correlation matrix to check whether the sample and test data were similar. But, it seems otherwise.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 318194,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/23/2018 11:22:01",
      "content": "<p>It depend on how you select your models.  If you use the public Lb score to select your models then you probably are overfitting.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 318198,
          "author_name": "maheshak04",
          "author_url": "",
          "post_date": "04/23/2018 11:35:40",
          "content": "<p>@CPMP many blended kernels i see are using LB score as the guiding principle.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318205,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/23/2018 11:48:15",
          "content": "<p>I know...  </p>\n\n<p>This happens in every competition. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318216,
          "author_name": "ug2409",
          "author_url": "",
          "post_date": "04/23/2018 12:18:10",
          "content": "<p>I learnt this the hard way in Mercedes-Benz competition. From 35th rank on Public to 1400 on private. Even though the size of the public set is substantial and much much greater than the Mercedes one, still it is only for an hour. Please don't rely totally on the Public LB. <a href=\"https://www.kaggle.com/c/mercedes-benz-greener-manufacturing/discussion/36136\">in CV you must trust</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318239,
          "author_name": "maheshak04",
          "author_url": "",
          "post_date": "04/23/2018 13:08:52",
          "content": "<p>I suffered a similar fate in Mercari contest :(</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 318274,
      "author_name": "bangdasun",
      "author_url": "",
      "post_date": "04/23/2018 14:14:22",
      "content": "<p>You can refer this kernel: <a href=\"https://www.kaggle.com/alexvonrass/how-to-get-0-5-public-auc/output\">https://www.kaggle.com/alexvonrass/how-to-get-0-5-public-auc/output</a>.</p>\n\n<p>That way I think we could get simulated public / private lb score in validation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 318624,
          "author_name": "maheshak04",
          "author_url": "",
          "post_date": "04/24/2018 06:41:39",
          "content": "<p>thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 318937,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "04/24/2018 19:44:59",
      "content": "<blockquote>\n  <p>Are we overfitting to 18% of test data?</p>\n</blockquote>\n\n<p>I think it depends on who you ask.</p>\n\n<p>From the comments I have read, it seems that a majority of users are training on an incomplete train dataset and doing simple validation rather than N-fold cross-validation. Normally, that would be a recipe for disaster, but it may work out OK here because the dataset is so big. It will depend on whether the (semi-)random portions of data used for training/validation capture well enough the properties of the whole dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 318941,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "04/24/2018 19:51:33",
          "content": "<p>I am quite curious how that one plays out - my validation is pretty consistent with the public lb, but the performance on the first hour of the holdout and the remaining period is vastly different (as in: 0.979 on the first hour, 0.9817 on the remaining ones). If this is the real deal, there might be some turbulence.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318945,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "04/24/2018 19:58:51",
          "content": "<blockquote>\n  <p>the performance on the first hour of the holdout and the remaining period is vastly different (as in: 0.979 on the first hour, 0.9817 on the remaining ones).</p>\n</blockquote>\n\n<p>Doesn't surprise me, though I haven't tried it directly. I'd rather train on full train data and get a lower score that I can understand than train on a part of data and get a higher score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318958,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "04/24/2018 21:17:13",
          "content": "<p>What if once we become confident on our validation setting and it is consistent with LB, then do a N-Fold average of your model or better train our model for each day and average the results. I may bet on that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318991,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/25/2018 01:27:04",
          "content": "<p>@Konrad, my local validation on full day 9 is also significantly higher than on hour 4 of that day.  I think others also reported it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 319001,
          "author_name": "maheshak04",
          "author_url": "",
          "post_date": "04/25/2018 02:10:29",
          "content": "<p>Completely agree with you.It's very important that the sample we are using for training is representattive of the whole training set. I took a simple correlation matrix to check whether the sample and test data were similar. But, it seems otherwise.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "318177": "Are we overfitting to 18% of test data?",
    "318194": "It depend on how you select your models.  If you use the public Lb score to select your models then you probably are overfitting.",
    "318198": "CPMP many blended kernels i see are using LB score as the guiding principle.",
    "318205": "I know...  \n\nThis happens in every competition.",
    "318216": "I learnt this the hard way in Mercedes-Benz competition. From 35th rank on Public to 1400 on private. Even though the size of the public set is substantial and much much greater than the Mercedes one, still it is only for an hour. Please don't rely totally on the Public LB. [in CV you must trust][1]\n\n\n  [1]: https://www.kaggle.com/c/mercedes-benz-greener-manufacturing/discussion/36136",
    "318239": "I suffered a similar fate in Mercari contest :(",
    "318274": "You can refer this kernel: https://www.kaggle.com/alexvonrass/how-to-get-0-5-public-auc/output.\n\nThat way I think we could get simulated public / private lb score in validation.",
    "318624": "thanks",
    "318937": "&gt; Are we overfitting to 18% of test data?\n\nI think it depends on who you ask.\n\nFrom the comments I have read, it seems that a majority of users are training on an incomplete train dataset and doing simple validation rather than N-fold cross-validation. Normally, that would be a recipe for disaster, but it may work out OK here because the dataset is so big. It will depend on whether the (semi-)random portions of data used for training/validation capture well enough the properties of the whole dataset.",
    "318941": "I am quite curious how that one plays out - my validation is pretty consistent with the public lb, but the performance on the first hour of the holdout and the remaining period is vastly different (as in: 0.979 on the first hour, 0.9817 on the remaining ones). If this is the real deal, there might be some turbulence.",
    "318945": "&gt; the performance on the first hour of the holdout and the remaining period is vastly different (as in: 0.979 on the first hour, 0.9817 on the remaining ones).\n\nDoesn't surprise me, though I haven't tried it directly. I'd rather train on full train data and get a lower score that I can understand than train on a part of data and get a higher score.",
    "318958": "What if once we become confident on our validation setting and it is consistent with LB, then do a N-Fold average of your model or better train our model for each day and average the results. I may bet on that.",
    "318991": "Konrad, my local validation on full day 9 is also significantly higher than on hour 4 of that day.  I think others also reported it.",
    "319001": "Completely agree with you.It's very important that the sample we are using for training is representattive of the whole training set. I took a simple correlation matrix to check whether the sample and test data were similar. But, it seems otherwise."
  },
  "source": "meta"
}