{
  "id": 338140,
  "title": "How overfit is the leaderboard?",
  "url": "/competitions/amex-default-prediction/discussion/338140",
  "author_name": "raddar",
  "post_date": "2022-07-19T10:56:56.552000",
  "votes": 49,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Going to share my take of how I feel about the current state of the leaderboard.</p>\n<p>I have 3 strong models which:</p>\n<ul>\n<li>ensemble scores 0.8001 on CV</li>\n<li>LB score 0.7999 (quite confident about 4th digit as its the best 0.799 model I have submitted when ranked in submission page)</li>\n</ul>\n<p>That's good enough for ~#30 place on the current state of LB.</p>\n<p>Then I average my best CV ensemble  with best public notebook submission - submit multiple times to optimize ensemble weights to maximize public LB score - current #11th place. <br>\nBased on current standings I would extrapolate that to LB score of 0.8007 public LB score.</p>\n<p>I am very confident that the public LB model did not improve my model, as I have reused it in my current models on my folds. </p>\n<p>What does this mean?</p>\n<ul>\n<li>I simply gained +0.0008 on LB by using the available public kernels. I am fairly confident that this gain will not hold on the private LB - in fact I am not going to choose this submission as my final submissions at the end of the competition</li>\n<li>If you use several models weights optimized on public LB ranking, you are essentially doing the very same thing as I did with my experiment. It's even worse as it is even much harder to measure how much you have overfitted and might make you overconfident of your LB rankings. </li>\n<li>Ignore available public submission files - they were not trained on your CV folds - you can still take the ideas from them, or retrain on your folds :)</li>\n<li>Optimize model weights based on CV score. Don't optimize model ensembles based on the leaderboard results.</li>\n</ul>",
  "messages": [
    {
      "id": 1861935,
      "postDate": "2022-07-19T10:56:56.553Z",
      "content": "<p>Going to share my take of how I feel about the current state of the leaderboard.</p>\n<p>I have 3 strong models which:</p>\n<ul>\n<li>ensemble scores 0.8001 on CV</li>\n<li>LB score 0.7999 (quite confident about 4th digit as its the best 0.799 model I have submitted when ranked in submission page)</li>\n</ul>\n<p>That's good enough for ~#30 place on the current state of LB.</p>\n<p>Then I average my best CV ensemble  with best public notebook submission - submit multiple times to optimize ensemble weights to maximize public LB score - current #11th place. <br>\nBased on current standings I would extrapolate that to LB score of 0.8007 public LB score.</p>\n<p>I am very confident that the public LB model did not improve my model, as I have reused it in my current models on my folds. </p>\n<p>What does this mean?</p>\n<ul>\n<li>I simply gained +0.0008 on LB by using the available public kernels. I am fairly confident that this gain will not hold on the private LB - in fact I am not going to choose this submission as my final submissions at the end of the competition</li>\n<li>If you use several models weights optimized on public LB ranking, you are essentially doing the very same thing as I did with my experiment. It's even worse as it is even much harder to measure how much you have overfitted and might make you overconfident of your LB rankings. </li>\n<li>Ignore available public submission files - they were not trained on your CV folds - you can still take the ideas from them, or retrain on your folds :)</li>\n<li>Optimize model weights based on CV score. Don't optimize model ensembles based on the leaderboard results.</li>\n</ul>",
      "rawMarkdown": "Going to share my take of how I feel about the current state of the leaderboard.\n\nI have 3 strong models which:\n- ensemble scores 0.8001 on CV\n- LB score 0.7999 (quite confident about 4th digit as its the best 0.799 model I have submitted when ranked in submission page)\n\nThat's good enough for ~#30 place on the current state of LB.\n\nThen I average my best CV ensemble  with best public notebook submission - submit multiple times to optimize ensemble weights to maximize public LB score - current #11th place. \nBased on current standings I would extrapolate that to LB score of 0.8007 public LB score.\n\nI am very confident that the public LB model did not improve my model, as I have reused it in my current models on my folds. \n\nWhat does this mean?\n- I simply gained +0.0008 on LB by using the available public kernels. I am fairly confident that this gain will not hold on the private LB - in fact I am not going to choose this submission as my final submissions at the end of the competition\n- If you use several models weights optimized on public LB ranking, you are essentially doing the very same thing as I did with my experiment. It's even worse as it is even much harder to measure how much you have overfitted and might make you overconfident of your LB rankings. \n- Ignore available public submission files - they were not trained on your CV folds - you can still take the ideas from them, or retrain on your folds :)\n- Optimize model weights based on CV score. Don't optimize model ensembles based on the leaderboard results.\n\n\n",
      "votes": 48
    },
    {
      "id": 1862702,
      "postDate": "2022-07-19T22:59:38.583Z",
      "content": "<p>In fact, if your PC has enough RAM or batch training for you is feasible, you can simply add all the useful features from the public notebook to your own model. If you can get to 0.800+ through blend, that should already give you a score of 0.800+ with extra features. As for blend, you must carefully analyze the correlation between the public notebook output and the output of your original model. Even so, we should be careful. </p>",
      "rawMarkdown": "In fact, if your PC has enough RAM or batch training for you is feasible, you can simply add all the useful features from the public notebook to your own model. If you can get to 0.800+ through blend, that should already give you a score of 0.800+ with extra features. As for blend, you must carefully analyze the correlation between the public notebook output and the output of your original model. Even so, we should be careful. ",
      "votes": 4
    },
    {
      "id": 1872937,
      "postDate": "2022-07-27T10:36:29.387Z",
      "content": "<p>Don’t forget about covid…. feels weird that everyone is fighting for the 4th decimal place on CV/LB while the private will be affected by the health crisis of the century.</p>",
      "rawMarkdown": "Don’t forget about covid.... feels weird that everyone is fighting for the 4th decimal place on CV/LB while the private will be affected by the health crisis of the century.",
      "votes": 1
    },
    {
      "id": 1868329,
      "postDate": "2022-07-23T23:10:59.710Z",
      "content": "<p>Great insights! I think you've really hit the nail on the head with regard to the current state of the leaderboard.<br>\nI think you're also right that it's very hard to measure how much you've overfit on the public LB, and that it's probably not worth the risk in the end.</p>",
      "rawMarkdown": "Great insights! I think you've really hit the nail on the head with regard to the current state of the leaderboard.\nI think you're also right that it's very hard to measure how much you've overfit on the public LB, and that it's probably not worth the risk in the end.\n\n",
      "votes": 1
    },
    {
      "id": 1865078,
      "postDate": "2022-07-21T13:59:44.970Z",
      "content": "<p>good job !!!</p>",
      "rawMarkdown": "good job !!!\n",
      "votes": 1
    },
    {
      "id": 1864957,
      "postDate": "2022-07-21T11:59:27.840Z",
      "content": "<p>yes yes yes !!!1</p>",
      "rawMarkdown": "yes yes yes !!!1",
      "votes": 1
    },
    {
      "id": 1862804,
      "postDate": "2022-07-20T02:39:16.737Z",
      "content": "<p>The best score in the public notebook is using DART, your three models, are you using DART?   😯</p>",
      "rawMarkdown": "The best score in the public notebook is using DART, your three models, are you using DART?   😯",
      "votes": 1,
      "replies": [
        {
          "id": 1863200,
          "postDate": "2022-07-20T08:00:15.910Z",
          "content": "<p>one of them - yes</p>",
          "rawMarkdown": "one of them - yes",
          "votes": 5
        }
      ]
    },
    {
      "id": 1864227,
      "postDate": "2022-07-20T21:08:58.217Z",
      "content": "<p>Do you think that optimizing models must be only by cv scores?</p>",
      "rawMarkdown": "Do you think that optimizing models must be only by cv scores?",
      "votes": 2,
      "replies": [
        {
          "id": 1864239,
          "postDate": "2022-07-20T21:14:51.117Z",
          "content": "<p>For this competition - 100% sure</p>",
          "rawMarkdown": "For this competition - 100% sure",
          "votes": 7
        },
        {
          "id": 1864254,
          "postDate": "2022-07-20T21:27:01.337Z",
          "content": "<p>I agree with you that if the proportion of positive and negative samples in the training and test sets is basically the same, and the data distribution of public and private LB are also similar. Then you should trust cv. But you may want to set a tolerance standard deviation of about 0.0001. One more thing, 51% of the data is in the public LB, it should not be ignored.</p>",
          "rawMarkdown": "I agree with you that if the proportion of positive and negative samples in the training and test sets is basically the same, and the data distribution of public and private LB are also similar. Then you should trust cv. But you may want to set a tolerance standard deviation of about 0.0001. One more thing, 51% of the data is in the public LB, it should not be ignored.",
          "votes": 6
        }
      ]
    },
    {
      "id": 1862176,
      "postDate": "2022-07-19T13:55:55.220Z",
      "content": "<p>We are are having something similar with our best CV ensemble. If I may ask based on OOF are you just doing a average ensemble or also trying weighted or stacking strategies (all based on oof of course) to boost CV .</p>",
      "rawMarkdown": "We are are having something similar with our best CV ensemble. If I may ask based on OOF are you just doing a average ensemble or also trying weighted or stacking strategies (all based on oof of course) to boost CV .",
      "votes": 2,
      "replies": [
        {
          "id": 1862357,
          "postDate": "2022-07-19T16:15:01.280Z",
          "content": "<p>simple weighted average</p>",
          "rawMarkdown": "simple weighted average",
          "votes": 8
        },
        {
          "id": 1862491,
          "postDate": "2022-07-19T18:33:30.133Z",
          "content": "<p>May I ask if you are averaging probabilities, ranked probabilities or the raw output from the models? 😸</p>",
          "rawMarkdown": "May I ask if you are averaging probabilities, ranked probabilities or the raw output from the models? 😸",
          "votes": 1
        },
        {
          "id": 1862537,
          "postDate": "2022-07-19T19:21:55.067Z",
          "content": "<p>ranked probabilities</p>",
          "rawMarkdown": "ranked probabilities",
          "votes": 7
        },
        {
          "id": 1864217,
          "postDate": "2022-07-20T20:57:59.390Z",
          "content": "<p>sorry, what do you mean by ranked probabilities?</p>",
          "rawMarkdown": "sorry, what do you mean by ranked probabilities?",
          "votes": 1
        },
        {
          "id": 1864237,
          "postDate": "2022-07-20T21:14:29.250Z",
          "content": "<p><a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.rankdata.html\" target=\"_blank\">https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.rankdata.html</a></p>",
          "rawMarkdown": "https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.rankdata.html",
          "votes": 5
        }
      ]
    },
    {
      "id": 1862003,
      "postDate": "2022-07-19T11:40:21.310Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1862702,
      "author_name": "Mengfei Li",
      "author_url": "",
      "post_date": "2022-07-19T22:59:38.583000",
      "content": "<p>In fact, if your PC has enough RAM or batch training for you is feasible, you can simply add all the useful features from the public notebook to your own model. If you can get to 0.800+ through blend, that should already give you a score of 0.800+ with extra features. As for blend, you must carefully analyze the correlation between the public notebook output and the output of your original model. Even so, we should be careful. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1872937,
      "author_name": "Lucas Morin",
      "author_url": "",
      "post_date": "2022-07-27T10:36:29.387000",
      "content": "<p>Don’t forget about covid…. feels weird that everyone is fighting for the 4th decimal place on CV/LB while the private will be affected by the health crisis of the century.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1868329,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-07-23T23:10:59.710000",
      "content": "<p>Great insights! I think you've really hit the nail on the head with regard to the current state of the leaderboard.<br>\nI think you're also right that it's very hard to measure how much you've overfit on the public LB, and that it's probably not worth the risk in the end.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1865078,
      "author_name": "Fares Abbas",
      "author_url": "",
      "post_date": "2022-07-21T13:59:44.970000",
      "content": "<p>good job !!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1864957,
      "author_name": "Pratik Bharamgude",
      "author_url": "",
      "post_date": "2022-07-21T11:59:27.840000",
      "content": "<p>yes yes yes !!!1</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1862804,
      "author_name": "kgxiao",
      "author_url": "",
      "post_date": "2022-07-20T02:39:16.737000",
      "content": "<p>The best score in the public notebook is using DART, your three models, are you using DART?   😯</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1863200,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-07-20T08:00:15.910000",
          "content": "<p>one of them - yes</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1864227,
      "author_name": "Oday Mourad",
      "author_url": "",
      "post_date": "2022-07-20T21:08:58.217000",
      "content": "<p>Do you think that optimizing models must be only by cv scores?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1864239,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-07-20T21:14:51.117000",
          "content": "<p>For this competition - 100% sure</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1864254,
          "author_name": "Mengfei Li",
          "author_url": "",
          "post_date": "2022-07-20T21:27:01.337000",
          "content": "<p>I agree with you that if the proportion of positive and negative samples in the training and test sets is basically the same, and the data distribution of public and private LB are also similar. Then you should trust cv. But you may want to set a tolerance standard deviation of about 0.0001. One more thing, 51% of the data is in the public LB, it should not be ignored.</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 1862176,
      "author_name": "Gaurav Rawat",
      "author_url": "",
      "post_date": "2022-07-19T13:55:55.220000",
      "content": "<p>We are are having something similar with our best CV ensemble. If I may ask based on OOF are you just doing a average ensemble or also trying weighted or stacking strategies (all based on oof of course) to boost CV .</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1862357,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-07-19T16:15:01.280000",
          "content": "<p>simple weighted average</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1862491,
          "author_name": "delai50",
          "author_url": "",
          "post_date": "2022-07-19T18:33:30.133000",
          "content": "<p>May I ask if you are averaging probabilities, ranked probabilities or the raw output from the models? 😸</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1862537,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-07-19T19:21:55.067000",
          "content": "<p>ranked probabilities</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1864217,
          "author_name": "mavillan",
          "author_url": "",
          "post_date": "2022-07-20T20:57:59.390000",
          "content": "<p>sorry, what do you mean by ranked probabilities?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1864237,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-07-20T21:14:29.250000",
          "content": "<p><a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.rankdata.html\" target=\"_blank\">https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.rankdata.html</a></p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1862003,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-19T11:40:21.310000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1861935": "Going to share my take of how I feel about the current state of the leaderboard.\n\nI have 3 strong models which:\n- ensemble scores 0.8001 on CV\n- LB score 0.7999 (quite confident about 4th digit as its the best 0.799 model I have submitted when ranked in submission page)\n\nThat's good enough for ~#30 place on the current state of LB.\n\nThen I average my best CV ensemble  with best public notebook submission - submit multiple times to optimize ensemble weights to maximize public LB score - current #11th place. \nBased on current standings I would extrapolate that to LB score of 0.8007 public LB score.\n\nI am very confident that the public LB model did not improve my model, as I have reused it in my current models on my folds. \n\nWhat does this mean?\n- I simply gained +0.0008 on LB by using the available public kernels. I am fairly confident that this gain will not hold on the private LB - in fact I am not going to choose this submission as my final submissions at the end of the competition\n- If you use several models weights optimized on public LB ranking, you are essentially doing the very same thing as I did with my experiment. It's even worse as it is even much harder to measure how much you have overfitted and might make you overconfident of your LB rankings. \n- Ignore available public submission files - they were not trained on your CV folds - you can still take the ideas from them, or retrain on your folds :)\n- Optimize model weights based on CV score. Don't optimize model ensembles based on the leaderboard results.\n\n\n",
    "1862702": "In fact, if your PC has enough RAM or batch training for you is feasible, you can simply add all the useful features from the public notebook to your own model. If you can get to 0.800+ through blend, that should already give you a score of 0.800+ with extra features. As for blend, you must carefully analyze the correlation between the public notebook output and the output of your original model. Even so, we should be careful. ",
    "1872937": "Don’t forget about covid.... feels weird that everyone is fighting for the 4th decimal place on CV/LB while the private will be affected by the health crisis of the century.",
    "1868329": "Great insights! I think you've really hit the nail on the head with regard to the current state of the leaderboard.\nI think you're also right that it's very hard to measure how much you've overfit on the public LB, and that it's probably not worth the risk in the end.\n\n",
    "1865078": "good job !!!\n",
    "1864957": "yes yes yes !!!1",
    "1862804": "The best score in the public notebook is using DART, your three models, are you using DART?   😯",
    "1864227": "Do you think that optimizing models must be only by cv scores?",
    "1862176": "We are are having something similar with our best CV ensemble. If I may ask based on OOF are you just doing a average ensemble or also trying weighted or stacking strategies (all based on oof of course) to boost CV .",
    "1862003": ""
  }
}