{
  "id": 329738,
  "title": "Big \"positional\" shakeup",
  "url": "/competitions/amex-default-prediction/discussion/329738",
  "author_name": "",
  "post_date": "2022-06-08T12:15:13.954603100Z",
  "votes": 47,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Even though this competition has only just started, I am already going to be so bold as to suggest that there will be a big \"positional\" shakeup in this competition.</p>\n<p>There are two types of shakeup, a <strong><em>score</em></strong> shakeup, and a <strong><em>positional</em></strong> or leaderboard shakeup. Naturally they are interlinked, however, one is manageable and the other, not so much. </p>\n<p>A score shakeup, a big difference between Public and Private Leaderboard scores, is either due to overfitting, or the distribution of the Private Leaderboard data being very different from either the training data and/or the Public Leaderboard data. There are countless topics on overfitting, and avoiding it is one of the hallmarks of the great kaggle competitors.</p>\n<p>However, a <em>positional</em> shakeup is a different creature; all data is noisy, and has an <em>irreducible error</em> component, and there is (almost by definition) nothing that can be done about it. In a positional shakeup one can have very comparable Public and Private scores, to with a reasonable and totally acceptable statistical difference. However, on kaggle this small difference could actually result in a huge jump in LB positions. It was pointed out by AmbrosM that the <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329088#1811782\" target=\"_blank\">metric is rather noisy</a>, and we are all subject to its vagaries. </p>\n<p>A look at the leaderboard after only two weeks shows the LB is already very compact;  there are already 400 teams in the range [0.793, 0.797] (a difference of only 0.005).  Add to that situation the possibility of <em>forkmitting</em>, which is where numerous people may submit the very same <code>submission.csv</code>, creating an <a href=\"https://www.kaggle.com/carlmcbrideellis/shakeup-scatterplots-boxes-strings-and-things\" target=\"_blank\">isoscore string</a> where a large block of people all have exactly the same both Public and Private scores, again it is possible to experience a large shakeup based on the most minuscule of differences: if one is just below the isoscore string value on the Public LB, but just above it on the Private LB, then one will leapfrog all of the people in the block at once. </p>\n<p>And finally, suffice to say, it is also very hard not to fall into the trap of fitting to the Public leaderboard; after all, who does not like to see themselves climb positions with each submission!</p>\n<p>So, all in all, in this competition I somewhat suspect, despite best efforts not to overfit,  we shall still be in for a quite substantial \"positional\" shakeup; many of the the medals (or no medals) will be awarded on the fourth significant figure, which is basically noise. In view of this, I think anyone whose model has the same Public and Private LB scores will have every right to be proud of their work, and the medals are what they are; just a part of kaggle.</p>\n<p>All the best,<br>\ncarl</p>\n<p>PS: By the way, if one has an idle five minutes whilst running their model that will definitely have a LB score of 0.800 <em>this</em> time, here are a couple of thought provoking papers perhaps worth perusing: <a href=\"https://proceedings.neurips.cc/paper/2019/file/ee39e503b6bedf0c98c388b7e8589aca-Paper.pdf\" target=\"_blank\">\"A Meta-Analysis of Overfitting in Machine Learning\"</a> where the authors analyze overfitting by looking at 120 kaggle competition leaderboards, and also the paper <a href=\"https://arxiv.org/pdf/2109.08203.pdf\" target=\"_blank\">\"<code>torch.manual seed(3407)</code> is all you need: On the influence of random seeds in deep learning architectures for computer vision\"</a> which, despite the title, is also applicable to tabular predictions.</p>",
  "messages": [
    {
      "id": "1814860",
      "postDate": "06/08/2022 12:15:13",
      "content": "<p>Even though this competition has only just started, I am already going to be so bold as to suggest that there will be a big \"positional\" shakeup in this competition.</p>\n<p>There are two types of shakeup, a <strong><em>score</em></strong> shakeup, and a <strong><em>positional</em></strong> or leaderboard shakeup. Naturally they are interlinked, however, one is manageable and the other, not so much. </p>\n<p>A score shakeup, a big difference between Public and Private Leaderboard scores, is either due to overfitting, or the distribution of the Private Leaderboard data being very different from either the training data and/or the Public Leaderboard data. There are countless topics on overfitting, and avoiding it is one of the hallmarks of the great kaggle competitors.</p>\n<p>However, a <em>positional</em> shakeup is a different creature; all data is noisy, and has an <em>irreducible error</em> component, and there is (almost by definition) nothing that can be done about it. In a positional shakeup one can have very comparable Public and Private scores, to with a reasonable and totally acceptable statistical difference. However, on kaggle this small difference could actually result in a huge jump in LB positions. It was pointed out by AmbrosM that the <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329088#1811782\" target=\"_blank\">metric is rather noisy</a>, and we are all subject to its vagaries. </p>\n<p>A look at the leaderboard after only two weeks shows the LB is already very compact;  there are already 400 teams in the range [0.793, 0.797] (a difference of only 0.005).  Add to that situation the possibility of <em>forkmitting</em>, which is where numerous people may submit the very same <code>submission.csv</code>, creating an <a href=\"https://www.kaggle.com/carlmcbrideellis/shakeup-scatterplots-boxes-strings-and-things\" target=\"_blank\">isoscore string</a> where a large block of people all have exactly the same both Public and Private scores, again it is possible to experience a large shakeup based on the most minuscule of differences: if one is just below the isoscore string value on the Public LB, but just above it on the Private LB, then one will leapfrog all of the people in the block at once. </p>\n<p>And finally, suffice to say, it is also very hard not to fall into the trap of fitting to the Public leaderboard; after all, who does not like to see themselves climb positions with each submission!</p>\n<p>So, all in all, in this competition I somewhat suspect, despite best efforts not to overfit,  we shall still be in for a quite substantial \"positional\" shakeup; many of the the medals (or no medals) will be awarded on the fourth significant figure, which is basically noise. In view of this, I think anyone whose model has the same Public and Private LB scores will have every right to be proud of their work, and the medals are what they are; just a part of kaggle.</p>\n<p>All the best,<br>\ncarl</p>\n<p>PS: By the way, if one has an idle five minutes whilst running their model that will definitely have a LB score of 0.800 <em>this</em> time, here are a couple of thought provoking papers perhaps worth perusing: <a href=\"https://proceedings.neurips.cc/paper/2019/file/ee39e503b6bedf0c98c388b7e8589aca-Paper.pdf\" target=\"_blank\">\"A Meta-Analysis of Overfitting in Machine Learning\"</a> where the authors analyze overfitting by looking at 120 kaggle competition leaderboards, and also the paper <a href=\"https://arxiv.org/pdf/2109.08203.pdf\" target=\"_blank\">\"<code>torch.manual seed(3407)</code> is all you need: On the influence of random seeds in deep learning architectures for computer vision\"</a> which, despite the title, is also applicable to tabular predictions.</p>",
      "rawMarkdown": "Even though this competition has only just started, I am already going to be so bold as to suggest that there will be a big \"positional\" shakeup in this competition.\n\nThere are two types of shakeup, a ***score*** shakeup, and a ***positional*** or leaderboard shakeup. Naturally they are interlinked, however, one is manageable and the other, not so much. \n\nA score shakeup, a big difference between Public and Private Leaderboard scores, is either due to overfitting, or the distribution of the Private Leaderboard data being very different from either the training data and/or the Public Leaderboard data. There are countless topics on overfitting, and avoiding it is one of the hallmarks of the great kaggle competitors.\n\nHowever, a *positional* shakeup is a different creature; all data is noisy, and has an *irreducible error* component, and there is (almost by definition) nothing that can be done about it. In a positional shakeup one can have very comparable Public and Private scores, to with a reasonable and totally acceptable statistical difference. However, on kaggle this small difference could actually result in a huge jump in LB positions. It was pointed out by AmbrosM that the [metric is rather noisy](https://www.kaggle.com/competitions/amex-default-prediction/discussion/329088#1811782), and we are all subject to its vagaries. \n\nA look at the leaderboard after only two weeks shows the LB is already very compact;  there are already 400 teams in the range [0.793, 0.797] (a difference of only 0.005).  Add to that situation the possibility of *forkmitting*, which is where numerous people may submit the very same `submission.csv`, creating an [isoscore string](https://www.kaggle.com/carlmcbrideellis/shakeup-scatterplots-boxes-strings-and-things) where a large block of people all have exactly the same both Public and Private scores, again it is possible to experience a large shakeup based on the most minuscule of differences: if one is just below the isoscore string value on the Public LB, but just above it on the Private LB, then one will leapfrog all of the people in the block at once. \n\nAnd finally, suffice to say, it is also very hard not to fall into the trap of fitting to the Public leaderboard; after all, who does not like to see themselves climb positions with each submission!\n\nSo, all in all, in this competition I somewhat suspect, despite best efforts not to overfit,  we shall still be in for a quite substantial \"positional\" shakeup; many of the the medals (or no medals) will be awarded on the fourth significant figure, which is basically noise. In view of this, I think anyone whose model has the same Public and Private LB scores will have every right to be proud of their work, and the medals are what they are; just a part of kaggle.\n\nAll the best,\ncarl\n\nPS: By the way, if one has an idle five minutes whilst running their model that will definitely have a LB score of 0.800 *this* time, here are a couple of thought provoking papers perhaps worth perusing: [\"A Meta-Analysis of Overfitting in Machine Learning\"](https://proceedings.neurips.cc/paper/2019/file/ee39e503b6bedf0c98c388b7e8589aca-Paper.pdf) where the authors analyze overfitting by looking at 120 kaggle competition leaderboards, and also the paper [\"`torch.manual seed(3407)` is all you need: On the influence of random seeds in deep learning architectures for computer vision\"](https://arxiv.org/pdf/2109.08203.pdf) which, despite the title, is also applicable to tabular predictions.",
      "votes": null
    },
    {
      "id": "1814934",
      "postDate": "06/08/2022 13:45:30",
      "content": "<p>Having a huge fall on the private leaderboard from the public leaderboard is almost a rite of passage on Kaggle. In the Porto Seguro competition, I fell 870 places. A large (but not the only) part of that was being ahead of a popular public kernel on public but behind it on private as you described. The key thing is to learn from it if it happens. I agree that there will be more instability than if we were using some other metrics (such as AUC only).</p>",
      "rawMarkdown": "Having a huge fall on the private leaderboard from the public leaderboard is almost a rite of passage on Kaggle. In the Porto Seguro competition, I fell 870 places. A large (but not the only) part of that was being ahead of a popular public kernel on public but behind it on private as you described. The key thing is to learn from it if it happens. I agree that there will be more instability than if we were using some other metrics (such as AUC only).",
      "votes": null
    },
    {
      "id": "1814952",
      "postDate": "06/08/2022 13:59:49",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/datahobbit\" target=\"_blank\">@datahobbit</a> </p>\n<blockquote>\n  <p>\"<em>The key thing is to learn from it if it happens.</em>\"</p>\n</blockquote>\n<p>When it comes to positional shakeups I think there is little to really learn; outside of a kaggle competition context I think such solutions would be more than acceptable. In this case I would rather look to Rudyard Kipling:</p>\n<p><a href=\"https://www.poetryfoundation.org/poems/46473/if---\" target=\"_blank\"><em>If you can meet with Triumph and Disaster <br>&nbsp; &nbsp; And treat those two impostors just the same;</em></a></p>\n<p>As for the metric, I get the feeling we are being pushed into a (4%) \"numerical corner\" by the <em>D</em> component, with diminishing returns, just as in the <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\">excellent image generated again by AmbrosM</a>.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @datahobbit \n\n> \"*The key thing is to learn from it if it happens.*\"\n\nWhen it comes to positional shakeups I think there is little to really learn; outside of a kaggle competition context I think such solutions would be more than acceptable. In this case I would rather look to Rudyard Kipling:\n\n\n[*If you can meet with Triumph and Disaster <br>&nbsp; &nbsp; And treat those two impostors just the same;*](https://www.poetryfoundation.org/poems/46473/if---)\n\nAs for the metric, I get the feeling we are being pushed into a (4%) \"numerical corner\" by the *D* component, with diminishing returns, just as in the [excellent image generated again by AmbrosM](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464).\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1815044",
      "postDate": "06/08/2022 15:51:56",
      "content": "<p>Hi Carl, I agree that some shakeups are just a matter of statistical chance but others are caused by chasing a high public score at the expense of a solid cv approach. From my personal experience, the 1 place shakeup fall (from gold to silver) in the Google Quest Q&amp;A competition could have gone either way and was partly bad luck. In that case, we did have a pretty solid cv approach but the metric was unstable. On the other hand, the Porto Seguro huge fall was because I was focusing too much on the public score and not enough on cv. There will probably be people in this competition who are unlucky and others who suffer from chasing the public score. There may be some who have used a solid cv and shake up slightly perhaps across a medal threshold. I personally think that anyone who uses robust cv is more likely to shake the right way above high-scoring public kernels than vice versa. </p>\n<p>Yes, I agree that there could be diminishing returns on the competition metric. Will be interesting to see.</p>",
      "rawMarkdown": "Hi Carl, I agree that some shakeups are just a matter of statistical chance but others are caused by chasing a high public score at the expense of a solid cv approach. From my personal experience, the 1 place shakeup fall (from gold to silver) in the Google Quest Q&A competition could have gone either way and was partly bad luck. In that case, we did have a pretty solid cv approach but the metric was unstable. On the other hand, the Porto Seguro huge fall was because I was focusing too much on the public score and not enough on cv. There will probably be people in this competition who are unlucky and others who suffer from chasing the public score. There may be some who have used a solid cv and shake up slightly perhaps across a medal threshold. I personally think that anyone who uses robust cv is more likely to shake the right way above high-scoring public kernels than vice versa. \n\nYes, I agree that there could be diminishing returns on the competition metric. Will be interesting to see.",
      "votes": null
    },
    {
      "id": "1817547",
      "postDate": "06/11/2022 14:28:34",
      "content": "<p>That's exactly what I was thinking.</p>",
      "rawMarkdown": "That's exactly what I was thinking.",
      "votes": null
    },
    {
      "id": "1818160",
      "postDate": "06/12/2022 10:52:00",
      "content": "<p>Do you think the custom evaluation metric will reduce the shakeup?</p>",
      "rawMarkdown": "Do you think the custom evaluation metric will reduce the shakeup?",
      "votes": null
    },
    {
      "id": "1818437",
      "postDate": "06/12/2022 17:31:50",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/heyspaceturtle\" target=\"_blank\">@heyspaceturtle</a> </p>\n<p>Having the custom evaluation metric should certainly help one to reduce overfitting. However, in this case the shakeup will occur even if one does not overfit at all, simply due to a combination of the density of the leaderboard in conjunction with random noise that starts appearing in the (hidden) fourth significant figure of the leaderboard scores. To better visualize the situation may I heartily recommend taking a look at the plots made by the user <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>  both <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329088#1811782\" target=\"_blank\">here</a> and in the topic <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329787\" target=\"_blank\">\"Can you find the best seed?\"</a>. </p>\n<p>Un saludo muy cordial,<br>\ncarl</p>",
      "rawMarkdown": "Dear @heyspaceturtle \n\nHaving the custom evaluation metric should certainly help one to reduce overfitting. However, in this case the shakeup will occur even if one does not overfit at all, simply due to a combination of the density of the leaderboard in conjunction with random noise that starts appearing in the (hidden) fourth significant figure of the leaderboard scores. To better visualize the situation may I heartily recommend taking a look at the plots made by the user @ambrosm  both [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/329088#1811782) and in the topic [\"Can you find the best seed?\"](https://www.kaggle.com/competitions/amex-default-prediction/discussion/329787). \n\nUn saludo muy cordial,\ncarl",
      "votes": null
    },
    {
      "id": "1818455",
      "postDate": "06/12/2022 17:59:03",
      "content": "<p>Wow, muchas gracias por la respuesta. </p>\n<p>This is my very first Kaggle competition and its exciting to see how all the leaderboard positions will change. I'm learning a lot on the way and discovering great data-driven stuff on the forums. </p>\n<p>What do you think is the best way to build a stable model? Even if in the public leaderboard is not on the top. Will a normal cross-validation do? </p>",
      "rawMarkdown": "Wow, muchas gracias por la respuesta. \n\nThis is my very first Kaggle competition and its exciting to see how all the leaderboard positions will change. I'm learning a lot on the way and discovering great data-driven stuff on the forums. \n\nWhat do you think is the best way to build a stable model? Even if in the public leaderboard is not on the top. Will a normal cross-validation do?",
      "votes": null
    },
    {
      "id": "1818468",
      "postDate": "06/12/2022 18:12:12",
      "content": "<blockquote>\n  <p>\"<em>What do you think is the best way to build a stable model?</em>\"</p>\n</blockquote>\n<p>There have been rivers of ink and a good percentage of kaggle topics dedicated to your question; in general terms simple precautions such as in addition to cross validation (or perhaps even <a href=\"https://www.kaggle.com/discussions/general/255783\" target=\"_blank\">nested cross-validation</a>) OOF scores, one can create an additional hold-out  dataset to test ones model on (imagine it to be ones own personal off-line leaderboard  score!). Also to reduce the possibility of overfitting, one can create a final ensemble of your models. For more on this see the fantastic <a href=\"https://web.archive.org/web/20160304031055/http://mlwave.com/kaggle-ensembling-guide/\" target=\"_blank\">\"Kaggle Ensembling Guide\"</a>.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "> \"*What do you think is the best way to build a stable model?*\"\n\nThere have been rivers of ink and a good percentage of kaggle topics dedicated to your question; in general terms simple precautions such as in addition to cross validation (or perhaps even [nested cross-validation](https://www.kaggle.com/discussions/general/255783)) OOF scores, one can create an additional hold-out  dataset to test ones model on (imagine it to be ones own personal off-line leaderboard  score!). Also to reduce the possibility of overfitting, one can create a final ensemble of your models. For more on this see the fantastic [\"Kaggle Ensembling Guide\"](https://web.archive.org/web/20160304031055/http://mlwave.com/kaggle-ensembling-guide/).\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1818569",
      "postDate": "06/12/2022 22:12:31",
      "content": "<p>This is just what I needed! Thank you so much🐢</p>",
      "rawMarkdown": "This is just what I needed! Thank you so much🐢",
      "votes": null
    },
    {
      "id": "1821944",
      "postDate": "06/16/2022 00:22:12",
      "content": "<p>It doesn't seem easy .. but reading constructive criticism is a real eye-opener. Thank you for sharing <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> </p>",
      "rawMarkdown": "It doesn't seem easy .. but reading constructive criticism is a real eye-opener. Thank you for sharing @carlmcbrideellis",
      "votes": null
    },
    {
      "id": "1872467",
      "postDate": "07/27/2022 02:09:40",
      "content": "<p>A huge positional shakeup is quite possible since the scores are so close on public LB. Also the variance in local CV (and presumably public and private LB) is std 0.0012 which is hundreds of LB ranks.</p>",
      "rawMarkdown": "A huge positional shakeup is quite possible since the scores are so close on public LB. Also the variance in local CV (and presumably public and private LB) is std 0.0012 which is hundreds of LB ranks.",
      "votes": null
    },
    {
      "id": "1872572",
      "postDate": "07/27/2022 05:18:20",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<p>Indeed; as of today, with a month still to go, the 0.799 score occupies the 54th position all the way down to the 880th position on the Public LB. I cannot help but think how much more well behaved the LB would be if the competition metric were the <em>G</em> component (basically the AUC) alone, without the <em>D</em> threshold (which is the source of most of the variance).<br>\nI suspect the 25th of August will be a day of omphaloskepsis for many.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @cdeotte \n\nIndeed; as of today, with a month still to go, the 0.799 score occupies the 54th position all the way down to the 880th position on the Public LB. I cannot help but think how much more well behaved the LB would be if the competition metric were the *G* component (basically the AUC) alone, without the *D* threshold (which is the source of most of the variance).\nI suspect the 25th of August will be a day of omphaloskepsis for many.\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1873273",
      "postDate": "07/27/2022 14:09:32",
      "content": "<p>( Apropos of nothing, I like your icon. )</p>",
      "rawMarkdown": "( Apropos of nothing, I like your icon. )",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1814934,
      "author_name": "datahobbit",
      "author_url": "",
      "post_date": "06/08/2022 13:45:30",
      "content": "<p>Having a huge fall on the private leaderboard from the public leaderboard is almost a rite of passage on Kaggle. In the Porto Seguro competition, I fell 870 places. A large (but not the only) part of that was being ahead of a popular public kernel on public but behind it on private as you described. The key thing is to learn from it if it happens. I agree that there will be more instability than if we were using some other metrics (such as AUC only).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1814952,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "06/08/2022 13:59:49",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/datahobbit\" target=\"_blank\">@datahobbit</a> </p>\n<blockquote>\n  <p>\"<em>The key thing is to learn from it if it happens.</em>\"</p>\n</blockquote>\n<p>When it comes to positional shakeups I think there is little to really learn; outside of a kaggle competition context I think such solutions would be more than acceptable. In this case I would rather look to Rudyard Kipling:</p>\n<p><a href=\"https://www.poetryfoundation.org/poems/46473/if---\" target=\"_blank\"><em>If you can meet with Triumph and Disaster <br>&nbsp; &nbsp; And treat those two impostors just the same;</em></a></p>\n<p>As for the metric, I get the feeling we are being pushed into a (4%) \"numerical corner\" by the <em>D</em> component, with diminishing returns, just as in the <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\">excellent image generated again by AmbrosM</a>.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1815044,
          "author_name": "datahobbit",
          "author_url": "",
          "post_date": "06/08/2022 15:51:56",
          "content": "<p>Hi Carl, I agree that some shakeups are just a matter of statistical chance but others are caused by chasing a high public score at the expense of a solid cv approach. From my personal experience, the 1 place shakeup fall (from gold to silver) in the Google Quest Q&amp;A competition could have gone either way and was partly bad luck. In that case, we did have a pretty solid cv approach but the metric was unstable. On the other hand, the Porto Seguro huge fall was because I was focusing too much on the public score and not enough on cv. There will probably be people in this competition who are unlucky and others who suffer from chasing the public score. There may be some who have used a solid cv and shake up slightly perhaps across a medal threshold. I personally think that anyone who uses robust cv is more likely to shake the right way above high-scoring public kernels than vice versa. </p>\n<p>Yes, I agree that there could be diminishing returns on the competition metric. Will be interesting to see.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1817547,
      "author_name": "wuuthraad",
      "author_url": "",
      "post_date": "06/11/2022 14:28:34",
      "content": "<p>That's exactly what I was thinking.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1818160,
      "author_name": "heyspaceturtle",
      "author_url": "",
      "post_date": "06/12/2022 10:52:00",
      "content": "<p>Do you think the custom evaluation metric will reduce the shakeup?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1818437,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "06/12/2022 17:31:50",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/heyspaceturtle\" target=\"_blank\">@heyspaceturtle</a> </p>\n<p>Having the custom evaluation metric should certainly help one to reduce overfitting. However, in this case the shakeup will occur even if one does not overfit at all, simply due to a combination of the density of the leaderboard in conjunction with random noise that starts appearing in the (hidden) fourth significant figure of the leaderboard scores. To better visualize the situation may I heartily recommend taking a look at the plots made by the user <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>  both <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329088#1811782\" target=\"_blank\">here</a> and in the topic <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329787\" target=\"_blank\">\"Can you find the best seed?\"</a>. </p>\n<p>Un saludo muy cordial,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1818455,
          "author_name": "heyspaceturtle",
          "author_url": "",
          "post_date": "06/12/2022 17:59:03",
          "content": "<p>Wow, muchas gracias por la respuesta. </p>\n<p>This is my very first Kaggle competition and its exciting to see how all the leaderboard positions will change. I'm learning a lot on the way and discovering great data-driven stuff on the forums. </p>\n<p>What do you think is the best way to build a stable model? Even if in the public leaderboard is not on the top. Will a normal cross-validation do? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1818468,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "06/12/2022 18:12:12",
          "content": "<blockquote>\n  <p>\"<em>What do you think is the best way to build a stable model?</em>\"</p>\n</blockquote>\n<p>There have been rivers of ink and a good percentage of kaggle topics dedicated to your question; in general terms simple precautions such as in addition to cross validation (or perhaps even <a href=\"https://www.kaggle.com/discussions/general/255783\" target=\"_blank\">nested cross-validation</a>) OOF scores, one can create an additional hold-out  dataset to test ones model on (imagine it to be ones own personal off-line leaderboard  score!). Also to reduce the possibility of overfitting, one can create a final ensemble of your models. For more on this see the fantastic <a href=\"https://web.archive.org/web/20160304031055/http://mlwave.com/kaggle-ensembling-guide/\" target=\"_blank\">\"Kaggle Ensembling Guide\"</a>.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1818569,
          "author_name": "heyspaceturtle",
          "author_url": "",
          "post_date": "06/12/2022 22:12:31",
          "content": "<p>This is just what I needed! Thank you so much🐢</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1821944,
      "author_name": "arti1117",
      "author_url": "",
      "post_date": "06/16/2022 00:22:12",
      "content": "<p>It doesn't seem easy .. but reading constructive criticism is a real eye-opener. Thank you for sharing <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1872467,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/27/2022 02:09:40",
      "content": "<p>A huge positional shakeup is quite possible since the scores are so close on public LB. Also the variance in local CV (and presumably public and private LB) is std 0.0012 which is hundreds of LB ranks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1872572,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "07/27/2022 05:18:20",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<p>Indeed; as of today, with a month still to go, the 0.799 score occupies the 54th position all the way down to the 880th position on the Public LB. I cannot help but think how much more well behaved the LB would be if the competition metric were the <em>G</em> component (basically the AUC) alone, without the <em>D</em> threshold (which is the source of most of the variance).<br>\nI suspect the 25th of August will be a day of omphaloskepsis for many.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1873273,
      "author_name": "bellapisani",
      "author_url": "",
      "post_date": "07/27/2022 14:09:32",
      "content": "<p>( Apropos of nothing, I like your icon. )</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1814860": "Even though this competition has only just started, I am already going to be so bold as to suggest that there will be a big \"positional\" shakeup in this competition.\n\nThere are two types of shakeup, a ***score*** shakeup, and a ***positional*** or leaderboard shakeup. Naturally they are interlinked, however, one is manageable and the other, not so much. \n\nA score shakeup, a big difference between Public and Private Leaderboard scores, is either due to overfitting, or the distribution of the Private Leaderboard data being very different from either the training data and/or the Public Leaderboard data. There are countless topics on overfitting, and avoiding it is one of the hallmarks of the great kaggle competitors.\n\nHowever, a *positional* shakeup is a different creature; all data is noisy, and has an *irreducible error* component, and there is (almost by definition) nothing that can be done about it. In a positional shakeup one can have very comparable Public and Private scores, to with a reasonable and totally acceptable statistical difference. However, on kaggle this small difference could actually result in a huge jump in LB positions. It was pointed out by AmbrosM that the [metric is rather noisy](https://www.kaggle.com/competitions/amex-default-prediction/discussion/329088#1811782), and we are all subject to its vagaries. \n\nA look at the leaderboard after only two weeks shows the LB is already very compact;  there are already 400 teams in the range [0.793, 0.797] (a difference of only 0.005).  Add to that situation the possibility of *forkmitting*, which is where numerous people may submit the very same `submission.csv`, creating an [isoscore string](https://www.kaggle.com/carlmcbrideellis/shakeup-scatterplots-boxes-strings-and-things) where a large block of people all have exactly the same both Public and Private scores, again it is possible to experience a large shakeup based on the most minuscule of differences: if one is just below the isoscore string value on the Public LB, but just above it on the Private LB, then one will leapfrog all of the people in the block at once. \n\nAnd finally, suffice to say, it is also very hard not to fall into the trap of fitting to the Public leaderboard; after all, who does not like to see themselves climb positions with each submission!\n\nSo, all in all, in this competition I somewhat suspect, despite best efforts not to overfit,  we shall still be in for a quite substantial \"positional\" shakeup; many of the the medals (or no medals) will be awarded on the fourth significant figure, which is basically noise. In view of this, I think anyone whose model has the same Public and Private LB scores will have every right to be proud of their work, and the medals are what they are; just a part of kaggle.\n\nAll the best,\ncarl\n\nPS: By the way, if one has an idle five minutes whilst running their model that will definitely have a LB score of 0.800 *this* time, here are a couple of thought provoking papers perhaps worth perusing: [\"A Meta-Analysis of Overfitting in Machine Learning\"](https://proceedings.neurips.cc/paper/2019/file/ee39e503b6bedf0c98c388b7e8589aca-Paper.pdf) where the authors analyze overfitting by looking at 120 kaggle competition leaderboards, and also the paper [\"`torch.manual seed(3407)` is all you need: On the influence of random seeds in deep learning architectures for computer vision\"](https://arxiv.org/pdf/2109.08203.pdf) which, despite the title, is also applicable to tabular predictions.",
    "1814934": "Having a huge fall on the private leaderboard from the public leaderboard is almost a rite of passage on Kaggle. In the Porto Seguro competition, I fell 870 places. A large (but not the only) part of that was being ahead of a popular public kernel on public but behind it on private as you described. The key thing is to learn from it if it happens. I agree that there will be more instability than if we were using some other metrics (such as AUC only).",
    "1814952": "Dear @datahobbit \n\n> \"*The key thing is to learn from it if it happens.*\"\n\nWhen it comes to positional shakeups I think there is little to really learn; outside of a kaggle competition context I think such solutions would be more than acceptable. In this case I would rather look to Rudyard Kipling:\n\n\n[*If you can meet with Triumph and Disaster <br>&nbsp; &nbsp; And treat those two impostors just the same;*](https://www.poetryfoundation.org/poems/46473/if---)\n\nAs for the metric, I get the feeling we are being pushed into a (4%) \"numerical corner\" by the *D* component, with diminishing returns, just as in the [excellent image generated again by AmbrosM](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464).\n\nAll the best,\ncarl",
    "1815044": "Hi Carl, I agree that some shakeups are just a matter of statistical chance but others are caused by chasing a high public score at the expense of a solid cv approach. From my personal experience, the 1 place shakeup fall (from gold to silver) in the Google Quest Q&A competition could have gone either way and was partly bad luck. In that case, we did have a pretty solid cv approach but the metric was unstable. On the other hand, the Porto Seguro huge fall was because I was focusing too much on the public score and not enough on cv. There will probably be people in this competition who are unlucky and others who suffer from chasing the public score. There may be some who have used a solid cv and shake up slightly perhaps across a medal threshold. I personally think that anyone who uses robust cv is more likely to shake the right way above high-scoring public kernels than vice versa. \n\nYes, I agree that there could be diminishing returns on the competition metric. Will be interesting to see.",
    "1817547": "That's exactly what I was thinking.",
    "1818160": "Do you think the custom evaluation metric will reduce the shakeup?",
    "1818437": "Dear @heyspaceturtle \n\nHaving the custom evaluation metric should certainly help one to reduce overfitting. However, in this case the shakeup will occur even if one does not overfit at all, simply due to a combination of the density of the leaderboard in conjunction with random noise that starts appearing in the (hidden) fourth significant figure of the leaderboard scores. To better visualize the situation may I heartily recommend taking a look at the plots made by the user @ambrosm  both [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/329088#1811782) and in the topic [\"Can you find the best seed?\"](https://www.kaggle.com/competitions/amex-default-prediction/discussion/329787). \n\nUn saludo muy cordial,\ncarl",
    "1818455": "Wow, muchas gracias por la respuesta. \n\nThis is my very first Kaggle competition and its exciting to see how all the leaderboard positions will change. I'm learning a lot on the way and discovering great data-driven stuff on the forums. \n\nWhat do you think is the best way to build a stable model? Even if in the public leaderboard is not on the top. Will a normal cross-validation do?",
    "1818468": "> \"*What do you think is the best way to build a stable model?*\"\n\nThere have been rivers of ink and a good percentage of kaggle topics dedicated to your question; in general terms simple precautions such as in addition to cross validation (or perhaps even [nested cross-validation](https://www.kaggle.com/discussions/general/255783)) OOF scores, one can create an additional hold-out  dataset to test ones model on (imagine it to be ones own personal off-line leaderboard  score!). Also to reduce the possibility of overfitting, one can create a final ensemble of your models. For more on this see the fantastic [\"Kaggle Ensembling Guide\"](https://web.archive.org/web/20160304031055/http://mlwave.com/kaggle-ensembling-guide/).\n\nAll the best,\ncarl",
    "1818569": "This is just what I needed! Thank you so much🐢",
    "1821944": "It doesn't seem easy .. but reading constructive criticism is a real eye-opener. Thank you for sharing @carlmcbrideellis",
    "1872467": "A huge positional shakeup is quite possible since the scores are so close on public LB. Also the variance in local CV (and presumably public and private LB) is std 0.0012 which is hundreds of LB ranks.",
    "1872572": "Dear @cdeotte \n\nIndeed; as of today, with a month still to go, the 0.799 score occupies the 54th position all the way down to the 880th position on the Public LB. I cannot help but think how much more well behaved the LB would be if the competition metric were the *G* component (basically the AUC) alone, without the *D* threshold (which is the source of most of the variance).\nI suspect the 25th of August will be a day of omphaloskepsis for many.\n\nAll the best,\ncarl",
    "1873273": "( Apropos of nothing, I like your icon. )"
  },
  "source": "meta"
}