{
  "id": 551758,
  "title": "Evidence of overfitting, or how to recognize bad public notebooks",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/551758",
  "author_name": "AmbrosM",
  "post_date": "2024-12-15T11:13:49.779000",
  "votes": 51,
  "comment_count": 26,
  "views": 0,
  "content": "<p>If we compare today's two highest-scoring public notebooks, we see that they differ in exactly one number:</p>\n<p><a href=\"https://www.kaggle.com/code/batprem/cmi-tuning-ensemble-of-solutions\" target=\"_blank\">CMI| Tuning | Ensemble of solutions</a> V111, lb score 0.497, 2 days ago:</p>\n<pre><code>    KappaOPtimizer = minimize(evaluate_predictions,\n                              x0=[, , ], args=(y, oof_non_rounded), \n                              ='-')\n</code></pre>\n<p><a href=\"https://www.kaggle.com/code/vitalykudelya/cmi-issues-with-the-metric-and-baseline\" target=\"_blank\">CMI: Issues with the Metric and Baseline</a> V3, lb score 0.495, 3 days ago:</p>\n<pre><code>    KappaOPtimizer = minimize(evaluate_predictions,\n                              x0=[, , ], args=(y, oof_non_rounded), \n                              ='-')\n</code></pre>\n<p>The output of the corresponding cells differs slightly:</p>\n<p>CMI| Tuning | Ensemble of solutions:</p>\n<pre><code>\n</code></pre>\n<p>CMI: Issues with the Metric and Baseline</p>\n<pre><code>\n</code></pre>\n<p>What does this mean? The newer notebook changes the initial guess of the <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.optimize.minimize.html\" target=\"_blank\">Nelder–Mead optimization</a> (not even a model hyperparameter), the Nelder–Mead optimization falls into another local minimum, the cv score goes down and the leaderboard score goes up. </p>\n<p>\"cv down and leaderboard up\" should ring a bell: This is called overfitting to the leaderboard. Don't expect that the newer model will keep its advantage on the private leaderboard!</p>\n<p>If you belong to the hundreds of people who have copied and submitted this notebook, try changing the seed and resubmit! If the notebook is good, it will keep its high score with other seeds. If the public leaderboard score drops when changing the seed, don't expect anything good from the private leaderboard. The transition from public to private is equivalent to changing the seed.</p>\n<p>P.S. In the meantime, another public notebook has appeared, <a href=\"https://www.kaggle.com/code/shodaifuruya/lb-497-multi-model-feature-importance-analysis\" target=\"_blank\">LB.497|Multi-Model Feature Importance Analysis</a> V15. It is an exact copy of the overfitting model with the addition of some feature importance bar charts.</p>",
  "messages": [
    {
      "id": 3072530,
      "postDate": "2024-12-15T11:13:49.780Z",
      "content": "<p>If we compare today's two highest-scoring public notebooks, we see that they differ in exactly one number:</p>\n<p><a href=\"https://www.kaggle.com/code/batprem/cmi-tuning-ensemble-of-solutions\" target=\"_blank\">CMI| Tuning | Ensemble of solutions</a> V111, lb score 0.497, 2 days ago:</p>\n<pre><code>    KappaOPtimizer = minimize(evaluate_predictions,\n                              x0=[, , ], args=(y, oof_non_rounded), \n                              ='-')\n</code></pre>\n<p><a href=\"https://www.kaggle.com/code/vitalykudelya/cmi-issues-with-the-metric-and-baseline\" target=\"_blank\">CMI: Issues with the Metric and Baseline</a> V3, lb score 0.495, 3 days ago:</p>\n<pre><code>    KappaOPtimizer = minimize(evaluate_predictions,\n                              x0=[, , ], args=(y, oof_non_rounded), \n                              ='-')\n</code></pre>\n<p>The output of the corresponding cells differs slightly:</p>\n<p>CMI| Tuning | Ensemble of solutions:</p>\n<pre><code>\n</code></pre>\n<p>CMI: Issues with the Metric and Baseline</p>\n<pre><code>\n</code></pre>\n<p>What does this mean? The newer notebook changes the initial guess of the <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.optimize.minimize.html\" target=\"_blank\">Nelder–Mead optimization</a> (not even a model hyperparameter), the Nelder–Mead optimization falls into another local minimum, the cv score goes down and the leaderboard score goes up. </p>\n<p>\"cv down and leaderboard up\" should ring a bell: This is called overfitting to the leaderboard. Don't expect that the newer model will keep its advantage on the private leaderboard!</p>\n<p>If you belong to the hundreds of people who have copied and submitted this notebook, try changing the seed and resubmit! If the notebook is good, it will keep its high score with other seeds. If the public leaderboard score drops when changing the seed, don't expect anything good from the private leaderboard. The transition from public to private is equivalent to changing the seed.</p>\n<p>P.S. In the meantime, another public notebook has appeared, <a href=\"https://www.kaggle.com/code/shodaifuruya/lb-497-multi-model-feature-importance-analysis\" target=\"_blank\">LB.497|Multi-Model Feature Importance Analysis</a> V15. It is an exact copy of the overfitting model with the addition of some feature importance bar charts.</p>",
      "rawMarkdown": "If we compare today's two highest-scoring public notebooks, we see that they differ in exactly one number:\n\n[CMI| Tuning | Ensemble of solutions](https://www.kaggle.com/code/batprem/cmi-tuning-ensemble-of-solutions) V111, lb score 0.497, 2 days ago:\n```\n    KappaOPtimizer = minimize(evaluate_predictions,\n                              x0=[0.5, 1.49, 2.5], args=(y, oof_non_rounded), \n                              method='Nelder-Mead')\n```\n\n[CMI: Issues with the Metric and Baseline](https://www.kaggle.com/code/vitalykudelya/cmi-issues-with-the-metric-and-baseline) V3, lb score 0.495, 3 days ago:\n```\n    KappaOPtimizer = minimize(evaluate_predictions,\n                              x0=[0.5, 1.5, 2.5], args=(y, oof_non_rounded), \n                              method='Nelder-Mead')\n```\n\nThe output of the corresponding cells differs slightly:\n\nCMI| Tuning | Ensemble of solutions:\n```\n----> || Optimized QWK SCORE ::  0.456\n```\n\nCMI: Issues with the Metric and Baseline\n```\n----> || Optimized QWK SCORE ::  0.457\n```\n\nWhat does this mean? The newer notebook changes the initial guess of the [Nelder–Mead optimization](https://docs.scipy.org/doc/scipy/reference/generated/scipy.optimize.minimize.html) (not even a model hyperparameter), the Nelder–Mead optimization falls into another local minimum, the cv score goes down and the leaderboard score goes up. \n\n\"cv down and leaderboard up\" should ring a bell: This is called overfitting to the leaderboard. Don't expect that the newer model will keep its advantage on the private leaderboard!\n\nIf you belong to the hundreds of people who have copied and submitted this notebook, try changing the seed and resubmit! If the notebook is good, it will keep its high score with other seeds. If the public leaderboard score drops when changing the seed, don't expect anything good from the private leaderboard. The transition from public to private is equivalent to changing the seed.\n\nP.S. In the meantime, another public notebook has appeared, [LB.497|Multi-Model Feature Importance Analysis](https://www.kaggle.com/code/shodaifuruya/lb-497-multi-model-feature-importance-analysis) V15. It is an exact copy of the overfitting model with the addition of some feature importance bar charts.",
      "votes": 51
    },
    {
      "id": 3072570,
      "postDate": "2024-12-15T11:57:32.287Z",
      "content": "<p>I checked some competitions with qwk metric and most of them had some kind of a shakeup. Since the penalty for farther off predictions (e.g., predicting 0 for a 3) is quadratic, a small change in the model can result in large swings. The dataset is also small here, so those swings are becoming larger. If I'm not mistaken, there aren't any class 3 in public test set, so it might boil down to how well you are predicting them in private test set.</p>",
      "rawMarkdown": "I checked some competitions with qwk metric and most of them had some kind of a shakeup. Since the penalty for farther off predictions (e.g., predicting 0 for a 3) is quadratic, a small change in the model can result in large swings. The dataset is also small here, so those swings are becoming larger. If I'm not mistaken, there aren't any class 3 in public test set, so it might boil down to how well you are predicting them in private test set.",
      "votes": 14,
      "replies": [
        {
          "id": 3073758,
          "postDate": "2024-12-16T21:29:03.300Z",
          "content": "<p><code>If I'm not mistaken, there aren't any class 3 in public test set</code></p>\n<p>This is REALLY interesting and, frankly, sounds like an oversight on the part of whoever's curating the data. How'd you reach this conclusion?</p>",
          "rawMarkdown": "`If I'm not mistaken, there aren't any class 3 in public test set`\n\nThis is REALLY interesting and, frankly, sounds like an oversight on the part of whoever's curating the data. How'd you reach this conclusion?",
          "votes": 1,
          "replies": [
            {
              "id": 3073833,
              "postDate": "2024-12-17T01:30:13.043Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 3074306,
          "postDate": "2024-12-17T14:47:14.833Z",
          "content": "<p>Hmm I doubt this is true. I tried changing all my sii=3 predictions to 2, and the LB decreased slightly </p>",
          "rawMarkdown": "Hmm I doubt this is true. I tried changing all my sii=3 predictions to 2, and the LB decreased slightly ",
          "votes": 1,
          "replies": [
            {
              "id": 3074316,
              "postDate": "2024-12-17T15:02:35.553Z",
              "content": "<p>It's my bad then. My last threshold was 1.5 or something like that and I tried changing it to a large number in order to change all 3 predictions to 2, but my score didn't change at all. I guess my model wasn't predicting anything beyond that threshold.</p>",
              "rawMarkdown": "It's my bad then. My last threshold was 1.5 or something like that and I tried changing it to a large number in order to change all 3 predictions to 2, but my score didn't change at all. I guess my model wasn't predicting anything beyond that threshold."
            }
          ]
        }
      ]
    },
    {
      "id": 3072596,
      "postDate": "2024-12-15T12:34:04.493Z",
      "content": "<p>The easiest way to get notebook medal may be to publish high score overfitting ensemble notebook near end.<br>\nUnfortunately, they may bring us almost nothing but copy-and-paste content.</p>",
      "rawMarkdown": "The easiest way to get notebook medal may be to publish high score overfitting ensemble notebook near end.\nUnfortunately, they may bring us almost nothing but copy-and-paste content.\n",
      "votes": 9,
      "replies": [
        {
          "id": 3073037,
          "postDate": "2024-12-16T00:40:44.900Z",
          "content": "<blockquote>\n  <p>Unfortunately, they may bring us almost nothing but copy-and-paste content.  </p>\n</blockquote>\n<p>There is actually a great value in knowing that the upper score limit can be achieved by random overfitting.  </p>",
          "rawMarkdown": ">Unfortunately, they may bring us almost nothing but copy-and-paste content.  \n\nThere is actually a great value in knowing that the upper score limit can be achieved by random overfitting.  ",
          "votes": 1,
          "replies": [
            {
              "id": 3073041,
              "postDate": "2024-12-16T01:01:51.247Z",
              "content": "<p>A kaggle friends say that when they have free time, they repeatedly change the seed of best public notebooks and submit them to observe the lb score variance. This may be good analysis for winning comp.</p>",
              "rawMarkdown": "A kaggle friends say that when they have free time, they repeatedly change the seed of best public notebooks and submit them to observe the lb score variance. This may be good analysis for winning comp.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3072582,
      "postDate": "2024-12-15T12:12:27.837Z",
      "content": "<p>I noticed such strange paradoxes with the Prem Chotepanit - CMI| Tuning | Ensemble of solutions V111 notebook, lb score 0.497. </p>\n<p>If you change the accelerator from GPU T4 x2 to GPU P100, the score drops to 0.488. </p>\n<ul>\n<li>GPU T4 x2 - Mean Train QWK: 0.7240, Mean Validation QWK: 0.4613 Optimized QWK SCORE: 0.5185 [Model 1]</li>\n<li>GPU P100 - Mean Train QWK: 0.7278, Mean Validation QWK: 0.4758 Optimized QWK SCORE: 0.5298 [Model 1]</li>\n</ul>\n<p>There is an explanation for this, and to test the idea, for example, I re-sorted the incoming sets, LB fell to 0.452. </p>\n<ul>\n<li>GPU T4 x2 - Mean Train QWK: 0.7293, Mean Validation QWK: 0.4752 Optimized QWK SCORE: 0.5324 [Model 1]</li>\n</ul>\n<p>imputer = KNNImputer(n_neighbors=4) LB=0.487<br>\nMean Train QWK: 0.7374, Mean Validation QWK: 0.4816 Optimized QWK SCORE: 0.5423 [Model 1]</p>\n<p>imputer = KNNImputer(n_neighbors=6) LB=0.489<br>\nMean Train QWK: 0.7358, Mean Validation QWK: 0.4888 Optimized QWK SCORE: 0.5555 [Model 1]</p>\n<p>imputer = KNNImputer(n_neighbors=52) LB=0.484<br>\nMean Train QWK: 0.7676, Mean Validation QWK: 0.5490 Optimized QWK SCORE: 0.6012 [Model 1]</p>\n<p>How is this possible?</p>\n<p>P.S.<br>\n— Is it possible to achieve a result without knowing the reasons for the phenomenon?<br>\n— I don’t know what to answer… The reason for the strength of damask steel was discovered only in the nineteenth century, but the best swords were forged in the eleventh.</p>",
      "rawMarkdown": "I noticed such strange paradoxes with the Prem Chotepanit - CMI| Tuning | Ensemble of solutions V111 notebook, lb score 0.497. \n\nIf you change the accelerator from GPU T4 x2 to GPU P100, the score drops to 0.488. \n- GPU T4 x2 - Mean Train QWK: 0.7240, Mean Validation QWK: 0.4613 Optimized QWK SCORE: 0.5185 [Model 1]\n- GPU P100 - Mean Train QWK: 0.7278, Mean Validation QWK: 0.4758 Optimized QWK SCORE: 0.5298 [Model 1]\n\nThere is an explanation for this, and to test the idea, for example, I re-sorted the incoming sets, LB fell to 0.452. \n- GPU T4 x2 - Mean Train QWK: 0.7293, Mean Validation QWK: 0.4752 Optimized QWK SCORE: 0.5324 [Model 1]\n\nimputer = KNNImputer(n_neighbors=4) LB=0.487\nMean Train QWK: 0.7374, Mean Validation QWK: 0.4816 Optimized QWK SCORE: 0.5423 [Model 1]\n\nimputer = KNNImputer(n_neighbors=6) LB=0.489\nMean Train QWK: 0.7358, Mean Validation QWK: 0.4888 Optimized QWK SCORE: 0.5555 [Model 1]\n\nimputer = KNNImputer(n_neighbors=52) LB=0.484\nMean Train QWK: 0.7676, Mean Validation QWK: 0.5490 Optimized QWK SCORE: 0.6012 [Model 1]\n\nHow is this possible?\n\nP.S.\n— Is it possible to achieve a result without knowing the reasons for the phenomenon?\n— I don’t know what to answer… The reason for the strength of damask steel was discovered only in the nineteenth century, but the best swords were forged in the eleventh.",
      "votes": 10
    },
    {
      "id": 3072768,
      "postDate": "2024-12-15T16:19:26.253Z",
      "content": "<p>Yep. Changed seed for this public NB &gt;&gt; LB score dropped by 23 points 🤣  <br>\nThis is very fun competition. Those are some great odds for winnings- much higher than regular lottery! May the odds will be with me 🤣  </p>",
      "rawMarkdown": "Yep. Changed seed for this public NB >> LB score dropped by 23 points 🤣  \nThis is very fun competition. Those are some great odds for winnings- much higher than regular lottery! May the odds will be with me 🤣  ",
      "votes": 8
    },
    {
      "id": 3072569,
      "postDate": "2024-12-15T11:57:08.907Z",
      "content": "<p><a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> yes you are right , for high scored public notebook , thee is a lot on flctatuations on the scores when we changes the SEED , THIS RED FLAG OF OVERFITTING </p>",
      "rawMarkdown": "@ambrosm yes you are right , for high scored public notebook , thee is a lot on flctatuations on the scores when we changes the SEED , THIS RED FLAG OF OVERFITTING ",
      "votes": 6,
      "replies": [
        {
          "id": 3072580,
          "postDate": "2024-12-15T12:08:40.580Z",
          "content": "<p>This is just 1 red flag - we have at least 5-6 more if you peruse them closely <a href=\"https://www.kaggle.com/saidkoussi\" target=\"_blank\">@saidkoussi</a> </p>",
          "rawMarkdown": "This is just 1 red flag - we have at least 5-6 more if you peruse them closely @saidkoussi ",
          "votes": 4
        }
      ]
    },
    {
      "id": 3072581,
      "postDate": "2024-12-15T12:09:10.390Z",
      "content": "<p>Mean Train QWK --&gt; 0.8612<br>\nMean Validation QWK ---&gt; 0.4910<br>\n----&gt; || Optimized QWK SCORE ::  0.543</p>\n<p><em>Combined with the High Score Program</em> got 0.491 pb score,I have good reason to believe that many of the programs are already overfitted and that there is something very wrong with the structure of their code，they overfitted the data by 38%</p>",
      "rawMarkdown": "Mean Train QWK --> 0.8612\nMean Validation QWK ---> 0.4910\n----> || Optimized QWK SCORE ::  0.543\n\n*Combined with the High Score Program* got 0.491 pb score,I have good reason to believe that many of the programs are already overfitted and that there is something very wrong with the structure of their code，they overfitted the data by 38%",
      "votes": 3,
      "replies": [
        {
          "id": 3072585,
          "postDate": "2024-12-15T12:16:12.573Z",
          "content": "<p>Imagine the state of the private LB on 20thDec <a href=\"https://www.kaggle.com/aristotlechen\" target=\"_blank\">@aristotlechen</a>!<br>\nI think this will be the new ICR! Whatever may be the result, a lot of us are going to remember this over the longest run!</p>",
          "rawMarkdown": "Imagine the state of the private LB on 20thDec @aristotlechen!\nI think this will be the new ICR! Whatever may be the result, a lot of us are going to remember this over the longest run!"
        }
      ]
    },
    {
      "id": 3074155,
      "postDate": "2024-12-17T11:34:30.507Z",
      "content": "<p>I manually changed the thresholds of the best public 0.497 notebook (up or down by less than 0.03 per value). The LB dropped to below 0.48</p>",
      "rawMarkdown": "I manually changed the thresholds of the best public 0.497 notebook (up or down by less than 0.03 per value). The LB dropped to below 0.48",
      "votes": 1
    },
    {
      "id": 3072552,
      "postDate": "2024-12-15T11:39:36.673Z",
      "content": "<p>Most of the high scoring public notebooks are like this - changing the seed either takes one up/ down on the public LB. I am unsure of what can be considered a good public kernel in this case <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "rawMarkdown": "Most of the high scoring public notebooks are like this - changing the seed either takes one up/ down on the public LB. I am unsure of what can be considered a good public kernel in this case @ambrosm ",
      "votes": 1,
      "replies": [
        {
          "id": 3073004,
          "postDate": "2024-12-15T22:59:13.983Z",
          "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> , i have one confusion, if we change seed value of a single model , if it score good on public by changing seed is this also count in overfitting ? </p>",
          "rawMarkdown": "@ravi20076 , i have one confusion, if we change seed value of a single model , if it score good on public by changing seed is this also count in overfitting ? ",
          "votes": 1,
          "replies": [
            {
              "id": 3073154,
              "postDate": "2024-12-16T05:38:05.320Z",
              "content": "<p>This is luck and not overfitting <a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> </p>",
              "rawMarkdown": "This is luck and not overfitting @abdmental01 ",
              "votes": 3
            }
          ]
        },
        {
          "id": 3073063,
          "postDate": "2024-12-16T02:14:10.280Z",
          "content": "<p>What are we looking to improve???</p>\n<p>Exactly Which metric???</p>\n<p>All i get my reading here is No of seeds in public nbks are pretty high</p>\n<p>Does high QWK Score mean the nbk is better?</p>",
          "rawMarkdown": "What are we looking to improve???\n\nExactly Which metric???\n\nAll i get my reading here is No of seeds in public nbks are pretty high\n\nDoes high QWK Score mean the nbk is better?",
          "votes": 1,
          "replies": [
            {
              "id": 3073155,
              "postDate": "2024-12-16T05:38:36.637Z",
              "content": "<p>Quadratic Cohen Kappa is the metric - we need to maximize this <a href=\"https://www.kaggle.com/vedantsinghthakur\" target=\"_blank\">@vedantsinghthakur</a> </p>",
              "rawMarkdown": "Quadratic Cohen Kappa is the metric - we need to maximize this @vedantsinghthakur ",
              "votes": 1
            },
            {
              "id": 3073272,
              "postDate": "2024-12-16T08:52:42.560Z",
              "content": "<p>Thanks soo much!</p>",
              "rawMarkdown": "Thanks soo much!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3073300,
      "postDate": "2024-12-16T10:02:05.253Z",
      "content": "<p>Many public notebooks have data leakage when using KNN.</p>",
      "rawMarkdown": "Many public notebooks have data leakage when using KNN.",
      "votes": 2
    },
    {
      "id": 3073122,
      "postDate": "2024-12-16T04:39:38.970Z",
      "content": "<p>Let's wait for the lottery</p>",
      "rawMarkdown": "Let's wait for the lottery",
      "votes": 2,
      "replies": [
        {
          "id": 3073329,
          "postDate": "2024-12-16T11:12:44.180Z",
          "content": "<p>haha, just like waiting \"double color balls\"</p>",
          "rawMarkdown": "haha, just like waiting \"double color balls\"",
          "votes": 1
        },
        {
          "id": 3074379,
          "postDate": "2024-12-17T15:47:14.730Z",
          "content": "<p>i also hope this lottery</p>",
          "rawMarkdown": "i also hope this lottery"
        },
        {
          "id": 3074783,
          "postDate": "2024-12-18T04:30:07.823Z",
          "content": "<p>definitely！</p>",
          "rawMarkdown": "definitely！"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3072570,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-12-15T11:57:32.287000",
      "content": "<p>I checked some competitions with qwk metric and most of them had some kind of a shakeup. Since the penalty for farther off predictions (e.g., predicting 0 for a 3) is quadratic, a small change in the model can result in large swings. The dataset is also small here, so those swings are becoming larger. If I'm not mistaken, there aren't any class 3 in public test set, so it might boil down to how well you are predicting them in private test set.</p>",
      "votes": 14,
      "replies": [
        {
          "id": 3073758,
          "author_name": "Taizhuo Tang",
          "author_url": "",
          "post_date": "2024-12-16T21:29:03.300000",
          "content": "<p><code>If I'm not mistaken, there aren't any class 3 in public test set</code></p>\n<p>This is REALLY interesting and, frankly, sounds like an oversight on the part of whoever's curating the data. How'd you reach this conclusion?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3073833,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-12-17T01:30:13.043000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3074306,
          "author_name": "Geremie Yeo",
          "author_url": "",
          "post_date": "2024-12-17T14:47:14.833000",
          "content": "<p>Hmm I doubt this is true. I tried changing all my sii=3 predictions to 2, and the LB decreased slightly </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3074316,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-12-17T15:02:35.553000",
              "content": "<p>It's my bad then. My last threshold was 1.5 or something like that and I tried changing it to a large number in order to change all 3 predictions to 2, but my score didn't change at all. I guess my model wasn't predicting anything beyond that threshold.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3072596,
      "author_name": "Aurora_blue",
      "author_url": "",
      "post_date": "2024-12-15T12:34:04.493000",
      "content": "<p>The easiest way to get notebook medal may be to publish high score overfitting ensemble notebook near end.<br>\nUnfortunately, they may bring us almost nothing but copy-and-paste content.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 3073037,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-12-16T00:40:44.900000",
          "content": "<blockquote>\n  <p>Unfortunately, they may bring us almost nothing but copy-and-paste content.  </p>\n</blockquote>\n<p>There is actually a great value in knowing that the upper score limit can be achieved by random overfitting.  </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3073041,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-12-16T01:01:51.247000",
              "content": "<p>A kaggle friends say that when they have free time, they repeatedly change the seed of best public notebooks and submit them to observe the lb score variance. This may be good analysis for winning comp.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3072582,
      "author_name": "AI-Cat",
      "author_url": "",
      "post_date": "2024-12-15T12:12:27.837000",
      "content": "<p>I noticed such strange paradoxes with the Prem Chotepanit - CMI| Tuning | Ensemble of solutions V111 notebook, lb score 0.497. </p>\n<p>If you change the accelerator from GPU T4 x2 to GPU P100, the score drops to 0.488. </p>\n<ul>\n<li>GPU T4 x2 - Mean Train QWK: 0.7240, Mean Validation QWK: 0.4613 Optimized QWK SCORE: 0.5185 [Model 1]</li>\n<li>GPU P100 - Mean Train QWK: 0.7278, Mean Validation QWK: 0.4758 Optimized QWK SCORE: 0.5298 [Model 1]</li>\n</ul>\n<p>There is an explanation for this, and to test the idea, for example, I re-sorted the incoming sets, LB fell to 0.452. </p>\n<ul>\n<li>GPU T4 x2 - Mean Train QWK: 0.7293, Mean Validation QWK: 0.4752 Optimized QWK SCORE: 0.5324 [Model 1]</li>\n</ul>\n<p>imputer = KNNImputer(n_neighbors=4) LB=0.487<br>\nMean Train QWK: 0.7374, Mean Validation QWK: 0.4816 Optimized QWK SCORE: 0.5423 [Model 1]</p>\n<p>imputer = KNNImputer(n_neighbors=6) LB=0.489<br>\nMean Train QWK: 0.7358, Mean Validation QWK: 0.4888 Optimized QWK SCORE: 0.5555 [Model 1]</p>\n<p>imputer = KNNImputer(n_neighbors=52) LB=0.484<br>\nMean Train QWK: 0.7676, Mean Validation QWK: 0.5490 Optimized QWK SCORE: 0.6012 [Model 1]</p>\n<p>How is this possible?</p>\n<p>P.S.<br>\n— Is it possible to achieve a result without knowing the reasons for the phenomenon?<br>\n— I don’t know what to answer… The reason for the strength of damask steel was discovered only in the nineteenth century, but the best swords were forged in the eleventh.</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 3072768,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-12-15T16:19:26.253000",
      "content": "<p>Yep. Changed seed for this public NB &gt;&gt; LB score dropped by 23 points 🤣  <br>\nThis is very fun competition. Those are some great odds for winnings- much higher than regular lottery! May the odds will be with me 🤣  </p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 3072569,
      "author_name": "work work",
      "author_url": "",
      "post_date": "2024-12-15T11:57:08.907000",
      "content": "<p><a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> yes you are right , for high scored public notebook , thee is a lot on flctatuations on the scores when we changes the SEED , THIS RED FLAG OF OVERFITTING </p>",
      "votes": 6,
      "replies": [
        {
          "id": 3072580,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-12-15T12:08:40.580000",
          "content": "<p>This is just 1 red flag - we have at least 5-6 more if you peruse them closely <a href=\"https://www.kaggle.com/saidkoussi\" target=\"_blank\">@saidkoussi</a> </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 3072581,
      "author_name": "Aristotle.Chen",
      "author_url": "",
      "post_date": "2024-12-15T12:09:10.390000",
      "content": "<p>Mean Train QWK --&gt; 0.8612<br>\nMean Validation QWK ---&gt; 0.4910<br>\n----&gt; || Optimized QWK SCORE ::  0.543</p>\n<p><em>Combined with the High Score Program</em> got 0.491 pb score,I have good reason to believe that many of the programs are already overfitted and that there is something very wrong with the structure of their code，they overfitted the data by 38%</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3072585,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-12-15T12:16:12.573000",
          "content": "<p>Imagine the state of the private LB on 20thDec <a href=\"https://www.kaggle.com/aristotlechen\" target=\"_blank\">@aristotlechen</a>!<br>\nI think this will be the new ICR! Whatever may be the result, a lot of us are going to remember this over the longest run!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3074155,
      "author_name": "Geremie Yeo",
      "author_url": "",
      "post_date": "2024-12-17T11:34:30.507000",
      "content": "<p>I manually changed the thresholds of the best public 0.497 notebook (up or down by less than 0.03 per value). The LB dropped to below 0.48</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3072552,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2024-12-15T11:39:36.673000",
      "content": "<p>Most of the high scoring public notebooks are like this - changing the seed either takes one up/ down on the public LB. I am unsure of what can be considered a good public kernel in this case <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3073004,
          "author_name": "Sheikh Muhammad Abdullah",
          "author_url": "",
          "post_date": "2024-12-15T22:59:13.983000",
          "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> , i have one confusion, if we change seed value of a single model , if it score good on public by changing seed is this also count in overfitting ? </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3073154,
              "author_name": "Ravi Ramakrishnan",
              "author_url": "",
              "post_date": "2024-12-16T05:38:05.320000",
              "content": "<p>This is luck and not overfitting <a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> </p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 3073063,
          "author_name": "Vedant Singh Thakur",
          "author_url": "",
          "post_date": "2024-12-16T02:14:10.280000",
          "content": "<p>What are we looking to improve???</p>\n<p>Exactly Which metric???</p>\n<p>All i get my reading here is No of seeds in public nbks are pretty high</p>\n<p>Does high QWK Score mean the nbk is better?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3073155,
              "author_name": "Ravi Ramakrishnan",
              "author_url": "",
              "post_date": "2024-12-16T05:38:36.637000",
              "content": "<p>Quadratic Cohen Kappa is the metric - we need to maximize this <a href=\"https://www.kaggle.com/vedantsinghthakur\" target=\"_blank\">@vedantsinghthakur</a> </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3073272,
              "author_name": "Vedant Singh Thakur",
              "author_url": "",
              "post_date": "2024-12-16T08:52:42.560000",
              "content": "<p>Thanks soo much!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3073300,
      "author_name": "SnakeWoodMan",
      "author_url": "",
      "post_date": "2024-12-16T10:02:05.253000",
      "content": "<p>Many public notebooks have data leakage when using KNN.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3073122,
      "author_name": "Luck is all you need",
      "author_url": "",
      "post_date": "2024-12-16T04:39:38.970000",
      "content": "<p>Let's wait for the lottery</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3073329,
          "author_name": "Xiaolei Lian",
          "author_url": "",
          "post_date": "2024-12-16T11:12:44.180000",
          "content": "<p>haha, just like waiting \"double color balls\"</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3074379,
          "author_name": "Đàm Văn Tài",
          "author_url": "",
          "post_date": "2024-12-17T15:47:14.730000",
          "content": "<p>i also hope this lottery</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3074783,
          "author_name": "yueming",
          "author_url": "",
          "post_date": "2024-12-18T04:30:07.823000",
          "content": "<p>definitely！</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3072530": "If we compare today's two highest-scoring public notebooks, we see that they differ in exactly one number:\n\n[CMI| Tuning | Ensemble of solutions](https://www.kaggle.com/code/batprem/cmi-tuning-ensemble-of-solutions) V111, lb score 0.497, 2 days ago:\n```\n    KappaOPtimizer = minimize(evaluate_predictions,\n                              x0=[0.5, 1.49, 2.5], args=(y, oof_non_rounded), \n                              method='Nelder-Mead')\n```\n\n[CMI: Issues with the Metric and Baseline](https://www.kaggle.com/code/vitalykudelya/cmi-issues-with-the-metric-and-baseline) V3, lb score 0.495, 3 days ago:\n```\n    KappaOPtimizer = minimize(evaluate_predictions,\n                              x0=[0.5, 1.5, 2.5], args=(y, oof_non_rounded), \n                              method='Nelder-Mead')\n```\n\nThe output of the corresponding cells differs slightly:\n\nCMI| Tuning | Ensemble of solutions:\n```\n----> || Optimized QWK SCORE ::  0.456\n```\n\nCMI: Issues with the Metric and Baseline\n```\n----> || Optimized QWK SCORE ::  0.457\n```\n\nWhat does this mean? The newer notebook changes the initial guess of the [Nelder–Mead optimization](https://docs.scipy.org/doc/scipy/reference/generated/scipy.optimize.minimize.html) (not even a model hyperparameter), the Nelder–Mead optimization falls into another local minimum, the cv score goes down and the leaderboard score goes up. \n\n\"cv down and leaderboard up\" should ring a bell: This is called overfitting to the leaderboard. Don't expect that the newer model will keep its advantage on the private leaderboard!\n\nIf you belong to the hundreds of people who have copied and submitted this notebook, try changing the seed and resubmit! If the notebook is good, it will keep its high score with other seeds. If the public leaderboard score drops when changing the seed, don't expect anything good from the private leaderboard. The transition from public to private is equivalent to changing the seed.\n\nP.S. In the meantime, another public notebook has appeared, [LB.497|Multi-Model Feature Importance Analysis](https://www.kaggle.com/code/shodaifuruya/lb-497-multi-model-feature-importance-analysis) V15. It is an exact copy of the overfitting model with the addition of some feature importance bar charts.",
    "3072570": "I checked some competitions with qwk metric and most of them had some kind of a shakeup. Since the penalty for farther off predictions (e.g., predicting 0 for a 3) is quadratic, a small change in the model can result in large swings. The dataset is also small here, so those swings are becoming larger. If I'm not mistaken, there aren't any class 3 in public test set, so it might boil down to how well you are predicting them in private test set.",
    "3072596": "The easiest way to get notebook medal may be to publish high score overfitting ensemble notebook near end.\nUnfortunately, they may bring us almost nothing but copy-and-paste content.\n",
    "3072582": "I noticed such strange paradoxes with the Prem Chotepanit - CMI| Tuning | Ensemble of solutions V111 notebook, lb score 0.497. \n\nIf you change the accelerator from GPU T4 x2 to GPU P100, the score drops to 0.488. \n- GPU T4 x2 - Mean Train QWK: 0.7240, Mean Validation QWK: 0.4613 Optimized QWK SCORE: 0.5185 [Model 1]\n- GPU P100 - Mean Train QWK: 0.7278, Mean Validation QWK: 0.4758 Optimized QWK SCORE: 0.5298 [Model 1]\n\nThere is an explanation for this, and to test the idea, for example, I re-sorted the incoming sets, LB fell to 0.452. \n- GPU T4 x2 - Mean Train QWK: 0.7293, Mean Validation QWK: 0.4752 Optimized QWK SCORE: 0.5324 [Model 1]\n\nimputer = KNNImputer(n_neighbors=4) LB=0.487\nMean Train QWK: 0.7374, Mean Validation QWK: 0.4816 Optimized QWK SCORE: 0.5423 [Model 1]\n\nimputer = KNNImputer(n_neighbors=6) LB=0.489\nMean Train QWK: 0.7358, Mean Validation QWK: 0.4888 Optimized QWK SCORE: 0.5555 [Model 1]\n\nimputer = KNNImputer(n_neighbors=52) LB=0.484\nMean Train QWK: 0.7676, Mean Validation QWK: 0.5490 Optimized QWK SCORE: 0.6012 [Model 1]\n\nHow is this possible?\n\nP.S.\n— Is it possible to achieve a result without knowing the reasons for the phenomenon?\n— I don’t know what to answer… The reason for the strength of damask steel was discovered only in the nineteenth century, but the best swords were forged in the eleventh.",
    "3072768": "Yep. Changed seed for this public NB >> LB score dropped by 23 points 🤣  \nThis is very fun competition. Those are some great odds for winnings- much higher than regular lottery! May the odds will be with me 🤣  ",
    "3072569": "@ambrosm yes you are right , for high scored public notebook , thee is a lot on flctatuations on the scores when we changes the SEED , THIS RED FLAG OF OVERFITTING ",
    "3072581": "Mean Train QWK --> 0.8612\nMean Validation QWK ---> 0.4910\n----> || Optimized QWK SCORE ::  0.543\n\n*Combined with the High Score Program* got 0.491 pb score,I have good reason to believe that many of the programs are already overfitted and that there is something very wrong with the structure of their code，they overfitted the data by 38%",
    "3074155": "I manually changed the thresholds of the best public 0.497 notebook (up or down by less than 0.03 per value). The LB dropped to below 0.48",
    "3072552": "Most of the high scoring public notebooks are like this - changing the seed either takes one up/ down on the public LB. I am unsure of what can be considered a good public kernel in this case @ambrosm ",
    "3073300": "Many public notebooks have data leakage when using KNN.",
    "3073122": "Let's wait for the lottery"
  }
}