{
  "id": 457753,
  "title": "Stop spoiling the competition with high scoring notebooks",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/457753",
  "author_name": "Nischay Dhankhar",
  "post_date": "2023-11-26T16:02:17.298000",
  "votes": 29,
  "comment_count": 18,
  "views": 0,
  "content": "<p>No wonder tuning every sample/target manually will most likely deteriorate private leaderboard score. Sharing such high scoring gold place notebooks in final week should be stopped, as it completely ruin the importance of public leaderboard and distract the new kagglers to spend time more time in wrong direction. </p>\n<p>This is a small snippet of code from top scoring notebook, <strong>massive overfitting</strong>:</p>\n<pre><code>\n\n[  :   +] =     .   + df[  :   +]  - .    #  good              .  #  .\n[ :  +] =     .    + df[ :  +]  - .     #  good               .  #  .\n[ :  +] =     .   #+ df[ :  +]  - .   #  good              df[ :  +] =   .   #.   #   .\n[ :  +] =     .    + df[ :  +]  - .   # good               .    #.  #\n[ :  +] =     .    + df[ :  +]  - .   # good               .    #. #\n[ :  +] =     .   + df[ :  +]  - .   # good               . .  #\n[ :  +] =     .  + df[ :  +]  - .   #  good              .  #   \n[ :  +] =     .  + df[ :  +]  - .   # good               #.  # .\n[ :  +] =     .  + df[ :  +]  - .    #   good                 .  #.  #\n[ :  +] =     .  + df[ :  +]  - . #  good  weekSelect.            .  #.  #\n[ :  +] =     .  + df[ :  +]  - . #  good  weekSelect.        .  #.  #\n[ :  +] =     .  + df[ :  +]  - . #   weekSelect. add  ???        .  #.  #\n</code></pre>",
  "messages": [
    {
      "id": 2539021,
      "postDate": "2023-11-26T16:02:17.297Z",
      "content": "<p>No wonder tuning every sample/target manually will most likely deteriorate private leaderboard score. Sharing such high scoring gold place notebooks in final week should be stopped, as it completely ruin the importance of public leaderboard and distract the new kagglers to spend time more time in wrong direction. </p>\n<p>This is a small snippet of code from top scoring notebook, <strong>massive overfitting</strong>:</p>\n<pre><code>\n\n[  :   +] =     .   + df[  :   +]  - .    #  good              .  #  .\n[ :  +] =     .    + df[ :  +]  - .     #  good               .  #  .\n[ :  +] =     .   #+ df[ :  +]  - .   #  good              df[ :  +] =   .   #.   #   .\n[ :  +] =     .    + df[ :  +]  - .   # good               .    #.  #\n[ :  +] =     .    + df[ :  +]  - .   # good               .    #. #\n[ :  +] =     .   + df[ :  +]  - .   # good               . .  #\n[ :  +] =     .  + df[ :  +]  - .   #  good              .  #   \n[ :  +] =     .  + df[ :  +]  - .   # good               #.  # .\n[ :  +] =     .  + df[ :  +]  - .    #   good                 .  #.  #\n[ :  +] =     .  + df[ :  +]  - . #  good  weekSelect.            .  #.  #\n[ :  +] =     .  + df[ :  +]  - . #  good  weekSelect.        .  #.  #\n[ :  +] =     .  + df[ :  +]  - . #   weekSelect. add  ???        .  #.  #\n</code></pre>",
      "rawMarkdown": "No wonder tuning every sample/target manually will most likely deteriorate private leaderboard score. Sharing such high scoring gold place notebooks in final week should be stopped, as it completely ruin the importance of public leaderboard and distract the new kagglers to spend time more time in wrong direction. \n\n\nThis is a small snippet of code from top scoring notebook, **massive overfitting**:\n\n```\n##############\n  \ndf[58  : 58  +1] =     7.2068508521732895   + df[58  : 58  +1]  - 4.4488508521732895    #  good              0.05505962617010652  #  4.4488508521732895\ndf[249 : 249 +1] =     4.019050299888815    + df[249 : 249 +1]  - 2.779050299888815     #  good               .04099198546854504  #  2.779050299888815\ndf[122 : 122 +1] =     1.2170550317747087   #+ df[122 : 122 +1]  - 3.8354043063329506   #  good  544            df[122 : 122 +1] =   2.5470550317747087   #0.0318830319364953   #   2.5570550317747087\ndf[185 : 185 +1] =     0.759059971979153    + df[185 : 185 +1]  - 2.11905997197915319   # good               2.119059971979153    #0.02817637204112935  #\ndf[215 : 215 +1] =     0.967704245720528    + df[215 : 215 +1]  - 1.96770424572052804   # good               1.967704245720528    #0.023496010800942546 #\ndf[179 : 179 +1] =     2.3142648679548018   + df[179 : 179 +1]  - 1.67426486795480178   # good               . 0.02173648680072903  #\ndf[209 : 209 +1] =     0.83435555596533819  + df[209 : 209 +1]  - 0.73435562905765783   #  good              0.02028904039707182  #  185 215\ndf[232 : 232 +1] =     0.80731463350279168  + df[232 : 232 +1]  - 1.09731483117745654   # good               #0.01531950341962440  # 1.09731463350279168\ndf[182 : 182 +1] =     0.57364752751639799  + df[182 : 182 +1]  - 1.3456692698176478    #   good                 0.78664752751639799  #0.01462655858637081  #\ndf[128 : 128 +1] =     2.44645498049249803  + df[128 : 128 +1]  - 0.7464551083726029 #  good  weekSelect.            0.74645498049249803  #0.01419926547415611  #\ndf[208 : 208 +1] =     1.78602244991658055  + df[208 : 208 +1]  - 0.7860227180528301 #  good  weekSelect.        0.78602244991658055  #0.01374043895474812  #\ndf[209 : 209 +1] =     0.83435555596533819  + df[209 : 209 +1]  - 0.8343555887327185 #   weekSelect. add  ???        0.73435555596533819  #0.01363631071050974  #\n```",
      "votes": 28
    },
    {
      "id": 2539227,
      "postDate": "2023-11-26T21:08:04.643Z",
      "content": "<p>Thank you for breaking the silence about this issue! I've been expecting a shakeup at the end of the competition all along, but I find it sad that the public leaderboard has now become entirely useless as an indicator for how well one's doing compared to the other competitors.</p>",
      "rawMarkdown": "Thank you for breaking the silence about this issue! I've been expecting a shakeup at the end of the competition all along, but I find it sad that the public leaderboard has now become entirely useless as an indicator for how well one's doing compared to the other competitors.",
      "votes": 11
    },
    {
      "id": 2539639,
      "postDate": "2023-11-27T08:28:29.240Z",
      "content": "<p>I actually like those kind of notebooks. It's the Kaggle version of natural selection. </p>",
      "rawMarkdown": "I actually like those kind of notebooks. It's the Kaggle version of natural selection. ",
      "votes": 12,
      "replies": [
        {
          "id": 2539651,
          "postDate": "2023-11-27T08:35:47.737Z",
          "content": "<p>Exactly my thoughts. </p>",
          "rawMarkdown": "Exactly my thoughts. ",
          "votes": 1
        },
        {
          "id": 2544454,
          "postDate": "2023-11-30T20:22:52.080Z",
          "content": "<p>Good point <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> !</p>",
          "rawMarkdown": "Good point @gunesevitan !"
        }
      ]
    },
    {
      "id": 2539052,
      "postDate": "2023-11-26T16:52:16.993Z",
      "content": "<p>Clearly the public LB has descended into self-parody and is now all-but completely useless. Maybe it adds a layer of mystery to have absolutely no idea where we stand? Enjoy the shake-up!</p>",
      "rawMarkdown": "Clearly the public LB has descended into self-parody and is now all-but completely useless. Maybe it adds a layer of mystery to have absolutely no idea where we stand? Enjoy the shake-up!",
      "votes": 7
    },
    {
      "id": 2539047,
      "postDate": "2023-11-26T16:45:06.567Z",
      "content": "<p>I'm crossing my fingers that the next test set of the next competitions would be so colossal that rookies will need a lifetime to unravel the mysteries of overfitting tuning 🤣 I only used 4 models with w1<em>df1+…+w4</em>df4 </p>",
      "rawMarkdown": "I'm crossing my fingers that the next test set of the next competitions would be so colossal that rookies will need a lifetime to unravel the mysteries of overfitting tuning 🤣 I only used 4 models with w1*df1+...+w4*df4 ",
      "votes": 7,
      "replies": [
        {
          "id": 2544457,
          "postDate": "2023-11-30T20:24:16.977Z",
          "content": "<p>Great work <a href=\"https://www.kaggle.com/eliork\" target=\"_blank\">@eliork</a> ! Hope you will get higher in the private LB</p>",
          "rawMarkdown": "Great work @eliork ! Hope you will get higher in the private LB"
        }
      ]
    },
    {
      "id": 2539554,
      "postDate": "2023-11-27T06:15:24.617Z",
      "content": "<p>I welcome it! The more people that submit their public leaderboard overfit 'predictions' the better chance people who actually submit models that generalize will do on the private leaderboard. And in the end, that's all that matters.</p>",
      "rawMarkdown": "I welcome it! The more people that submit their public leaderboard overfit 'predictions' the better chance people who actually submit models that generalize will do on the private leaderboard. And in the end, that's all that matters.",
      "votes": 3,
      "replies": [
        {
          "id": 2542826,
          "postDate": "2023-11-29T14:46:03.387Z",
          "content": "<p>Participants should feel the shake up on the Private LB themselves.</p>",
          "rawMarkdown": "Participants should feel the shake up on the Private LB themselves."
        }
      ]
    },
    {
      "id": 2541137,
      "postDate": "2023-11-28T08:57:20.120Z",
      "content": "<p>I agree that this kind of LB probing is rather disorderly conduct.</p>\n<p>However, this already happened, and I think one can just blend their solution with the top-scored notebook of these <strong>using only rows that belong to the public test</strong>. In such a way the predictions for the private part of the test are untouched and this will not affect the final private LB score. At the same time, with such a blend, one would see its public LB position more accurately. </p>",
      "rawMarkdown": "I agree that this kind of LB probing is rather disorderly conduct.\n\nHowever, this already happened, and I think one can just blend their solution with the top-scored notebook of these **using only rows that belong to the public test**. In such a way the predictions for the private part of the test are untouched and this will not affect the final private LB score. At the same time, with such a blend, one would see its public LB position more accurately. ",
      "votes": 3
    },
    {
      "id": 2543670,
      "postDate": "2023-11-30T09:01:18.897Z",
      "content": "<p>Agreed, sharing high scoring notebooks close to competition end can mislead new Kagglers. It's crucial to focus on genuine model improvement rather than overfitting to public LB. Building a robust model that generalizes well to unseen data is key, rather than chasing leaderboard scores with overfit methods.</p>",
      "rawMarkdown": "Agreed, sharing high scoring notebooks close to competition end can mislead new Kagglers. It's crucial to focus on genuine model improvement rather than overfitting to public LB. Building a robust model that generalizes well to unseen data is key, rather than chasing leaderboard scores with overfit methods.",
      "votes": 1
    },
    {
      "id": 2543378,
      "postDate": "2023-11-30T03:23:25.333Z",
      "content": "<p>This indeed is a great issue! Why doesn't Kaggle do anything about it? People copy others' notebook and paste in verge of getting top scores. Learning is different thing but just copying blindly is just the waste of time.</p>",
      "rawMarkdown": "This indeed is a great issue! Why doesn't Kaggle do anything about it? People copy others' notebook and paste in verge of getting top scores. Learning is different thing but just copying blindly is just the waste of time.",
      "votes": 1
    },
    {
      "id": 2540681,
      "postDate": "2023-11-27T22:27:09.213Z",
      "content": "<p>pure overfitting!</p>",
      "rawMarkdown": "pure overfitting!",
      "votes": 1
    },
    {
      "id": 2539197,
      "postDate": "2023-11-26T20:04:49.797Z",
      "content": "<p>idk how, but I knew this would start. The competition become a joke and I feel embarrassed I'm a passive part of this circus…</p>",
      "rawMarkdown": "idk how, but I knew this would start. The competition become a joke and I feel embarrassed I'm a passive part of this circus...",
      "votes": 1
    },
    {
      "id": 2544496,
      "postDate": "2023-11-30T21:17:57.823Z",
      "content": "<p>I'm interested in a model which is fitted with submission data except above rows. If that model won, this competition would be a complete failure.</p>",
      "rawMarkdown": "I'm interested in a model which is fitted with submission data except above rows. If that model won, this competition would be a complete failure."
    },
    {
      "id": 2543384,
      "postDate": "2023-11-30T03:35:26.900Z",
      "content": "<p>I agree. 0.531 is too much in the leaderboard.</p>",
      "rawMarkdown": "I agree. 0.531 is too much in the leaderboard."
    },
    {
      "id": 2543328,
      "postDate": "2023-11-30T01:50:16.583Z",
      "content": "<p>I agree with your opinion. It seems likely that there will be a big shake-up/down in this competition. I hope that the model-based method has high generalizability and can get good scores…</p>",
      "rawMarkdown": "I agree with your opinion. It seems likely that there will be a big shake-up/down in this competition. I hope that the model-based method has high generalizability and can get good scores..."
    },
    {
      "id": 2539369,
      "postDate": "2023-11-27T02:33:32.553Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2539227,
      "author_name": "Frenio Redeker",
      "author_url": "",
      "post_date": "2023-11-26T21:08:04.643000",
      "content": "<p>Thank you for breaking the silence about this issue! I've been expecting a shakeup at the end of the competition all along, but I find it sad that the public leaderboard has now become entirely useless as an indicator for how well one's doing compared to the other competitors.</p>",
      "votes": 11,
      "replies": []
    },
    {
      "id": 2539639,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-11-27T08:28:29.240000",
      "content": "<p>I actually like those kind of notebooks. It's the Kaggle version of natural selection. </p>",
      "votes": 12,
      "replies": [
        {
          "id": 2539651,
          "author_name": "Joakim Arvidsson",
          "author_url": "",
          "post_date": "2023-11-27T08:35:47.737000",
          "content": "<p>Exactly my thoughts. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2544454,
          "author_name": "Octavi Grau",
          "author_url": "",
          "post_date": "2023-11-30T20:22:52.080000",
          "content": "<p>Good point <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2539052,
      "author_name": "John Mitchell",
      "author_url": "",
      "post_date": "2023-11-26T16:52:16.993000",
      "content": "<p>Clearly the public LB has descended into self-parody and is now all-but completely useless. Maybe it adds a layer of mystery to have absolutely no idea where we stand? Enjoy the shake-up!</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 2539047,
      "author_name": "EliKal",
      "author_url": "",
      "post_date": "2023-11-26T16:45:06.567000",
      "content": "<p>I'm crossing my fingers that the next test set of the next competitions would be so colossal that rookies will need a lifetime to unravel the mysteries of overfitting tuning 🤣 I only used 4 models with w1<em>df1+…+w4</em>df4 </p>",
      "votes": 7,
      "replies": [
        {
          "id": 2544457,
          "author_name": "Octavi Grau",
          "author_url": "",
          "post_date": "2023-11-30T20:24:16.977000",
          "content": "<p>Great work <a href=\"https://www.kaggle.com/eliork\" target=\"_blank\">@eliork</a> ! Hope you will get higher in the private LB</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2539554,
      "author_name": "Joakim Arvidsson",
      "author_url": "",
      "post_date": "2023-11-27T06:15:24.617000",
      "content": "<p>I welcome it! The more people that submit their public leaderboard overfit 'predictions' the better chance people who actually submit models that generalize will do on the private leaderboard. And in the end, that's all that matters.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2542826,
          "author_name": "Zhuravlev Viktor",
          "author_url": "",
          "post_date": "2023-11-29T14:46:03.387000",
          "content": "<p>Participants should feel the shake up on the Private LB themselves.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2541137,
      "author_name": "Antonina Dolgorukova",
      "author_url": "",
      "post_date": "2023-11-28T08:57:20.120000",
      "content": "<p>I agree that this kind of LB probing is rather disorderly conduct.</p>\n<p>However, this already happened, and I think one can just blend their solution with the top-scored notebook of these <strong>using only rows that belong to the public test</strong>. In such a way the predictions for the private part of the test are untouched and this will not affect the final private LB score. At the same time, with such a blend, one would see its public LB position more accurately. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2543670,
      "author_name": "Nikolenko_Sergei",
      "author_url": "",
      "post_date": "2023-11-30T09:01:18.897000",
      "content": "<p>Agreed, sharing high scoring notebooks close to competition end can mislead new Kagglers. It's crucial to focus on genuine model improvement rather than overfitting to public LB. Building a robust model that generalizes well to unseen data is key, rather than chasing leaderboard scores with overfit methods.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2543378,
      "author_name": "Abhas Malguri",
      "author_url": "",
      "post_date": "2023-11-30T03:23:25.333000",
      "content": "<p>This indeed is a great issue! Why doesn't Kaggle do anything about it? People copy others' notebook and paste in verge of getting top scores. Learning is different thing but just copying blindly is just the waste of time.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2540681,
      "author_name": "JK-Piece",
      "author_url": "",
      "post_date": "2023-11-27T22:27:09.213000",
      "content": "<p>pure overfitting!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2539197,
      "author_name": "Yurnero",
      "author_url": "",
      "post_date": "2023-11-26T20:04:49.797000",
      "content": "<p>idk how, but I knew this would start. The competition become a joke and I feel embarrassed I'm a passive part of this circus…</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2544496,
      "author_name": "Daisy",
      "author_url": "",
      "post_date": "2023-11-30T21:17:57.823000",
      "content": "<p>I'm interested in a model which is fitted with submission data except above rows. If that model won, this competition would be a complete failure.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2543384,
      "author_name": "FanghuiZh",
      "author_url": "",
      "post_date": "2023-11-30T03:35:26.900000",
      "content": "<p>I agree. 0.531 is too much in the leaderboard.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2543328,
      "author_name": "shisa07",
      "author_url": "",
      "post_date": "2023-11-30T01:50:16.583000",
      "content": "<p>I agree with your opinion. It seems likely that there will be a big shake-up/down in this competition. I hope that the model-based method has high generalizability and can get good scores…</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2539369,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-11-27T02:33:32.553000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2539021": "No wonder tuning every sample/target manually will most likely deteriorate private leaderboard score. Sharing such high scoring gold place notebooks in final week should be stopped, as it completely ruin the importance of public leaderboard and distract the new kagglers to spend time more time in wrong direction. \n\n\nThis is a small snippet of code from top scoring notebook, **massive overfitting**:\n\n```\n##############\n  \ndf[58  : 58  +1] =     7.2068508521732895   + df[58  : 58  +1]  - 4.4488508521732895    #  good              0.05505962617010652  #  4.4488508521732895\ndf[249 : 249 +1] =     4.019050299888815    + df[249 : 249 +1]  - 2.779050299888815     #  good               .04099198546854504  #  2.779050299888815\ndf[122 : 122 +1] =     1.2170550317747087   #+ df[122 : 122 +1]  - 3.8354043063329506   #  good  544            df[122 : 122 +1] =   2.5470550317747087   #0.0318830319364953   #   2.5570550317747087\ndf[185 : 185 +1] =     0.759059971979153    + df[185 : 185 +1]  - 2.11905997197915319   # good               2.119059971979153    #0.02817637204112935  #\ndf[215 : 215 +1] =     0.967704245720528    + df[215 : 215 +1]  - 1.96770424572052804   # good               1.967704245720528    #0.023496010800942546 #\ndf[179 : 179 +1] =     2.3142648679548018   + df[179 : 179 +1]  - 1.67426486795480178   # good               . 0.02173648680072903  #\ndf[209 : 209 +1] =     0.83435555596533819  + df[209 : 209 +1]  - 0.73435562905765783   #  good              0.02028904039707182  #  185 215\ndf[232 : 232 +1] =     0.80731463350279168  + df[232 : 232 +1]  - 1.09731483117745654   # good               #0.01531950341962440  # 1.09731463350279168\ndf[182 : 182 +1] =     0.57364752751639799  + df[182 : 182 +1]  - 1.3456692698176478    #   good                 0.78664752751639799  #0.01462655858637081  #\ndf[128 : 128 +1] =     2.44645498049249803  + df[128 : 128 +1]  - 0.7464551083726029 #  good  weekSelect.            0.74645498049249803  #0.01419926547415611  #\ndf[208 : 208 +1] =     1.78602244991658055  + df[208 : 208 +1]  - 0.7860227180528301 #  good  weekSelect.        0.78602244991658055  #0.01374043895474812  #\ndf[209 : 209 +1] =     0.83435555596533819  + df[209 : 209 +1]  - 0.8343555887327185 #   weekSelect. add  ???        0.73435555596533819  #0.01363631071050974  #\n```",
    "2539227": "Thank you for breaking the silence about this issue! I've been expecting a shakeup at the end of the competition all along, but I find it sad that the public leaderboard has now become entirely useless as an indicator for how well one's doing compared to the other competitors.",
    "2539639": "I actually like those kind of notebooks. It's the Kaggle version of natural selection. ",
    "2539052": "Clearly the public LB has descended into self-parody and is now all-but completely useless. Maybe it adds a layer of mystery to have absolutely no idea where we stand? Enjoy the shake-up!",
    "2539047": "I'm crossing my fingers that the next test set of the next competitions would be so colossal that rookies will need a lifetime to unravel the mysteries of overfitting tuning 🤣 I only used 4 models with w1*df1+...+w4*df4 ",
    "2539554": "I welcome it! The more people that submit their public leaderboard overfit 'predictions' the better chance people who actually submit models that generalize will do on the private leaderboard. And in the end, that's all that matters.",
    "2541137": "I agree that this kind of LB probing is rather disorderly conduct.\n\nHowever, this already happened, and I think one can just blend their solution with the top-scored notebook of these **using only rows that belong to the public test**. In such a way the predictions for the private part of the test are untouched and this will not affect the final private LB score. At the same time, with such a blend, one would see its public LB position more accurately. ",
    "2543670": "Agreed, sharing high scoring notebooks close to competition end can mislead new Kagglers. It's crucial to focus on genuine model improvement rather than overfitting to public LB. Building a robust model that generalizes well to unseen data is key, rather than chasing leaderboard scores with overfit methods.",
    "2543378": "This indeed is a great issue! Why doesn't Kaggle do anything about it? People copy others' notebook and paste in verge of getting top scores. Learning is different thing but just copying blindly is just the waste of time.",
    "2540681": "pure overfitting!",
    "2539197": "idk how, but I knew this would start. The competition become a joke and I feel embarrassed I'm a passive part of this circus...",
    "2544496": "I'm interested in a model which is fitted with submission data except above rows. If that model won, this competition would be a complete failure.",
    "2543384": "I agree. 0.531 is too much in the leaderboard.",
    "2543328": "I agree with your opinion. It seems likely that there will be a big shake-up/down in this competition. I hope that the model-based method has high generalizability and can get good scores...",
    "2539369": ""
  }
}