{
  "id": 487615,
  "title": "Be careful with public models! ",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/487615",
  "author_name": "Cody_Null",
  "post_date": "2024-03-29T18:25:44.410000",
  "votes": 11,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi all, I do see a new high-scoring public model has been released. Though I personally believe it is inappropriate to be posting high-scoring public notebooks this close to the competition end, regardless of its probability of success in the private LB, I wanted to bring additional attention to this. Just because this notebook may have a higher LB than an existing notebook of yours does not necessarily mean it will be a better submission. Please do your due diligence and be extra careful when selecting submissions. Think about what could affect the final scores and what is required to make each model succeed. Be careful and good luck out there!</p>",
  "messages": [
    {
      "id": 2722712,
      "postDate": "2024-03-29T18:25:44.410Z",
      "content": "<p>Hi all, I do see a new high-scoring public model has been released. Though I personally believe it is inappropriate to be posting high-scoring public notebooks this close to the competition end, regardless of its probability of success in the private LB, I wanted to bring additional attention to this. Just because this notebook may have a higher LB than an existing notebook of yours does not necessarily mean it will be a better submission. Please do your due diligence and be extra careful when selecting submissions. Think about what could affect the final scores and what is required to make each model succeed. Be careful and good luck out there!</p>",
      "rawMarkdown": "Hi all, I do see a new high-scoring public model has been released. Though I personally believe it is inappropriate to be posting high-scoring public notebooks this close to the competition end, regardless of its probability of success in the private LB, I wanted to bring additional attention to this. Just because this notebook may have a higher LB than an existing notebook of yours does not necessarily mean it will be a better submission. Please do your due diligence and be extra careful when selecting submissions. Think about what could affect the final scores and what is required to make each model succeed. Be careful and good luck out there!",
      "votes": 10
    },
    {
      "id": 2744504,
      "postDate": "2024-04-10T00:17:54.750Z",
      "content": "<p>I checked the public code and found that ensemble of public models get 0.35X in private LB, bronze zone. <br>\nIndeed that each single model did not get top of public LB, but ensembles of it destroyed LB.<br>\nFinally, as discussed again, it would be not recommended to share high score notebook near ends.</p>",
      "rawMarkdown": "I checked the public code and found that ensemble of public models get 0.35X in private LB, bronze zone. \nIndeed that each single model did not get top of public LB, but ensembles of it destroyed LB.\nFinally, as discussed again, it would be not recommended to share high score notebook near ends.\n",
      "votes": 1
    },
    {
      "id": 2735318,
      "postDate": "2024-04-04T16:56:42.060Z",
      "content": "<p>That is true I see in public ensembles even weaker models so far mixed the right way or probing LB are producing nice scores but that could be due to the distribution on public test data . so unless the model generalizes to cv or most evaluators it could shake up in private . Better to play both ways to be safe at least .</p>",
      "rawMarkdown": "That is true I see in public ensembles even weaker models so far mixed the right way or probing LB are producing nice scores but that could be due to the distribution on public test data . so unless the model generalizes to cv or most evaluators it could shake up in private . Better to play both ways to be safe at least .",
      "votes": 1,
      "replies": [
        {
          "id": 2735348,
          "postDate": "2024-04-04T17:21:35.317Z",
          "content": "<p>I totally agree. I normally like to stick to cv. But so many people have worked on high votes only that if we don’t have a sub with it and it is correct we will fall behind on the private LB. I am hoping it is closer to the training CV rather than high votes only</p>",
          "rawMarkdown": "I totally agree. I normally like to stick to cv. But so many people have worked on high votes only that if we don’t have a sub with it and it is correct we will fall behind on the private LB. I am hoping it is closer to the training CV rather than high votes only",
          "votes": 1,
          "replies": [
            {
              "id": 2735928,
              "postDate": "2024-04-05T00:46:31.857Z",
              "content": "<p>There was a competition in which the solution ignoring CV with all data and fitting to high quality label won, so in some situation trusting CV with all training data does not win.<br>\n<a href=\"https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/leaderboard\" target=\"_blank\">Prostate cANcer graDe Assessment (PANDA) Challenge</a><br>\n6th solution</p>\n<pre><code>Validation strategy\nAs written  task description, the label quality  train data differs a lot  those  test data. So  the very beginning we assumed this part would be critical  this comp. Roughly speaking,\ntrain data: noisy  big\npublic test data: clean but small\nprivate test data: clean but small\nSo our strategy  IGNORE CV, CARE ABOUT PUBLIC LB, AND TRUST METHODOLOGY.\nFor us, the results were unstable due to small size of test data, but  so ‘lottery’.\n</code></pre>",
              "rawMarkdown": "There was a competition in which the solution ignoring CV with all data and fitting to high quality label won, so in some situation trusting CV with all training data does not win.\n[Prostate cANcer graDe Assessment (PANDA) Challenge](https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/leaderboard)\n6th solution\n```python\nValidation strategy\nAs written in task description, the label quality in train data differs a lot from those in test data. So from the very beginning we assumed this part would be critical in this comp. Roughly speaking,\ntrain data: noisy and big\npublic test data: clean but small\nprivate test data: clean but small\nSo our strategy is IGNORE CV, CARE ABOUT PUBLIC LB, AND TRUST METHODOLOGY.\nFor us, the results were unstable due to small size of test data, but not so ‘lottery’.\n```",
              "votes": 3
            },
            {
              "id": 2735948,
              "postDate": "2024-04-05T01:16:42.520Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2723109,
      "postDate": "2024-03-30T00:58:31.757Z",
      "content": "<p>If we have OOF pred and know how to train model, It may be possible to optimize ensemble weights, but changing weights of public models with trial and errors to public LB score may cause overfitting.</p>",
      "rawMarkdown": "If we have OOF pred and know how to train model, It may be possible to optimize ensemble weights, but changing weights of public models with trial and errors to public LB score may cause overfitting.",
      "votes": 1
    },
    {
      "id": 2733294,
      "postDate": "2024-04-03T15:14:50.480Z",
      "content": "<p>Thank you for the reminder, <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a>. It's important to prioritize quality over simply chasing high scores</p>",
      "rawMarkdown": "Thank you for the reminder, @cody11null. It's important to prioritize quality over simply chasing high scores"
    },
    {
      "id": 2724676,
      "postDate": "2024-03-31T03:28:55.993Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2744504,
      "author_name": "Aurora_blue",
      "author_url": "",
      "post_date": "2024-04-10T00:17:54.750000",
      "content": "<p>I checked the public code and found that ensemble of public models get 0.35X in private LB, bronze zone. <br>\nIndeed that each single model did not get top of public LB, but ensembles of it destroyed LB.<br>\nFinally, as discussed again, it would be not recommended to share high score notebook near ends.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2735318,
      "author_name": "Gaurav Rawat",
      "author_url": "",
      "post_date": "2024-04-04T16:56:42.060000",
      "content": "<p>That is true I see in public ensembles even weaker models so far mixed the right way or probing LB are producing nice scores but that could be due to the distribution on public test data . so unless the model generalizes to cv or most evaluators it could shake up in private . Better to play both ways to be safe at least .</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2735348,
          "author_name": "Cody_Null",
          "author_url": "",
          "post_date": "2024-04-04T17:21:35.317000",
          "content": "<p>I totally agree. I normally like to stick to cv. But so many people have worked on high votes only that if we don’t have a sub with it and it is correct we will fall behind on the private LB. I am hoping it is closer to the training CV rather than high votes only</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2735928,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-04-05T00:46:31.857000",
              "content": "<p>There was a competition in which the solution ignoring CV with all data and fitting to high quality label won, so in some situation trusting CV with all training data does not win.<br>\n<a href=\"https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/leaderboard\" target=\"_blank\">Prostate cANcer graDe Assessment (PANDA) Challenge</a><br>\n6th solution</p>\n<pre><code>Validation strategy\nAs written  task description, the label quality  train data differs a lot  those  test data. So  the very beginning we assumed this part would be critical  this comp. Roughly speaking,\ntrain data: noisy  big\npublic test data: clean but small\nprivate test data: clean but small\nSo our strategy  IGNORE CV, CARE ABOUT PUBLIC LB, AND TRUST METHODOLOGY.\nFor us, the results were unstable due to small size of test data, but  so ‘lottery’.\n</code></pre>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2735948,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-05T01:16:42.520000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2723109,
      "author_name": "Aurora_blue",
      "author_url": "",
      "post_date": "2024-03-30T00:58:31.757000",
      "content": "<p>If we have OOF pred and know how to train model, It may be possible to optimize ensemble weights, but changing weights of public models with trial and errors to public LB score may cause overfitting.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2733294,
      "author_name": "Sahir Maharaj",
      "author_url": "",
      "post_date": "2024-04-03T15:14:50.480000",
      "content": "<p>Thank you for the reminder, <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a>. It's important to prioritize quality over simply chasing high scores</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2724676,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-31T03:28:55.993000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2722712": "Hi all, I do see a new high-scoring public model has been released. Though I personally believe it is inappropriate to be posting high-scoring public notebooks this close to the competition end, regardless of its probability of success in the private LB, I wanted to bring additional attention to this. Just because this notebook may have a higher LB than an existing notebook of yours does not necessarily mean it will be a better submission. Please do your due diligence and be extra careful when selecting submissions. Think about what could affect the final scores and what is required to make each model succeed. Be careful and good luck out there!",
    "2744504": "I checked the public code and found that ensemble of public models get 0.35X in private LB, bronze zone. \nIndeed that each single model did not get top of public LB, but ensembles of it destroyed LB.\nFinally, as discussed again, it would be not recommended to share high score notebook near ends.\n",
    "2735318": "That is true I see in public ensembles even weaker models so far mixed the right way or probing LB are producing nice scores but that could be due to the distribution on public test data . so unless the model generalizes to cv or most evaluators it could shake up in private . Better to play both ways to be safe at least .",
    "2723109": "If we have OOF pred and know how to train model, It may be possible to optimize ensemble weights, but changing weights of public models with trial and errors to public LB score may cause overfitting.",
    "2733294": "Thank you for the reminder, @cody11null. It's important to prioritize quality over simply chasing high scores",
    "2724676": ""
  }
}