{
  "id": 389554,
  "title": "Is best threshold for test different from that for train? ",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/389554",
  "author_name": "",
  "post_date": "2023-02-22T06:10:36.621858300Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Is it possible to increase LB by using a different threshold during test than during train?</p>",
  "messages": [
    {
      "id": "2154606",
      "postDate": "02/22/2023 06:10:36",
      "content": "<p>Is it possible to increase LB by using a different threshold during test than during train?</p>",
      "rawMarkdown": "Is it possible to increase LB by using a different threshold during test than during train?",
      "votes": null
    },
    {
      "id": "2154621",
      "postDate": "02/22/2023 06:33:20",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yurimaeda\" target=\"_blank\">@yurimaeda</a>,</p>\n<p>It's possible to find a better threshold for public LB. However, I think that's not a good idea to tune threshold based on test set, which can lead to <strong>overfitting on the public LB</strong> now. Then, the over-tuned model will have worse performance on the <strong>private LB</strong> (<em>i.e.,</em> unseen test set), whose distribution remains unknown.</p>\n<p>Hope this helps, thanks!</p>",
      "rawMarkdown": "Hi @yurimaeda,\n\nIt's possible to find a better threshold for public LB. However, I think that's not a good idea to tune threshold based on test set, which can lead to **overfitting on the public LB** now. Then, the over-tuned model will have worse performance on the **private LB** (*i.e.,* unseen test set), whose distribution remains unknown.\n\nHope this helps, thanks!",
      "votes": null
    },
    {
      "id": "2156040",
      "postDate": "02/23/2023 04:16:28",
      "content": "<p>Thank you very much. Since CV went up but LB went down, I wondered if it might be because the optimal value of the threshold for LB is different from that for CV.</p>",
      "rawMarkdown": "Thank you very much. Since CV went up but LB went down, I wondered if it might be because the optimal value of the threshold for LB is different from that for CV.",
      "votes": null
    },
    {
      "id": "2157292",
      "postDate": "02/23/2023 21:55:25",
      "content": "<p>I think it is better to go with the best cv threshold. This may save you from shakes in the private lb.</p>",
      "rawMarkdown": "I think it is better to go with the best cv threshold. This may save you from shakes in the private lb.",
      "votes": null
    },
    {
      "id": "2160409",
      "postDate": "02/26/2023 16:44:50",
      "content": "<p>CV up vs LB down can happen for many many reasons. </p>\n<p>If it's a third digit issue (example: CV up +0.003 vs LB down -0.002), consider whether this could happen if you just changed random seed. If it's within normal variance, then there are so so many possible reasons it could happen. Asd most likely is that you made a decision based on CV improvement (aka hyperparameter tuning), but the apparent improvement was partially or entirely random noise and not a model improvement. And even the model improvement, if any, could be smaller than the overall random seed variance, so LB could randomly go down. </p>\n<p>Bigger CV vs LB deltas have fewer plausible explanations, while simultaneously being much more concerning. But threshold tuning is also very unlikely to cause a big CV vs LB deltas by itself, I think. I would expect only third digit impact.</p>",
      "rawMarkdown": "CV up vs LB down can happen for many many reasons. \n\nIf it's a third digit issue (example: CV up +0.003 vs LB down -0.002), consider whether this could happen if you just changed random seed. If it's within normal variance, then there are so so many possible reasons it could happen. Asd most likely is that you made a decision based on CV improvement (aka hyperparameter tuning), but the apparent improvement was partially or entirely random noise and not a model improvement. And even the model improvement, if any, could be smaller than the overall random seed variance, so LB could randomly go down. \n\nBigger CV vs LB deltas have fewer plausible explanations, while simultaneously being much more concerning. But threshold tuning is also very unlikely to cause a big CV vs LB deltas by itself, I think. I would expect only third digit impact.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2154621,
      "author_name": "abaojiang",
      "author_url": "",
      "post_date": "02/22/2023 06:33:20",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yurimaeda\" target=\"_blank\">@yurimaeda</a>,</p>\n<p>It's possible to find a better threshold for public LB. However, I think that's not a good idea to tune threshold based on test set, which can lead to <strong>overfitting on the public LB</strong> now. Then, the over-tuned model will have worse performance on the <strong>private LB</strong> (<em>i.e.,</em> unseen test set), whose distribution remains unknown.</p>\n<p>Hope this helps, thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2156040,
          "author_name": "yurimaeda",
          "author_url": "",
          "post_date": "02/23/2023 04:16:28",
          "content": "<p>Thank you very much. Since CV went up but LB went down, I wondered if it might be because the optimal value of the threshold for LB is different from that for CV.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2157292,
              "author_name": "mohammad2012191",
              "author_url": "",
              "post_date": "02/23/2023 21:55:25",
              "content": "<p>I think it is better to go with the best cv threshold. This may save you from shakes in the private lb.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2160409,
              "author_name": "roberthatch",
              "author_url": "",
              "post_date": "02/26/2023 16:44:50",
              "content": "<p>CV up vs LB down can happen for many many reasons. </p>\n<p>If it's a third digit issue (example: CV up +0.003 vs LB down -0.002), consider whether this could happen if you just changed random seed. If it's within normal variance, then there are so so many possible reasons it could happen. Asd most likely is that you made a decision based on CV improvement (aka hyperparameter tuning), but the apparent improvement was partially or entirely random noise and not a model improvement. And even the model improvement, if any, could be smaller than the overall random seed variance, so LB could randomly go down. </p>\n<p>Bigger CV vs LB deltas have fewer plausible explanations, while simultaneously being much more concerning. But threshold tuning is also very unlikely to cause a big CV vs LB deltas by itself, I think. I would expect only third digit impact.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2154606": "Is it possible to increase LB by using a different threshold during test than during train?",
    "2154621": "Hi @yurimaeda,\n\nIt's possible to find a better threshold for public LB. However, I think that's not a good idea to tune threshold based on test set, which can lead to **overfitting on the public LB** now. Then, the over-tuned model will have worse performance on the **private LB** (*i.e.,* unseen test set), whose distribution remains unknown.\n\nHope this helps, thanks!",
    "2156040": "Thank you very much. Since CV went up but LB went down, I wondered if it might be because the optimal value of the threshold for LB is different from that for CV.",
    "2157292": "I think it is better to go with the best cv threshold. This may save you from shakes in the private lb.",
    "2160409": "CV up vs LB down can happen for many many reasons. \n\nIf it's a third digit issue (example: CV up +0.003 vs LB down -0.002), consider whether this could happen if you just changed random seed. If it's within normal variance, then there are so so many possible reasons it could happen. Asd most likely is that you made a decision based on CV improvement (aka hyperparameter tuning), but the apparent improvement was partially or entirely random noise and not a model improvement. And even the model improvement, if any, could be smaller than the overall random seed variance, so LB could randomly go down. \n\nBigger CV vs LB deltas have fewer plausible explanations, while simultaneously being much more concerning. But threshold tuning is also very unlikely to cause a big CV vs LB deltas by itself, I think. I would expect only third digit impact."
  },
  "source": "meta"
}