{
  "id": 389217,
  "title": "Different Thresholds for Different Question Models",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/389217",
  "author_name": "",
  "post_date": "2023-02-21T05:12:33.705993600Z",
  "votes": 13,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Hii,<br>\nI tried using different thresholds for the different question models. This boosts the CV Score for each question significantly but reduces the LB score.<br>\nAm I doing something wrong?</p>",
  "messages": [
    {
      "id": "2152924",
      "postDate": "02/21/2023 05:12:33",
      "content": "<p>Hii,<br>\nI tried using different thresholds for the different question models. This boosts the CV Score for each question significantly but reduces the LB score.<br>\nAm I doing something wrong?</p>",
      "rawMarkdown": "Hii,\nI tried using different thresholds for the different question models. This boosts the CV Score for each question significantly but reduces the LB score.\nAm I doing something wrong?",
      "votes": null
    },
    {
      "id": "2152931",
      "postDate": "02/21/2023 05:21:13",
      "content": "<p>I also tried that and found different Thresholds boost each question in f1 but reduce the overall f1.I think the reason may fall on the imbalance of some questions like q2 and q3.Best local thresholds on these questions (approximately 0.9) make more negative predictions to boost negative recall but harm the overall precision and recall. <br>\nI think different thresholds works but we need to find out how to optimize the overall f1 </p>",
      "rawMarkdown": "I also tried that and found different Thresholds boost each question in f1 but reduce the overall f1.I think the reason may fall on the imbalance of some questions like q2 and q3.Best local thresholds on these questions (approximately 0.9) make more negative predictions to boost negative recall but harm the overall precision and recall. \nI think different thresholds works but we need to find out how to optimize the overall f1",
      "votes": null
    },
    {
      "id": "2152932",
      "postDate": "02/21/2023 05:24:36",
      "content": "<p>Yes. I think then maybe it's just better to get the best threshold for the overall f1.</p>",
      "rawMarkdown": "Yes. I think then maybe it's just better to get the best threshold for the overall f1.",
      "votes": null
    },
    {
      "id": "2152942",
      "postDate": "02/21/2023 05:33:05",
      "content": "<p>I tried optimizing the f1 score for each question individually, but it turned out to harm the overall performnce.<br>\nYou can check this discussion:<br>\n<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384677\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384677</a></p>",
      "rawMarkdown": "I tried optimizing the f1 score for each question individually, but it turned out to harm the overall performnce.\nYou can check this discussion:\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384677",
      "votes": null
    },
    {
      "id": "2152945",
      "postDate": "02/21/2023 05:33:27",
      "content": "<p>Yes, you r right, coz greedy optimization per question doesn't give global optima. Remember we r scoring our submission as global predictions, not per question.</p>",
      "rawMarkdown": "Yes, you r right, coz greedy optimization per question doesn't give global optima. Remember we r scoring our submission as global predictions, not per question.",
      "votes": null
    },
    {
      "id": "2152954",
      "postDate": "02/21/2023 05:39:34",
      "content": "<p>Yeah, I get it now. Thanks.</p>",
      "rawMarkdown": "Yeah, I get it now. Thanks.",
      "votes": null
    },
    {
      "id": "2152956",
      "postDate": "02/21/2023 05:39:50",
      "content": "<p>Thank you.</p>",
      "rawMarkdown": "Thank you.",
      "votes": null
    },
    {
      "id": "2153386",
      "postDate": "02/21/2023 11:56:44",
      "content": "<p>I tried to look for the overall optimal f1, instead of optimizing each question seperately. But it is difficult to achieve this by something like gridsearch, since there are 18 parameters. </p>\n<p>I just optimized one threshold for a question and then fixed it to optimize another one by one, trying to get local optimal f1. It works much better than optimizing the f1 score for each question individually, but still not as well as setting one threshold(~0.63).</p>",
      "rawMarkdown": "I tried to look for the overall optimal f1, instead of optimizing each question seperately. But it is difficult to achieve this by something like gridsearch, since there are 18 parameters. \n\nI just optimized one threshold for a question and then fixed it to optimize another one by one, trying to get local optimal f1. It works much better than optimizing the f1 score for each question individually, but still not as well as setting one threshold(~0.63).",
      "votes": null
    },
    {
      "id": "2154177",
      "postDate": "02/21/2023 21:10:18",
      "content": "<p>I tried this as well and had the same issue. You can use optuna to try optimising for global F1 like you would a hyperparameter search by using the thresholds for each question as a hyperparameter which gives a bit of a boost, especially if you narrow down the range you consider</p>\n<pre><code>def objective(trial):\n    # Suggest thresholds for each question\n    params = {\n        k: trial.suggest_float(k, 0.5, 0.7, step=0.01) for k in range(18)\n    }\n\n    # Calculate global F1 score\n    return f1_score(true.values.reshape((-1)), np.array([oof[k].values &gt; params[k] for k in range(18)]).T.reshape((-1)).astype('int'), average='macro')\n\nstudy = optuna.create_study(direction='maximize', study_name='Best threshold for each question')\nstudy.optimize(objective, n_trials=500)\nprint('Number of finished trials:', len(study.trials))\nprint('Best trial:', study.best_trial.params)\nbest_thresholds = study.best_trial.params\nprint('Best F1 score:', study.best_trial.value)\n</code></pre>",
      "rawMarkdown": "I tried this as well and had the same issue. You can use optuna to try optimising for global F1 like you would a hyperparameter search by using the thresholds for each question as a hyperparameter which gives a bit of a boost, especially if you narrow down the range you consider\n\n```\ndef objective(trial):\n    # Suggest thresholds for each question\n    params = {\n        k: trial.suggest_float(k, 0.5, 0.7, step=0.01) for k in range(18)\n    }\n    \n    # Calculate global F1 score\n    return f1_score(true.values.reshape((-1)), np.array([oof[k].values > params[k] for k in range(18)]).T.reshape((-1)).astype('int'), average='macro')\n\nstudy = optuna.create_study(direction='maximize', study_name='Best threshold for each question')\nstudy.optimize(objective, n_trials=500)\nprint('Number of finished trials:', len(study.trials))\nprint('Best trial:', study.best_trial.params)\nbest_thresholds = study.best_trial.params\nprint('Best F1 score:', study.best_trial.value)\n```",
      "votes": null
    },
    {
      "id": "2154622",
      "postDate": "02/22/2023 06:33:35",
      "content": "<p>I made a notebook for this very particular thing. You may have a look here: <a href=\"https://www.kaggle.com/code/mayukh18/optimize-for-multiple-thresholds\" target=\"_blank\">Optimize for Multiple Thresholds</a></p>",
      "rawMarkdown": "I made a notebook for this very particular thing. You may have a look here: [Optimize for Multiple Thresholds](https://www.kaggle.com/code/mayukh18/optimize-for-multiple-thresholds)",
      "votes": null
    },
    {
      "id": "2155030",
      "postDate": "02/22/2023 11:54:06",
      "content": "<p>Yes, I guess 0.63 is the magical number for this competiton ;) LoL</p>",
      "rawMarkdown": "Yes, I guess 0.63 is the magical number for this competiton ;) LoL",
      "votes": null
    },
    {
      "id": "2155043",
      "postDate": "02/22/2023 12:03:19",
      "content": "<p>I'll try this. Thanks :) </p>",
      "rawMarkdown": "I'll try this. Thanks :)",
      "votes": null
    },
    {
      "id": "2155278",
      "postDate": "02/22/2023 14:36:34",
      "content": "<p>It's a great job. Thanks a lot.</p>\n<p>Although I have tried 10000 trials, I cannot find a set of thresholds better than a uniform threshold.</p>\n<p>I think it's quite strange. I'm wondering whether there's any internal law over the choice of thresholds, something like a uniform threshold must perform better than different threshold under certain circumstance…</p>",
      "rawMarkdown": "It's a great job. Thanks a lot.\n\nAlthough I have tried 10000 trials, I cannot find a set of thresholds better than a uniform threshold.\n\nI think it's quite strange. I'm wondering whether there's any internal law over the choice of thresholds, something like a uniform threshold must perform better than different threshold under certain circumstance...",
      "votes": null
    },
    {
      "id": "2155287",
      "postDate": "02/22/2023 14:41:58",
      "content": "<p>Updated. After I narrawed down the search range, I found some sets of thresholds boosting ~0.001 CV.</p>",
      "rawMarkdown": "Updated. After I narrawed down the search range, I found some sets of thresholds boosting ~0.001 CV.",
      "votes": null
    },
    {
      "id": "2155401",
      "postDate": "02/22/2023 16:07:40",
      "content": "<p>I should have pointed that out! It's a good idea to run it once as a rough attempt to find an approximate range of thresholds to search through, and then reducing that range (and possibly the step as well) to zone in on the best thresholds. It's great to hear you got a boost!</p>",
      "rawMarkdown": "I should have pointed that out! It's a good idea to run it once as a rough attempt to find an approximate range of thresholds to search through, and then reducing that range (and possibly the step as well) to zone in on the best thresholds. It's great to hear you got a boost!",
      "votes": null
    },
    {
      "id": "2156240",
      "postDate": "02/23/2023 07:27:42",
      "content": "<p>It's amazing. Thanks.</p>",
      "rawMarkdown": "It's amazing. Thanks.",
      "votes": null
    },
    {
      "id": "2254222",
      "postDate": "05/10/2023 18:23:53",
      "content": "<p>Different threshold give different LB and CV spreads , we tried with optuna but seems CV,LB correlation is not that reliable and effected by threshold a lot . quite confusing not sure its due to train and test data difference</p>",
      "rawMarkdown": "Different threshold give different LB and CV spreads , we tried with optuna but seems CV,LB correlation is not that reliable and effected by threshold a lot . quite confusing not sure its due to train and test data difference",
      "votes": null
    },
    {
      "id": "2254489",
      "postDate": "05/11/2023 03:38:32",
      "content": "<p>Optimizing on validation set results in better CV worse LB.</p>\n<p>I also tried optimizing different threshold on training set base on macro F1 score, which only increase macro F1 for training set and not CV. I guess trying to optimize thresholds only result in overfitting.</p>",
      "rawMarkdown": "Optimizing on validation set results in better CV worse LB.\n\nI also tried optimizing different threshold on training set base on macro F1 score, which only increase macro F1 for training set and not CV. I guess trying to optimize thresholds only result in overfitting.",
      "votes": null
    },
    {
      "id": "2254916",
      "postDate": "05/11/2023 10:43:50",
      "content": "<p>Seeing same not sure tuning threshold results in lb overfitting. Trusting cv hopefully should generalize. The difference seems to be for us in range of 0.001-0.003</p>",
      "rawMarkdown": "Seeing same not sure tuning threshold results in lb overfitting. Trusting cv hopefully should generalize. The difference seems to be for us in range of 0.001-0.003",
      "votes": null
    },
    {
      "id": "2254925",
      "postDate": "05/11/2023 10:49:10",
      "content": "<p>Maybe you can try to have a held-out test set, and then optimizing thresholds base on CV on rest of the data, and see whether F1 for the held-out test set increase. For me trying to optimize the threshold on CV does not translate to the held-out set, therefore I was confident that it was overfitting. </p>",
      "rawMarkdown": "Maybe you can try to have a held-out test set, and then optimizing thresholds base on CV on rest of the data, and see whether F1 for the held-out test set increase. For me trying to optimize the threshold on CV does not translate to the held-out set, therefore I was confident that it was overfitting.",
      "votes": null
    },
    {
      "id": "2286647",
      "postDate": "06/03/2023 15:56:52",
      "content": "<p>Check out my genetic algorithm for threshold optimization:<br>\n<a href=\"https://www.kaggle.com/code/jakubzeman/genetic-alogrithm-for-threshold-optimization\" target=\"_blank\">https://www.kaggle.com/code/jakubzeman/genetic-alogrithm-for-threshold-optimization</a></p>",
      "rawMarkdown": "Check out my genetic algorithm for threshold optimization:\nhttps://www.kaggle.com/code/jakubzeman/genetic-alogrithm-for-threshold-optimization",
      "votes": null
    },
    {
      "id": "2316381",
      "postDate": "06/24/2023 20:56:59",
      "content": "<p>If still interested, here is some math explaining this:</p>\n<p><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417478\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417478</a></p>",
      "rawMarkdown": "If still interested, here is some math explaining this:\n\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417478",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2152931,
      "author_name": "kobeinmemory",
      "author_url": "",
      "post_date": "02/21/2023 05:21:13",
      "content": "<p>I also tried that and found different Thresholds boost each question in f1 but reduce the overall f1.I think the reason may fall on the imbalance of some questions like q2 and q3.Best local thresholds on these questions (approximately 0.9) make more negative predictions to boost negative recall but harm the overall precision and recall. <br>\nI think different thresholds works but we need to find out how to optimize the overall f1 </p>",
      "votes": null,
      "replies": [
        {
          "id": 2152932,
          "author_name": "shashwatraman",
          "author_url": "",
          "post_date": "02/21/2023 05:24:36",
          "content": "<p>Yes. I think then maybe it's just better to get the best threshold for the overall f1.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2152945,
              "author_name": "sggpls",
              "author_url": "",
              "post_date": "02/21/2023 05:33:27",
              "content": "<p>Yes, you r right, coz greedy optimization per question doesn't give global optima. Remember we r scoring our submission as global predictions, not per question.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2152954,
                  "author_name": "shashwatraman",
                  "author_url": "",
                  "post_date": "02/21/2023 05:39:34",
                  "content": "<p>Yeah, I get it now. Thanks.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 2152942,
          "author_name": "mohammad2012191",
          "author_url": "",
          "post_date": "02/21/2023 05:33:05",
          "content": "<p>I tried optimizing the f1 score for each question individually, but it turned out to harm the overall performnce.<br>\nYou can check this discussion:<br>\n<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384677\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384677</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2152956,
              "author_name": "shashwatraman",
              "author_url": "",
              "post_date": "02/21/2023 05:39:50",
              "content": "<p>Thank you.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2153386,
      "author_name": "yunqicao",
      "author_url": "",
      "post_date": "02/21/2023 11:56:44",
      "content": "<p>I tried to look for the overall optimal f1, instead of optimizing each question seperately. But it is difficult to achieve this by something like gridsearch, since there are 18 parameters. </p>\n<p>I just optimized one threshold for a question and then fixed it to optimize another one by one, trying to get local optimal f1. It works much better than optimizing the f1 score for each question individually, but still not as well as setting one threshold(~0.63).</p>",
      "votes": null,
      "replies": [
        {
          "id": 2155030,
          "author_name": "shashwatraman",
          "author_url": "",
          "post_date": "02/22/2023 11:54:06",
          "content": "<p>Yes, I guess 0.63 is the magical number for this competiton ;) LoL</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2154177,
      "author_name": "judehunt23",
      "author_url": "",
      "post_date": "02/21/2023 21:10:18",
      "content": "<p>I tried this as well and had the same issue. You can use optuna to try optimising for global F1 like you would a hyperparameter search by using the thresholds for each question as a hyperparameter which gives a bit of a boost, especially if you narrow down the range you consider</p>\n<pre><code>def objective(trial):\n    # Suggest thresholds for each question\n    params = {\n        k: trial.suggest_float(k, 0.5, 0.7, step=0.01) for k in range(18)\n    }\n\n    # Calculate global F1 score\n    return f1_score(true.values.reshape((-1)), np.array([oof[k].values &gt; params[k] for k in range(18)]).T.reshape((-1)).astype('int'), average='macro')\n\nstudy = optuna.create_study(direction='maximize', study_name='Best threshold for each question')\nstudy.optimize(objective, n_trials=500)\nprint('Number of finished trials:', len(study.trials))\nprint('Best trial:', study.best_trial.params)\nbest_thresholds = study.best_trial.params\nprint('Best F1 score:', study.best_trial.value)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2155043,
          "author_name": "shashwatraman",
          "author_url": "",
          "post_date": "02/22/2023 12:03:19",
          "content": "<p>I'll try this. Thanks :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2155278,
          "author_name": "yunqicao",
          "author_url": "",
          "post_date": "02/22/2023 14:36:34",
          "content": "<p>It's a great job. Thanks a lot.</p>\n<p>Although I have tried 10000 trials, I cannot find a set of thresholds better than a uniform threshold.</p>\n<p>I think it's quite strange. I'm wondering whether there's any internal law over the choice of thresholds, something like a uniform threshold must perform better than different threshold under certain circumstance…</p>",
          "votes": null,
          "replies": [
            {
              "id": 2155287,
              "author_name": "yunqicao",
              "author_url": "",
              "post_date": "02/22/2023 14:41:58",
              "content": "<p>Updated. After I narrawed down the search range, I found some sets of thresholds boosting ~0.001 CV.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2155401,
                  "author_name": "judehunt23",
                  "author_url": "",
                  "post_date": "02/22/2023 16:07:40",
                  "content": "<p>I should have pointed that out! It's a good idea to run it once as a rough attempt to find an approximate range of thresholds to search through, and then reducing that range (and possibly the step as well) to zone in on the best thresholds. It's great to hear you got a boost!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2154622,
      "author_name": "mayukh18",
      "author_url": "",
      "post_date": "02/22/2023 06:33:35",
      "content": "<p>I made a notebook for this very particular thing. You may have a look here: <a href=\"https://www.kaggle.com/code/mayukh18/optimize-for-multiple-thresholds\" target=\"_blank\">Optimize for Multiple Thresholds</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2156240,
          "author_name": "shashwatraman",
          "author_url": "",
          "post_date": "02/23/2023 07:27:42",
          "content": "<p>It's amazing. Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2254222,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "05/10/2023 18:23:53",
      "content": "<p>Different threshold give different LB and CV spreads , we tried with optuna but seems CV,LB correlation is not that reliable and effected by threshold a lot . quite confusing not sure its due to train and test data difference</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2254489,
      "author_name": "woprime",
      "author_url": "",
      "post_date": "05/11/2023 03:38:32",
      "content": "<p>Optimizing on validation set results in better CV worse LB.</p>\n<p>I also tried optimizing different threshold on training set base on macro F1 score, which only increase macro F1 for training set and not CV. I guess trying to optimize thresholds only result in overfitting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2254916,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "05/11/2023 10:43:50",
          "content": "<p>Seeing same not sure tuning threshold results in lb overfitting. Trusting cv hopefully should generalize. The difference seems to be for us in range of 0.001-0.003</p>",
          "votes": null,
          "replies": [
            {
              "id": 2254925,
              "author_name": "woprime",
              "author_url": "",
              "post_date": "05/11/2023 10:49:10",
              "content": "<p>Maybe you can try to have a held-out test set, and then optimizing thresholds base on CV on rest of the data, and see whether F1 for the held-out test set increase. For me trying to optimize the threshold on CV does not translate to the held-out set, therefore I was confident that it was overfitting. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2286647,
      "author_name": "jakubzeman",
      "author_url": "",
      "post_date": "06/03/2023 15:56:52",
      "content": "<p>Check out my genetic algorithm for threshold optimization:<br>\n<a href=\"https://www.kaggle.com/code/jakubzeman/genetic-alogrithm-for-threshold-optimization\" target=\"_blank\">https://www.kaggle.com/code/jakubzeman/genetic-alogrithm-for-threshold-optimization</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2316381,
      "author_name": "kononenko",
      "author_url": "",
      "post_date": "06/24/2023 20:56:59",
      "content": "<p>If still interested, here is some math explaining this:</p>\n<p><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417478\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417478</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2152924": "Hii,\nI tried using different thresholds for the different question models. This boosts the CV Score for each question significantly but reduces the LB score.\nAm I doing something wrong?",
    "2152931": "I also tried that and found different Thresholds boost each question in f1 but reduce the overall f1.I think the reason may fall on the imbalance of some questions like q2 and q3.Best local thresholds on these questions (approximately 0.9) make more negative predictions to boost negative recall but harm the overall precision and recall. \nI think different thresholds works but we need to find out how to optimize the overall f1",
    "2152932": "Yes. I think then maybe it's just better to get the best threshold for the overall f1.",
    "2152942": "I tried optimizing the f1 score for each question individually, but it turned out to harm the overall performnce.\nYou can check this discussion:\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384677",
    "2152945": "Yes, you r right, coz greedy optimization per question doesn't give global optima. Remember we r scoring our submission as global predictions, not per question.",
    "2152954": "Yeah, I get it now. Thanks.",
    "2152956": "Thank you.",
    "2153386": "I tried to look for the overall optimal f1, instead of optimizing each question seperately. But it is difficult to achieve this by something like gridsearch, since there are 18 parameters. \n\nI just optimized one threshold for a question and then fixed it to optimize another one by one, trying to get local optimal f1. It works much better than optimizing the f1 score for each question individually, but still not as well as setting one threshold(~0.63).",
    "2154177": "I tried this as well and had the same issue. You can use optuna to try optimising for global F1 like you would a hyperparameter search by using the thresholds for each question as a hyperparameter which gives a bit of a boost, especially if you narrow down the range you consider\n\n```\ndef objective(trial):\n    # Suggest thresholds for each question\n    params = {\n        k: trial.suggest_float(k, 0.5, 0.7, step=0.01) for k in range(18)\n    }\n    \n    # Calculate global F1 score\n    return f1_score(true.values.reshape((-1)), np.array([oof[k].values > params[k] for k in range(18)]).T.reshape((-1)).astype('int'), average='macro')\n\nstudy = optuna.create_study(direction='maximize', study_name='Best threshold for each question')\nstudy.optimize(objective, n_trials=500)\nprint('Number of finished trials:', len(study.trials))\nprint('Best trial:', study.best_trial.params)\nbest_thresholds = study.best_trial.params\nprint('Best F1 score:', study.best_trial.value)\n```",
    "2154622": "I made a notebook for this very particular thing. You may have a look here: [Optimize for Multiple Thresholds](https://www.kaggle.com/code/mayukh18/optimize-for-multiple-thresholds)",
    "2155030": "Yes, I guess 0.63 is the magical number for this competiton ;) LoL",
    "2155043": "I'll try this. Thanks :)",
    "2155278": "It's a great job. Thanks a lot.\n\nAlthough I have tried 10000 trials, I cannot find a set of thresholds better than a uniform threshold.\n\nI think it's quite strange. I'm wondering whether there's any internal law over the choice of thresholds, something like a uniform threshold must perform better than different threshold under certain circumstance...",
    "2155287": "Updated. After I narrawed down the search range, I found some sets of thresholds boosting ~0.001 CV.",
    "2155401": "I should have pointed that out! It's a good idea to run it once as a rough attempt to find an approximate range of thresholds to search through, and then reducing that range (and possibly the step as well) to zone in on the best thresholds. It's great to hear you got a boost!",
    "2156240": "It's amazing. Thanks.",
    "2254222": "Different threshold give different LB and CV spreads , we tried with optuna but seems CV,LB correlation is not that reliable and effected by threshold a lot . quite confusing not sure its due to train and test data difference",
    "2254489": "Optimizing on validation set results in better CV worse LB.\n\nI also tried optimizing different threshold on training set base on macro F1 score, which only increase macro F1 for training set and not CV. I guess trying to optimize thresholds only result in overfitting.",
    "2254916": "Seeing same not sure tuning threshold results in lb overfitting. Trusting cv hopefully should generalize. The difference seems to be for us in range of 0.001-0.003",
    "2254925": "Maybe you can try to have a held-out test set, and then optimizing thresholds base on CV on rest of the data, and see whether F1 for the held-out test set increase. For me trying to optimize the threshold on CV does not translate to the held-out set, therefore I was confident that it was overfitting.",
    "2286647": "Check out my genetic algorithm for threshold optimization:\nhttps://www.kaggle.com/code/jakubzeman/genetic-alogrithm-for-threshold-optimization",
    "2316381": "If still interested, here is some math explaining this:\n\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417478"
  },
  "source": "meta"
}