{
  "id": 502437,
  "title": "Calibration : Right way to do it?",
  "url": "/competitions/leash-BELKA/discussion/502437",
  "author_name": "",
  "post_date": "2024-05-13T12:53:06.249751700Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Last couple of days made a significant change in the leaderboard, where calibration resulted in higher jumps at the public LB. But is it the right choice? If the distribution of samples in private set is not similar to that of public set, are we all directed in the wrong way and stick to non-calibration approaches ? </p>",
  "messages": [
    {
      "id": "2810780",
      "postDate": "05/13/2024 12:53:06",
      "content": "<p>Last couple of days made a significant change in the leaderboard, where calibration resulted in higher jumps at the public LB. But is it the right choice? If the distribution of samples in private set is not similar to that of public set, are we all directed in the wrong way and stick to non-calibration approaches ? </p>",
      "rawMarkdown": "Last couple of days made a significant change in the leaderboard, where calibration resulted in higher jumps at the public LB. But is it the right choice? If the distribution of samples in private set is not similar to that of public set, are we all directed in the wrong way and stick to non-calibration approaches ?",
      "votes": null
    },
    {
      "id": "2810834",
      "postDate": "05/13/2024 13:18:43",
      "content": "<p>for the private test, you still can apply \"shifting trick\". but because share and nonshare distribution are different, \"where, what and how much\" to shift will be different from the public one</p>\n<p>if you make the wrong shift, you will get very bad results, worse than non-shifting.<br>\n(how much you can gain will be how much you can lose)</p>\n<hr>\n<p>please understand the statistics and reason why it work (and the risk why it doesn't). <br>\ndo not blindly adjust.</p>\n<hr>\n<p>finally you do not need to apply methods if test and train distributions are the same.<br>\nlearning is better than betting. <br>\nif you want to bet, the chance of winning should be bigger than random and you have a backup plan</p>\n<hr>\n<p>basically for this trick to work, you must have:</p>\n<ul>\n<li>identify one subset(s) that you know the truth pos/neg ratio. </li>\n<li>at the current subset location the measured predicted pos/neg  is very different from what you \"know\".</li>\n<li>so you shifted the subset to the correct location.</li>\n<li>after shifting, it won't worsen the original results of that new location <br>\n(that is the meaning of calibration)</li>\n</ul>\n<p>in the public test data, in particular, the nonshare  predictions are stuck at zero. we know this is wrong location because we know there are some nonshare pos samples (via probing). so just shift it a little towards the correct location will get good results.This will not affect the share prediction if they are already well predicted and concentrated the high probability location.</p>\n<p>you can treat this as a game of ranking. how to swap locations (ranks) to make pos samples move to the right side.</p>",
      "rawMarkdown": "for the private test, you still can apply \"shifting trick\". but because share and nonshare distribution are different, \"where, what and how much\" to shift will be different from the public one\n\nif you make the wrong shift, you will get very bad results, worse than non-shifting.\n(how much you can gain will be how much you can lose)\n\n---\n\nplease understand the statistics and reason why it work (and the risk why it doesn't). \ndo not blindly adjust.\n\n---\n\nfinally you do not need to apply methods if test and train distributions are the same.\nlearning is better than betting. \nif you want to bet, the chance of winning should be bigger than random and you have a backup plan\n\n\n---\n\nbasically for this trick to work, you must have:\n- identify one subset(s) that you know the truth pos/neg ratio. \n- at the current subset location the measured predicted pos/neg  is very different from what you \"know\".\n- so you shifted the subset to the correct location.\n- after shifting, it won't worsen the original results of that new location \n(that is the meaning of calibration)\n\nin the public test data, in particular, the nonshare  predictions are stuck at zero. we know this is wrong location because we know there are some nonshare pos samples (via probing). so just shift it a little towards the correct location will get good results.This will not affect the share prediction if they are already well predicted and concentrated the high probability location.\n\nyou can treat this as a game of ranking. how to swap locations (ranks) to make pos samples move to the right side.",
      "votes": null
    },
    {
      "id": "2810856",
      "postDate": "05/13/2024 13:33:32",
      "content": "<p>Calibration of models can indeed lead to significant changes in leaderboard standing, specially in competitions where the evaluation metric is sensitive to the confidence or probability estimates of predictions. However, whether calibration is the right choice depends on various factors, including the natur of the data, the evaluation metric, and the goals of the competition.<br>\nIf the distributon of samples in the private set differs significantly from that of the public set, calibration based solely on the public set may not generalize well to the private set. In such cases, relying solely on calibration approaches might indeed lead participants in the wrong direction. Strategies such as CrossValidation, Ensemble Methods, Post-Competition Analysis can be helpful. Calibrayution can lead to improvements in leaderboard standings, it's essential to consider the potential pitfalls, especially if the distribution of samples in the private set differs significantly from that of the public set. By understanding the data distribution, employing cross-validation techniques, using ensemble methods, and conducting post-competition analysis, participants can make more informed decisions and develop models that generalize well beyond the public leaderboard.</p>",
      "rawMarkdown": "Calibration of models can indeed lead to significant changes in leaderboard standing, specially in competitions where the evaluation metric is sensitive to the confidence or probability estimates of predictions. However, whether calibration is the right choice depends on various factors, including the natur of the data, the evaluation metric, and the goals of the competition.\nIf the distributon of samples in the private set differs significantly from that of the public set, calibration based solely on the public set may not generalize well to the private set. In such cases, relying solely on calibration approaches might indeed lead participants in the wrong direction. Strategies such as CrossValidation, Ensemble Methods, Post-Competition Analysis can be helpful. Calibrayution can lead to improvements in leaderboard standings, it's essential to consider the potential pitfalls, especially if the distribution of samples in the private set differs significantly from that of the public set. By understanding the data distribution, employing cross-validation techniques, using ensemble methods, and conducting post-competition analysis, participants can make more informed decisions and develop models that generalize well beyond the public leaderboard.",
      "votes": null
    },
    {
      "id": "2810885",
      "postDate": "05/13/2024 13:44:51",
      "content": "<p>Hey! Thank you for such a detailed explanation! Much appreciated. It was the post-processing technique you shared in one of the threads that helped me understand what was truly happening and about calibration. I have a question, is there a possibility for a non-calibration technique to outperform on the public lb or it doesn't stand a chance at all ? </p>",
      "rawMarkdown": "Hey! Thank you for such a detailed explanation! Much appreciated. It was the post-processing technique you shared in one of the threads that helped me understand what was truly happening and about calibration. I have a question, is there a possibility for a non-calibration technique to outperform on the public lb or it doesn't stand a chance at all ?",
      "votes": null
    },
    {
      "id": "2810902",
      "postDate": "05/13/2024 13:55:39",
      "content": "<p>\" non-calibration technique to outperform on the public l\" </p>\n<p>i would say you can actually include calibration in the learning process. (not trial and error or manual adjust).<br>\nthat should perform very well.</p>\n<p>becuase the share and nonshare molecule are very \"different\", currnetly i cannot find a method that predict the nonshare moelcule well. in that sense, i don't think non-calibrated methods work better than pure prediction.</p>\n<p>but we haven't explore all methods yet, e.g. 3d graph, docking and MD, foundation models, SSL …<br>\nif you can find some very good molecular property prediction, you maybe do not need calibration.<br>\nwhen i say \"very good \" it means smiliar (or less than 5% performance drop) validation/public AP for share or nonshare molecules</p>",
      "rawMarkdown": "\" non-calibration technique to outperform on the public l\" \n\ni would say you can actually include calibration in the learning process. (not trial and error or manual adjust).\nthat should perform very well.\n\nbecuase the share and nonshare molecule are very \"different\", currnetly i cannot find a method that predict the nonshare moelcule well. in that sense, i don't think non-calibrated methods work better than pure prediction.\n\nbut we haven't explore all methods yet, e.g. 3d graph, docking and MD, foundation models, SSL ...\nif you can find some very good molecular property prediction, you maybe do not need calibration.\nwhen i say \"very good \" it means smiliar (or less than 5% performance drop) validation/public AP for share or nonshare molecules",
      "votes": null
    },
    {
      "id": "2810944",
      "postDate": "05/13/2024 14:30:32",
      "content": "<p>Makes sense. Fair enough. For me, I haven't separately tried to test them (share and non-share dataset), but used the entire dataset instead,  the results don't include calibration technique. Need to see how well/worse the model performs on both the shared/non-shared molecules and will retrain the model including calibration while training and see. Thanks again for taking out time and replying! :)</p>",
      "rawMarkdown": "Makes sense. Fair enough. For me, I haven't separately tried to test them (share and non-share dataset), but used the entire dataset instead,  the results don't include calibration technique. Need to see how well/worse the model performs on both the shared/non-shared molecules and will retrain the model including calibration while training and see. Thanks again for taking out time and replying! :)",
      "votes": null
    },
    {
      "id": "2812456",
      "postDate": "05/14/2024 08:32:53",
      "content": "<p>you should do something like this<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff67b6f787485745a2de7af1407201037%2FSelection_114.png?generation=1715679072660828&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F18efc9c7ff01f5956f8c7df5286ae475%2FSelection_122.png?generation=1715682158180679&amp;alt=media\"></p>\n<p>your homework: plot these curves after shifting or retrain with more neg weight in loss, etc<br>\nmight be easier to analyse if if add \"recall vs p\" and \"precision vs p\" in the plots as well</p>",
      "rawMarkdown": "you should do something like this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff67b6f787485745a2de7af1407201037%2FSelection_114.png?generation=1715679072660828&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F18efc9c7ff01f5956f8c7df5286ae475%2FSelection_122.png?generation=1715682158180679&alt=media)\n\nyour homework: plot these curves after shifting or retrain with more neg weight in loss, etc\nmight be easier to analyse if if add \"recall vs p\" and \"precision vs p\" in the plots as well",
      "votes": null
    },
    {
      "id": "2813818",
      "postDate": "05/15/2024 04:08:46",
      "content": "<p>The simplest approach, which I haven't tried yet, is train a model and validate that model against some shared and some non shared blocks. Then, since it's validation data, directly calibrate the (lessened) accuracy of the non-share predictions. Then save the calibration adjustment(s) and use it for your LB non-share predictions. See if you get approximately the same benefit simply from predicting how much the model predictions have degraded on non share even without explicitly giving yourself any outside knowledge about average binding frequency on test non share. </p>",
      "rawMarkdown": "The simplest approach, which I haven't tried yet, is train a model and validate that model against some shared and some non shared blocks. Then, since it's validation data, directly calibrate the (lessened) accuracy of the non-share predictions. Then save the calibration adjustment(s) and use it for your LB non-share predictions. See if you get approximately the same benefit simply from predicting how much the model predictions have degraded on non share even without explicitly giving yourself any outside knowledge about average binding frequency on test non share.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2810834,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/13/2024 13:18:43",
      "content": "<p>for the private test, you still can apply \"shifting trick\". but because share and nonshare distribution are different, \"where, what and how much\" to shift will be different from the public one</p>\n<p>if you make the wrong shift, you will get very bad results, worse than non-shifting.<br>\n(how much you can gain will be how much you can lose)</p>\n<hr>\n<p>please understand the statistics and reason why it work (and the risk why it doesn't). <br>\ndo not blindly adjust.</p>\n<hr>\n<p>finally you do not need to apply methods if test and train distributions are the same.<br>\nlearning is better than betting. <br>\nif you want to bet, the chance of winning should be bigger than random and you have a backup plan</p>\n<hr>\n<p>basically for this trick to work, you must have:</p>\n<ul>\n<li>identify one subset(s) that you know the truth pos/neg ratio. </li>\n<li>at the current subset location the measured predicted pos/neg  is very different from what you \"know\".</li>\n<li>so you shifted the subset to the correct location.</li>\n<li>after shifting, it won't worsen the original results of that new location <br>\n(that is the meaning of calibration)</li>\n</ul>\n<p>in the public test data, in particular, the nonshare  predictions are stuck at zero. we know this is wrong location because we know there are some nonshare pos samples (via probing). so just shift it a little towards the correct location will get good results.This will not affect the share prediction if they are already well predicted and concentrated the high probability location.</p>\n<p>you can treat this as a game of ranking. how to swap locations (ranks) to make pos samples move to the right side.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2810885,
          "author_name": "ahsuna123",
          "author_url": "",
          "post_date": "05/13/2024 13:44:51",
          "content": "<p>Hey! Thank you for such a detailed explanation! Much appreciated. It was the post-processing technique you shared in one of the threads that helped me understand what was truly happening and about calibration. I have a question, is there a possibility for a non-calibration technique to outperform on the public lb or it doesn't stand a chance at all ? </p>",
          "votes": null,
          "replies": [
            {
              "id": 2810902,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "05/13/2024 13:55:39",
              "content": "<p>\" non-calibration technique to outperform on the public l\" </p>\n<p>i would say you can actually include calibration in the learning process. (not trial and error or manual adjust).<br>\nthat should perform very well.</p>\n<p>becuase the share and nonshare molecule are very \"different\", currnetly i cannot find a method that predict the nonshare moelcule well. in that sense, i don't think non-calibrated methods work better than pure prediction.</p>\n<p>but we haven't explore all methods yet, e.g. 3d graph, docking and MD, foundation models, SSL …<br>\nif you can find some very good molecular property prediction, you maybe do not need calibration.<br>\nwhen i say \"very good \" it means smiliar (or less than 5% performance drop) validation/public AP for share or nonshare molecules</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2810944,
                  "author_name": "ahsuna123",
                  "author_url": "",
                  "post_date": "05/13/2024 14:30:32",
                  "content": "<p>Makes sense. Fair enough. For me, I haven't separately tried to test them (share and non-share dataset), but used the entire dataset instead,  the results don't include calibration technique. Need to see how well/worse the model performs on both the shared/non-shared molecules and will retrain the model including calibration while training and see. Thanks again for taking out time and replying! :)</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2810856,
      "author_name": "masroormohajerani",
      "author_url": "",
      "post_date": "05/13/2024 13:33:32",
      "content": "<p>Calibration of models can indeed lead to significant changes in leaderboard standing, specially in competitions where the evaluation metric is sensitive to the confidence or probability estimates of predictions. However, whether calibration is the right choice depends on various factors, including the natur of the data, the evaluation metric, and the goals of the competition.<br>\nIf the distributon of samples in the private set differs significantly from that of the public set, calibration based solely on the public set may not generalize well to the private set. In such cases, relying solely on calibration approaches might indeed lead participants in the wrong direction. Strategies such as CrossValidation, Ensemble Methods, Post-Competition Analysis can be helpful. Calibrayution can lead to improvements in leaderboard standings, it's essential to consider the potential pitfalls, especially if the distribution of samples in the private set differs significantly from that of the public set. By understanding the data distribution, employing cross-validation techniques, using ensemble methods, and conducting post-competition analysis, participants can make more informed decisions and develop models that generalize well beyond the public leaderboard.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2812456,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/14/2024 08:32:53",
      "content": "<p>you should do something like this<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff67b6f787485745a2de7af1407201037%2FSelection_114.png?generation=1715679072660828&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F18efc9c7ff01f5956f8c7df5286ae475%2FSelection_122.png?generation=1715682158180679&amp;alt=media\"></p>\n<p>your homework: plot these curves after shifting or retrain with more neg weight in loss, etc<br>\nmight be easier to analyse if if add \"recall vs p\" and \"precision vs p\" in the plots as well</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2813818,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "05/15/2024 04:08:46",
      "content": "<p>The simplest approach, which I haven't tried yet, is train a model and validate that model against some shared and some non shared blocks. Then, since it's validation data, directly calibrate the (lessened) accuracy of the non-share predictions. Then save the calibration adjustment(s) and use it for your LB non-share predictions. See if you get approximately the same benefit simply from predicting how much the model predictions have degraded on non share even without explicitly giving yourself any outside knowledge about average binding frequency on test non share. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2810780": "Last couple of days made a significant change in the leaderboard, where calibration resulted in higher jumps at the public LB. But is it the right choice? If the distribution of samples in private set is not similar to that of public set, are we all directed in the wrong way and stick to non-calibration approaches ?",
    "2810834": "for the private test, you still can apply \"shifting trick\". but because share and nonshare distribution are different, \"where, what and how much\" to shift will be different from the public one\n\nif you make the wrong shift, you will get very bad results, worse than non-shifting.\n(how much you can gain will be how much you can lose)\n\n---\n\nplease understand the statistics and reason why it work (and the risk why it doesn't). \ndo not blindly adjust.\n\n---\n\nfinally you do not need to apply methods if test and train distributions are the same.\nlearning is better than betting. \nif you want to bet, the chance of winning should be bigger than random and you have a backup plan\n\n\n---\n\nbasically for this trick to work, you must have:\n- identify one subset(s) that you know the truth pos/neg ratio. \n- at the current subset location the measured predicted pos/neg  is very different from what you \"know\".\n- so you shifted the subset to the correct location.\n- after shifting, it won't worsen the original results of that new location \n(that is the meaning of calibration)\n\nin the public test data, in particular, the nonshare  predictions are stuck at zero. we know this is wrong location because we know there are some nonshare pos samples (via probing). so just shift it a little towards the correct location will get good results.This will not affect the share prediction if they are already well predicted and concentrated the high probability location.\n\nyou can treat this as a game of ranking. how to swap locations (ranks) to make pos samples move to the right side.",
    "2810856": "Calibration of models can indeed lead to significant changes in leaderboard standing, specially in competitions where the evaluation metric is sensitive to the confidence or probability estimates of predictions. However, whether calibration is the right choice depends on various factors, including the natur of the data, the evaluation metric, and the goals of the competition.\nIf the distributon of samples in the private set differs significantly from that of the public set, calibration based solely on the public set may not generalize well to the private set. In such cases, relying solely on calibration approaches might indeed lead participants in the wrong direction. Strategies such as CrossValidation, Ensemble Methods, Post-Competition Analysis can be helpful. Calibrayution can lead to improvements in leaderboard standings, it's essential to consider the potential pitfalls, especially if the distribution of samples in the private set differs significantly from that of the public set. By understanding the data distribution, employing cross-validation techniques, using ensemble methods, and conducting post-competition analysis, participants can make more informed decisions and develop models that generalize well beyond the public leaderboard.",
    "2810885": "Hey! Thank you for such a detailed explanation! Much appreciated. It was the post-processing technique you shared in one of the threads that helped me understand what was truly happening and about calibration. I have a question, is there a possibility for a non-calibration technique to outperform on the public lb or it doesn't stand a chance at all ?",
    "2810902": "\" non-calibration technique to outperform on the public l\" \n\ni would say you can actually include calibration in the learning process. (not trial and error or manual adjust).\nthat should perform very well.\n\nbecuase the share and nonshare molecule are very \"different\", currnetly i cannot find a method that predict the nonshare moelcule well. in that sense, i don't think non-calibrated methods work better than pure prediction.\n\nbut we haven't explore all methods yet, e.g. 3d graph, docking and MD, foundation models, SSL ...\nif you can find some very good molecular property prediction, you maybe do not need calibration.\nwhen i say \"very good \" it means smiliar (or less than 5% performance drop) validation/public AP for share or nonshare molecules",
    "2810944": "Makes sense. Fair enough. For me, I haven't separately tried to test them (share and non-share dataset), but used the entire dataset instead,  the results don't include calibration technique. Need to see how well/worse the model performs on both the shared/non-shared molecules and will retrain the model including calibration while training and see. Thanks again for taking out time and replying! :)",
    "2812456": "you should do something like this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff67b6f787485745a2de7af1407201037%2FSelection_114.png?generation=1715679072660828&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F18efc9c7ff01f5956f8c7df5286ae475%2FSelection_122.png?generation=1715682158180679&alt=media)\n\nyour homework: plot these curves after shifting or retrain with more neg weight in loss, etc\nmight be easier to analyse if if add \"recall vs p\" and \"precision vs p\" in the plots as well",
    "2813818": "The simplest approach, which I haven't tried yet, is train a model and validate that model against some shared and some non shared blocks. Then, since it's validation data, directly calibrate the (lessened) accuracy of the non-share predictions. Then save the calibration adjustment(s) and use it for your LB non-share predictions. See if you get approximately the same benefit simply from predicting how much the model predictions have degraded on non share even without explicitly giving yourself any outside knowledge about average binding frequency on test non share."
  },
  "source": "meta"
}