{
  "id": 501901,
  "title": "Score Calibration Problem",
  "url": "/competitions/leash-BELKA/discussion/501901",
  "author_name": "Ruby",
  "post_date": "2024-05-11T07:32:55.339000",
  "votes": 36,
  "comment_count": 33,
  "views": 0,
  "content": "<p><strong>In short</strong>: my score boost from 0.639 to 0.668 (before metric update) comes from training model for noshare part with smaller weight for negative samples.<br>\nBased on <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a>'s amazing work-<a href=\"https://www.kaggle.com/code/junkoda/unknowns-in-public-test-set\" target=\"_blank\">https://www.kaggle.com/code/junkoda/unknowns-in-public-test-set</a>, it seems LB consists of several groups of data with different positive rate. I want to known if such difference comes from different filtering rules related to molecule property only or it is filtered directly based on target value? (minor structure changes based past positive experiment should also belongs to this case)<br>\nIf it is the latter case there is no way we can predict the unseen experiment design/filtering rules only based on molecule and this shouldn’t be what host really cares. Given specific molecule the binding should be deterministic but unless we have perfect model the negative rate matters when compare predicted probability.<br>\nTo more concentrate on the problem itself, I think it is better to compute MAP within each design group and then average the scores or provide us the exact positive rate of each groups such that we can do calibration without gambling.<br>\nHope for host reply. <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a><br>\nEdit: Maybe we can do calibration based on predicted average probability but this involves unnecessary complexity.</p>",
  "messages": [
    {
      "id": 2806614,
      "postDate": "2024-05-11T07:32:55.340Z",
      "content": "<p><strong>In short</strong>: my score boost from 0.639 to 0.668 (before metric update) comes from training model for noshare part with smaller weight for negative samples.<br>\nBased on <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a>'s amazing work-<a href=\"https://www.kaggle.com/code/junkoda/unknowns-in-public-test-set\" target=\"_blank\">https://www.kaggle.com/code/junkoda/unknowns-in-public-test-set</a>, it seems LB consists of several groups of data with different positive rate. I want to known if such difference comes from different filtering rules related to molecule property only or it is filtered directly based on target value? (minor structure changes based past positive experiment should also belongs to this case)<br>\nIf it is the latter case there is no way we can predict the unseen experiment design/filtering rules only based on molecule and this shouldn’t be what host really cares. Given specific molecule the binding should be deterministic but unless we have perfect model the negative rate matters when compare predicted probability.<br>\nTo more concentrate on the problem itself, I think it is better to compute MAP within each design group and then average the scores or provide us the exact positive rate of each groups such that we can do calibration without gambling.<br>\nHope for host reply. <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a><br>\nEdit: Maybe we can do calibration based on predicted average probability but this involves unnecessary complexity.</p>",
      "rawMarkdown": "**In short**: my score boost from 0.639 to 0.668 (before metric update) comes from training model for noshare part with smaller weight for negative samples.\nBased on @junkoda's amazing work-https://www.kaggle.com/code/junkoda/unknowns-in-public-test-set, it seems LB consists of several groups of data with different positive rate. I want to known if such difference comes from different filtering rules related to molecule property only or it is filtered directly based on target value? (minor structure changes based past positive experiment should also belongs to this case)\nIf it is the latter case there is no way we can predict the unseen experiment design/filtering rules only based on molecule and this shouldn’t be what host really cares. Given specific molecule the binding should be deterministic but unless we have perfect model the negative rate matters when compare predicted probability.\nTo more concentrate on the problem itself, I think it is better to compute MAP within each design group and then average the scores or provide us the exact positive rate of each groups such that we can do calibration without gambling.\nHope for host reply. @andrewdblevins\nEdit: Maybe we can do calibration based on predicted average probability but this involves unnecessary complexity.",
      "votes": 36
    },
    {
      "id": 2807674,
      "postDate": "2024-05-11T19:22:52.697Z",
      "content": "<p>instead of retraining noshare part with smaller weight for negative samples, i implement in another way. post process the results to give boost to non-share prediction</p>\n<p>(mathematically, loss weight in bce loss is power in probability, aka temperature scaling)</p>\n<pre><code>        nonshare_submit_df = submit_df()\n        nonshare_submit_df **= \n        nonshare_submit_df(ensemble_file(, ), index=False)\n        (nonshare_submit_df())\n</code></pre>\n<p>lb increase from 0.620 to 0.643.</p>\n<hr>\n<p>so it is not only calibration for different protein, it can also be calibration for different group (share vs non share)</p>",
      "rawMarkdown": "instead of retraining noshare part with smaller weight for negative samples, i implement in another way. post process the results to give boost to non-share prediction\n\n\n(mathematically, loss weight in bce loss is power in probability, aka temperature scaling)\n\n```\n\t\tnonshare_submit_df = submit_df.copy()\n\t\tnonshare_submit_df.loc[sharing_df.is_share == 0, 'binds'] **= 0.75\n\t\tnonshare_submit_df.to_csv(ensemble_file.replace('.csv', '.nonshare-boost0.75.csv'), index=False)\n\t\tprint(nonshare_submit_df.sum())\n```\n\nlb increase from 0.620 to 0.643.\n\n---\n\nso it is not only calibration for different protein, it can also be calibration for different group (share vs non share)",
      "votes": 10,
      "replies": [
        {
          "id": 2809133,
          "postDate": "2024-05-12T14:46:07.303Z",
          "content": "<p>So this is the mysterious 'pow(p,n)'? 🤣<br>\nThis competition is too much tricks and micro-management…the opposite of what I thought it would be.</p>",
          "rawMarkdown": "So this is the mysterious 'pow(p,n)'? 🤣\nThis competition is too much tricks and micro-management...the opposite of what I thought it would be.",
          "votes": 2,
          "replies": [
            {
              "id": 2809180,
              "postDate": "2024-05-12T15:13:28.277Z",
              "content": "<p>actually calibration is part of model building in bioinformatics.<br>\n(they need this to do hypothesis tests)</p>\n<p>but it would be incredibly difficult here because of different domain for share and nonshare molecules. But it isn't impossible to do … now non-ML methods like MD and docking would become important (and with active probing) to provide guidance for ML calibration.</p>\n<p>once we have full excess to the test data, anything is possible</p>",
              "rawMarkdown": "actually calibration is part of model building in bioinformatics.\n(they need this to do hypothesis tests)\n\nbut it would be incredibly difficult here because of different domain for share and nonshare molecules. But it isn't impossible to do ... now non-ML methods like MD and docking would become important (and with active probing) to provide guidance for ML calibration.\n\nonce we have full excess to the test data, anything is possible",
              "votes": 3
            },
            {
              "id": 2809198,
              "postDate": "2024-05-12T15:23:15.010Z",
              "content": "<p>I thought it would be like Ribonanza, where we had a massive train set that generalized to test, and we could just build the best model.<br>\nBut it starts to look more like Novozymes…</p>",
              "rawMarkdown": "I thought it would be like Ribonanza, where we had a massive train set that generalized to test, and we could just build the best model.\nBut it starts to look more like Novozymes...",
              "votes": 5
            },
            {
              "id": 2809453,
              "postDate": "2024-05-12T18:43:02.630Z",
              "content": "<p>I've been thinking of it as half and half for awhile now. Half Novozymes has interesting implications, this thread is touching on a problem I've been thinking about for weeks, actually. If half your predictions are great, and half are an unknown level of terrible… How do you blend? Is docking a subset a way of determining how bad your second model is (and thus calibrating)? Is there some trustworthy way to quantify the comparative predictive power? Do you just guess?</p>\n<p>I see it as three areas. The \"Ribonanza\", the \"Novozymes\", but also the blending/docking/calibrating area of focus. </p>\n<p>But the better the final models, the less the last area matters!</p>",
              "rawMarkdown": "I've been thinking of it as half and half for awhile now. Half Novozymes has interesting implications, this thread is touching on a problem I've been thinking about for weeks, actually. If half your predictions are great, and half are an unknown level of terrible... How do you blend? Is docking a subset a way of determining how bad your second model is (and thus calibrating)? Is there some trustworthy way to quantify the comparative predictive power? Do you just guess?\n\nI see it as three areas. The \"Ribonanza\", the \"Novozymes\", but also the blending/docking/calibrating area of focus. \n\nBut the better the final models, the less the last area matters!",
              "votes": 1
            },
            {
              "id": 2809463,
              "postDate": "2024-05-12T19:06:15.803Z",
              "content": "<p>The problem is…[Accurate (Ribonanza]]+[Random (Novozymes)] = Random hahaha<br>\nI mean…yeah you will have a slightly better chance to finish at top if you are top in the first half. But it will still be only a (small as far as we can tell now) chance…sigh. I hate random competitions. It's good I have the LEAP competition; now THAT is a great one.</p>",
              "rawMarkdown": "The problem is...[Accurate (Ribonanza]]+[Random (Novozymes)] = Random hahaha\nI mean...yeah you will have a slightly better chance to finish at top if you are top in the first half. But it will still be only a (small as far as we can tell now) chance...sigh. I hate random competitions. It's good I have the LEAP competition; now THAT is a great one."
            },
            {
              "id": 2809480,
              "postDate": "2024-05-12T19:24:17.380Z",
              "content": "<p>Oh it's very problematic, yeah. I just enjoy trying to overcome that sort of thing. Lol</p>",
              "rawMarkdown": "Oh it's very problematic, yeah. I just enjoy trying to overcome that sort of thing. Lol",
              "votes": 1
            },
            {
              "id": 2809893,
              "postDate": "2024-05-13T03:44:14.470Z",
              "content": "<p>Might have to give LEAP a shot, quite unfortunate that this seems to be a \"random\" competition as I was looking for something closer to Ribonanza too.</p>",
              "rawMarkdown": "Might have to give LEAP a shot, quite unfortunate that this seems to be a \"random\" competition as I was looking for something closer to Ribonanza too.",
              "votes": 2
            },
            {
              "id": 2809908,
              "postDate": "2024-05-13T04:03:22.080Z",
              "content": "<p>LEAP is hard core DL. You just need to build a really good model and be able to train it on large data. It's like Ribonanza on steroids in that aspect. The difference- no bio, no external data/special features like BPP/CapR etc. Also no edges data. So it's even more simple/pure. IDK why it's so unpopular…well outside the fact that the total size of available data is ~490 GB hahaha</p>",
              "rawMarkdown": "LEAP is hard core DL. You just need to build a really good model and be able to train it on large data. It's like Ribonanza on steroids in that aspect. The difference- no bio, no external data/special features like BPP/CapR etc. Also no edges data. So it's even more simple/pure. IDK why it's so unpopular...well outside the fact that the total size of available data is ~490 GB hahaha",
              "votes": 1
            },
            {
              "id": 2809924,
              "postDate": "2024-05-13T04:15:15.743Z",
              "content": "<p>i don't mind nonshare block. But the distribution of share and nonshare should be similar in public and private test split. This would make our life easier.</p>\n<hr>\n<p>\"Ideas for calibration: Do 10, 20, 30+ folds of CV with this same problem.\" …<br>\nThere is alternative solution. Someone have to standardize the split and probing. all kagglers can disclose some non-share statistics  and submission score. I think we can find some pattern.</p>\n<p>NOTE: the current way to mask out share (set to zero) to probe nonshare is wrong.<br>\nThis is because if you submit (share=0, nonshare) and (share=0, monotonic function(share)), they would have the same results. </p>\n<p>AP is a ranking problem. you basically need to submit share = some good prediction (e.g from public notebook) as a reference, and see how nonshare rank relative to this reference.  </p>",
              "rawMarkdown": "i don't mind nonshare block. But the distribution of share and nonshare should be similar in public and private test split. This would make our life easier.\n\n---\n\n\"Ideas for calibration: Do 10, 20, 30+ folds of CV with this same problem.\" ...\nThere is alternative solution. Someone have to standardize the split and probing. all kagglers can disclose some non-share statistics  and submission score. I think we can find some pattern.\n\nNOTE: the current way to mask out share (set to zero) to probe nonshare is wrong.\nThis is because if you submit (share=0, nonshare) and (share=0, monotonic function(share)), they would have the same results. \n\nAP is a ranking problem. you basically need to submit share = some good prediction (e.g from public notebook) as a reference, and see how nonshare rank relative to this reference.  \n\n",
              "votes": 1
            },
            {
              "id": 2810024,
              "postDate": "2024-05-13T05:29:16.713Z",
              "content": "<p>There's a lot of ways to do it. But I would caution against focusing mainly on public LB. Or even a (single) standard split. We have a large training dataset, and can look at what generally works on a variety of target non share blocks and what doesn't, instead of just over fitting public LB. </p>",
              "rawMarkdown": "There's a lot of ways to do it. But I would caution against focusing mainly on public LB. Or even a (single) standard split. We have a large training dataset, and can look at what generally works on a variety of target non share blocks and what doesn't, instead of just over fitting public LB. "
            },
            {
              "id": 2810618,
              "postDate": "2024-05-13T11:01:05.623Z",
              "content": "<p>i just realise one thing interesting<br>\nthere about unalbelled 880k test. my training is batch size = 5k<br>\nthis means 176 batches. </p>\n<p>I did a \"thought experiment\". it takes 5 weeks to submit all batches, of which 35% (public) will return with batch-level label (pos fraction) .</p>\n<hr>\n<p>then actually we don't need to submit and probe at all.<br>\nif the batch size is large enough, we definitely have some positive samples.<br>\nso actually we <strong>do have some batch-level label for the unlabelled dataset</strong>.</p>\n<p>so if we are training with <strong>large batch size</strong>,</p>\n<ol>\n<li>all train label pos must rank better than all train labelled neg</li>\n<li>some train unlabelled must rank better than all train labelled neg</li>\n</ol>\n<p>we have sample loss for labelled and batch level loss for unlabelled<br>\n(something like multiple instant ranking … concept of bag loss)</p>\n<hr>\n<p>i haven't did experiments, but i can imaging, this forces the unlabelled score to stretch out during training.<br>\nnow if we imagine some parametric distribution, e.g. poisson as target, batch loss can be just kl fitting.</p>\n<p>currently, my resources can go up to batch-size = 10k . so we can train and calibrate at the same time</p>",
              "rawMarkdown": "i just realise one thing interesting\nthere about unalbelled 880k test. my training is batch size = 5k\nthis means 176 batches. \n\nI did a \"thought experiment\". it takes 5 weeks to submit all batches, of which 35% (public) will return with batch-level label (pos fraction) .\n\n---\n\nthen actually we don't need to submit and probe at all.\nif the batch size is large enough, we definitely have some positive samples.\nso actually we **do have some batch-level label for the unlabelled dataset**.\n\nso if we are training with **large batch size**,\n1. all train label pos must rank better than all train labelled neg\n2. some train unlabelled must rank better than all train labelled neg\n\nwe have sample loss for labelled and batch level loss for unlabelled\n(something like multiple instant ranking ... concept of bag loss)\n\n---\ni haven't did experiments, but i can imaging, this forces the unlabelled score to stretch out during training.\nnow if we imagine some parametric distribution, e.g. poisson as target, batch loss can be just kl fitting.\n\ncurrently, my resources can go up to batch-size = 10k . so we can train and calibrate at the same time",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2807698,
      "postDate": "2024-05-11T19:45:35Z",
      "content": "<p>updated my diagram<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F97f5a323b292a958c07110e61dac6764%2FSelection_103.png?generation=1715456726680517&amp;alt=media\"></p>\n<p>so we can \"cheat\" by \"stretching the prediction score\" in post processing in the second case.</p>",
      "rawMarkdown": "updated my diagram\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F97f5a323b292a958c07110e61dac6764%2FSelection_103.png?generation=1715456726680517&alt=media)\n\nso we can \"cheat\" by \"stretching the prediction score\" in post processing in the second case.",
      "votes": 7
    },
    {
      "id": 2809236,
      "postDate": "2024-05-12T16:05:37.370Z",
      "content": "<p>I'm not too happy with the scoring mechanics for this competition. I think it would've been better to score groups and then average those scores. And/or score top 1% or in some way emphasize the importance of getting as many positives into your high confidence region as possible, rather than trying to detangle which of hundreds of thousands of low probability are 'less low'. </p>\n<p>HOWEVER, we don't know whether this calibration issue is mainly a solvable problem or not. Mostly people didn't  realize the issue existed so haven't even tried. Also the private LB will frankly have much bigger problems then merely calibration, heh. </p>\n<p>Ideas for calibration: Do 10, 20, 30+ folds of CV with this same problem. Do we get similar clumps of (lower and) higher numbers of positive samples? Can we train a model to have any confidence in distinguishing GROUP positive rate? Even if not, can we <em>directly</em> calibrate a large amount of many folds of CV \"non-share\" vs \"share\" predictions and get appropriately less confident predictions on nonshare, and thus the same LB scoring bump? Or does it not work, and the LB bump is only from knowing there's more positive samples in that group?</p>\n<p>It's discouraging when you can get a huge LB increase with \"tricks\", but let's not give up hope it can be handled by a smart model :)</p>",
      "rawMarkdown": "I'm not too happy with the scoring mechanics for this competition. I think it would've been better to score groups and then average those scores. And/or score top 1% or in some way emphasize the importance of getting as many positives into your high confidence region as possible, rather than trying to detangle which of hundreds of thousands of low probability are 'less low'. \n\nHOWEVER, we don't know whether this calibration issue is mainly a solvable problem or not. Mostly people didn't  realize the issue existed so haven't even tried. Also the private LB will frankly have much bigger problems then merely calibration, heh. \n\nIdeas for calibration: Do 10, 20, 30+ folds of CV with this same problem. Do we get similar clumps of (lower and) higher numbers of positive samples? Can we train a model to have any confidence in distinguishing GROUP positive rate? Even if not, can we *directly* calibrate a large amount of many folds of CV \"non-share\" vs \"share\" predictions and get appropriately less confident predictions on nonshare, and thus the same LB scoring bump? Or does it not work, and the LB bump is only from knowing there's more positive samples in that group?\n\nIt's discouraging when you can get a huge LB increase with \"tricks\", but let's not give up hope it can be handled by a smart model :)",
      "votes": 5,
      "replies": [
        {
          "id": 2809326,
          "postDate": "2024-05-12T17:11:14.847Z",
          "content": "<p>I completely agree. In a realistic drug discovery setting, what we really care about is to maximize the number of true actives in the top 1%. After all, due to cost constraints, only less than a hundred of these compounds would be routinely synthesized off-DNA for pharmacological assays. </p>\n<p>Furthermore, the winning model will mostly likely be used to screen virtual libraries. Having a model that can rank the best molecules in the top1% would indeed be the most useful, rather than a model that can distinguish a weak binder from a very weak binder. </p>",
          "rawMarkdown": "I completely agree. In a realistic drug discovery setting, what we really care about is to maximize the number of true actives in the top 1%. After all, due to cost constraints, only less than a hundred of these compounds would be routinely synthesized off-DNA for pharmacological assays. \n\nFurthermore, the winning model will mostly likely be used to screen virtual libraries. Having a model that can rank the best molecules in the top1% would indeed be the most useful, rather than a model that can distinguish a weak binder from a very weak binder. ",
          "votes": 8
        }
      ]
    },
    {
      "id": 2806722,
      "postDate": "2024-05-11T08:35:09.187Z",
      "content": "<p>to put into picture, the issue is:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F67f9f78321b75c65228ebbe841d635d1%2FSelection_093.png?generation=1715416504126885&amp;alt=media\"></p>",
      "rawMarkdown": "to put into picture, the issue is:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F67f9f78321b75c65228ebbe841d635d1%2FSelection_093.png?generation=1715416504126885&alt=media)\n\n",
      "votes": 4,
      "replies": [
        {
          "id": 2806735,
          "postDate": "2024-05-11T08:46:15.017Z",
          "content": "<p>I think it is a bit different from difficulty. Even all proteins has same difficulty, score calibration will also boost the score by give higher probability to those proteins with more positive molecule candidates. By the way your example are all correctly ordered so score in picture can't represent meaningful probability, if your score is 0.9 then there should be 10% blue dot and 90% red dot nearby.</p>",
          "rawMarkdown": "I think it is a bit different from difficulty. Even all proteins has same difficulty, score calibration will also boost the score by give higher probability to those proteins with more positive molecule candidates. By the way your example are all correctly ordered so score in picture can't represent meaningful probability, if your score is 0.9 then there should be 10% blue dot and 90% red dot nearby.",
          "replies": [
            {
              "id": 2806893,
              "postDate": "2024-05-11T11:28:49.700Z",
              "content": "<p>\"so score in picture can't represent meaningful probability, if your score is 0.9 then there should be 10% blue dot and 90% red dot nearby.\"</p>\n<p>you can imagine after training a model, you take 10 test samples and make the plot.</p>\n<p>\"10% blue dot and 90% red dot nearby.\" . this is for train samples.</p>\n<hr>\n<p>it really depends on how good are your model.</p>\n<p>I did experiments on dummy data. we can simulate difficulty by adding label noise.<br>\ni measure ap(each and average micro) on the train data as train learning iteration proceeds.<br>\nyou can monitor how the ap changes (and also the predict prob distributions) as the model get stronger and become perfect (since it is train data)</p>",
              "rawMarkdown": "\"so score in picture can't represent meaningful probability, if your score is 0.9 then there should be 10% blue dot and 90% red dot nearby.\"\n\nyou can imagine after training a model, you take 10 test samples and make the plot.\n\n\"10% blue dot and 90% red dot nearby.\" . this is for train samples.\n\n---\n\nit really depends on how good are your model.\n\nI did experiments on dummy data. we can simulate difficulty by adding label noise.\ni measure ap(each and average micro) on the train data as train learning iteration proceeds.\nyou can monitor how the ap changes (and also the predict prob distributions) as the model get stronger and become perfect (since it is train data)"
            }
          ]
        }
      ]
    },
    {
      "id": 2815040,
      "postDate": "2024-05-15T16:51:51.497Z",
      "content": "<p>I also think that it is better to use the average score of MAPs calculated separately for each target protein as the competition metric. These proteins have somewhat different properties, and thus complex calibration might sometimes be required to optimize the metric currently used. Is there any benefit to calculating these together in this sort of virtual screening?</p>",
      "rawMarkdown": "I also think that it is better to use the average score of MAPs calculated separately for each target protein as the competition metric. These proteins have somewhat different properties, and thus complex calibration might sometimes be required to optimize the metric currently used. Is there any benefit to calculating these together in this sort of virtual screening?",
      "votes": 1
    },
    {
      "id": 2810887,
      "postDate": "2024-05-13T13:46:25.127Z",
      "content": "<p>So if the unknown ratio is larger in private lb, there could be very large shake up?</p>",
      "rawMarkdown": "So if the unknown ratio is larger in private lb, there could be very large shake up?",
      "votes": 1,
      "replies": [
        {
          "id": 2810895,
          "postDate": "2024-05-13T13:51:23.550Z",
          "content": "<p>Exactly, what I'm skeptical of as well. </p>",
          "rawMarkdown": "Exactly, what I'm skeptical of as well. ",
          "replies": [
            {
              "id": 2810912,
              "postDate": "2024-05-13T14:00:44.173Z",
              "content": "<p>shake up is what i call \"kaggle version of short selling\" 卖空 做空</p>",
              "rawMarkdown": "shake up is what i call \"kaggle version of short selling\" 卖空 做空",
              "votes": 2
            },
            {
              "id": 2810951,
              "postDate": "2024-05-13T14:32:51.937Z",
              "content": "<p>Hilarious analogy! On point thou! 😀</p>",
              "rawMarkdown": "Hilarious analogy! On point thou! 😀"
            }
          ]
        },
        {
          "id": 2810933,
          "postDate": "2024-05-13T14:14:49.047Z",
          "content": "<p>Yes if it is this case. But based on competition data description- \"To test generalizability, the test set contains building blocks that are not in the training set. These datasets are very imbalanced: <strong>roughly 0.5% of examples are classified as binders</strong>\", maybe public noshare part is just the result of an accident. If test positive rate is very different I don't think host will release such misleading information. Personally I will try to collect more hints before make decision.</p>",
          "rawMarkdown": "Yes if it is this case. But based on competition data description- \"To test generalizability, the test set contains building blocks that are not in the training set. These datasets are very imbalanced: **roughly 0.5% of examples are classified as binders**\", maybe public noshare part is just the result of an accident. If test positive rate is very different I don't think host will release such misleading information. Personally I will try to collect more hints before make decision.\n",
          "votes": 1,
          "replies": [
            {
              "id": 2811070,
              "postDate": "2024-05-13T15:41:57.983Z",
              "content": "<p>It would be better asking host directly for clarification upon it.</p>",
              "rawMarkdown": "It would be better asking host directly for clarification upon it."
            }
          ]
        }
      ]
    },
    {
      "id": 2808802,
      "postDate": "2024-05-12T12:05:23.933Z",
      "content": "<p>I have a very silly question to ask but I'll ask it anyway, what are you referring to <strong>no-share</strong> here? The building blocks that are not in the training set? Or I am getting it all wrong? </p>",
      "rawMarkdown": "I have a very silly question to ask but I'll ask it anyway, what are you referring to **no-share** here? The building blocks that are not in the training set? Or I am getting it all wrong? ",
      "votes": 1,
      "replies": [
        {
          "id": 2808806,
          "postDate": "2024-05-12T12:07:47.657Z",
          "content": "<p>Yes,some test samples with BB not seen in training set.</p>",
          "rawMarkdown": "Yes,some test samples with BB not seen in training set.",
          "votes": 2,
          "replies": [
            {
              "id": 2808817,
              "postDate": "2024-05-12T12:12:32.057Z",
              "content": "<p>Thank you so much for clarifying! And kudos to you for currently being at 1st! :)🥳</p>",
              "rawMarkdown": "Thank you so much for clarifying! And kudos to you for currently being at 1st! :)🥳"
            }
          ]
        }
      ]
    },
    {
      "id": 2816314,
      "postDate": "2024-05-16T08:49:52.930Z",
      "content": "<p>Would also prefer to use average score</p>",
      "rawMarkdown": "Would also prefer to use average score",
      "votes": -1
    },
    {
      "id": 2811712,
      "postDate": "2024-05-13T20:28:37.463Z",
      "content": "<p>Let me check my understanding, I've been doing research on noshare in silo for quite a while. Tuning the predictions to share part is what seem to happen here (downplaying the noshare predictions) and it works well for public LB. But this will be a mistake to submit as a final prediction as \"true\" private LB is actually 34% noshare and the good generalization there will be important. I assume that competition authors included noshare part in the public test set, but not randomly sampled from noshare part, but mostly sampled positive signals (reason for this is unknown). If 5% of noshare in public has conversion ratio of 4%, when when it will be expanded to private test set - it will get diluted (lets assume only negative samples for no share) to about 0.5%. Which means that average binding ratio is still 0.5% for (private + public) sets. WDYT?</p>",
      "rawMarkdown": "Let me check my understanding, I've been doing research on noshare in silo for quite a while. Tuning the predictions to share part is what seem to happen here (downplaying the noshare predictions) and it works well for public LB. But this will be a mistake to submit as a final prediction as \"true\" private LB is actually 34% noshare and the good generalization there will be important. I assume that competition authors included noshare part in the public test set, but not randomly sampled from noshare part, but mostly sampled positive signals (reason for this is unknown). If 5% of noshare in public has conversion ratio of 4%, when when it will be expanded to private test set - it will get diluted (lets assume only negative samples for no share) to about 0.5%. Which means that average binding ratio is still 0.5% for (private + public) sets. WDYT?",
      "replies": [
        {
          "id": 2811727,
          "postDate": "2024-05-13T20:55:40.523Z",
          "content": "<p>Couple quick comments:</p>\n<ul>\n<li>Re: downplaying: It's more like increasing noshare ptedictions, not downplaying, that is working well on public LB</li>\n<li>Re: submit as final prediction: we don't have much way of knowing if it will be a mistake or not. We aren't turning off any predictions. With or without these methods or ideas, the noshare predictions will be ordered as best as the model can do. So it won't impact score to scale them up or down. EXCEPT we need to blend noshare (model is less trusted, but model doesn't know that it is predicting noshare) with share preds (model is good and is trusted. Model is calibrated fine. )</li>\n<li>We have no way to know the private LB avg binding ratio. Unless we think the organizers were giving that info in one particular quote. That's one of the questions of this thread. </li>\n</ul>",
          "rawMarkdown": "Couple quick comments:\n* Re: downplaying: It's more like increasing noshare ptedictions, not downplaying, that is working well on public LB\n* Re: submit as final prediction: we don't have much way of knowing if it will be a mistake or not. We aren't turning off any predictions. With or without these methods or ideas, the noshare predictions will be ordered as best as the model can do. So it won't impact score to scale them up or down. EXCEPT we need to blend noshare (model is less trusted, but model doesn't know that it is predicting noshare) with share preds (model is good and is trusted. Model is calibrated fine. )\n* We have no way to know the private LB avg binding ratio. Unless we think the organizers were giving that info in one particular quote. That's one of the questions of this thread. ",
          "votes": 1,
          "replies": [
            {
              "id": 2811819,
              "postDate": "2024-05-13T23:12:35.607Z",
              "content": "<p>i have a question. 0.5% of 890k test molecules is 4450.<br>\nThis is not a large number. would it be easy to verify these 4450 cases using more time consuming methods like MD or docking software?</p>",
              "rawMarkdown": "i have a question. 0.5% of 890k test molecules is 4450.\nThis is not a large number. would it be easy to verify these 4450 cases using more time consuming methods like MD or docking software?"
            }
          ]
        }
      ]
    },
    {
      "id": 2811658,
      "postDate": "2024-05-13T19:59:56.200Z",
      "rawMarkdown": "",
      "votes": -6,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2807674,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-11T19:22:52.697000",
      "content": "<p>instead of retraining noshare part with smaller weight for negative samples, i implement in another way. post process the results to give boost to non-share prediction</p>\n<p>(mathematically, loss weight in bce loss is power in probability, aka temperature scaling)</p>\n<pre><code>        nonshare_submit_df = submit_df()\n        nonshare_submit_df **= \n        nonshare_submit_df(ensemble_file(, ), index=False)\n        (nonshare_submit_df())\n</code></pre>\n<p>lb increase from 0.620 to 0.643.</p>\n<hr>\n<p>so it is not only calibration for different protein, it can also be calibration for different group (share vs non share)</p>",
      "votes": 10,
      "replies": [
        {
          "id": 2809133,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-05-12T14:46:07.303000",
          "content": "<p>So this is the mysterious 'pow(p,n)'? 🤣<br>\nThis competition is too much tricks and micro-management…the opposite of what I thought it would be.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2809180,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-05-12T15:13:28.277000",
              "content": "<p>actually calibration is part of model building in bioinformatics.<br>\n(they need this to do hypothesis tests)</p>\n<p>but it would be incredibly difficult here because of different domain for share and nonshare molecules. But it isn't impossible to do … now non-ML methods like MD and docking would become important (and with active probing) to provide guidance for ML calibration.</p>\n<p>once we have full excess to the test data, anything is possible</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2809198,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-05-12T15:23:15.010000",
              "content": "<p>I thought it would be like Ribonanza, where we had a massive train set that generalized to test, and we could just build the best model.<br>\nBut it starts to look more like Novozymes…</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2809453,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2024-05-12T18:43:02.630000",
              "content": "<p>I've been thinking of it as half and half for awhile now. Half Novozymes has interesting implications, this thread is touching on a problem I've been thinking about for weeks, actually. If half your predictions are great, and half are an unknown level of terrible… How do you blend? Is docking a subset a way of determining how bad your second model is (and thus calibrating)? Is there some trustworthy way to quantify the comparative predictive power? Do you just guess?</p>\n<p>I see it as three areas. The \"Ribonanza\", the \"Novozymes\", but also the blending/docking/calibrating area of focus. </p>\n<p>But the better the final models, the less the last area matters!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2809463,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-05-12T19:06:15.803000",
              "content": "<p>The problem is…[Accurate (Ribonanza]]+[Random (Novozymes)] = Random hahaha<br>\nI mean…yeah you will have a slightly better chance to finish at top if you are top in the first half. But it will still be only a (small as far as we can tell now) chance…sigh. I hate random competitions. It's good I have the LEAP competition; now THAT is a great one.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2809480,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2024-05-12T19:24:17.380000",
              "content": "<p>Oh it's very problematic, yeah. I just enjoy trying to overcome that sort of thing. Lol</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2809893,
              "author_name": "sroger",
              "author_url": "",
              "post_date": "2024-05-13T03:44:14.470000",
              "content": "<p>Might have to give LEAP a shot, quite unfortunate that this seems to be a \"random\" competition as I was looking for something closer to Ribonanza too.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2809908,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-05-13T04:03:22.080000",
              "content": "<p>LEAP is hard core DL. You just need to build a really good model and be able to train it on large data. It's like Ribonanza on steroids in that aspect. The difference- no bio, no external data/special features like BPP/CapR etc. Also no edges data. So it's even more simple/pure. IDK why it's so unpopular…well outside the fact that the total size of available data is ~490 GB hahaha</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2809924,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-05-13T04:15:15.743000",
              "content": "<p>i don't mind nonshare block. But the distribution of share and nonshare should be similar in public and private test split. This would make our life easier.</p>\n<hr>\n<p>\"Ideas for calibration: Do 10, 20, 30+ folds of CV with this same problem.\" …<br>\nThere is alternative solution. Someone have to standardize the split and probing. all kagglers can disclose some non-share statistics  and submission score. I think we can find some pattern.</p>\n<p>NOTE: the current way to mask out share (set to zero) to probe nonshare is wrong.<br>\nThis is because if you submit (share=0, nonshare) and (share=0, monotonic function(share)), they would have the same results. </p>\n<p>AP is a ranking problem. you basically need to submit share = some good prediction (e.g from public notebook) as a reference, and see how nonshare rank relative to this reference.  </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2810024,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2024-05-13T05:29:16.713000",
              "content": "<p>There's a lot of ways to do it. But I would caution against focusing mainly on public LB. Or even a (single) standard split. We have a large training dataset, and can look at what generally works on a variety of target non share blocks and what doesn't, instead of just over fitting public LB. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2810618,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-05-13T11:01:05.623000",
              "content": "<p>i just realise one thing interesting<br>\nthere about unalbelled 880k test. my training is batch size = 5k<br>\nthis means 176 batches. </p>\n<p>I did a \"thought experiment\". it takes 5 weeks to submit all batches, of which 35% (public) will return with batch-level label (pos fraction) .</p>\n<hr>\n<p>then actually we don't need to submit and probe at all.<br>\nif the batch size is large enough, we definitely have some positive samples.<br>\nso actually we <strong>do have some batch-level label for the unlabelled dataset</strong>.</p>\n<p>so if we are training with <strong>large batch size</strong>,</p>\n<ol>\n<li>all train label pos must rank better than all train labelled neg</li>\n<li>some train unlabelled must rank better than all train labelled neg</li>\n</ol>\n<p>we have sample loss for labelled and batch level loss for unlabelled<br>\n(something like multiple instant ranking … concept of bag loss)</p>\n<hr>\n<p>i haven't did experiments, but i can imaging, this forces the unlabelled score to stretch out during training.<br>\nnow if we imagine some parametric distribution, e.g. poisson as target, batch loss can be just kl fitting.</p>\n<p>currently, my resources can go up to batch-size = 10k . so we can train and calibrate at the same time</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2807698,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-11T19:45:35",
      "content": "<p>updated my diagram<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F97f5a323b292a958c07110e61dac6764%2FSelection_103.png?generation=1715456726680517&amp;alt=media\"></p>\n<p>so we can \"cheat\" by \"stretching the prediction score\" in post processing in the second case.</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 2809236,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2024-05-12T16:05:37.370000",
      "content": "<p>I'm not too happy with the scoring mechanics for this competition. I think it would've been better to score groups and then average those scores. And/or score top 1% or in some way emphasize the importance of getting as many positives into your high confidence region as possible, rather than trying to detangle which of hundreds of thousands of low probability are 'less low'. </p>\n<p>HOWEVER, we don't know whether this calibration issue is mainly a solvable problem or not. Mostly people didn't  realize the issue existed so haven't even tried. Also the private LB will frankly have much bigger problems then merely calibration, heh. </p>\n<p>Ideas for calibration: Do 10, 20, 30+ folds of CV with this same problem. Do we get similar clumps of (lower and) higher numbers of positive samples? Can we train a model to have any confidence in distinguishing GROUP positive rate? Even if not, can we <em>directly</em> calibrate a large amount of many folds of CV \"non-share\" vs \"share\" predictions and get appropriately less confident predictions on nonshare, and thus the same LB scoring bump? Or does it not work, and the LB bump is only from knowing there's more positive samples in that group?</p>\n<p>It's discouraging when you can get a huge LB increase with \"tricks\", but let's not give up hope it can be handled by a smart model :)</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2809326,
          "author_name": "Rced_AI",
          "author_url": "",
          "post_date": "2024-05-12T17:11:14.847000",
          "content": "<p>I completely agree. In a realistic drug discovery setting, what we really care about is to maximize the number of true actives in the top 1%. After all, due to cost constraints, only less than a hundred of these compounds would be routinely synthesized off-DNA for pharmacological assays. </p>\n<p>Furthermore, the winning model will mostly likely be used to screen virtual libraries. Having a model that can rank the best molecules in the top1% would indeed be the most useful, rather than a model that can distinguish a weak binder from a very weak binder. </p>",
          "votes": 8,
          "replies": []
        }
      ]
    },
    {
      "id": 2806722,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-11T08:35:09.187000",
      "content": "<p>to put into picture, the issue is:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F67f9f78321b75c65228ebbe841d635d1%2FSelection_093.png?generation=1715416504126885&amp;alt=media\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 2806735,
          "author_name": "Ruby",
          "author_url": "",
          "post_date": "2024-05-11T08:46:15.017000",
          "content": "<p>I think it is a bit different from difficulty. Even all proteins has same difficulty, score calibration will also boost the score by give higher probability to those proteins with more positive molecule candidates. By the way your example are all correctly ordered so score in picture can't represent meaningful probability, if your score is 0.9 then there should be 10% blue dot and 90% red dot nearby.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2806893,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-05-11T11:28:49.700000",
              "content": "<p>\"so score in picture can't represent meaningful probability, if your score is 0.9 then there should be 10% blue dot and 90% red dot nearby.\"</p>\n<p>you can imagine after training a model, you take 10 test samples and make the plot.</p>\n<p>\"10% blue dot and 90% red dot nearby.\" . this is for train samples.</p>\n<hr>\n<p>it really depends on how good are your model.</p>\n<p>I did experiments on dummy data. we can simulate difficulty by adding label noise.<br>\ni measure ap(each and average micro) on the train data as train learning iteration proceeds.<br>\nyou can monitor how the ap changes (and also the predict prob distributions) as the model get stronger and become perfect (since it is train data)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2815040,
      "author_name": "rikein12",
      "author_url": "",
      "post_date": "2024-05-15T16:51:51.497000",
      "content": "<p>I also think that it is better to use the average score of MAPs calculated separately for each target protein as the competition metric. These proteins have somewhat different properties, and thus complex calibration might sometimes be required to optimize the metric currently used. Is there any benefit to calculating these together in this sort of virtual screening?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2810887,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2024-05-13T13:46:25.127000",
      "content": "<p>So if the unknown ratio is larger in private lb, there could be very large shake up?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2810895,
          "author_name": "AC",
          "author_url": "",
          "post_date": "2024-05-13T13:51:23.550000",
          "content": "<p>Exactly, what I'm skeptical of as well. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2810912,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-05-13T14:00:44.173000",
              "content": "<p>shake up is what i call \"kaggle version of short selling\" 卖空 做空</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2810951,
              "author_name": "AC",
              "author_url": "",
              "post_date": "2024-05-13T14:32:51.937000",
              "content": "<p>Hilarious analogy! On point thou! 😀</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2810933,
          "author_name": "Ruby",
          "author_url": "",
          "post_date": "2024-05-13T14:14:49.047000",
          "content": "<p>Yes if it is this case. But based on competition data description- \"To test generalizability, the test set contains building blocks that are not in the training set. These datasets are very imbalanced: <strong>roughly 0.5% of examples are classified as binders</strong>\", maybe public noshare part is just the result of an accident. If test positive rate is very different I don't think host will release such misleading information. Personally I will try to collect more hints before make decision.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2811070,
              "author_name": "yuanzhe zhou",
              "author_url": "",
              "post_date": "2024-05-13T15:41:57.983000",
              "content": "<p>It would be better asking host directly for clarification upon it.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2808802,
      "author_name": "AC",
      "author_url": "",
      "post_date": "2024-05-12T12:05:23.933000",
      "content": "<p>I have a very silly question to ask but I'll ask it anyway, what are you referring to <strong>no-share</strong> here? The building blocks that are not in the training set? Or I am getting it all wrong? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2808806,
          "author_name": "Ruby",
          "author_url": "",
          "post_date": "2024-05-12T12:07:47.657000",
          "content": "<p>Yes,some test samples with BB not seen in training set.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2808817,
              "author_name": "AC",
              "author_url": "",
              "post_date": "2024-05-12T12:12:32.057000",
              "content": "<p>Thank you so much for clarifying! And kudos to you for currently being at 1st! :)🥳</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2816314,
      "author_name": "Natarajan Vijaikumar",
      "author_url": "",
      "post_date": "2024-05-16T08:49:52.930000",
      "content": "<p>Would also prefer to use average score</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2811712,
      "author_name": "yamu_duck",
      "author_url": "",
      "post_date": "2024-05-13T20:28:37.463000",
      "content": "<p>Let me check my understanding, I've been doing research on noshare in silo for quite a while. Tuning the predictions to share part is what seem to happen here (downplaying the noshare predictions) and it works well for public LB. But this will be a mistake to submit as a final prediction as \"true\" private LB is actually 34% noshare and the good generalization there will be important. I assume that competition authors included noshare part in the public test set, but not randomly sampled from noshare part, but mostly sampled positive signals (reason for this is unknown). If 5% of noshare in public has conversion ratio of 4%, when when it will be expanded to private test set - it will get diluted (lets assume only negative samples for no share) to about 0.5%. Which means that average binding ratio is still 0.5% for (private + public) sets. WDYT?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2811727,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2024-05-13T20:55:40.523000",
          "content": "<p>Couple quick comments:</p>\n<ul>\n<li>Re: downplaying: It's more like increasing noshare ptedictions, not downplaying, that is working well on public LB</li>\n<li>Re: submit as final prediction: we don't have much way of knowing if it will be a mistake or not. We aren't turning off any predictions. With or without these methods or ideas, the noshare predictions will be ordered as best as the model can do. So it won't impact score to scale them up or down. EXCEPT we need to blend noshare (model is less trusted, but model doesn't know that it is predicting noshare) with share preds (model is good and is trusted. Model is calibrated fine. )</li>\n<li>We have no way to know the private LB avg binding ratio. Unless we think the organizers were giving that info in one particular quote. That's one of the questions of this thread. </li>\n</ul>",
          "votes": 1,
          "replies": [
            {
              "id": 2811819,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-05-13T23:12:35.607000",
              "content": "<p>i have a question. 0.5% of 890k test molecules is 4450.<br>\nThis is not a large number. would it be easy to verify these 4450 cases using more time consuming methods like MD or docking software?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2811658,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-13T19:59:56.200000",
      "content": "",
      "votes": -6,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2806614": "**In short**: my score boost from 0.639 to 0.668 (before metric update) comes from training model for noshare part with smaller weight for negative samples.\nBased on @junkoda's amazing work-https://www.kaggle.com/code/junkoda/unknowns-in-public-test-set, it seems LB consists of several groups of data with different positive rate. I want to known if such difference comes from different filtering rules related to molecule property only or it is filtered directly based on target value? (minor structure changes based past positive experiment should also belongs to this case)\nIf it is the latter case there is no way we can predict the unseen experiment design/filtering rules only based on molecule and this shouldn’t be what host really cares. Given specific molecule the binding should be deterministic but unless we have perfect model the negative rate matters when compare predicted probability.\nTo more concentrate on the problem itself, I think it is better to compute MAP within each design group and then average the scores or provide us the exact positive rate of each groups such that we can do calibration without gambling.\nHope for host reply. @andrewdblevins\nEdit: Maybe we can do calibration based on predicted average probability but this involves unnecessary complexity.",
    "2807674": "instead of retraining noshare part with smaller weight for negative samples, i implement in another way. post process the results to give boost to non-share prediction\n\n\n(mathematically, loss weight in bce loss is power in probability, aka temperature scaling)\n\n```\n\t\tnonshare_submit_df = submit_df.copy()\n\t\tnonshare_submit_df.loc[sharing_df.is_share == 0, 'binds'] **= 0.75\n\t\tnonshare_submit_df.to_csv(ensemble_file.replace('.csv', '.nonshare-boost0.75.csv'), index=False)\n\t\tprint(nonshare_submit_df.sum())\n```\n\nlb increase from 0.620 to 0.643.\n\n---\n\nso it is not only calibration for different protein, it can also be calibration for different group (share vs non share)",
    "2807698": "updated my diagram\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F97f5a323b292a958c07110e61dac6764%2FSelection_103.png?generation=1715456726680517&alt=media)\n\nso we can \"cheat\" by \"stretching the prediction score\" in post processing in the second case.",
    "2809236": "I'm not too happy with the scoring mechanics for this competition. I think it would've been better to score groups and then average those scores. And/or score top 1% or in some way emphasize the importance of getting as many positives into your high confidence region as possible, rather than trying to detangle which of hundreds of thousands of low probability are 'less low'. \n\nHOWEVER, we don't know whether this calibration issue is mainly a solvable problem or not. Mostly people didn't  realize the issue existed so haven't even tried. Also the private LB will frankly have much bigger problems then merely calibration, heh. \n\nIdeas for calibration: Do 10, 20, 30+ folds of CV with this same problem. Do we get similar clumps of (lower and) higher numbers of positive samples? Can we train a model to have any confidence in distinguishing GROUP positive rate? Even if not, can we *directly* calibrate a large amount of many folds of CV \"non-share\" vs \"share\" predictions and get appropriately less confident predictions on nonshare, and thus the same LB scoring bump? Or does it not work, and the LB bump is only from knowing there's more positive samples in that group?\n\nIt's discouraging when you can get a huge LB increase with \"tricks\", but let's not give up hope it can be handled by a smart model :)",
    "2806722": "to put into picture, the issue is:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F67f9f78321b75c65228ebbe841d635d1%2FSelection_093.png?generation=1715416504126885&alt=media)\n\n",
    "2815040": "I also think that it is better to use the average score of MAPs calculated separately for each target protein as the competition metric. These proteins have somewhat different properties, and thus complex calibration might sometimes be required to optimize the metric currently used. Is there any benefit to calculating these together in this sort of virtual screening?",
    "2810887": "So if the unknown ratio is larger in private lb, there could be very large shake up?",
    "2808802": "I have a very silly question to ask but I'll ask it anyway, what are you referring to **no-share** here? The building blocks that are not in the training set? Or I am getting it all wrong? ",
    "2816314": "Would also prefer to use average score",
    "2811712": "Let me check my understanding, I've been doing research on noshare in silo for quite a while. Tuning the predictions to share part is what seem to happen here (downplaying the noshare predictions) and it works well for public LB. But this will be a mistake to submit as a final prediction as \"true\" private LB is actually 34% noshare and the good generalization there will be important. I assume that competition authors included noshare part in the public test set, but not randomly sampled from noshare part, but mostly sampled positive signals (reason for this is unknown). If 5% of noshare in public has conversion ratio of 4%, when when it will be expanded to private test set - it will get diluted (lets assume only negative samples for no share) to about 0.5%. Which means that average binding ratio is still 0.5% for (private + public) sets. WDYT?",
    "2811658": ""
  }
}