{
  "id": 503232,
  "title": "Announcement: Changes to Scoring Metric in the Competition",
  "url": "/competitions/leash-BELKA/discussion/503232",
  "author_name": "Andrew D. Blevins",
  "post_date": "2024-05-16T15:30:32.053000",
  "votes": 68,
  "comment_count": 44,
  "views": 0,
  "content": "<p>Dear Participants,</p>\n<p>We hope you are enjoying the competition so far. With nearly 900 teams participating, the level of engagement and enthusiasm has been incredible! Your diligence has identified an issue with the scoring metric currently in use and we will be making adjustments to ensure a fair and accurate evaluation of your models.</p>\n<p><strong>Issue Description:</strong></p>\n<p>Our dataset, derived from DEL screens containing 133 million molecules against 3 proteins, was split using multiple techniques:</p>\n<ol>\n<li><strong>Reserved Building Blocks:</strong> Some building blocks were reserved for the test set, ensuring they did not appear in the training set.</li>\n<li><strong>Random and Chemical Scaffold Separation:</strong> Certain Murcko scaffolds were made exclusive to the test set, then we additionally sampled random molecules.</li>\n<li><strong>Proprietary Library Addition:</strong> A proprietary chemical library was included solely in the private test set.</li>\n</ol>\n<p>Each of these sampling techniques resulted in different proportions of hits. When holding out building blocks, all molecules containing those pieces were removed from the training set. Given that single building blocks drive a lot of activity, we had to do some careful stratified sampling to get enough positive examples into the test sets, and this led to varying levels of difficulty and base rates across splits and proteins.</p>\n<p>A number of teams have pointed out these differences and their effects on scoring. Chemdatafarmer and Robert Hatch identified the presence of novel chemistry in the test set early in the competition (<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/493294\" target=\"_blank\">1</a>, <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/496576\" target=\"_blank\">2</a>). Several competitors probed CV/LB results in a thread started by Darek Kłeczek (<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/498983\" target=\"_blank\">3</a>). Jun Koda described the differences in known/unknown building block proportions in the sets more precisely (<a href=\"https://www.kaggle.com/code/junkoda/unknowns-in-public-test-set\" target=\"_blank\">4</a>), leading to the discovery that tweaking calibration gave a private test score boost by Ruby (<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/501901\" target=\"_blank\">5</a>). Hengck23 beautifully summarized the issue with the below diagrams from (<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/501901#2806722\" target=\"_blank\">6</a>) and (<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/501901#2807698\" target=\"_blank\">7</a>):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F57970%2F80570cda23eeca6845ea100bb0588a40%2FScreen%20Shot%202024-05-16%20at%209.31.41%20AM.png?generation=1715873422843117&amp;alt=media\"></p>\n<p>Using a single average precision (AP) number as the metric inadvertently emphasized the importance of understanding these base rates, which was not our intention. Our goal is to focus on identifying the molecule representations and model architectures that provide the greatest generalization.</p>\n<p><strong>New Scoring Metric:</strong></p>\n<p>To address this issue, we will calculate the average precision for each protein and each split group individually and then average these scores. This method ensures a balanced evaluation across different groups. The revised scoring metric is as follows:</p>\n<p><strong>Public Metric:</strong><br>\nThe Average Precision will be calculated across 6 groups:</p>\n<ul>\n<li>sEH<ul>\n<li>Shared-bb</li>\n<li>Nonshared-bb</li></ul></li>\n<li>HSA<ul>\n<li>Shared-bb</li>\n<li>Nonshared-bb</li></ul></li>\n<li>BRD4<ul>\n<li>Shared-bb</li>\n<li>Nonshared-bb</li></ul></li>\n</ul>\n<p>Those 6 values will be averaged with equal weighting into the final public score</p>\n<p><strong>Private Metric:</strong></p>\n<p>The Average Precision will be calculated across 9 groups:</p>\n<ul>\n<li>sEH<ul>\n<li>Shared-bb</li>\n<li>Nonshared-bb</li>\n<li>New library</li></ul></li>\n<li>HSA<ul>\n<li>Shared-bb</li>\n<li>Nonshared-bb</li>\n<li>New library</li></ul></li>\n<li>BRD4<ul>\n<li>Shared-bb</li>\n<li>Nonshared-bb</li>\n<li>New library</li></ul></li>\n</ul>\n<p>Those 9 values will be averaged with equal weighting into the final public score. This means that 2/3rds the final score will come from molecules outside the training distribution, and we know predicting out of distribution is a hard challenge here. Still, the kinds of physical interactions that drive binding (hydrogen bonds, shape, pi-stacking, Van der Waals, charge distribution, etc.) are universal. We hope that progress can be made towards generalizing.</p>\n<p><strong>Nefarious New Library</strong><br>\nAs many have already noticed, we have added a set of molecules into the private set that are completely different from the triazine molecules in the test set.</p>\n<p>This new library was constructed by attaching DNA to a trifunctional core, and reacting fmoc/boc protected amide bond formation on one side and a Suzuki reaction on the other side. We suspect predicting these will be difficult but being able to predict these will show that you truly have trained a model that generalizes across chemical space.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F57970%2F8021959caea9c31363f400870b5f65ac%2FScreen%20Shot%202024-05-16%20at%209.28.27%20AM.png?generation=1715873251490099&amp;alt=media\"></p>\n<p><strong>Implementation Timeline:</strong></p>\n<p>These changes will be implemented immediately, and the leaderboard will be in a temporary state of flux as previous submissions are rescored. An announcement will be made when that process is complete. We appreciate your understanding and cooperation as we strive to maintain the integrity and fairness of the competition.</p>\n<p>Thank you for your participation and dedication. We look forward to seeing your continued progress!</p>\n<p>Best regards,  <br>\nLeash Biosciences</p>",
  "messages": [
    {
      "id": 2816896,
      "postDate": "2024-05-16T15:30:32.053Z",
      "content": "<p>Dear Participants,</p>\n<p>We hope you are enjoying the competition so far. With nearly 900 teams participating, the level of engagement and enthusiasm has been incredible! Your diligence has identified an issue with the scoring metric currently in use and we will be making adjustments to ensure a fair and accurate evaluation of your models.</p>\n<p><strong>Issue Description:</strong></p>\n<p>Our dataset, derived from DEL screens containing 133 million molecules against 3 proteins, was split using multiple techniques:</p>\n<ol>\n<li><strong>Reserved Building Blocks:</strong> Some building blocks were reserved for the test set, ensuring they did not appear in the training set.</li>\n<li><strong>Random and Chemical Scaffold Separation:</strong> Certain Murcko scaffolds were made exclusive to the test set, then we additionally sampled random molecules.</li>\n<li><strong>Proprietary Library Addition:</strong> A proprietary chemical library was included solely in the private test set.</li>\n</ol>\n<p>Each of these sampling techniques resulted in different proportions of hits. When holding out building blocks, all molecules containing those pieces were removed from the training set. Given that single building blocks drive a lot of activity, we had to do some careful stratified sampling to get enough positive examples into the test sets, and this led to varying levels of difficulty and base rates across splits and proteins.</p>\n<p>A number of teams have pointed out these differences and their effects on scoring. Chemdatafarmer and Robert Hatch identified the presence of novel chemistry in the test set early in the competition (<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/493294\" target=\"_blank\">1</a>, <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/496576\" target=\"_blank\">2</a>). Several competitors probed CV/LB results in a thread started by Darek Kłeczek (<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/498983\" target=\"_blank\">3</a>). Jun Koda described the differences in known/unknown building block proportions in the sets more precisely (<a href=\"https://www.kaggle.com/code/junkoda/unknowns-in-public-test-set\" target=\"_blank\">4</a>), leading to the discovery that tweaking calibration gave a private test score boost by Ruby (<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/501901\" target=\"_blank\">5</a>). Hengck23 beautifully summarized the issue with the below diagrams from (<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/501901#2806722\" target=\"_blank\">6</a>) and (<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/501901#2807698\" target=\"_blank\">7</a>):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F57970%2F80570cda23eeca6845ea100bb0588a40%2FScreen%20Shot%202024-05-16%20at%209.31.41%20AM.png?generation=1715873422843117&amp;alt=media\"></p>\n<p>Using a single average precision (AP) number as the metric inadvertently emphasized the importance of understanding these base rates, which was not our intention. Our goal is to focus on identifying the molecule representations and model architectures that provide the greatest generalization.</p>\n<p><strong>New Scoring Metric:</strong></p>\n<p>To address this issue, we will calculate the average precision for each protein and each split group individually and then average these scores. This method ensures a balanced evaluation across different groups. The revised scoring metric is as follows:</p>\n<p><strong>Public Metric:</strong><br>\nThe Average Precision will be calculated across 6 groups:</p>\n<ul>\n<li>sEH<ul>\n<li>Shared-bb</li>\n<li>Nonshared-bb</li></ul></li>\n<li>HSA<ul>\n<li>Shared-bb</li>\n<li>Nonshared-bb</li></ul></li>\n<li>BRD4<ul>\n<li>Shared-bb</li>\n<li>Nonshared-bb</li></ul></li>\n</ul>\n<p>Those 6 values will be averaged with equal weighting into the final public score</p>\n<p><strong>Private Metric:</strong></p>\n<p>The Average Precision will be calculated across 9 groups:</p>\n<ul>\n<li>sEH<ul>\n<li>Shared-bb</li>\n<li>Nonshared-bb</li>\n<li>New library</li></ul></li>\n<li>HSA<ul>\n<li>Shared-bb</li>\n<li>Nonshared-bb</li>\n<li>New library</li></ul></li>\n<li>BRD4<ul>\n<li>Shared-bb</li>\n<li>Nonshared-bb</li>\n<li>New library</li></ul></li>\n</ul>\n<p>Those 9 values will be averaged with equal weighting into the final public score. This means that 2/3rds the final score will come from molecules outside the training distribution, and we know predicting out of distribution is a hard challenge here. Still, the kinds of physical interactions that drive binding (hydrogen bonds, shape, pi-stacking, Van der Waals, charge distribution, etc.) are universal. We hope that progress can be made towards generalizing.</p>\n<p><strong>Nefarious New Library</strong><br>\nAs many have already noticed, we have added a set of molecules into the private set that are completely different from the triazine molecules in the test set.</p>\n<p>This new library was constructed by attaching DNA to a trifunctional core, and reacting fmoc/boc protected amide bond formation on one side and a Suzuki reaction on the other side. We suspect predicting these will be difficult but being able to predict these will show that you truly have trained a model that generalizes across chemical space.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F57970%2F8021959caea9c31363f400870b5f65ac%2FScreen%20Shot%202024-05-16%20at%209.28.27%20AM.png?generation=1715873251490099&amp;alt=media\"></p>\n<p><strong>Implementation Timeline:</strong></p>\n<p>These changes will be implemented immediately, and the leaderboard will be in a temporary state of flux as previous submissions are rescored. An announcement will be made when that process is complete. We appreciate your understanding and cooperation as we strive to maintain the integrity and fairness of the competition.</p>\n<p>Thank you for your participation and dedication. We look forward to seeing your continued progress!</p>\n<p>Best regards,  <br>\nLeash Biosciences</p>",
      "rawMarkdown": "Dear Participants,\n\nWe hope you are enjoying the competition so far. With nearly 900 teams participating, the level of engagement and enthusiasm has been incredible! Your diligence has identified an issue with the scoring metric currently in use and we will be making adjustments to ensure a fair and accurate evaluation of your models.\n\n**Issue Description:**\n\nOur dataset, derived from DEL screens containing 133 million molecules against 3 proteins, was split using multiple techniques:\n\n1. **Reserved Building Blocks:** Some building blocks were reserved for the test set, ensuring they did not appear in the training set.\n2. **Random and Chemical Scaffold Separation:** Certain Murcko scaffolds were made exclusive to the test set, then we additionally sampled random molecules.\n3. **Proprietary Library Addition:** A proprietary chemical library was included solely in the private test set.\n\nEach of these sampling techniques resulted in different proportions of hits. When holding out building blocks, all molecules containing those pieces were removed from the training set. Given that single building blocks drive a lot of activity, we had to do some careful stratified sampling to get enough positive examples into the test sets, and this led to varying levels of difficulty and base rates across splits and proteins.\n\nA number of teams have pointed out these differences and their effects on scoring. Chemdatafarmer and Robert Hatch identified the presence of novel chemistry in the test set early in the competition ([1](https://www.kaggle.com/competitions/leash-BELKA/discussion/493294), [2](https://www.kaggle.com/competitions/leash-BELKA/discussion/496576)). Several competitors probed CV/LB results in a thread started by Darek Kłeczek ([3](https://www.kaggle.com/competitions/leash-BELKA/discussion/498983)). Jun Koda described the differences in known/unknown building block proportions in the sets more precisely ([4](https://www.kaggle.com/code/junkoda/unknowns-in-public-test-set)), leading to the discovery that tweaking calibration gave a private test score boost by Ruby ([5](https://www.kaggle.com/competitions/leash-BELKA/discussion/501901)). Hengck23 beautifully summarized the issue with the below diagrams from ([6](https://www.kaggle.com/competitions/leash-BELKA/discussion/501901#2806722)) and ([7](https://www.kaggle.com/competitions/leash-BELKA/discussion/501901#2807698)):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F57970%2F80570cda23eeca6845ea100bb0588a40%2FScreen%20Shot%202024-05-16%20at%209.31.41%20AM.png?generation=1715873422843117&alt=media)\n\nUsing a single average precision (AP) number as the metric inadvertently emphasized the importance of understanding these base rates, which was not our intention. Our goal is to focus on identifying the molecule representations and model architectures that provide the greatest generalization.\n\n**New Scoring Metric:**\n\nTo address this issue, we will calculate the average precision for each protein and each split group individually and then average these scores. This method ensures a balanced evaluation across different groups. The revised scoring metric is as follows:\n\n**Public Metric:**\nThe Average Precision will be calculated across 6 groups:\n* sEH\n  * Shared-bb\n  * Nonshared-bb\n* HSA\n  * Shared-bb\n  * Nonshared-bb\n* BRD4\n  * Shared-bb\n  * Nonshared-bb\n\n\nThose 6 values will be averaged with equal weighting into the final public score\n\n**Private Metric:**\n\nThe Average Precision will be calculated across 9 groups:\n* sEH\n  * Shared-bb\n  * Nonshared-bb\n  * New library\n* HSA\n  * Shared-bb\n  * Nonshared-bb\n  * New library\n* BRD4\n  * Shared-bb\n  * Nonshared-bb\n  * New library\n\nThose 9 values will be averaged with equal weighting into the final public score. This means that 2/3rds the final score will come from molecules outside the training distribution, and we know predicting out of distribution is a hard challenge here. Still, the kinds of physical interactions that drive binding (hydrogen bonds, shape, pi-stacking, Van der Waals, charge distribution, etc.) are universal. We hope that progress can be made towards generalizing.\n\n**Nefarious New Library**\nAs many have already noticed, we have added a set of molecules into the private set that are completely different from the triazine molecules in the test set.\n\nThis new library was constructed by attaching DNA to a trifunctional core, and reacting fmoc/boc protected amide bond formation on one side and a Suzuki reaction on the other side. We suspect predicting these will be difficult but being able to predict these will show that you truly have trained a model that generalizes across chemical space.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F57970%2F8021959caea9c31363f400870b5f65ac%2FScreen%20Shot%202024-05-16%20at%209.28.27%20AM.png?generation=1715873251490099&alt=media)\n\n**Implementation Timeline:**\n\nThese changes will be implemented immediately, and the leaderboard will be in a temporary state of flux as previous submissions are rescored. An announcement will be made when that process is complete. We appreciate your understanding and cooperation as we strive to maintain the integrity and fairness of the competition.\n\nThank you for your participation and dedication. We look forward to seeing your continued progress!\n\nBest regards,  \nLeash Biosciences\n",
      "votes": 67
    },
    {
      "id": 2816964,
      "postDate": "2024-05-16T16:17:20.370Z",
      "content": "<p>Thanks for the update.</p>\n<p>here i highlight the the important changes from machine learning point of view. Kagglers should take note note of these to make sure your validation consider this!</p>\n<ol>\n<li>\"different proportions of hits\"</li>\n<li>\"single building blocks drive a lot of activity, …. stratified ….\"</li>\n<li>\"enough positive examples into the test sets\"</li>\n<li>\"identifying the molecule representations that provide the greatest generalization.\"</li>\n<li>\" This means that 2/3rds the final score will come from molecules outside the training distribution,\"</li>\n<li>\" 9 … averaged with equal weighting\"</li>\n</ol>",
      "rawMarkdown": "Thanks for the update.\n\nhere i highlight the the important changes from machine learning point of view. Kagglers should take note note of these to make sure your validation consider this!\n1. \"different proportions of hits\"\n2. \"single building blocks drive a lot of activity, .... stratified ....\"\n3. \"enough positive examples into the test sets\"\n4. \"identifying the molecule representations that provide the greatest generalization.\"\n5. \" This means that 2/3rds the final score will come from molecules outside the training distribution,\"\n6. \" 9 ... averaged with equal weighting\"\n",
      "votes": 5
    },
    {
      "id": 2821136,
      "postDate": "2024-05-18T00:20:35.743Z",
      "content": "<p>Overall this seems good. However, if it's equal 1/6th and 1/9th, then I am pretty concerned this could over-emphasize the smaller number of samples of \"non-shared bb\"? Would it be reasonable to have all the scores be weighted by number of samples in that group?</p>\n<p>e.g. Public LB, approximately :<br>\n180,000 * shared sEH mean average precision + 11,000 * nonshared sEH + …  + 180k * HSA share + 11k * HSA nonshare. All divided by total number of public LB samples.</p>\n<p>And similarly for the more even distribution of the 9 groups in private LB?</p>",
      "rawMarkdown": "Overall this seems good. However, if it's equal 1/6th and 1/9th, then I am pretty concerned this could over-emphasize the smaller number of samples of \"non-shared bb\"? Would it be reasonable to have all the scores be weighted by number of samples in that group?\n\ne.g. Public LB, approximately :\n180,000 * shared sEH mean average precision + 11,000 * nonshared sEH + ...  + 180k * HSA share + 11k * HSA nonshare. All divided by total number of public LB samples.\n\nAnd similarly for the more even distribution of the 9 groups in private LB?",
      "votes": 6,
      "replies": [
        {
          "id": 2821139,
          "postDate": "2024-05-18T00:23:56.930Z",
          "content": "<p>Just to add now that I see my score: I don't care that the current metric hurts my current score, since I had intentionally focused solely on shared BB. But I do worry that there will just be higher variance even between good models on the small 11 thousand sample groups.</p>",
          "rawMarkdown": "Just to add now that I see my score: I don't care that the current metric hurts my current score, since I had intentionally focused solely on shared BB. But I do worry that there will just be higher variance even between good models on the small 11 thousand sample groups.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2913491,
      "postDate": "2024-07-09T13:47:29.240Z",
      "content": "<p>I've been trying to understand the public-private leaderboards shakeup: with this metric, in very simplified case, if we assume that most models failed to generalize onto non-triazine / new library molecules (and so score -&gt; 0), a simple rule of thumb for the scores relations between the public and private would've have to be something like this:</p>\n<p><code>Public LB ~= Private LB * 3 / 2</code></p>\n<p>However it doesn't seem to be the case? So I'm starting to think - how were non-shared building blocks distributed between private and public parts of the test set?<br>\nIf that's ok to ask, <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> - was that uniform random, or were these two independent sets, so public and private parts had different non-shared building blocks - non-shared not only with the train set, but also non-shared between each other? Thanks!</p>",
      "rawMarkdown": "I've been trying to understand the public-private leaderboards shakeup: with this metric, in very simplified case, if we assume that most models failed to generalize onto non-triazine / new library molecules (and so score -> 0), a simple rule of thumb for the scores relations between the public and private would've have to be something like this:\n\n`Public LB ~= Private LB * 3 / 2`\n\nHowever it doesn't seem to be the case? So I'm starting to think - how were non-shared building blocks distributed between private and public parts of the test set?\nIf that's ok to ask, @andrewdblevins - was that uniform random, or were these two independent sets, so public and private parts had different non-shared building blocks - non-shared not only with the train set, but also non-shared between each other? Thanks!",
      "votes": 1,
      "replies": [
        {
          "id": 2913518,
          "postDate": "2024-07-09T13:59:09.217Z",
          "content": "<p>with so few building blocks we were afraid of probing, so the non-shared bbs were different in the public and private set.</p>",
          "rawMarkdown": "with so few building blocks we were afraid of probing, so the non-shared bbs were different in the public and private set.",
          "replies": [
            {
              "id": 2913530,
              "postDate": "2024-07-09T14:04:26.983Z",
              "content": "<p>That makes sense, thank you!</p>",
              "rawMarkdown": "That makes sense, thank you!"
            }
          ]
        },
        {
          "id": 2913574,
          "postDate": "2024-07-09T14:58:04.530Z",
          "content": "<p>The \"real\" public LB scores were ~400 with private ~266. Public no share might happen to be a little easier on avg so slightly above 400 vs slightly below 266. </p>\n<p>But then you have the small sample size of each non share block adding large variation. So most public LB scores is matching to that not CV, so big increase possible, won't generalize well and probably further decrease private score.  And as seen by posts here, huge variation on private with single models single folds, different seeds, so it's possible to get large increase to private score way above the ~260 that a good generalized model can get. </p>",
          "rawMarkdown": "The \"real\" public LB scores were ~400 with private ~266. Public no share might happen to be a little easier on avg so slightly above 400 vs slightly below 266. \n\nBut then you have the small sample size of each non share block adding large variation. So most public LB scores is matching to that not CV, so big increase possible, won't generalize well and probably further decrease private score.  And as seen by posts here, huge variation on private with single models single folds, different seeds, so it's possible to get large increase to private score way above the ~260 that a good generalized model can get. ",
          "replies": [
            {
              "id": 2913619,
              "postDate": "2024-07-09T15:25:12.243Z",
              "content": "<p>Yes, that definitely makes sense - however, whilst <strong>drops</strong> in ranking are quite expected here, I was rather a bit puzzled by models that scored high on Private but low on Public, jumping <strong>up</strong>. This means that these models are also not generalizing well - they just landed in a better spot in the bbs-performance space!<br>\nGenerally given that bbs in private and public are different - it's all really obvious now, but before I was mostly focused on thinking about triazine - non-triazine, so I only expected more systematic shift down (like from 400 -&gt; 266, yes)</p>",
              "rawMarkdown": "Yes, that definitely makes sense - however, whilst **drops** in ranking are quite expected here, I was rather a bit puzzled by models that scored high on Private but low on Public, jumping **up**. This means that these models are also not generalizing well - they just landed in a better spot in the bbs-performance space!\nGenerally given that bbs in private and public are different - it's all really obvious now, but before I was mostly focused on thinking about triazine - non-triazine, so I only expected more systematic shift down (like from 400 -> 266, yes)",
              "votes": 2
            },
            {
              "id": 2913633,
              "postDate": "2024-07-09T15:31:59.557Z",
              "content": "<p>I actually brought up how noisy the new scoring metric would be in this very thread, lol. The intent was good - necessary! - but making 40000 rows 1/3 of the score is too much variance (1/2 if no one figures out nontriazine). </p>",
              "rawMarkdown": "I actually brought up how noisy the new scoring metric would be in this very thread, lol. The intent was good - necessary! - but making 40000 rows 1/3 of the score is too much variance (1/2 if no one figures out nontriazine). "
            },
            {
              "id": 2913669,
              "postDate": "2024-07-09T15:43:56.067Z",
              "content": "<p>As the targets were well defined, I think it might have been possible to use external binding/non-binding data to push the model on non-triazines more to the top. Modelling might also have helped, but modelling molecules with attached DNA is somewhat difficult.</p>",
              "rawMarkdown": "As the targets were well defined, I think it might have been possible to use external binding/non-binding data to push the model on non-triazines more to the top. Modelling might also have helped, but modelling molecules with attached DNA is somewhat difficult."
            }
          ]
        }
      ]
    },
    {
      "id": 2817539,
      "postDate": "2024-05-17T01:41:56.283Z",
      "content": "<p>Thank you for paying such close attention to the discussions and the courage to implement feedback and improve the competition on the fly!</p>",
      "rawMarkdown": "Thank you for paying such close attention to the discussions and the courage to implement feedback and improve the competition on the fly!",
      "votes": 1
    },
    {
      "id": 2816990,
      "postDate": "2024-05-16T16:34:52.760Z",
      "content": "<p>Thanks for quick update! But I note some notebook has rescoring error, which succeed in the past. Do you have any hints on this? I think there is nothing else constraints except for binds value should lies in [0,1]?</p>",
      "rawMarkdown": "Thanks for quick update! But I note some notebook has rescoring error, which succeed in the past. Do you have any hints on this? I think there is nothing else constraints except for binds value should lies in [0,1]?",
      "votes": 1,
      "replies": [
        {
          "id": 2816995,
          "postDate": "2024-05-16T16:37:00.900Z",
          "content": "<p>I am facing them too. Alot of failed submissions, the error it says is : <strong>Internal scoring error</strong></p>",
          "rawMarkdown": "I am facing them too. Alot of failed submissions, the error it says is : **Internal scoring error**",
          "replies": [
            {
              "id": 2817014,
              "postDate": "2024-05-16T16:48:51.293Z",
              "content": "<p>We are looking into it, but rescoring will happen over the next couple hours and things will be in flux until it finishes.</p>",
              "rawMarkdown": "We are looking into it, but rescoring will happen over the next couple hours and things will be in flux until it finishes.",
              "votes": 5
            }
          ]
        },
        {
          "id": 2817013,
          "postDate": "2024-05-16T16:48:23.543Z",
          "content": "<p>I'll investigate all of the scoring errors, and rescore, if necessary.</p>",
          "rawMarkdown": "I'll investigate all of the scoring errors, and rescore, if necessary.",
          "votes": 6,
          "replies": [
            {
              "id": 2817015,
              "postDate": "2024-05-16T16:48:55.340Z",
              "content": "<p>Thanks! :)</p>",
              "rawMarkdown": "Thanks! :)"
            },
            {
              "id": 2817276,
              "postDate": "2024-05-16T19:49:02.600Z",
              "content": "<p>i have many re-scoring error and many pending (even after 3 hours from the metric change is made)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc6c9e4b77d99de340560c428edfaac86%2FSelection_128.png?generation=1715888923935188&amp;alt=media\"></p>",
              "rawMarkdown": "i have many re-scoring error and many pending (even after 3 hours from the metric change is made)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc6c9e4b77d99de340560c428edfaac86%2FSelection_128.png?generation=1715888923935188&alt=media)"
            },
            {
              "id": 2817307,
              "postDate": "2024-05-16T20:11:55.327Z",
              "content": "<p>Yeah same here</p>",
              "rawMarkdown": "Yeah same here"
            },
            {
              "id": 2817326,
              "postDate": "2024-05-16T20:20:59.317Z",
              "content": "<p>It appears the rescoring error is general. I have many of it at as well</p>",
              "rawMarkdown": "It appears the rescoring error is general. I have many of it at as well"
            },
            {
              "id": 2817353,
              "postDate": "2024-05-16T20:57:22.260Z",
              "content": "<p>Yeah, there's something wrong on the eng side. Meeting with them in a few minutes to discuss.</p>",
              "rawMarkdown": "Yeah, there's something wrong on the eng side. Meeting with them in a few minutes to discuss."
            },
            {
              "id": 2817509,
              "postDate": "2024-05-17T00:48:53.403Z",
              "content": "<p>Update: For me, the rescoring error is resolved. As of now, 2 of my submissions are still processing but all the rest that were throwing internal scoring error have been submitted successfully. Thank you! :)</p>",
              "rawMarkdown": "Update: For me, the rescoring error is resolved. As of now, 2 of my submissions are still processing but all the rest that were throwing internal scoring error have been submitted successfully. Thank you! :)"
            }
          ]
        }
      ]
    },
    {
      "id": 2816975,
      "postDate": "2024-05-16T16:23:54.927Z",
      "content": "<p>Thanks to everybody who has put in the efforts. Appreciated :) <br>\n2/3rd is huge, focusing on better generalization is the goal here. 😄</p>",
      "rawMarkdown": "Thanks to everybody who has put in the efforts. Appreciated :) \n2/3rd is huge, focusing on better generalization is the goal here. 😄",
      "votes": 1
    },
    {
      "id": 2861667,
      "postDate": "2024-06-08T09:52:56.237Z",
      "content": "<p>Thanks to everybody who has put in the efforts. </p>",
      "rawMarkdown": "Thanks to everybody who has put in the efforts. "
    },
    {
      "id": 2819579,
      "postDate": "2024-05-17T15:34:17.827Z",
      "content": "<p>Thanks to everybody who has put in the efforts. Appreciated :)!</p>",
      "rawMarkdown": "Thanks to everybody who has put in the efforts. Appreciated :)!"
    },
    {
      "id": 2859405,
      "postDate": "2024-06-07T02:50:55.213Z",
      "content": "<p>Thanks to everybody who has put in the efforts. </p>",
      "rawMarkdown": "Thanks to everybody who has put in the efforts. ",
      "votes": -1
    },
    {
      "id": 2909581,
      "postDate": "2024-07-07T05:43:58.333Z",
      "content": "<p>Thanks for this competition. I learned a lot.</p>",
      "rawMarkdown": "Thanks for this competition. I learned a lot."
    },
    {
      "id": 2899971,
      "postDate": "2024-07-02T02:03:50.293Z",
      "content": "<p>good but can be more</p>",
      "rawMarkdown": "good but can be more"
    },
    {
      "id": 2893684,
      "postDate": "2024-06-28T02:26:57.650Z",
      "content": "<p>I know it is a bit late in the game, and I am just trying to learn. But when I try to copy &amp; edit this notebook I am running in to several errors, currently got stuck with a 'DEPRECATION WARNING: USE MORGAN GENERATOR', and it never finishes running. I tried adding a warning ignore snippet but with no luck so far. Also needed to comment to become a contributor to get into a team! Thanks for reading.</p>",
      "rawMarkdown": "I know it is a bit late in the game, and I am just trying to learn. But when I try to copy & edit this notebook I am running in to several errors, currently got stuck with a 'DEPRECATION WARNING: USE MORGAN GENERATOR', and it never finishes running. I tried adding a warning ignore snippet but with no luck so far. Also needed to comment to become a contributor to get into a team! Thanks for reading."
    },
    {
      "id": 2889664,
      "postDate": "2024-06-25T16:00:21.323Z",
      "content": "<p>Thanks to everybody who has put in the efforts.</p>",
      "rawMarkdown": "Thanks to everybody who has put in the efforts."
    },
    {
      "id": 2879978,
      "postDate": "2024-06-20T00:47:33.933Z",
      "content": "<p>Is there any relationship among the three proteins should be taken into consideration？</p>",
      "rawMarkdown": "Is there any relationship among the three proteins should be taken into consideration？"
    },
    {
      "id": 2870298,
      "postDate": "2024-06-13T13:57:16.783Z",
      "content": "<p>Hi there,</p>\n<p>Just looking for clarification here - </p>\n<p>Does \"Shared - bb\" mean that <em>all or some</em> of the building blocks in the test set are seen in the training set?</p>\n<p>Does \"Non-shared - bb\" mean that <em>all or some</em> of the building blocks in the test set are not seen in the training set?</p>",
      "rawMarkdown": "Hi there,\n\nJust looking for clarification here - \n\nDoes \"Shared - bb\" mean that *all or some* of the building blocks in the test set are seen in the training set?\n\nDoes \"Non-shared - bb\" mean that *all or some* of the building blocks in the test set are not seen in the training set?",
      "replies": [
        {
          "id": 2870394,
          "postDate": "2024-06-13T15:02:35.180Z",
          "content": "<p>I think <strong>all</strong> shared bb in the test test are seen in the training set and <strong>all</strong> non-shared bb are not seen in the training set</p>",
          "rawMarkdown": "I think **all** shared bb in the test test are seen in the training set and **all** non-shared bb are not seen in the training set",
          "votes": 1
        },
        {
          "id": 2871541,
          "postDate": "2024-06-14T09:08:22.760Z",
          "content": "<p>The test dataset comprises 878,022 unique molecules. Their building blocks 1 - 3 are either all shared or not shared at all. Additionally, all the molecules with non-triazine cores are made of non-shared building blocks. (my analysis <a href=\"https://www.kaggle.com/code/hideakiogasawara/chemspace-visualization\" target=\"_blank\">here</a>)</p>\n<table>\n<thead>\n<tr>\n<th>BB1 shared</th>\n<th>BB2 shared</th>\n<th>BB3s hared</th>\n<th>Triazine</th>\n<th>Number</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>True</td>\n<td>True</td>\n<td>True</td>\n<td>True</td>\n<td>369,039</td>\n</tr>\n<tr>\n<td>False</td>\n<td>False</td>\n<td>False</td>\n<td>True</td>\n<td>22,593</td>\n</tr>\n<tr>\n<td>False</td>\n<td>False</td>\n<td>False</td>\n<td>False</td>\n<td>486,390</td>\n</tr>\n<tr>\n<td></td>\n<td></td>\n<td></td>\n<td><strong>Total</strong></td>\n<td>878,022</td>\n</tr>\n</tbody>\n</table>",
          "rawMarkdown": "The test dataset comprises 878,022 unique molecules. Their building blocks 1 - 3 are either all shared or not shared at all. Additionally, all the molecules with non-triazine cores are made of non-shared building blocks. (my analysis [here](https://www.kaggle.com/code/hideakiogasawara/chemspace-visualization))\n\n| BB1 shared | BB2 shared | BB3s hared | Triazine | Number |\n| --- | --- | --- | --- | --- |\n| True | True | True | True | 369,039 |\n| False | False | False | True | 22,593 |\n| False  | False  | False  | False  | 486,390 |\n|    |    |  | **Total**  | 878,022  |",
          "votes": 3
        }
      ]
    },
    {
      "id": 2868732,
      "postDate": "2024-06-12T15:28:21.923Z",
      "content": "<p>Hi,</p>\n<p>I can't find new library required for Private Metric . Can anyone tell me where it is?</p>\n<p>Thank's.</p>",
      "rawMarkdown": "Hi,\n\nI can't find new library required for Private Metric . Can anyone tell me where it is?\n\nThank's."
    },
    {
      "id": 2842527,
      "postDate": "2024-05-29T06:11:17.670Z",
      "content": "<p>Since we now know better the structure of training and public and private test datasets, it would probably also be interesting to learn about the exact procedure of hit calling in the subsets. E.g.: is a single read in the third enrichment sufficient to call a hit or is the percentage of hits constant and the read threshold is picked accordingly.</p>\n<p>This should help in scaling the predictions for the subsets, although it might possibly not affect the final score as the three sets are scored separately.</p>",
      "rawMarkdown": "Since we now know better the structure of training and public and private test datasets, it would probably also be interesting to learn about the exact procedure of hit calling in the subsets. E.g.: is a single read in the third enrichment sufficient to call a hit or is the percentage of hits constant and the read threshold is picked accordingly.\n\nThis should help in scaling the predictions for the subsets, although it might possibly not affect the final score as the three sets are scored separately."
    },
    {
      "id": 2824259,
      "postDate": "2024-05-19T16:57:05.033Z",
      "content": "<p>I still couldn't understand the new library part.</p>",
      "rawMarkdown": "I still couldn't understand the new library part.",
      "replies": [
        {
          "id": 2828186,
          "postDate": "2024-05-21T22:58:01.397Z",
          "content": "<p>me too.<br>\naccording to the illustration, it seems to be only two bbs.<br>\ncan anyone help clarify this?</p>",
          "rawMarkdown": "me too.\naccording to the illustration, it seems to be only two bbs.\ncan anyone help clarify this?",
          "replies": [
            {
              "id": 2828232,
              "postDate": "2024-05-22T01:12:09.653Z",
              "content": "<p>38 cores -&gt; bb1. (I counted 36, in my post that's footnote 2, above, but either way). Then bb2 binds to bb1 via boronate (I think), and bb3 to bb1 via the other binder in their diagram. </p>",
              "rawMarkdown": "38 cores -> bb1. (I counted 36, in my post that's footnote 2, above, but either way). Then bb2 binds to bb1 via boronate (I think), and bb3 to bb1 via the other binder in their diagram. ",
              "votes": 1
            },
            {
              "id": 2830010,
              "postDate": "2024-05-23T00:36:32.100Z",
              "content": "<p>thanks(✺ω✺)</p>",
              "rawMarkdown": "thanks(✺ω✺)"
            }
          ]
        }
      ]
    },
    {
      "id": 2916597,
      "postDate": "2024-07-11T06:02:56.717Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2886121,
      "postDate": "2024-06-23T13:14:54.763Z",
      "content": "<p>Thanks to everybody who has put in the efforts.</p>",
      "rawMarkdown": "Thanks to everybody who has put in the efforts.",
      "isDeleted": true
    },
    {
      "id": 2862152,
      "postDate": "2024-06-08T15:43:53.563Z",
      "content": "<p>Thanks for update</p>",
      "rawMarkdown": "Thanks for update"
    },
    {
      "id": 2861876,
      "postDate": "2024-06-08T12:13:19.243Z",
      "content": "<p>Thanks for update!</p>",
      "rawMarkdown": "Thanks for update!"
    },
    {
      "id": 2822142,
      "postDate": "2024-05-18T13:13:33.060Z",
      "content": "<p>Thanks for the update.</p>",
      "rawMarkdown": "Thanks for the update."
    }
  ],
  "comments": [
    {
      "id": 2816964,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-16T16:17:20.370000",
      "content": "<p>Thanks for the update.</p>\n<p>here i highlight the the important changes from machine learning point of view. Kagglers should take note note of these to make sure your validation consider this!</p>\n<ol>\n<li>\"different proportions of hits\"</li>\n<li>\"single building blocks drive a lot of activity, …. stratified ….\"</li>\n<li>\"enough positive examples into the test sets\"</li>\n<li>\"identifying the molecule representations that provide the greatest generalization.\"</li>\n<li>\" This means that 2/3rds the final score will come from molecules outside the training distribution,\"</li>\n<li>\" 9 … averaged with equal weighting\"</li>\n</ol>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2821136,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2024-05-18T00:20:35.743000",
      "content": "<p>Overall this seems good. However, if it's equal 1/6th and 1/9th, then I am pretty concerned this could over-emphasize the smaller number of samples of \"non-shared bb\"? Would it be reasonable to have all the scores be weighted by number of samples in that group?</p>\n<p>e.g. Public LB, approximately :<br>\n180,000 * shared sEH mean average precision + 11,000 * nonshared sEH + …  + 180k * HSA share + 11k * HSA nonshare. All divided by total number of public LB samples.</p>\n<p>And similarly for the more even distribution of the 9 groups in private LB?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2821139,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2024-05-18T00:23:56.930000",
          "content": "<p>Just to add now that I see my score: I don't care that the current metric hurts my current score, since I had intentionally focused solely on shared BB. But I do worry that there will just be higher variance even between good models on the small 11 thousand sample groups.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2913491,
      "author_name": "Mikhail Pershin",
      "author_url": "",
      "post_date": "2024-07-09T13:47:29.240000",
      "content": "<p>I've been trying to understand the public-private leaderboards shakeup: with this metric, in very simplified case, if we assume that most models failed to generalize onto non-triazine / new library molecules (and so score -&gt; 0), a simple rule of thumb for the scores relations between the public and private would've have to be something like this:</p>\n<p><code>Public LB ~= Private LB * 3 / 2</code></p>\n<p>However it doesn't seem to be the case? So I'm starting to think - how were non-shared building blocks distributed between private and public parts of the test set?<br>\nIf that's ok to ask, <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> - was that uniform random, or were these two independent sets, so public and private parts had different non-shared building blocks - non-shared not only with the train set, but also non-shared between each other? Thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2913518,
          "author_name": "Andrew D. Blevins",
          "author_url": "",
          "post_date": "2024-07-09T13:59:09.217000",
          "content": "<p>with so few building blocks we were afraid of probing, so the non-shared bbs were different in the public and private set.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2913530,
              "author_name": "Mikhail Pershin",
              "author_url": "",
              "post_date": "2024-07-09T14:04:26.983000",
              "content": "<p>That makes sense, thank you!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2913574,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2024-07-09T14:58:04.530000",
          "content": "<p>The \"real\" public LB scores were ~400 with private ~266. Public no share might happen to be a little easier on avg so slightly above 400 vs slightly below 266. </p>\n<p>But then you have the small sample size of each non share block adding large variation. So most public LB scores is matching to that not CV, so big increase possible, won't generalize well and probably further decrease private score.  And as seen by posts here, huge variation on private with single models single folds, different seeds, so it's possible to get large increase to private score way above the ~260 that a good generalized model can get. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2913619,
              "author_name": "Mikhail Pershin",
              "author_url": "",
              "post_date": "2024-07-09T15:25:12.243000",
              "content": "<p>Yes, that definitely makes sense - however, whilst <strong>drops</strong> in ranking are quite expected here, I was rather a bit puzzled by models that scored high on Private but low on Public, jumping <strong>up</strong>. This means that these models are also not generalizing well - they just landed in a better spot in the bbs-performance space!<br>\nGenerally given that bbs in private and public are different - it's all really obvious now, but before I was mostly focused on thinking about triazine - non-triazine, so I only expected more systematic shift down (like from 400 -&gt; 266, yes)</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2913633,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2024-07-09T15:31:59.557000",
              "content": "<p>I actually brought up how noisy the new scoring metric would be in this very thread, lol. The intent was good - necessary! - but making 40000 rows 1/3 of the score is too much variance (1/2 if no one figures out nontriazine). </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2913669,
              "author_name": "Bernhard Rohde",
              "author_url": "",
              "post_date": "2024-07-09T15:43:56.067000",
              "content": "<p>As the targets were well defined, I think it might have been possible to use external binding/non-binding data to push the model on non-triazines more to the top. Modelling might also have helped, but modelling molecules with attached DNA is somewhat difficult.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2817539,
      "author_name": "Frenio Redeker",
      "author_url": "",
      "post_date": "2024-05-17T01:41:56.283000",
      "content": "<p>Thank you for paying such close attention to the discussions and the courage to implement feedback and improve the competition on the fly!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2816990,
      "author_name": "Ruby",
      "author_url": "",
      "post_date": "2024-05-16T16:34:52.760000",
      "content": "<p>Thanks for quick update! But I note some notebook has rescoring error, which succeed in the past. Do you have any hints on this? I think there is nothing else constraints except for binds value should lies in [0,1]?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2816995,
          "author_name": "AC",
          "author_url": "",
          "post_date": "2024-05-16T16:37:00.900000",
          "content": "<p>I am facing them too. Alot of failed submissions, the error it says is : <strong>Internal scoring error</strong></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2817014,
              "author_name": "Andrew D. Blevins",
              "author_url": "",
              "post_date": "2024-05-16T16:48:51.293000",
              "content": "<p>We are looking into it, but rescoring will happen over the next couple hours and things will be in flux until it finishes.</p>",
              "votes": 5,
              "replies": []
            }
          ]
        },
        {
          "id": 2817013,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2024-05-16T16:48:23.543000",
          "content": "<p>I'll investigate all of the scoring errors, and rescore, if necessary.</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2817015,
              "author_name": "AC",
              "author_url": "",
              "post_date": "2024-05-16T16:48:55.340000",
              "content": "<p>Thanks! :)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2817276,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-05-16T19:49:02.600000",
              "content": "<p>i have many re-scoring error and many pending (even after 3 hours from the metric change is made)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc6c9e4b77d99de340560c428edfaac86%2FSelection_128.png?generation=1715888923935188&amp;alt=media\"></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2817307,
              "author_name": "Steven_Y",
              "author_url": "",
              "post_date": "2024-05-16T20:11:55.327000",
              "content": "<p>Yeah same here</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2817326,
              "author_name": "HungryLearner",
              "author_url": "",
              "post_date": "2024-05-16T20:20:59.317000",
              "content": "<p>It appears the rescoring error is general. I have many of it at as well</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2817353,
              "author_name": "inversion",
              "author_url": "",
              "post_date": "2024-05-16T20:57:22.260000",
              "content": "<p>Yeah, there's something wrong on the eng side. Meeting with them in a few minutes to discuss.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2817509,
              "author_name": "AC",
              "author_url": "",
              "post_date": "2024-05-17T00:48:53.403000",
              "content": "<p>Update: For me, the rescoring error is resolved. As of now, 2 of my submissions are still processing but all the rest that were throwing internal scoring error have been submitted successfully. Thank you! :)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2816975,
      "author_name": "AC",
      "author_url": "",
      "post_date": "2024-05-16T16:23:54.927000",
      "content": "<p>Thanks to everybody who has put in the efforts. Appreciated :) <br>\n2/3rd is huge, focusing on better generalization is the goal here. 😄</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2861667,
      "author_name": "AngelatSwezey",
      "author_url": "",
      "post_date": "2024-06-08T09:52:56.237000",
      "content": "<p>Thanks to everybody who has put in the efforts. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2819579,
      "author_name": "Sheema Zain",
      "author_url": "",
      "post_date": "2024-05-17T15:34:17.827000",
      "content": "<p>Thanks to everybody who has put in the efforts. Appreciated :)!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2859405,
      "author_name": "Bakery",
      "author_url": "",
      "post_date": "2024-06-07T02:50:55.213000",
      "content": "<p>Thanks to everybody who has put in the efforts. </p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2909581,
      "author_name": "Terry Park",
      "author_url": "",
      "post_date": "2024-07-07T05:43:58.333000",
      "content": "<p>Thanks for this competition. I learned a lot.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2899971,
      "author_name": "Ziyue Jiang0322",
      "author_url": "",
      "post_date": "2024-07-02T02:03:50.293000",
      "content": "<p>good but can be more</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2893684,
      "author_name": "Yash. S",
      "author_url": "",
      "post_date": "2024-06-28T02:26:57.650000",
      "content": "<p>I know it is a bit late in the game, and I am just trying to learn. But when I try to copy &amp; edit this notebook I am running in to several errors, currently got stuck with a 'DEPRECATION WARNING: USE MORGAN GENERATOR', and it never finishes running. I tried adding a warning ignore snippet but with no luck so far. Also needed to comment to become a contributor to get into a team! Thanks for reading.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2889664,
      "author_name": "sciuromorphsm",
      "author_url": "",
      "post_date": "2024-06-25T16:00:21.323000",
      "content": "<p>Thanks to everybody who has put in the efforts.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2879978,
      "author_name": "Colin Wu",
      "author_url": "",
      "post_date": "2024-06-20T00:47:33.933000",
      "content": "<p>Is there any relationship among the three proteins should be taken into consideration？</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2870298,
      "author_name": "PaulFAClarke",
      "author_url": "",
      "post_date": "2024-06-13T13:57:16.783000",
      "content": "<p>Hi there,</p>\n<p>Just looking for clarification here - </p>\n<p>Does \"Shared - bb\" mean that <em>all or some</em> of the building blocks in the test set are seen in the training set?</p>\n<p>Does \"Non-shared - bb\" mean that <em>all or some</em> of the building blocks in the test set are not seen in the training set?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2870394,
          "author_name": "Steven_Y",
          "author_url": "",
          "post_date": "2024-06-13T15:02:35.180000",
          "content": "<p>I think <strong>all</strong> shared bb in the test test are seen in the training set and <strong>all</strong> non-shared bb are not seen in the training set</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2871541,
          "author_name": "Hideaki Ogasawara",
          "author_url": "",
          "post_date": "2024-06-14T09:08:22.760000",
          "content": "<p>The test dataset comprises 878,022 unique molecules. Their building blocks 1 - 3 are either all shared or not shared at all. Additionally, all the molecules with non-triazine cores are made of non-shared building blocks. (my analysis <a href=\"https://www.kaggle.com/code/hideakiogasawara/chemspace-visualization\" target=\"_blank\">here</a>)</p>\n<table>\n<thead>\n<tr>\n<th>BB1 shared</th>\n<th>BB2 shared</th>\n<th>BB3s hared</th>\n<th>Triazine</th>\n<th>Number</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>True</td>\n<td>True</td>\n<td>True</td>\n<td>True</td>\n<td>369,039</td>\n</tr>\n<tr>\n<td>False</td>\n<td>False</td>\n<td>False</td>\n<td>True</td>\n<td>22,593</td>\n</tr>\n<tr>\n<td>False</td>\n<td>False</td>\n<td>False</td>\n<td>False</td>\n<td>486,390</td>\n</tr>\n<tr>\n<td></td>\n<td></td>\n<td></td>\n<td><strong>Total</strong></td>\n<td>878,022</td>\n</tr>\n</tbody>\n</table>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2868732,
      "author_name": "joan pedragosa ai",
      "author_url": "",
      "post_date": "2024-06-12T15:28:21.923000",
      "content": "<p>Hi,</p>\n<p>I can't find new library required for Private Metric . Can anyone tell me where it is?</p>\n<p>Thank's.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2842527,
      "author_name": "Bernhard Rohde",
      "author_url": "",
      "post_date": "2024-05-29T06:11:17.670000",
      "content": "<p>Since we now know better the structure of training and public and private test datasets, it would probably also be interesting to learn about the exact procedure of hit calling in the subsets. E.g.: is a single read in the third enrichment sufficient to call a hit or is the percentage of hits constant and the read threshold is picked accordingly.</p>\n<p>This should help in scaling the predictions for the subsets, although it might possibly not affect the final score as the three sets are scored separately.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2824259,
      "author_name": "AC",
      "author_url": "",
      "post_date": "2024-05-19T16:57:05.033000",
      "content": "<p>I still couldn't understand the new library part.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2828186,
          "author_name": "henryzhaowong",
          "author_url": "",
          "post_date": "2024-05-21T22:58:01.397000",
          "content": "<p>me too.<br>\naccording to the illustration, it seems to be only two bbs.<br>\ncan anyone help clarify this?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2828232,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2024-05-22T01:12:09.653000",
              "content": "<p>38 cores -&gt; bb1. (I counted 36, in my post that's footnote 2, above, but either way). Then bb2 binds to bb1 via boronate (I think), and bb3 to bb1 via the other binder in their diagram. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2830010,
              "author_name": "henryzhaowong",
              "author_url": "",
              "post_date": "2024-05-23T00:36:32.100000",
              "content": "<p>thanks(✺ω✺)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2916597,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-11T06:02:56.717000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2886121,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-23T13:14:54.763000",
      "content": "<p>Thanks to everybody who has put in the efforts.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2862152,
      "author_name": "Luciean",
      "author_url": "",
      "post_date": "2024-06-08T15:43:53.563000",
      "content": "<p>Thanks for update</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2861876,
      "author_name": "repurika",
      "author_url": "",
      "post_date": "2024-06-08T12:13:19.243000",
      "content": "<p>Thanks for update!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2822142,
      "author_name": "Jayesh Kothavale",
      "author_url": "",
      "post_date": "2024-05-18T13:13:33.060000",
      "content": "<p>Thanks for the update.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2816896": "Dear Participants,\n\nWe hope you are enjoying the competition so far. With nearly 900 teams participating, the level of engagement and enthusiasm has been incredible! Your diligence has identified an issue with the scoring metric currently in use and we will be making adjustments to ensure a fair and accurate evaluation of your models.\n\n**Issue Description:**\n\nOur dataset, derived from DEL screens containing 133 million molecules against 3 proteins, was split using multiple techniques:\n\n1. **Reserved Building Blocks:** Some building blocks were reserved for the test set, ensuring they did not appear in the training set.\n2. **Random and Chemical Scaffold Separation:** Certain Murcko scaffolds were made exclusive to the test set, then we additionally sampled random molecules.\n3. **Proprietary Library Addition:** A proprietary chemical library was included solely in the private test set.\n\nEach of these sampling techniques resulted in different proportions of hits. When holding out building blocks, all molecules containing those pieces were removed from the training set. Given that single building blocks drive a lot of activity, we had to do some careful stratified sampling to get enough positive examples into the test sets, and this led to varying levels of difficulty and base rates across splits and proteins.\n\nA number of teams have pointed out these differences and their effects on scoring. Chemdatafarmer and Robert Hatch identified the presence of novel chemistry in the test set early in the competition ([1](https://www.kaggle.com/competitions/leash-BELKA/discussion/493294), [2](https://www.kaggle.com/competitions/leash-BELKA/discussion/496576)). Several competitors probed CV/LB results in a thread started by Darek Kłeczek ([3](https://www.kaggle.com/competitions/leash-BELKA/discussion/498983)). Jun Koda described the differences in known/unknown building block proportions in the sets more precisely ([4](https://www.kaggle.com/code/junkoda/unknowns-in-public-test-set)), leading to the discovery that tweaking calibration gave a private test score boost by Ruby ([5](https://www.kaggle.com/competitions/leash-BELKA/discussion/501901)). Hengck23 beautifully summarized the issue with the below diagrams from ([6](https://www.kaggle.com/competitions/leash-BELKA/discussion/501901#2806722)) and ([7](https://www.kaggle.com/competitions/leash-BELKA/discussion/501901#2807698)):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F57970%2F80570cda23eeca6845ea100bb0588a40%2FScreen%20Shot%202024-05-16%20at%209.31.41%20AM.png?generation=1715873422843117&alt=media)\n\nUsing a single average precision (AP) number as the metric inadvertently emphasized the importance of understanding these base rates, which was not our intention. Our goal is to focus on identifying the molecule representations and model architectures that provide the greatest generalization.\n\n**New Scoring Metric:**\n\nTo address this issue, we will calculate the average precision for each protein and each split group individually and then average these scores. This method ensures a balanced evaluation across different groups. The revised scoring metric is as follows:\n\n**Public Metric:**\nThe Average Precision will be calculated across 6 groups:\n* sEH\n  * Shared-bb\n  * Nonshared-bb\n* HSA\n  * Shared-bb\n  * Nonshared-bb\n* BRD4\n  * Shared-bb\n  * Nonshared-bb\n\n\nThose 6 values will be averaged with equal weighting into the final public score\n\n**Private Metric:**\n\nThe Average Precision will be calculated across 9 groups:\n* sEH\n  * Shared-bb\n  * Nonshared-bb\n  * New library\n* HSA\n  * Shared-bb\n  * Nonshared-bb\n  * New library\n* BRD4\n  * Shared-bb\n  * Nonshared-bb\n  * New library\n\nThose 9 values will be averaged with equal weighting into the final public score. This means that 2/3rds the final score will come from molecules outside the training distribution, and we know predicting out of distribution is a hard challenge here. Still, the kinds of physical interactions that drive binding (hydrogen bonds, shape, pi-stacking, Van der Waals, charge distribution, etc.) are universal. We hope that progress can be made towards generalizing.\n\n**Nefarious New Library**\nAs many have already noticed, we have added a set of molecules into the private set that are completely different from the triazine molecules in the test set.\n\nThis new library was constructed by attaching DNA to a trifunctional core, and reacting fmoc/boc protected amide bond formation on one side and a Suzuki reaction on the other side. We suspect predicting these will be difficult but being able to predict these will show that you truly have trained a model that generalizes across chemical space.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F57970%2F8021959caea9c31363f400870b5f65ac%2FScreen%20Shot%202024-05-16%20at%209.28.27%20AM.png?generation=1715873251490099&alt=media)\n\n**Implementation Timeline:**\n\nThese changes will be implemented immediately, and the leaderboard will be in a temporary state of flux as previous submissions are rescored. An announcement will be made when that process is complete. We appreciate your understanding and cooperation as we strive to maintain the integrity and fairness of the competition.\n\nThank you for your participation and dedication. We look forward to seeing your continued progress!\n\nBest regards,  \nLeash Biosciences\n",
    "2816964": "Thanks for the update.\n\nhere i highlight the the important changes from machine learning point of view. Kagglers should take note note of these to make sure your validation consider this!\n1. \"different proportions of hits\"\n2. \"single building blocks drive a lot of activity, .... stratified ....\"\n3. \"enough positive examples into the test sets\"\n4. \"identifying the molecule representations that provide the greatest generalization.\"\n5. \" This means that 2/3rds the final score will come from molecules outside the training distribution,\"\n6. \" 9 ... averaged with equal weighting\"\n",
    "2821136": "Overall this seems good. However, if it's equal 1/6th and 1/9th, then I am pretty concerned this could over-emphasize the smaller number of samples of \"non-shared bb\"? Would it be reasonable to have all the scores be weighted by number of samples in that group?\n\ne.g. Public LB, approximately :\n180,000 * shared sEH mean average precision + 11,000 * nonshared sEH + ...  + 180k * HSA share + 11k * HSA nonshare. All divided by total number of public LB samples.\n\nAnd similarly for the more even distribution of the 9 groups in private LB?",
    "2913491": "I've been trying to understand the public-private leaderboards shakeup: with this metric, in very simplified case, if we assume that most models failed to generalize onto non-triazine / new library molecules (and so score -> 0), a simple rule of thumb for the scores relations between the public and private would've have to be something like this:\n\n`Public LB ~= Private LB * 3 / 2`\n\nHowever it doesn't seem to be the case? So I'm starting to think - how were non-shared building blocks distributed between private and public parts of the test set?\nIf that's ok to ask, @andrewdblevins - was that uniform random, or were these two independent sets, so public and private parts had different non-shared building blocks - non-shared not only with the train set, but also non-shared between each other? Thanks!",
    "2817539": "Thank you for paying such close attention to the discussions and the courage to implement feedback and improve the competition on the fly!",
    "2816990": "Thanks for quick update! But I note some notebook has rescoring error, which succeed in the past. Do you have any hints on this? I think there is nothing else constraints except for binds value should lies in [0,1]?",
    "2816975": "Thanks to everybody who has put in the efforts. Appreciated :) \n2/3rd is huge, focusing on better generalization is the goal here. 😄",
    "2861667": "Thanks to everybody who has put in the efforts. ",
    "2819579": "Thanks to everybody who has put in the efforts. Appreciated :)!",
    "2859405": "Thanks to everybody who has put in the efforts. ",
    "2909581": "Thanks for this competition. I learned a lot.",
    "2899971": "good but can be more",
    "2893684": "I know it is a bit late in the game, and I am just trying to learn. But when I try to copy & edit this notebook I am running in to several errors, currently got stuck with a 'DEPRECATION WARNING: USE MORGAN GENERATOR', and it never finishes running. I tried adding a warning ignore snippet but with no luck so far. Also needed to comment to become a contributor to get into a team! Thanks for reading.",
    "2889664": "Thanks to everybody who has put in the efforts.",
    "2879978": "Is there any relationship among the three proteins should be taken into consideration？",
    "2870298": "Hi there,\n\nJust looking for clarification here - \n\nDoes \"Shared - bb\" mean that *all or some* of the building blocks in the test set are seen in the training set?\n\nDoes \"Non-shared - bb\" mean that *all or some* of the building blocks in the test set are not seen in the training set?",
    "2868732": "Hi,\n\nI can't find new library required for Private Metric . Can anyone tell me where it is?\n\nThank's.",
    "2842527": "Since we now know better the structure of training and public and private test datasets, it would probably also be interesting to learn about the exact procedure of hit calling in the subsets. E.g.: is a single read in the third enrichment sufficient to call a hit or is the percentage of hits constant and the read threshold is picked accordingly.\n\nThis should help in scaling the predictions for the subsets, although it might possibly not affect the final score as the three sets are scored separately.",
    "2824259": "I still couldn't understand the new library part.",
    "2916597": "",
    "2886121": "Thanks to everybody who has put in the efforts.",
    "2862152": "Thanks for update",
    "2861876": "Thanks for update!",
    "2822142": "Thanks for the update."
  }
}