{
  "id": 519068,
  "title": "Models performace variance between the leaderboards, visual explanation take",
  "url": "/competitions/leash-BELKA/discussion/519068",
  "author_name": "",
  "post_date": "2024-07-09T15:08:15.913571700Z",
  "votes": 7,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I've been contemplating on what caused such a shakeup in the leaderboards and, of course, most importantly what might've been done to produce more robust submissions?</p>\n<p>I found diagrams of the building blocks space helpful for my understanding, so decided to draw these rough sketches in case they might be helpful for someone else.</p>\n<p>So, if we abstract for a second from the target proteins (they could be imagined as channels / layers on these diagrams) and imagine some 2d space of the molecule building blocks, then train / public test / private test might land in that space something like that:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F3c63bb8adc5fb2354e772e7369a5043d%2Fbbs_space.drawio.png?generation=1720537526950444&amp;alt=media\" alt=\"bbs_space\"></p>\n<p>Where our test set would be a union of [1], [2], [3] and [4], but while public LB part of this set would be based on the union of [1] and [2], the private (and final) LB set seems to be the union of [1], [3] and [4]. It should probably already strongly suggest what caused such a shakeup, right :)</p>\n<p>The ideal model for this competition would've been something like this one, where dashed area is a rough depiction of bbs space that model mastered and makes good predictions in:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F4c9a234cdb8d10defc679b3348347975%2Fideal_model.drawio.png?generation=1720537551794763&amp;alt=media\" alt=\"ideal_model\"></p>\n<p>However given the complexity of the field, actual dashed areas were likely much smaller, barely reaching the non-triazine sector at all! This, combined with the fact that sets [2] and [3] are different ones, presents the issue: some models would generalize a little bit in one \"direction\", whilst other models - in other one. And so we saw some submissions / models score really high on the Public LB, but drop thousands of positions on the Private LB - these were likely models like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F69ab6bdaa66aa69fd06340f3379e9967%2Fmodel_high_public_low_private.drawio.png?generation=1720537578127954&amp;alt=media\" alt=\"high_public_low_private\"></p>\n<p>Other models, on opposite, gained quite a bit - jumping high up between the Public and Private. These could've been models like that:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F843cfd992d5e0a878d78eabeaa84df81%2Fmodel_low_public_high_private.drawio.png?generation=1720537598741299&amp;alt=media\" alt=\"low_public_high_private\"></p>\n<p>The issue with such models, of course, is that there was no way of knowing reliably whether your model is generalizes in the set [3] direction, or if it's a model with poor generalization / overfitted on train or [1]. The only source of additional signal was Public LB, so performance on [2] - but, of course, that only separated models that overfitted on train from models that generalize in the direction of [2], not indicative of generalization on [3].</p>\n<p>In other words, main question here - how to \"steer\" the model into \"right\" direction in this bbs space (be it direction of set [3] to win the competition or just direction of more useful molecules, for example), while there's no data to guide that? My best attempt to answer that would be that we had to invest more time into better local validation (or cross-validation) setup - ideally using some chemical similarity, trying to carve the part of the train set most closest to the building blocks in the test set, something like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2Fb39bb7a7fadd6bc34cd8b380f10e33f7%2Fideal_cv.drawio.png?generation=1720537686296622&amp;alt=media\" alt=\"ideal_cv\"></p>\n<p>So, I would be really curious to hear if someone attempted some nuanced approaches to validation like that, and if you had any luck? Have any of the top-scoring submissions on the private LB had any tactics that helped? Please let us know!</p>\n<p>Technically, this dataset is a monumental contribution - and whilst the competition officially ended, it can be used to tune different molecules representations / embeddings / etc, so I want to join saying a massive thank you to the organizers - and I am looking forward where it all takes us! </p>",
  "messages": [
    {
      "id": "2913591",
      "postDate": "07/09/2024 15:08:15",
      "content": "<p>I've been contemplating on what caused such a shakeup in the leaderboards and, of course, most importantly what might've been done to produce more robust submissions?</p>\n<p>I found diagrams of the building blocks space helpful for my understanding, so decided to draw these rough sketches in case they might be helpful for someone else.</p>\n<p>So, if we abstract for a second from the target proteins (they could be imagined as channels / layers on these diagrams) and imagine some 2d space of the molecule building blocks, then train / public test / private test might land in that space something like that:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F3c63bb8adc5fb2354e772e7369a5043d%2Fbbs_space.drawio.png?generation=1720537526950444&amp;alt=media\" alt=\"bbs_space\"></p>\n<p>Where our test set would be a union of [1], [2], [3] and [4], but while public LB part of this set would be based on the union of [1] and [2], the private (and final) LB set seems to be the union of [1], [3] and [4]. It should probably already strongly suggest what caused such a shakeup, right :)</p>\n<p>The ideal model for this competition would've been something like this one, where dashed area is a rough depiction of bbs space that model mastered and makes good predictions in:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F4c9a234cdb8d10defc679b3348347975%2Fideal_model.drawio.png?generation=1720537551794763&amp;alt=media\" alt=\"ideal_model\"></p>\n<p>However given the complexity of the field, actual dashed areas were likely much smaller, barely reaching the non-triazine sector at all! This, combined with the fact that sets [2] and [3] are different ones, presents the issue: some models would generalize a little bit in one \"direction\", whilst other models - in other one. And so we saw some submissions / models score really high on the Public LB, but drop thousands of positions on the Private LB - these were likely models like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F69ab6bdaa66aa69fd06340f3379e9967%2Fmodel_high_public_low_private.drawio.png?generation=1720537578127954&amp;alt=media\" alt=\"high_public_low_private\"></p>\n<p>Other models, on opposite, gained quite a bit - jumping high up between the Public and Private. These could've been models like that:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F843cfd992d5e0a878d78eabeaa84df81%2Fmodel_low_public_high_private.drawio.png?generation=1720537598741299&amp;alt=media\" alt=\"low_public_high_private\"></p>\n<p>The issue with such models, of course, is that there was no way of knowing reliably whether your model is generalizes in the set [3] direction, or if it's a model with poor generalization / overfitted on train or [1]. The only source of additional signal was Public LB, so performance on [2] - but, of course, that only separated models that overfitted on train from models that generalize in the direction of [2], not indicative of generalization on [3].</p>\n<p>In other words, main question here - how to \"steer\" the model into \"right\" direction in this bbs space (be it direction of set [3] to win the competition or just direction of more useful molecules, for example), while there's no data to guide that? My best attempt to answer that would be that we had to invest more time into better local validation (or cross-validation) setup - ideally using some chemical similarity, trying to carve the part of the train set most closest to the building blocks in the test set, something like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2Fb39bb7a7fadd6bc34cd8b380f10e33f7%2Fideal_cv.drawio.png?generation=1720537686296622&amp;alt=media\" alt=\"ideal_cv\"></p>\n<p>So, I would be really curious to hear if someone attempted some nuanced approaches to validation like that, and if you had any luck? Have any of the top-scoring submissions on the private LB had any tactics that helped? Please let us know!</p>\n<p>Technically, this dataset is a monumental contribution - and whilst the competition officially ended, it can be used to tune different molecules representations / embeddings / etc, so I want to join saying a massive thank you to the organizers - and I am looking forward where it all takes us! </p>",
      "rawMarkdown": "I've been contemplating on what caused such a shakeup in the leaderboards and, of course, most importantly what might've been done to produce more robust submissions?\n\nI found diagrams of the building blocks space helpful for my understanding, so decided to draw these rough sketches in case they might be helpful for someone else.\n\nSo, if we abstract for a second from the target proteins (they could be imagined as channels / layers on these diagrams) and imagine some 2d space of the molecule building blocks, then train / public test / private test might land in that space something like that:\n![bbs_space](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F3c63bb8adc5fb2354e772e7369a5043d%2Fbbs_space.drawio.png?generation=1720537526950444&alt=media)\n\n\nWhere our test set would be a union of [1], [2], [3] and [4], but while public LB part of this set would be based on the union of [1] and [2], the private (and final) LB set seems to be the union of [1], [3] and [4]. It should probably already strongly suggest what caused such a shakeup, right :)\n\nThe ideal model for this competition would've been something like this one, where dashed area is a rough depiction of bbs space that model mastered and makes good predictions in:\n![ideal_model](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F4c9a234cdb8d10defc679b3348347975%2Fideal_model.drawio.png?generation=1720537551794763&alt=media)\n\nHowever given the complexity of the field, actual dashed areas were likely much smaller, barely reaching the non-triazine sector at all! This, combined with the fact that sets [2] and [3] are different ones, presents the issue: some models would generalize a little bit in one \"direction\", whilst other models - in other one. And so we saw some submissions / models score really high on the Public LB, but drop thousands of positions on the Private LB - these were likely models like this:\n![high_public_low_private](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F69ab6bdaa66aa69fd06340f3379e9967%2Fmodel_high_public_low_private.drawio.png?generation=1720537578127954&alt=media)\n\nOther models, on opposite, gained quite a bit - jumping high up between the Public and Private. These could've been models like that:\n![low_public_high_private](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F843cfd992d5e0a878d78eabeaa84df81%2Fmodel_low_public_high_private.drawio.png?generation=1720537598741299&alt=media)\n\nThe issue with such models, of course, is that there was no way of knowing reliably whether your model is generalizes in the set [3] direction, or if it's a model with poor generalization / overfitted on train or [1]. The only source of additional signal was Public LB, so performance on [2] - but, of course, that only separated models that overfitted on train from models that generalize in the direction of [2], not indicative of generalization on [3].\n\nIn other words, main question here - how to \"steer\" the model into \"right\" direction in this bbs space (be it direction of set [3] to win the competition or just direction of more useful molecules, for example), while there's no data to guide that? My best attempt to answer that would be that we had to invest more time into better local validation (or cross-validation) setup - ideally using some chemical similarity, trying to carve the part of the train set most closest to the building blocks in the test set, something like this:\n![ideal_cv](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2Fb39bb7a7fadd6bc34cd8b380f10e33f7%2Fideal_cv.drawio.png?generation=1720537686296622&alt=media)\n\n\nSo, I would be really curious to hear if someone attempted some nuanced approaches to validation like that, and if you had any luck? Have any of the top-scoring submissions on the private LB had any tactics that helped? Please let us know!\n\nTechnically, this dataset is a monumental contribution - and whilst the competition officially ended, it can be used to tune different molecules representations / embeddings / etc, so I want to join saying a massive thank you to the organizers - and I am looking forward where it all takes us!",
      "votes": null
    },
    {
      "id": "2913746",
      "postDate": "07/09/2024 16:12:37",
      "content": "<p>Pretty vizualization! </p>\n<blockquote>\n  <p>ideally using some chemical similarity, trying to carve the part of the train set most closest to the building blocks in the test set</p>\n</blockquote>\n<p>I actually have some doubts about this, as I have done several analyses that lead me to conclude that structural similarity has no relationship with the binding capabilities of molecules (at least as captured by different molecular fingerprints, <a href=\"https://www.kaggle.com/code/antoninadolgorukova/belka-bb-similarity\" target=\"_blank\">this notebook</a>), nor with the quality of the model predictions (at least for boostings, <a href=\"https://www.kaggle.com/code/antoninadolgorukova/belka-generalizability-similarity#Generalizability-of-a-model\" target=\"_blank\">this notebook</a>). If you have suggestions for other methods to verify this, please let me know, I'd be happy to investigate this further!</p>\n<p>Also, for us there seems to be a good correlation between private and public scores for all triazine molecules, including those with non-shared BBs. We tend to think that it is true not only for our submits and the huge shakeup is due to small differences in the non-triazine score, which seem to be unpredictable at all and thus just a random number for most teams, but contribute to a third of the total score.</p>",
      "rawMarkdown": "Pretty vizualization! \n\n>ideally using some chemical similarity, trying to carve the part of the train set most closest to the building blocks in the test set\n\nI actually have some doubts about this, as I have done several analyses that lead me to conclude that structural similarity has no relationship with the binding capabilities of molecules (at least as captured by different molecular fingerprints, [this notebook](https://www.kaggle.com/code/antoninadolgorukova/belka-bb-similarity)), nor with the quality of the model predictions (at least for boostings, [this notebook](https://www.kaggle.com/code/antoninadolgorukova/belka-generalizability-similarity#Generalizability-of-a-model)). If you have suggestions for other methods to verify this, please let me know, I'd be happy to investigate this further!\n\nAlso, for us there seems to be a good correlation between private and public scores for all triazine molecules, including those with non-shared BBs. We tend to think that it is true not only for our submits and the huge shakeup is due to small differences in the non-triazine score, which seem to be unpredictable at all and thus just a random number for most teams, but contribute to a third of the total score.",
      "votes": null
    },
    {
      "id": "2913806",
      "postDate": "07/09/2024 16:31:26",
      "content": "<blockquote>\n  <p>structural similarity has no relationship with the binding capabilities of molecules (at least as captured by different molecular fingerprints, this notebook), nor with the quality of the model predictions (at least for boostings, this notebook)</p>\n</blockquote>\n<p>Thank you for sharing these! I'll have to dive deeper - but first wanted to elaborate a little that I've been speculating that bbs similarity could've been useful for the construction of the local validation set, not necessarily as a feature for the model! The reason for this though is that my local scores poorly translated into the test scores - which seems like an indication that the bbs that I used as a holdout from training might've been \"steering\" the model in the wrong direction (ex down on the pictures in the post, away from test set bbs). So I am curious if with more nuances selection of bbs to run the local validation - was it possible to have a better correlation with the test set (and private LB part of the test set specifically)?</p>\n<blockquote>\n  <p>for us there seems to be a good correlation between private and public scores for all triazine molecules, including those with non-shared BBs. </p>\n</blockquote>\n<p>That just because you folks have an awesome model :) Maybe not reaching the non-triazine region, yes - but expanding / generalizing beyond the train in all directions, sufficiently to cover bbs both in public and private parts of the test!</p>\n<blockquote>\n  <p>the huge shakeup is due to small differences in the non-triazine score</p>\n</blockquote>\n<p>Here where my intuition differs - from the posts in the discussions, and from my probing, my theory that non-triazine is -&gt; 0 for most models, so it's effect is closer to be a <strong>constant <code>2/3</code> multiplier</strong>, with minor variance introduction. The fact that private test triazine non-shared bbs were different from public test non-shared bbs - that what I think introduced the shakeup!</p>",
      "rawMarkdown": "> structural similarity has no relationship with the binding capabilities of molecules (at least as captured by different molecular fingerprints, this notebook), nor with the quality of the model predictions (at least for boostings, this notebook)\n\nThank you for sharing these! I'll have to dive deeper - but first wanted to elaborate a little that I've been speculating that bbs similarity could've been useful for the construction of the local validation set, not necessarily as a feature for the model! The reason for this though is that my local scores poorly translated into the test scores - which seems like an indication that the bbs that I used as a holdout from training might've been \"steering\" the model in the wrong direction (ex down on the pictures in the post, away from test set bbs). So I am curious if with more nuances selection of bbs to run the local validation - was it possible to have a better correlation with the test set (and private LB part of the test set specifically)?\n\n> for us there seems to be a good correlation between private and public scores for all triazine molecules, including those with non-shared BBs. \n\nThat just because you folks have an awesome model :) Maybe not reaching the non-triazine region, yes - but expanding / generalizing beyond the train in all directions, sufficiently to cover bbs both in public and private parts of the test!\n\n\n> the huge shakeup is due to small differences in the non-triazine score\n\nHere where my intuition differs - from the posts in the discussions, and from my probing, my theory that non-triazine is -> 0 for most models, so it's effect is closer to be a **constant `2/3` multiplier**, with minor variance introduction. The fact that private test triazine non-shared bbs were different from public test non-shared bbs - that what I think introduced the shakeup!",
      "votes": null
    },
    {
      "id": "2913868",
      "postDate": "07/09/2024 16:48:44",
      "content": "<p>Thank you)</p>\n<blockquote>\n  <p>The fact that private test triazine non-shared bbs were different from public test non-shared bbs</p>\n</blockquote>\n<p>Yep, I just caught that)) But even if they were similar, I've written myself before (<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/509015#2848577\" target=\"_blank\">here</a>) that with such small sample sizes neither the public nor the private LB scores for the non-shared part can accurately reflect the true performance of the model.</p>\n<p>Even if we found a model that generalizes well, we might get randomly biased scores on the public/private tests due to its small size and specific classes balance for each target. This is just what happened + small random variations in non-triazine scores.</p>",
      "rawMarkdown": "Thank you)\n\n> The fact that private test triazine non-shared bbs were different from public test non-shared bbs\n\nYep, I just caught that)) But even if they were similar, I've written myself before ([here](https://www.kaggle.com/competitions/leash-BELKA/discussion/509015#2848577)) that with such small sample sizes neither the public nor the private LB scores for the non-shared part can accurately reflect the true performance of the model.\n\nEven if we found a model that generalizes well, we might get randomly biased scores on the public/private tests due to its small size and specific classes balance for each target. This is just what happened + small random variations in non-triazine scores.",
      "votes": null
    },
    {
      "id": "2913931",
      "postDate": "07/09/2024 17:30:26",
      "content": "<p>Hmm, and about this</p>\n<blockquote>\n  <p>with more nuances selection of bbs to run the local validation - was it possible to have a better correlation with the test set  (and private LB part of the test set specifically)?</p>\n</blockquote>\n<p>I did some experiments with fast XGB, trained it on different subsets, each without 20% of random BBs, and with validation on subsets with shared and nonshared BBs. The local scores were highly variable form fold to fold. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F5e6db8bf0458e92b0909fffeafb5e741%2FScreenshot%202024-07-09%20200441.png?generation=1720544703496699&amp;alt=media\"></p>\n<p>But I got your idea, I will check what the difference will be if for validation use folds e.g. most similar to test non-shared BBs.</p>",
      "rawMarkdown": "Hmm, and about this\n\n>with more nuances selection of bbs to run the local validation - was it possible to have a better correlation with the test set  (and private LB part of the test set specifically)?\n\nI did some experiments with fast XGB, trained it on different subsets, each without 20% of random BBs, and with validation on subsets with shared and nonshared BBs. The local scores were highly variable form fold to fold. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F5e6db8bf0458e92b0909fffeafb5e741%2FScreenshot%202024-07-09%20200441.png?generation=1720544703496699&alt=media)\n\nBut I got your idea, I will check what the difference will be if for validation use folds e.g. most similar to test non-shared BBs.",
      "votes": null
    },
    {
      "id": "2914240",
      "postDate": "07/09/2024 19:29:02",
      "content": "<p>Beautiful visualization! Easy to understand, thanks for sharing.</p>\n<p>People in this competition mostly used DL models trained on SMILES features. Since they focus on SMILES, the fitting is primarily based on building blocks.  I'm guessing if people use structural info (3d conformer, pharmocophore), it's possible to see the generalization as that can ignore the building blocks. </p>",
      "rawMarkdown": "Beautiful visualization! Easy to understand, thanks for sharing.\n\nPeople in this competition mostly used DL models trained on SMILES features. Since they focus on SMILES, the fitting is primarily based on building blocks.  I'm guessing if people use structural info (3d conformer, pharmocophore), it's possible to see the generalization as that can ignore the building blocks.",
      "votes": null
    },
    {
      "id": "2914245",
      "postDate": "07/09/2024 19:37:37",
      "content": "<p>Thank you for sharing this thought. I'm very surprised by the significant changes in this competition. If I understand correctly, the ideal model for this problem would produce predictions in [1], [2], [3], and [4]. May I infer that the final results should also consider the performance of [2], not just [3] on the private leaderboard? This approach would be fairer to the Kaggle participants, in my opinion.</p>",
      "rawMarkdown": "Thank you for sharing this thought. I'm very surprised by the significant changes in this competition. If I understand correctly, the ideal model for this problem would produce predictions in [1], [2], [3], and [4]. May I infer that the final results should also consider the performance of [2], not just [3] on the private leaderboard? This approach would be fairer to the Kaggle participants, in my opinion.",
      "votes": null
    },
    {
      "id": "2914391",
      "postDate": "07/09/2024 22:32:40",
      "content": "<p>That's the thing! The private leaderboard part of the test appears to be solely [1], [3] and [4] - nothing from [2], as measure to prevent probing through the submissions to the public leaderboard, see: <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/503232#2913518\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/discussion/503232#2913518</a></p>",
      "rawMarkdown": "That's the thing! The private leaderboard part of the test appears to be solely [1], [3] and [4] - nothing from [2], as measure to prevent probing through the submissions to the public leaderboard, see: https://www.kaggle.com/competitions/leash-BELKA/discussion/503232#2913518",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2913746,
      "author_name": "antoninadolgorukova",
      "author_url": "",
      "post_date": "07/09/2024 16:12:37",
      "content": "<p>Pretty vizualization! </p>\n<blockquote>\n  <p>ideally using some chemical similarity, trying to carve the part of the train set most closest to the building blocks in the test set</p>\n</blockquote>\n<p>I actually have some doubts about this, as I have done several analyses that lead me to conclude that structural similarity has no relationship with the binding capabilities of molecules (at least as captured by different molecular fingerprints, <a href=\"https://www.kaggle.com/code/antoninadolgorukova/belka-bb-similarity\" target=\"_blank\">this notebook</a>), nor with the quality of the model predictions (at least for boostings, <a href=\"https://www.kaggle.com/code/antoninadolgorukova/belka-generalizability-similarity#Generalizability-of-a-model\" target=\"_blank\">this notebook</a>). If you have suggestions for other methods to verify this, please let me know, I'd be happy to investigate this further!</p>\n<p>Also, for us there seems to be a good correlation between private and public scores for all triazine molecules, including those with non-shared BBs. We tend to think that it is true not only for our submits and the huge shakeup is due to small differences in the non-triazine score, which seem to be unpredictable at all and thus just a random number for most teams, but contribute to a third of the total score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2913806,
          "author_name": "pershinmr",
          "author_url": "",
          "post_date": "07/09/2024 16:31:26",
          "content": "<blockquote>\n  <p>structural similarity has no relationship with the binding capabilities of molecules (at least as captured by different molecular fingerprints, this notebook), nor with the quality of the model predictions (at least for boostings, this notebook)</p>\n</blockquote>\n<p>Thank you for sharing these! I'll have to dive deeper - but first wanted to elaborate a little that I've been speculating that bbs similarity could've been useful for the construction of the local validation set, not necessarily as a feature for the model! The reason for this though is that my local scores poorly translated into the test scores - which seems like an indication that the bbs that I used as a holdout from training might've been \"steering\" the model in the wrong direction (ex down on the pictures in the post, away from test set bbs). So I am curious if with more nuances selection of bbs to run the local validation - was it possible to have a better correlation with the test set (and private LB part of the test set specifically)?</p>\n<blockquote>\n  <p>for us there seems to be a good correlation between private and public scores for all triazine molecules, including those with non-shared BBs. </p>\n</blockquote>\n<p>That just because you folks have an awesome model :) Maybe not reaching the non-triazine region, yes - but expanding / generalizing beyond the train in all directions, sufficiently to cover bbs both in public and private parts of the test!</p>\n<blockquote>\n  <p>the huge shakeup is due to small differences in the non-triazine score</p>\n</blockquote>\n<p>Here where my intuition differs - from the posts in the discussions, and from my probing, my theory that non-triazine is -&gt; 0 for most models, so it's effect is closer to be a <strong>constant <code>2/3</code> multiplier</strong>, with minor variance introduction. The fact that private test triazine non-shared bbs were different from public test non-shared bbs - that what I think introduced the shakeup!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2913868,
              "author_name": "antoninadolgorukova",
              "author_url": "",
              "post_date": "07/09/2024 16:48:44",
              "content": "<p>Thank you)</p>\n<blockquote>\n  <p>The fact that private test triazine non-shared bbs were different from public test non-shared bbs</p>\n</blockquote>\n<p>Yep, I just caught that)) But even if they were similar, I've written myself before (<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/509015#2848577\" target=\"_blank\">here</a>) that with such small sample sizes neither the public nor the private LB scores for the non-shared part can accurately reflect the true performance of the model.</p>\n<p>Even if we found a model that generalizes well, we might get randomly biased scores on the public/private tests due to its small size and specific classes balance for each target. This is just what happened + small random variations in non-triazine scores.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2913931,
              "author_name": "antoninadolgorukova",
              "author_url": "",
              "post_date": "07/09/2024 17:30:26",
              "content": "<p>Hmm, and about this</p>\n<blockquote>\n  <p>with more nuances selection of bbs to run the local validation - was it possible to have a better correlation with the test set  (and private LB part of the test set specifically)?</p>\n</blockquote>\n<p>I did some experiments with fast XGB, trained it on different subsets, each without 20% of random BBs, and with validation on subsets with shared and nonshared BBs. The local scores were highly variable form fold to fold. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F5e6db8bf0458e92b0909fffeafb5e741%2FScreenshot%202024-07-09%20200441.png?generation=1720544703496699&amp;alt=media\"></p>\n<p>But I got your idea, I will check what the difference will be if for validation use folds e.g. most similar to test non-shared BBs.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2914240,
      "author_name": "lililycai",
      "author_url": "",
      "post_date": "07/09/2024 19:29:02",
      "content": "<p>Beautiful visualization! Easy to understand, thanks for sharing.</p>\n<p>People in this competition mostly used DL models trained on SMILES features. Since they focus on SMILES, the fitting is primarily based on building blocks.  I'm guessing if people use structural info (3d conformer, pharmocophore), it's possible to see the generalization as that can ignore the building blocks. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2914245,
      "author_name": "yukifuji9",
      "author_url": "",
      "post_date": "07/09/2024 19:37:37",
      "content": "<p>Thank you for sharing this thought. I'm very surprised by the significant changes in this competition. If I understand correctly, the ideal model for this problem would produce predictions in [1], [2], [3], and [4]. May I infer that the final results should also consider the performance of [2], not just [3] on the private leaderboard? This approach would be fairer to the Kaggle participants, in my opinion.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2914391,
          "author_name": "pershinmr",
          "author_url": "",
          "post_date": "07/09/2024 22:32:40",
          "content": "<p>That's the thing! The private leaderboard part of the test appears to be solely [1], [3] and [4] - nothing from [2], as measure to prevent probing through the submissions to the public leaderboard, see: <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/503232#2913518\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/discussion/503232#2913518</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2913591": "I've been contemplating on what caused such a shakeup in the leaderboards and, of course, most importantly what might've been done to produce more robust submissions?\n\nI found diagrams of the building blocks space helpful for my understanding, so decided to draw these rough sketches in case they might be helpful for someone else.\n\nSo, if we abstract for a second from the target proteins (they could be imagined as channels / layers on these diagrams) and imagine some 2d space of the molecule building blocks, then train / public test / private test might land in that space something like that:\n![bbs_space](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F3c63bb8adc5fb2354e772e7369a5043d%2Fbbs_space.drawio.png?generation=1720537526950444&alt=media)\n\n\nWhere our test set would be a union of [1], [2], [3] and [4], but while public LB part of this set would be based on the union of [1] and [2], the private (and final) LB set seems to be the union of [1], [3] and [4]. It should probably already strongly suggest what caused such a shakeup, right :)\n\nThe ideal model for this competition would've been something like this one, where dashed area is a rough depiction of bbs space that model mastered and makes good predictions in:\n![ideal_model](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F4c9a234cdb8d10defc679b3348347975%2Fideal_model.drawio.png?generation=1720537551794763&alt=media)\n\nHowever given the complexity of the field, actual dashed areas were likely much smaller, barely reaching the non-triazine sector at all! This, combined with the fact that sets [2] and [3] are different ones, presents the issue: some models would generalize a little bit in one \"direction\", whilst other models - in other one. And so we saw some submissions / models score really high on the Public LB, but drop thousands of positions on the Private LB - these were likely models like this:\n![high_public_low_private](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F69ab6bdaa66aa69fd06340f3379e9967%2Fmodel_high_public_low_private.drawio.png?generation=1720537578127954&alt=media)\n\nOther models, on opposite, gained quite a bit - jumping high up between the Public and Private. These could've been models like that:\n![low_public_high_private](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F843cfd992d5e0a878d78eabeaa84df81%2Fmodel_low_public_high_private.drawio.png?generation=1720537598741299&alt=media)\n\nThe issue with such models, of course, is that there was no way of knowing reliably whether your model is generalizes in the set [3] direction, or if it's a model with poor generalization / overfitted on train or [1]. The only source of additional signal was Public LB, so performance on [2] - but, of course, that only separated models that overfitted on train from models that generalize in the direction of [2], not indicative of generalization on [3].\n\nIn other words, main question here - how to \"steer\" the model into \"right\" direction in this bbs space (be it direction of set [3] to win the competition or just direction of more useful molecules, for example), while there's no data to guide that? My best attempt to answer that would be that we had to invest more time into better local validation (or cross-validation) setup - ideally using some chemical similarity, trying to carve the part of the train set most closest to the building blocks in the test set, something like this:\n![ideal_cv](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2Fb39bb7a7fadd6bc34cd8b380f10e33f7%2Fideal_cv.drawio.png?generation=1720537686296622&alt=media)\n\n\nSo, I would be really curious to hear if someone attempted some nuanced approaches to validation like that, and if you had any luck? Have any of the top-scoring submissions on the private LB had any tactics that helped? Please let us know!\n\nTechnically, this dataset is a monumental contribution - and whilst the competition officially ended, it can be used to tune different molecules representations / embeddings / etc, so I want to join saying a massive thank you to the organizers - and I am looking forward where it all takes us!",
    "2913746": "Pretty vizualization! \n\n>ideally using some chemical similarity, trying to carve the part of the train set most closest to the building blocks in the test set\n\nI actually have some doubts about this, as I have done several analyses that lead me to conclude that structural similarity has no relationship with the binding capabilities of molecules (at least as captured by different molecular fingerprints, [this notebook](https://www.kaggle.com/code/antoninadolgorukova/belka-bb-similarity)), nor with the quality of the model predictions (at least for boostings, [this notebook](https://www.kaggle.com/code/antoninadolgorukova/belka-generalizability-similarity#Generalizability-of-a-model)). If you have suggestions for other methods to verify this, please let me know, I'd be happy to investigate this further!\n\nAlso, for us there seems to be a good correlation between private and public scores for all triazine molecules, including those with non-shared BBs. We tend to think that it is true not only for our submits and the huge shakeup is due to small differences in the non-triazine score, which seem to be unpredictable at all and thus just a random number for most teams, but contribute to a third of the total score.",
    "2913806": "> structural similarity has no relationship with the binding capabilities of molecules (at least as captured by different molecular fingerprints, this notebook), nor with the quality of the model predictions (at least for boostings, this notebook)\n\nThank you for sharing these! I'll have to dive deeper - but first wanted to elaborate a little that I've been speculating that bbs similarity could've been useful for the construction of the local validation set, not necessarily as a feature for the model! The reason for this though is that my local scores poorly translated into the test scores - which seems like an indication that the bbs that I used as a holdout from training might've been \"steering\" the model in the wrong direction (ex down on the pictures in the post, away from test set bbs). So I am curious if with more nuances selection of bbs to run the local validation - was it possible to have a better correlation with the test set (and private LB part of the test set specifically)?\n\n> for us there seems to be a good correlation between private and public scores for all triazine molecules, including those with non-shared BBs. \n\nThat just because you folks have an awesome model :) Maybe not reaching the non-triazine region, yes - but expanding / generalizing beyond the train in all directions, sufficiently to cover bbs both in public and private parts of the test!\n\n\n> the huge shakeup is due to small differences in the non-triazine score\n\nHere where my intuition differs - from the posts in the discussions, and from my probing, my theory that non-triazine is -> 0 for most models, so it's effect is closer to be a **constant `2/3` multiplier**, with minor variance introduction. The fact that private test triazine non-shared bbs were different from public test non-shared bbs - that what I think introduced the shakeup!",
    "2913868": "Thank you)\n\n> The fact that private test triazine non-shared bbs were different from public test non-shared bbs\n\nYep, I just caught that)) But even if they were similar, I've written myself before ([here](https://www.kaggle.com/competitions/leash-BELKA/discussion/509015#2848577)) that with such small sample sizes neither the public nor the private LB scores for the non-shared part can accurately reflect the true performance of the model.\n\nEven if we found a model that generalizes well, we might get randomly biased scores on the public/private tests due to its small size and specific classes balance for each target. This is just what happened + small random variations in non-triazine scores.",
    "2913931": "Hmm, and about this\n\n>with more nuances selection of bbs to run the local validation - was it possible to have a better correlation with the test set  (and private LB part of the test set specifically)?\n\nI did some experiments with fast XGB, trained it on different subsets, each without 20% of random BBs, and with validation on subsets with shared and nonshared BBs. The local scores were highly variable form fold to fold. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F5e6db8bf0458e92b0909fffeafb5e741%2FScreenshot%202024-07-09%20200441.png?generation=1720544703496699&alt=media)\n\nBut I got your idea, I will check what the difference will be if for validation use folds e.g. most similar to test non-shared BBs.",
    "2914240": "Beautiful visualization! Easy to understand, thanks for sharing.\n\nPeople in this competition mostly used DL models trained on SMILES features. Since they focus on SMILES, the fitting is primarily based on building blocks.  I'm guessing if people use structural info (3d conformer, pharmocophore), it's possible to see the generalization as that can ignore the building blocks.",
    "2914245": "Thank you for sharing this thought. I'm very surprised by the significant changes in this competition. If I understand correctly, the ideal model for this problem would produce predictions in [1], [2], [3], and [4]. May I infer that the final results should also consider the performance of [2], not just [3] on the private leaderboard? This approach would be fairer to the Kaggle participants, in my opinion.",
    "2914391": "That's the thing! The private leaderboard part of the test appears to be solely [1], [3] and [4] - nothing from [2], as measure to prevent probing through the submissions to the public leaderboard, see: https://www.kaggle.com/competitions/leash-BELKA/discussion/503232#2913518"
  },
  "source": "meta"
}