{
  "id": 523779,
  "title": "Average precision scores by protein and subgroup",
  "url": "/competitions/leash-BELKA/discussion/523779",
  "author_name": "",
  "post_date": "2024-08-02T16:35:40.911296700Z",
  "votes": 9,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Thanks to <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> for posting the <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/523644\" target=\"_blank\">answer key</a>.  </p>\n<p>It would be very interesting to see exactly how different models performed across proteins and subsets of the data.  I've created <a href=\"https://www.kaggle.com/code/kirkdco/precision-by-protein-and-group?scriptVersionId=190931649\" target=\"_blank\">this notebook</a> to calculate the Average Precision across those 9 different groups. </p>\n<p>I think it would be incredibly informative to see how the different models performed especially in the non-triazine sets of test compounds.  The conclusion was that models did not generalize well to that group, but I'm curious how any of the models generalize if at all, especially the top performing ones.  Please run the notebook and post your results here.</p>\n<p>For me, I tried to be clever and use <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/519152\" target=\"_blank\">pharmacophore similarities to represent the R-groups</a> which didn't work as hoped.  Here are my results.  I notice that my Mean Average Precision does not match that posted for the Private LB.  It isn't that different but I'm curious if there are any ideas why they don't match.  I'm using the same SciKit-Learn function as shown in the <a href=\"https://www.kaggle.com/code/metric/leash-average-map\" target=\"_blank\">evaluation notebook</a>.  <strong>EDIT - I realized that this is due to the fact that I did not separate the Public and Private sets, but used the entire dataset.</strong></p>\n<p>Thanks again to <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a>, the Kaggle organizers and all the participants.  This has been a great experience.</p>\n<p>BTW - in the table produced by the notebook, <code>kin0</code> refers to the non-triazine set, non_share refers to the triazine core but with non-shared BBs, and share refers to the triazine core with shared BBs.  </p>\n<p><strong>Update 24.08.05</strong>  <a href=\"https://www.kaggle.com/datasets/kirkdco/leashbio-competition-subs/data\" target=\"_blank\">A dataset</a> was created with the 21 submission files gathered from the Code section, and a .csv with performances (shown below in table form).</p>\n<pre><code> Average Precision:  .\n          non_share   share\n    .    .    .\n    .    .    .\n    .    .    .\n</code></pre>",
  "messages": [
    {
      "id": "2944629",
      "postDate": "08/02/2024 16:35:40",
      "content": "<p>Thanks to <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> for posting the <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/523644\" target=\"_blank\">answer key</a>.  </p>\n<p>It would be very interesting to see exactly how different models performed across proteins and subsets of the data.  I've created <a href=\"https://www.kaggle.com/code/kirkdco/precision-by-protein-and-group?scriptVersionId=190931649\" target=\"_blank\">this notebook</a> to calculate the Average Precision across those 9 different groups. </p>\n<p>I think it would be incredibly informative to see how the different models performed especially in the non-triazine sets of test compounds.  The conclusion was that models did not generalize well to that group, but I'm curious how any of the models generalize if at all, especially the top performing ones.  Please run the notebook and post your results here.</p>\n<p>For me, I tried to be clever and use <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/519152\" target=\"_blank\">pharmacophore similarities to represent the R-groups</a> which didn't work as hoped.  Here are my results.  I notice that my Mean Average Precision does not match that posted for the Private LB.  It isn't that different but I'm curious if there are any ideas why they don't match.  I'm using the same SciKit-Learn function as shown in the <a href=\"https://www.kaggle.com/code/metric/leash-average-map\" target=\"_blank\">evaluation notebook</a>.  <strong>EDIT - I realized that this is due to the fact that I did not separate the Public and Private sets, but used the entire dataset.</strong></p>\n<p>Thanks again to <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a>, the Kaggle organizers and all the participants.  This has been a great experience.</p>\n<p>BTW - in the table produced by the notebook, <code>kin0</code> refers to the non-triazine set, non_share refers to the triazine core but with non-shared BBs, and share refers to the triazine core with shared BBs.  </p>\n<p><strong>Update 24.08.05</strong>  <a href=\"https://www.kaggle.com/datasets/kirkdco/leashbio-competition-subs/data\" target=\"_blank\">A dataset</a> was created with the 21 submission files gathered from the Code section, and a .csv with performances (shown below in table form).</p>\n<pre><code> Average Precision:  .\n          non_share   share\n    .    .    .\n    .    .    .\n    .    .    .\n</code></pre>",
      "rawMarkdown": "Thanks to @andrewdblevins for posting the [answer key](https://www.kaggle.com/competitions/leash-BELKA/discussion/523644).  \n \nIt would be very interesting to see exactly how different models performed across proteins and subsets of the data.  I've created [this notebook](https://www.kaggle.com/code/kirkdco/precision-by-protein-and-group?scriptVersionId=190931649) to calculate the Average Precision across those 9 different groups. \n\nI think it would be incredibly informative to see how the different models performed especially in the non-triazine sets of test compounds.  The conclusion was that models did not generalize well to that group, but I'm curious how any of the models generalize if at all, especially the top performing ones.  Please run the notebook and post your results here.\n\nFor me, I tried to be clever and use [pharmacophore similarities to represent the R-groups](https://www.kaggle.com/competitions/leash-BELKA/discussion/519152) which didn't work as hoped.  Here are my results.  I notice that my Mean Average Precision does not match that posted for the Private LB.  It isn't that different but I'm curious if there are any ideas why they don't match.  I'm using the same SciKit-Learn function as shown in the [evaluation notebook](https://www.kaggle.com/code/metric/leash-average-map).  **EDIT - I realized that this is due to the fact that I did not separate the Public and Private sets, but used the entire dataset.**\n\nThanks again to @andrewdblevins, the Kaggle organizers and all the participants.  This has been a great experience.\n\nBTW - in the table produced by the notebook, `kin0` refers to the non-triazine set, non_share refers to the triazine core but with non-shared BBs, and share refers to the triazine core with shared BBs.  \n\n**Update 24.08.05**  [A dataset](https://www.kaggle.com/datasets/kirkdco/leashbio-competition-subs/data) was created with the 21 submission files gathered from the Code section, and a .csv with performances (shown below in table form).\n\n```\nMean Average Precision:  0.21924157357132887\n      kin0\tnon_share\tshare\nBRD4\t0.001740\t0.175586\t0.517075\nHSA\t0.001165\t0.088662\t0.270757\nsEH\t0.000782\t0.059996\t0.857411\n```",
      "votes": null
    },
    {
      "id": "2945893",
      "postDate": "08/03/2024 20:18:21",
      "content": "<p>Thanks for the notebook! Just to confirm, \"kin0\" is non-triazine, \"share\" is triazine with shared BBs, and \"non-share\" is triazine core with non-shared BBs?</p>",
      "rawMarkdown": "Thanks for the notebook! Just to confirm, \"kin0\" is non-triazine, \"share\" is triazine with shared BBs, and \"non-share\" is triazine core with non-shared BBs?",
      "votes": null
    },
    {
      "id": "2945896",
      "postDate": "08/03/2024 20:20:29",
      "content": "<p>Yes, that is correct.  That is how they are labelled in the solution file, but I should make a note of that on the notebook itself.</p>",
      "rawMarkdown": "Yes, that is correct.  That is how they are labelled in the solution file, but I should make a note of that on the notebook itself.",
      "votes": null
    },
    {
      "id": "2946010",
      "postDate": "08/03/2024 22:11:57",
      "content": "<p>From the Code section I gathered 20 submissions with Private LB scores noted from 0.210 to 0.310 and ran my evaluation notebook to get a comparison of how they all did across the various proteins and groups. I also added one of my own submissions which had the highest Private LB score, but which I hadn't selected as a possible submission.  🫠  Below is the table of results.</p>\n<p>Note that <a href=\"https://www.kaggle.com/code/ahsuna123/belka1dcnn-0-310-private-based-on-ah-s-notebook\" target=\"_blank\">one of these submissions, from Ah,</a> had a higher Private LB score than the 1st place solution.  It appears to have been a late submission - perhaps one that hadn't been selected for judging.</p>\n<p>The general conclusion is the same and described by the organizers, that Private LB scores were largely driven by the triazines with shared BBs, and some submissions had comparable scores for the triazine non-shared BB group.  They all did poorly on the non-triazine group, however, even the high scoring submission by Ah I mentioned above.  I had hoped than some of the better scoring algorithms showed some degree of generalization to the non-triazine compounds, alas, it was not to be.</p>\n<p>This result is troubling to me as this is exactly what computational chemistry and cheminformatics is supposed to do - translate to previously unseen structures.  While there is certainly a benefit and enrichment for the triazine non-shared BBs, the non-triazine set shows no enrichment for hits based on predictions.  It is an interesting result and certainly an interesting problem that remains to be solved.</p>\n<p>I will create a dataset shortly containing all the submissions I hand scraped from the Code section along with a file with tabulated results. I'll update this thread with a link to the dataset.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2F65285bf03f53048398a89bc2e8647e6d%2FPerformanceTable.jpeg?generation=1722723111614211&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "From the Code section I gathered 20 submissions with Private LB scores noted from 0.210 to 0.310 and ran my evaluation notebook to get a comparison of how they all did across the various proteins and groups. I also added one of my own submissions which had the highest Private LB score, but which I hadn't selected as a possible submission.  🫠  Below is the table of results.\n\nNote that [one of these submissions, from Ah,](https://www.kaggle.com/code/ahsuna123/belka1dcnn-0-310-private-based-on-ah-s-notebook) had a higher Private LB score than the 1st place solution.  It appears to have been a late submission - perhaps one that hadn't been selected for judging.\n\nThe general conclusion is the same and described by the organizers, that Private LB scores were largely driven by the triazines with shared BBs, and some submissions had comparable scores for the triazine non-shared BB group.  They all did poorly on the non-triazine group, however, even the high scoring submission by Ah I mentioned above.  I had hoped than some of the better scoring algorithms showed some degree of generalization to the non-triazine compounds, alas, it was not to be.\n\nThis result is troubling to me as this is exactly what computational chemistry and cheminformatics is supposed to do - translate to previously unseen structures.  While there is certainly a benefit and enrichment for the triazine non-shared BBs, the non-triazine set shows no enrichment for hits based on predictions.  It is an interesting result and certainly an interesting problem that remains to be solved.\n\nI will create a dataset shortly containing all the submissions I hand scraped from the Code section along with a file with tabulated results. I'll update this thread with a link to the dataset.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2F65285bf03f53048398a89bc2e8647e6d%2FPerformanceTable.jpeg?generation=1722723111614211&alt=media)",
      "votes": null
    },
    {
      "id": "2946056",
      "postDate": "08/03/2024 23:29:19",
      "content": "<p>Use your methods to check <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's submissions <img src=\"https://i.ibb.co/bdH20zd/1.png\" alt=\"image\">, check this <a href=\"https://www.kaggle.com/code/lililycai/ap-score-by-protein-and-group\" target=\"_blank\">notebook</a></p>",
      "rawMarkdown": "Use your methods to check @hengck23 's submissions ![image](https://i.ibb.co/bdH20zd/1.png), check this [notebook] (https://www.kaggle.com/code/lililycai/ap-score-by-protein-and-group)",
      "votes": null
    },
    {
      "id": "2946089",
      "postDate": "08/04/2024 00:22:39",
      "content": "<p>Same story as the others.</p>\n<p>I had thought that the use of any models trained using the SMILES strings would inherently have difficulty in generalization.  To me, the SMILES representation will lead to over-training by its very nature.  Even a GNN is biased due to the specific representations in SMILES strings.  It almost seems that the best performing models were those that were the most over-trained.  🫠</p>\n<p>That's why I went with the 2D pharmacophore, but unfortunately that also didn't generalize.  I'm curious as to what kind of representation would allow generalization.  🤔</p>\n<p>I've said for a long time (decades?) that we don't need better modeling methods, we need better representations.  Assuming the data is accurate given the complexities of the DEL assay, this is a beautiful set to search for new representations, IMO.</p>",
      "rawMarkdown": "Same story as the others.\n\nI had thought that the use of any models trained using the SMILES strings would inherently have difficulty in generalization.  To me, the SMILES representation will lead to over-training by its very nature.  Even a GNN is biased due to the specific representations in SMILES strings.  It almost seems that the best performing models were those that were the most over-trained.  🫠\n\nThat's why I went with the 2D pharmacophore, but unfortunately that also didn't generalize.  I'm curious as to what kind of representation would allow generalization.  🤔\n\nI've said for a long time (decades?) that we don't need better modeling methods, we need better representations.  Assuming the data is accurate given the complexities of the DEL assay, this is a beautiful set to search for new representations, IMO.",
      "votes": null
    },
    {
      "id": "2946127",
      "postDate": "08/04/2024 01:54:34",
      "content": "<p>Agree. This is a good dataset to explore new representations for generalizations. Some physics based models should work, but needs lots of computation resources.</p>",
      "rawMarkdown": "Agree. This is a good dataset to explore new representations for generalizations. Some physics based models should work, but needs lots of computation resources.",
      "votes": null
    },
    {
      "id": "2955534",
      "postDate": "08/11/2024 08:47:58",
      "content": "<p><a href=\"https://www.kaggle.com/kirkdco\" target=\"_blank\">@kirkdco</a> A little correction : It wasn't a late submisison 😬, it was during the submission timeline just didn't end up selecting as one of the final two submissions! :)<br>\nKudos to you for compiling the results! Very helpful! :)</p>",
      "rawMarkdown": "kirkdco A little correction : It wasn't a late submisison 😬, it was during the submission timeline just didn't end up selecting as one of the final two submissions! :)\nKudos to you for compiling the results! Very helpful! :)",
      "votes": null
    },
    {
      "id": "2955681",
      "postDate": "08/11/2024 12:41:55",
      "content": "<p>Thanks for the clarification, <a href=\"https://www.kaggle.com/ahsuna123\" target=\"_blank\">@ahsuna123</a>, and for the kind words.  Kudos for the great work and contributions during the competition!!</p>",
      "rawMarkdown": "Thanks for the clarification, @ahsuna123, and for the kind words.  Kudos for the great work and contributions during the competition!!",
      "votes": null
    },
    {
      "id": "2975566",
      "postDate": "09/01/2024 03:08:41",
      "content": "<p>Hi! Sorry for the late reply, was busy with another competition :)<br>\nI scored &gt;90 submissions, which I saved, and the top models are here: <a href=\"https://www.kaggle.com/code/antoninadolgorukova/belka-score-submissions\" target=\"_blank\">https://www.kaggle.com/code/antoninadolgorukova/belka-score-submissions</a>.</p>\n<p>I have come to the following conclusions:</p>\n<ol>\n<li>No Single Best Model Across All Contexts: The analysis highlights that there is no single best model that performs optimally across all proteins and test groups:</li>\n</ol>\n<ul>\n<li>CNNs are the best models for all proteins but only for molecules with shared building blocks  </li>\n<li>ChemBerta models and XGBoost models showed the strongest performance on molecules with unseen building blocks, though their success depends on specific parameters and training settings  </li>\n<li>Chemprop models perform relatively better than other models in the highly challenging non-triazines group, although all models struggle in this category</li>\n</ul>\n<p>2.Limitations of Private mAP Scores: When considering private mAP scores averaged across both share and non-share groups, the overall reliability of these scores as indicators of generalization ability may be misleading. It might overestimate generalization ability if dominated by the more consistent performance in share groups or underestimate it if heavily influenced by variability in non-share groups.</p>",
      "rawMarkdown": "Hi! Sorry for the late reply, was busy with another competition :)\nI scored >90 submissions, which I saved, and the top models are here: https://www.kaggle.com/code/antoninadolgorukova/belka-score-submissions.\n\nI have come to the following conclusions:\n\n1. No Single Best Model Across All Contexts: The analysis highlights that there is no single best model that performs optimally across all proteins and test groups:\n- CNNs are the best models for all proteins but only for molecules with shared building blocks  \n- ChemBerta models and XGBoost models showed the strongest performance on molecules with unseen building blocks, though their success depends on specific parameters and training settings  \n- Chemprop models perform relatively better than other models in the highly challenging non-triazines group, although all models struggle in this category\n\n2.Limitations of Private mAP Scores: When considering private mAP scores averaged across both share and non-share groups, the overall reliability of these scores as indicators of generalization ability may be misleading. It might overestimate generalization ability if dominated by the more consistent performance in share groups or underestimate it if heavily influenced by variability in non-share groups.",
      "votes": null
    },
    {
      "id": "2976219",
      "postDate": "09/01/2024 16:11:55",
      "content": "<p>This is an amazing amount of excellent work.  Thank you!</p>",
      "rawMarkdown": "This is an amazing amount of excellent work.  Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2945893,
      "author_name": "lililycai",
      "author_url": "",
      "post_date": "08/03/2024 20:18:21",
      "content": "<p>Thanks for the notebook! Just to confirm, \"kin0\" is non-triazine, \"share\" is triazine with shared BBs, and \"non-share\" is triazine core with non-shared BBs?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2945896,
          "author_name": "kirkdco",
          "author_url": "",
          "post_date": "08/03/2024 20:20:29",
          "content": "<p>Yes, that is correct.  That is how they are labelled in the solution file, but I should make a note of that on the notebook itself.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2946010,
      "author_name": "kirkdco",
      "author_url": "",
      "post_date": "08/03/2024 22:11:57",
      "content": "<p>From the Code section I gathered 20 submissions with Private LB scores noted from 0.210 to 0.310 and ran my evaluation notebook to get a comparison of how they all did across the various proteins and groups. I also added one of my own submissions which had the highest Private LB score, but which I hadn't selected as a possible submission.  🫠  Below is the table of results.</p>\n<p>Note that <a href=\"https://www.kaggle.com/code/ahsuna123/belka1dcnn-0-310-private-based-on-ah-s-notebook\" target=\"_blank\">one of these submissions, from Ah,</a> had a higher Private LB score than the 1st place solution.  It appears to have been a late submission - perhaps one that hadn't been selected for judging.</p>\n<p>The general conclusion is the same and described by the organizers, that Private LB scores were largely driven by the triazines with shared BBs, and some submissions had comparable scores for the triazine non-shared BB group.  They all did poorly on the non-triazine group, however, even the high scoring submission by Ah I mentioned above.  I had hoped than some of the better scoring algorithms showed some degree of generalization to the non-triazine compounds, alas, it was not to be.</p>\n<p>This result is troubling to me as this is exactly what computational chemistry and cheminformatics is supposed to do - translate to previously unseen structures.  While there is certainly a benefit and enrichment for the triazine non-shared BBs, the non-triazine set shows no enrichment for hits based on predictions.  It is an interesting result and certainly an interesting problem that remains to be solved.</p>\n<p>I will create a dataset shortly containing all the submissions I hand scraped from the Code section along with a file with tabulated results. I'll update this thread with a link to the dataset.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2F65285bf03f53048398a89bc2e8647e6d%2FPerformanceTable.jpeg?generation=1722723111614211&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2955534,
          "author_name": "ahsuna123",
          "author_url": "",
          "post_date": "08/11/2024 08:47:58",
          "content": "<p><a href=\"https://www.kaggle.com/kirkdco\" target=\"_blank\">@kirkdco</a> A little correction : It wasn't a late submisison 😬, it was during the submission timeline just didn't end up selecting as one of the final two submissions! :)<br>\nKudos to you for compiling the results! Very helpful! :)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2955681,
              "author_name": "kirkdco",
              "author_url": "",
              "post_date": "08/11/2024 12:41:55",
              "content": "<p>Thanks for the clarification, <a href=\"https://www.kaggle.com/ahsuna123\" target=\"_blank\">@ahsuna123</a>, and for the kind words.  Kudos for the great work and contributions during the competition!!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2946056,
      "author_name": "lililycai",
      "author_url": "",
      "post_date": "08/03/2024 23:29:19",
      "content": "<p>Use your methods to check <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's submissions <img src=\"https://i.ibb.co/bdH20zd/1.png\" alt=\"image\">, check this <a href=\"https://www.kaggle.com/code/lililycai/ap-score-by-protein-and-group\" target=\"_blank\">notebook</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2946089,
          "author_name": "kirkdco",
          "author_url": "",
          "post_date": "08/04/2024 00:22:39",
          "content": "<p>Same story as the others.</p>\n<p>I had thought that the use of any models trained using the SMILES strings would inherently have difficulty in generalization.  To me, the SMILES representation will lead to over-training by its very nature.  Even a GNN is biased due to the specific representations in SMILES strings.  It almost seems that the best performing models were those that were the most over-trained.  🫠</p>\n<p>That's why I went with the 2D pharmacophore, but unfortunately that also didn't generalize.  I'm curious as to what kind of representation would allow generalization.  🤔</p>\n<p>I've said for a long time (decades?) that we don't need better modeling methods, we need better representations.  Assuming the data is accurate given the complexities of the DEL assay, this is a beautiful set to search for new representations, IMO.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2946127,
              "author_name": "lililycai",
              "author_url": "",
              "post_date": "08/04/2024 01:54:34",
              "content": "<p>Agree. This is a good dataset to explore new representations for generalizations. Some physics based models should work, but needs lots of computation resources.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2975566,
      "author_name": "antoninadolgorukova",
      "author_url": "",
      "post_date": "09/01/2024 03:08:41",
      "content": "<p>Hi! Sorry for the late reply, was busy with another competition :)<br>\nI scored &gt;90 submissions, which I saved, and the top models are here: <a href=\"https://www.kaggle.com/code/antoninadolgorukova/belka-score-submissions\" target=\"_blank\">https://www.kaggle.com/code/antoninadolgorukova/belka-score-submissions</a>.</p>\n<p>I have come to the following conclusions:</p>\n<ol>\n<li>No Single Best Model Across All Contexts: The analysis highlights that there is no single best model that performs optimally across all proteins and test groups:</li>\n</ol>\n<ul>\n<li>CNNs are the best models for all proteins but only for molecules with shared building blocks  </li>\n<li>ChemBerta models and XGBoost models showed the strongest performance on molecules with unseen building blocks, though their success depends on specific parameters and training settings  </li>\n<li>Chemprop models perform relatively better than other models in the highly challenging non-triazines group, although all models struggle in this category</li>\n</ul>\n<p>2.Limitations of Private mAP Scores: When considering private mAP scores averaged across both share and non-share groups, the overall reliability of these scores as indicators of generalization ability may be misleading. It might overestimate generalization ability if dominated by the more consistent performance in share groups or underestimate it if heavily influenced by variability in non-share groups.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2976219,
          "author_name": "kirkdco",
          "author_url": "",
          "post_date": "09/01/2024 16:11:55",
          "content": "<p>This is an amazing amount of excellent work.  Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2944629": "Thanks to @andrewdblevins for posting the [answer key](https://www.kaggle.com/competitions/leash-BELKA/discussion/523644).  \n \nIt would be very interesting to see exactly how different models performed across proteins and subsets of the data.  I've created [this notebook](https://www.kaggle.com/code/kirkdco/precision-by-protein-and-group?scriptVersionId=190931649) to calculate the Average Precision across those 9 different groups. \n\nI think it would be incredibly informative to see how the different models performed especially in the non-triazine sets of test compounds.  The conclusion was that models did not generalize well to that group, but I'm curious how any of the models generalize if at all, especially the top performing ones.  Please run the notebook and post your results here.\n\nFor me, I tried to be clever and use [pharmacophore similarities to represent the R-groups](https://www.kaggle.com/competitions/leash-BELKA/discussion/519152) which didn't work as hoped.  Here are my results.  I notice that my Mean Average Precision does not match that posted for the Private LB.  It isn't that different but I'm curious if there are any ideas why they don't match.  I'm using the same SciKit-Learn function as shown in the [evaluation notebook](https://www.kaggle.com/code/metric/leash-average-map).  **EDIT - I realized that this is due to the fact that I did not separate the Public and Private sets, but used the entire dataset.**\n\nThanks again to @andrewdblevins, the Kaggle organizers and all the participants.  This has been a great experience.\n\nBTW - in the table produced by the notebook, `kin0` refers to the non-triazine set, non_share refers to the triazine core but with non-shared BBs, and share refers to the triazine core with shared BBs.  \n\n**Update 24.08.05**  [A dataset](https://www.kaggle.com/datasets/kirkdco/leashbio-competition-subs/data) was created with the 21 submission files gathered from the Code section, and a .csv with performances (shown below in table form).\n\n```\nMean Average Precision:  0.21924157357132887\n      kin0\tnon_share\tshare\nBRD4\t0.001740\t0.175586\t0.517075\nHSA\t0.001165\t0.088662\t0.270757\nsEH\t0.000782\t0.059996\t0.857411\n```",
    "2945893": "Thanks for the notebook! Just to confirm, \"kin0\" is non-triazine, \"share\" is triazine with shared BBs, and \"non-share\" is triazine core with non-shared BBs?",
    "2945896": "Yes, that is correct.  That is how they are labelled in the solution file, but I should make a note of that on the notebook itself.",
    "2946010": "From the Code section I gathered 20 submissions with Private LB scores noted from 0.210 to 0.310 and ran my evaluation notebook to get a comparison of how they all did across the various proteins and groups. I also added one of my own submissions which had the highest Private LB score, but which I hadn't selected as a possible submission.  🫠  Below is the table of results.\n\nNote that [one of these submissions, from Ah,](https://www.kaggle.com/code/ahsuna123/belka1dcnn-0-310-private-based-on-ah-s-notebook) had a higher Private LB score than the 1st place solution.  It appears to have been a late submission - perhaps one that hadn't been selected for judging.\n\nThe general conclusion is the same and described by the organizers, that Private LB scores were largely driven by the triazines with shared BBs, and some submissions had comparable scores for the triazine non-shared BB group.  They all did poorly on the non-triazine group, however, even the high scoring submission by Ah I mentioned above.  I had hoped than some of the better scoring algorithms showed some degree of generalization to the non-triazine compounds, alas, it was not to be.\n\nThis result is troubling to me as this is exactly what computational chemistry and cheminformatics is supposed to do - translate to previously unseen structures.  While there is certainly a benefit and enrichment for the triazine non-shared BBs, the non-triazine set shows no enrichment for hits based on predictions.  It is an interesting result and certainly an interesting problem that remains to be solved.\n\nI will create a dataset shortly containing all the submissions I hand scraped from the Code section along with a file with tabulated results. I'll update this thread with a link to the dataset.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2F65285bf03f53048398a89bc2e8647e6d%2FPerformanceTable.jpeg?generation=1722723111614211&alt=media)",
    "2946056": "Use your methods to check @hengck23 's submissions ![image](https://i.ibb.co/bdH20zd/1.png), check this [notebook] (https://www.kaggle.com/code/lililycai/ap-score-by-protein-and-group)",
    "2946089": "Same story as the others.\n\nI had thought that the use of any models trained using the SMILES strings would inherently have difficulty in generalization.  To me, the SMILES representation will lead to over-training by its very nature.  Even a GNN is biased due to the specific representations in SMILES strings.  It almost seems that the best performing models were those that were the most over-trained.  🫠\n\nThat's why I went with the 2D pharmacophore, but unfortunately that also didn't generalize.  I'm curious as to what kind of representation would allow generalization.  🤔\n\nI've said for a long time (decades?) that we don't need better modeling methods, we need better representations.  Assuming the data is accurate given the complexities of the DEL assay, this is a beautiful set to search for new representations, IMO.",
    "2946127": "Agree. This is a good dataset to explore new representations for generalizations. Some physics based models should work, but needs lots of computation resources.",
    "2955534": "kirkdco A little correction : It wasn't a late submisison 😬, it was during the submission timeline just didn't end up selecting as one of the final two submissions! :)\nKudos to you for compiling the results! Very helpful! :)",
    "2955681": "Thanks for the clarification, @ahsuna123, and for the kind words.  Kudos for the great work and contributions during the competition!!",
    "2975566": "Hi! Sorry for the late reply, was busy with another competition :)\nI scored >90 submissions, which I saved, and the top models are here: https://www.kaggle.com/code/antoninadolgorukova/belka-score-submissions.\n\nI have come to the following conclusions:\n\n1. No Single Best Model Across All Contexts: The analysis highlights that there is no single best model that performs optimally across all proteins and test groups:\n- CNNs are the best models for all proteins but only for molecules with shared building blocks  \n- ChemBerta models and XGBoost models showed the strongest performance on molecules with unseen building blocks, though their success depends on specific parameters and training settings  \n- Chemprop models perform relatively better than other models in the highly challenging non-triazines group, although all models struggle in this category\n\n2.Limitations of Private mAP Scores: When considering private mAP scores averaged across both share and non-share groups, the overall reliability of these scores as indicators of generalization ability may be misleading. It might overestimate generalization ability if dominated by the more consistent performance in share groups or underestimate it if heavily influenced by variability in non-share groups.",
    "2976219": "This is an amazing amount of excellent work.  Thank you!"
  },
  "source": "meta"
}