{
  "id": 518943,
  "title": "How about we share our best private LB results here? (current best: 0.347 from @Maciej Sypetkowski)",
  "url": "/competitions/leash-BELKA/discussion/518943",
  "author_name": "Yifan Wu",
  "post_date": "2024-07-09T00:52:23.598000",
  "votes": 13,
  "comment_count": 68,
  "views": 0,
  "content": "<p>This is really a fucking crazy shuffle in private LB. And we can check all of our private LB results now. My best private LB is from my customed pretrained LM using only SMILES. In public LB, it is only 0.391 but in private LB, it's 0.294!. (SHIT, I should trust my dream! Why not trust my own pretrained model!!! )</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1029068%2F9227f421039af5fdc45d8c33e8bb5893%2F1720486391365.jpg?generation=1720486410872878&amp;alt=media\" alt=\"fk\"></p>",
  "messages": [
    {
      "id": 2912532,
      "postDate": "2024-07-09T00:52:23.600Z",
      "content": "<p>This is really a fucking crazy shuffle in private LB. And we can check all of our private LB results now. My best private LB is from my customed pretrained LM using only SMILES. In public LB, it is only 0.391 but in private LB, it's 0.294!. (SHIT, I should trust my dream! Why not trust my own pretrained model!!! )</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1029068%2F9227f421039af5fdc45d8c33e8bb5893%2F1720486391365.jpg?generation=1720486410872878&amp;alt=media\" alt=\"fk\"></p>",
      "rawMarkdown": "This is really a fucking crazy shuffle in private LB. And we can check all of our private LB results now. My best private LB is from my customed pretrained LM using only SMILES. In public LB, it is only 0.391 but in private LB, it's 0.294!. (SHIT, I should trust my dream! Why not trust my own pretrained model!!! )\n \n![fk](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1029068%2F9227f421039af5fdc45d8c33e8bb5893%2F1720486391365.jpg?generation=1720486410872878&alt=media)",
      "votes": 13
    },
    {
      "id": 2913022,
      "postDate": "2024-07-09T07:26:21.470Z",
      "content": "<p>All this scores mean nothing until we will see separate metrics for each of 9 parts averaged in total LB score. Just small random (!) variation in scores for less predictable (or almost unpredictable) non-triazine part can lead to diferences like 0.03 or larger in total scores.</p>",
      "rawMarkdown": "All this scores mean nothing until we will see separate metrics for each of 9 parts averaged in total LB score. Just small random (!) variation in scores for less predictable (or almost unpredictable) non-triazine part can lead to diferences like 0.03 or larger in total scores.",
      "votes": 8
    },
    {
      "id": 2912552,
      "postDate": "2024-07-09T01:14:39.230Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F75c4c2c30cf8343ee1ef0e09127d1855%2FSelection_024.png?generation=1720487665043888&amp;alt=media\"></p>\n<p>tx= transformer too (fa is flash attention)</p>\n<p>transformer is the king … actually</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F75c4c2c30cf8343ee1ef0e09127d1855%2FSelection_024.png?generation=1720487665043888&alt=media)\n\ntx= transformer too (fa is flash attention)\n\ntransformer is the king ... actually",
      "votes": 8,
      "replies": [
        {
          "id": 2912567,
          "postDate": "2024-07-09T01:19:06.213Z",
          "content": "<p>What's the input to your transformer？ Is it only SMILES?</p>",
          "rawMarkdown": "What's the input to your transformer？ Is it only SMILES?",
          "replies": [
            {
              "id": 2912571,
              "postDate": "2024-07-09T01:20:42.720Z",
              "content": "<p>input charcters. it is based on the code i share</p>",
              "rawMarkdown": "input charcters. it is based on the code i share\n"
            },
            {
              "id": 2912586,
              "postDate": "2024-07-09T01:27:20.777Z",
              "content": "<p>input characters/SMILES is the king!</p>",
              "rawMarkdown": "input characters/SMILES is the king!"
            }
          ]
        },
        {
          "id": 2912569,
          "postDate": "2024-07-09T01:20:31.917Z",
          "content": "<p>0.315! you're the champion！</p>",
          "rawMarkdown": "0.315! you're the champion！"
        },
        {
          "id": 2912573,
          "postDate": "2024-07-09T01:21:16.590Z",
          "content": "<p>GNN and mamba and cnn1d score<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5121bf2f956d9d50ec163f74dea12300%2FSelection_025.png?generation=1720488206406902&amp;alt=media\"></p>",
          "rawMarkdown": "GNN and mamba and cnn1d score\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5121bf2f956d9d50ec163f74dea12300%2FSelection_025.png?generation=1720488206406902&alt=media)"
        },
        {
          "id": 2912583,
          "postDate": "2024-07-09T01:26:32.163Z",
          "content": "<p>Congrats on 5th <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> !<br>\nWhat was your transformer's valid(share) and valid(non-share) scores?</p>\n<p>How did you pick your two submissions?</p>",
          "rawMarkdown": "Congrats on 5th @hengck23 !\nWhat was your transformer's valid(share) and valid(non-share) scores?\n\nHow did you pick your two submissions?",
          "replies": [
            {
              "id": 2912601,
              "postDate": "2024-07-09T01:40:15.807Z",
              "content": "<p>the private 0.315/public0.482 solution is actually public -solution.<br>\ni will private more details (on share nonshare scores) later.<br>\ni need some sleep first.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3c5aaf52bf89941d8beafe1e30adb2de%2FSelection_026.png?generation=1720489124135543&amp;alt=media\"></p>",
              "rawMarkdown": "the private 0.315/public0.482 solution is actually public -solution.\ni will private more details (on share nonshare scores) later.\ni need some sleep first.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3c5aaf52bf89941d8beafe1e30adb2de%2FSelection_026.png?generation=1720489124135543&alt=media)",
              "votes": 2
            },
            {
              "id": 2912606,
              "postDate": "2024-07-09T01:46:10.540Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb491118540076973b5def4528aa4e4c7%2FSelection_027.png?generation=1720489561920264&amp;alt=media\"></p>\n<p>top solution.</p>",
              "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb491118540076973b5def4528aa4e4c7%2FSelection_027.png?generation=1720489561920264&alt=media)\n\ntop solution."
            },
            {
              "id": 2912639,
              "postDate": "2024-07-09T02:20:43.387Z",
              "content": "<p>Go sleep🤣, Looking forward to your solution!</p>",
              "rawMarkdown": "Go sleep🤣, Looking forward to your solution!"
            }
          ]
        },
        {
          "id": 2913130,
          "postDate": "2024-07-09T09:24:10.780Z",
          "content": "<p>Congratulations! <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 🥳</p>",
          "rawMarkdown": "Congratulations! @hengck23 🥳"
        },
        {
          "id": 2915601,
          "postDate": "2024-07-10T15:24:21.797Z",
          "content": "<p>Could you give details on the exact transformer architecture of your top-scoring model? <br>\nThe #1 winning solution used a seemingly under-parameterized transformer encoder, which is why (in my opinion) it physically could not overfit the training dataset and ended up generalizing well.</p>",
          "rawMarkdown": "Could you give details on the exact transformer architecture of your top-scoring model? \nThe #1 winning solution used a seemingly under-parameterized transformer encoder, which is why (in my opinion) it physically could not overfit the training dataset and ended up generalizing well."
        }
      ]
    },
    {
      "id": 2913328,
      "postDate": "2024-07-09T12:21:29.897Z",
      "content": "<p>I have a model (not ensemble!) that got <strong>0.347</strong> on private and 0.352 on public. Basically it was a MLP on concatenation of Morgan fingerprints, MACCS and embeddings from a pretrained ChemBERT_chEMBL. It was also pretty bad on a local evaluation, especially on splits with unseen blocks (compared to my other models). That competition was so noisy…<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2631834%2F3243f56a961c037b54c6477ddf9e10ae%2Fscreenshot.png?generation=1720527753507162&amp;alt=media\"></p>",
      "rawMarkdown": "I have a model (not ensemble!) that got **0.347** on private and 0.352 on public. Basically it was a MLP on concatenation of Morgan fingerprints, MACCS and embeddings from a pretrained ChemBERT_chEMBL. It was also pretty bad on a local evaluation, especially on splits with unseen blocks (compared to my other models). That competition was so noisy...\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2631834%2F3243f56a961c037b54c6477ddf9e10ae%2Fscreenshot.png?generation=1720527753507162&alt=media)",
      "votes": 5,
      "replies": [
        {
          "id": 2913498,
          "postDate": "2024-07-09T13:52:38.697Z",
          "content": "<p>We were watching this submission wondering what it was. It was head and shoulders above everything else we saw, we were hoping it was the beginning of a breakthrough, but alas it probably just got lucky on a couple of the of the building blocks. Most of the performance gains come from an unusually high performance on non-share for sEH and HSA. Kin0 is what we call the non-triazine, and the performance for this model is 0 :(<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F57970%2Fd39acf25211362a20daef30e2b460829%2FScreen%20Shot%202024-07-09%20at%207.50.55%20AM.png?generation=1720532964733655&amp;alt=media\"></p>\n<p>I will work on generating more of these plots for other submissions, and release a notebook so that people can do this themselves.</p>",
          "rawMarkdown": "We were watching this submission wondering what it was. It was head and shoulders above everything else we saw, we were hoping it was the beginning of a breakthrough, but alas it probably just got lucky on a couple of the of the building blocks. Most of the performance gains come from an unusually high performance on non-share for sEH and HSA. Kin0 is what we call the non-triazine, and the performance for this model is 0 :(\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F57970%2Fd39acf25211362a20daef30e2b460829%2FScreen%20Shot%202024-07-09%20at%207.50.55%20AM.png?generation=1720532964733655&alt=media)\n\nI will work on generating more of these plots for other submissions, and release a notebook so that people can do this themselves.",
          "votes": 14,
          "replies": [
            {
              "id": 2913581,
              "postDate": "2024-07-09T15:02:35.757Z",
              "content": "<p>Looking forward for the notebook!</p>",
              "rawMarkdown": "Looking forward for the notebook!",
              "votes": 1
            },
            {
              "id": 2913638,
              "postDate": "2024-07-09T15:32:35.573Z",
              "content": "<p>How ironic that a competition designed to test generalization could be won by a solution that does not generalize at all 😅</p>",
              "rawMarkdown": "How ironic that a competition designed to test generalization could be won by a solution that does not generalize at all 😅",
              "votes": 4
            },
            {
              "id": 2913655,
              "postDate": "2024-07-09T15:36:25.333Z",
              "content": "<p>generalizing to new building blocks is still something we usually do not see. But yeah, we were hoping for better</p>",
              "rawMarkdown": "generalizing to new building blocks is still something we usually do not see. But yeah, we were hoping for better"
            },
            {
              "id": 2914566,
              "postDate": "2024-07-10T03:32:56.953Z",
              "content": "<p>You probably should put some non-triazine samples in the public LB set.   It's already hard enough for them not being in the training set.  It would be almost impossible if their scores are also hidden. Eventually, a fuzzy winning model with large inferencing data space due to model yet to define its inferencing space via under training (like one epoch) or other ways.  If you train the model a few more epochs, the model quickly shrinking its inferencing space close to the real training data space.  Then, the score starts to drop.  Honestly, there is almost no practical values in these fuzzy models. </p>",
              "rawMarkdown": "You probably should put some non-triazine samples in the public LB set.   It's already hard enough for them not being in the training set.  It would be almost impossible if their scores are also hidden. Eventually, a fuzzy winning model with large inferencing data space due to model yet to define its inferencing space via under training (like one epoch) or other ways.  If you train the model a few more epochs, the model quickly shrinking its inferencing space close to the real training data space.  Then, the score starts to drop.  Honestly, there is almost no practical values in these fuzzy models. ",
              "votes": 2
            },
            {
              "id": 2955899,
              "postDate": "2024-08-11T15:13:03.127Z",
              "content": "<p>Created a notebook to visualize these plots for everybody to use it. You can find it <a href=\"https://www.kaggle.com/code/ahsuna123/visualizing-the-protein-group-wise-results?scriptVersionId=192147258\" target=\"_blank\">here</a>. Hope, it's helpful! :)<br>\n<a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> <a href=\"https://www.kaggle.com/lililycai\" target=\"_blank\">@lililycai</a> </p>",
              "rawMarkdown": "Created a notebook to visualize these plots for everybody to use it. You can find it [here](https://www.kaggle.com/code/ahsuna123/visualizing-the-protein-group-wise-results?scriptVersionId=192147258). Hope, it's helpful! :)\n@andrewdblevins @lililycai ",
              "votes": 1
            },
            {
              "id": 3382839,
              "postDate": "2025-12-28T17:16:31.370Z",
              "content": "<p><a href=\"https://www.kaggle.com/ahsuna123\" target=\"_blank\">@ahsuna123</a> is this notebook still available somewhere?</p>",
              "rawMarkdown": "@ahsuna123 is this notebook still available somewhere?"
            }
          ]
        },
        {
          "id": 2914558,
          "postDate": "2024-07-10T03:10:09.557Z",
          "content": "<p>cool! I also have a model which try to use MLP with 7 types fingerprint and 3types pretrained molecular representation model. In public LB, I found the fingerprint is very usefull (the case is training model on whole dataset for some epochs). Finally the public LB is about 0.430 but private LB is only 0.220. So I'm curious about how many epochs or how many data do you use? Just like \"all you need is frequent checkpoint\", I think the training steps is very important to control the model's performance on unseen blocks.</p>",
          "rawMarkdown": "cool! I also have a model which try to use MLP with 7 types fingerprint and 3types pretrained molecular representation model. In public LB, I found the fingerprint is very usefull (the case is training model on whole dataset for some epochs). Finally the public LB is about 0.430 but private LB is only 0.220. So I'm curious about how many epochs or how many data do you use? Just like \"all you need is frequent checkpoint\", I think the training steps is very important to control the model's performance on unseen blocks.",
          "replies": [
            {
              "id": 2914569,
              "postDate": "2024-07-10T03:38:28.397Z",
              "content": "<p>I think you already input too much information with \"7 types fingerprint and 3types pretrained molecular representation model\".  </p>",
              "rawMarkdown": "I think you already input too much information with \"7 types fingerprint and 3types pretrained molecular representation model\".  ",
              "votes": 1
            },
            {
              "id": 2914595,
              "postDate": "2024-07-10T04:15:43.067Z",
              "content": "<p>agree, too much information can easily lead to overfitting</p>",
              "rawMarkdown": "agree, too much information can easily lead to overfitting",
              "votes": 1
            },
            {
              "id": 2914697,
              "postDate": "2024-07-10T06:02:45.297Z",
              "content": "<p>You're right. Too much information will introduce many redundance to the model, but I still think the model can choose which part are chosen to use by the gradient descent. When we are talking about \"overfitting\", I think there are two aspect of meanning. One is to overfit the training dataset. When we are training the model, more useful features will lead to a faster convergence speed, what we need to do is to stop the training at an appropriate step. That's why we need the validation set. Another one is to generalize to other totally new data (new distribution). This part I think mainly depends on the features or how we represent the input data. In this competition, the case is the distribution of data in private LB is totally different from the data in public LB. But if we select a subset from the training set, or stop at some specific steps, at this monment, the model will have the best performance in the distribution of data in private LB. <br>\nAnother thought I have is that I don't think the AI model can really generalize to some totally new distribution of data, in CV or NLP, the training/validation/test set of most of the tasks belongs to a similar distribution (near to the distribution in native). However, in biological field, the difference of distribution between different target/molecules are even larger than different tasks in CV/NLP. Maybe some zero-shot/few-shot learning techniques will work better in biological field. </p>",
              "rawMarkdown": "You're right. Too much information will introduce many redundance to the model, but I still think the model can choose which part are chosen to use by the gradient descent. When we are talking about \"overfitting\", I think there are two aspect of meanning. One is to overfit the training dataset. When we are training the model, more useful features will lead to a faster convergence speed, what we need to do is to stop the training at an appropriate step. That's why we need the validation set. Another one is to generalize to other totally new data (new distribution). This part I think mainly depends on the features or how we represent the input data. In this competition, the case is the distribution of data in private LB is totally different from the data in public LB. But if we select a subset from the training set, or stop at some specific steps, at this monment, the model will have the best performance in the distribution of data in private LB. \nAnother thought I have is that I don't think the AI model can really generalize to some totally new distribution of data, in CV or NLP, the training/validation/test set of most of the tasks belongs to a similar distribution (near to the distribution in native). However, in biological field, the difference of distribution between different target/molecules are even larger than different tasks in CV/NLP. Maybe some zero-shot/few-shot learning techniques will work better in biological field. ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2913021,
      "postDate": "2024-07-09T07:26:05.880Z",
      "content": "<p>Had one model that scored <strong>0.330</strong> on private, though only 0.355 on public - so it wasn't selected as a final submission 🥲!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F7a868e57aace0fb83d5eb38c62fab31a%2Finbox_283587_cacabaea5900e2c2d1800eb40257db38_Screenshot%20from%202024-07-09%2008-23-53.png?generation=1720510160991417&amp;alt=media\" alt=\"330\"></p>",
      "rawMarkdown": "Had one model that scored **0.330** on private, though only 0.355 on public - so it wasn't selected as a final submission 🥲!\n\n![330](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F7a868e57aace0fb83d5eb38c62fab31a%2Finbox_283587_cacabaea5900e2c2d1800eb40257db38_Screenshot%20from%202024-07-09%2008-23-53.png?generation=1720510160991417&alt=media)",
      "votes": 5,
      "replies": [
        {
          "id": 2913046,
          "postDate": "2024-07-09T07:41:59.037Z",
          "content": "<p>Great!!!! You are the best now. So what special tricks or features did you use in this model?</p>",
          "rawMarkdown": "Great!!!! You are the best now. So what special tricks or features did you use in this model?",
          "replies": [
            {
              "id": 2913076,
              "postDate": "2024-07-09T08:24:17.030Z",
              "content": "<p>Honestly, nothing really interesting about this one: the model itself is just an MLP classifier on top of ECFP6 fingerprints and GIN GNN embeddings (from <a href=\"https://molfeat.datamol.io/featurizers/gin_supervised_masking\" target=\"_blank\">here</a>), trained on ~1/4 of the training set (24M train rows in the reduced version).</p>\n<p>Generally I was trying to find a point where models start to overfit on non-share - this submission was just probing into how well my validation setup is set and calibrated. I had pushed this model even further without any degradation on the non-shared part - but it seems this particular submission might've been a sweet spot for the non-triazine maybe?</p>",
              "rawMarkdown": "Honestly, nothing really interesting about this one: the model itself is just an MLP classifier on top of ECFP6 fingerprints and GIN GNN embeddings (from [here](https://molfeat.datamol.io/featurizers/gin_supervised_masking)), trained on ~1/4 of the training set (24M train rows in the reduced version).\n\nGenerally I was trying to find a point where models start to overfit on non-share - this submission was just probing into how well my validation setup is set and calibrated. I had pushed this model even further without any degradation on the non-shared part - but it seems this particular submission might've been a sweet spot for the non-triazine maybe?",
              "votes": 1
            },
            {
              "id": 2913184,
              "postDate": "2024-07-09T10:03:54.290Z",
              "content": "<p>really interesting! could yo umake a submit to check it? (mask all triazines with 0 and see th escores)</p>",
              "rawMarkdown": "really interesting! could yo umake a submit to check it? (mask all triazines with 0 and see th escores)",
              "votes": 1
            },
            {
              "id": 2913290,
              "postDate": "2024-07-09T11:40:27.707Z",
              "content": "<p>Sure thing, did a check - and it doesn't seem like it scores anything interesting on non-triazine either, giving only 0.015 private LB score<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F8c9ba36b33a29b1d5ba12039626528f5%2FScreenshot%20from%202024-07-09%2012-37-58.png?generation=1720525193036236&amp;alt=media\" alt=\"0.015\"></p>",
              "rawMarkdown": "Sure thing, did a check - and it doesn't seem like it scores anything interesting on non-triazine either, giving only 0.015 private LB score\n![0.015](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F8c9ba36b33a29b1d5ba12039626528f5%2FScreenshot%20from%202024-07-09%2012-37-58.png?generation=1720525193036236&alt=media)",
              "votes": 2
            },
            {
              "id": 2913300,
              "postDate": "2024-07-09T11:48:59.163Z",
              "content": "<p>Essentially it means that non-triazine had absolutely no contribution to these scores (my hypothesis above is wrong).<br>\nI even tried now a submission with non-triazine masked - and the scores are exactly the same: 0.330 and 0.355    </p>",
              "rawMarkdown": "Essentially it means that non-triazine had absolutely no contribution to these scores (my hypothesis above is wrong).\nI even tried now a submission with non-triazine masked - and the scores are exactly the same: 0.330 and 0.355    ",
              "votes": 1
            },
            {
              "id": 2913309,
              "postDate": "2024-07-09T11:55:52.703Z",
              "content": "<p>yes, I have scored a lot of models, it's always 0.015 (and once 0.016) 🧐, and i am getting same score for each separate protein</p>",
              "rawMarkdown": "yes, I have scored a lot of models, it's always 0.015 (and once 0.016) 🧐, and i am getting same score for each separate protein"
            },
            {
              "id": 2913359,
              "postDate": "2024-07-09T12:49:58.607Z",
              "content": "<p>Well submitting zeros is not exacly the same as masking / filtering - as you'd still get an AUC-PR for that part. In fact you'd get AUC-PR ~= positive rate, so 0.015 might be just an average positive rate in the test set</p>",
              "rawMarkdown": "Well submitting zeros is not exacly the same as masking / filtering - as you'd still get an AUC-PR for that part. In fact you'd get AUC-PR ~= positive rate, so 0.015 might be just an average positive rate in the test set",
              "votes": 1
            },
            {
              "id": 2913383,
              "postDate": "2024-07-09T12:59:54.960Z",
              "content": "<p>Yes, in fact submitting all zeros gives exactly same - 0.015 private 0.023 public.</p>\n<p>Significant difference from 0.005 positive rate in train!!</p>",
              "rawMarkdown": "Yes, in fact submitting all zeros gives exactly same - 0.015 private 0.023 public.\n\nSignificant difference from 0.005 positive rate in train!!",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2912692,
      "postDate": "2024-07-09T03:18:37.743Z",
      "content": "<p>I submitted a total of 65 entries. But after all, I got my best score of 0.275 (public score 0.403) when I just ran the shared code <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">BELKA 1DCNN Starter with all data</a> using a different random seed. I’m left with mixed feelings.</p>",
      "rawMarkdown": "I submitted a total of 65 entries. But after all, I got my best score of 0.275 (public score 0.403) when I just ran the shared code [BELKA 1DCNN Starter with all data](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data) using a different random seed. I’m left with mixed feelings.",
      "votes": 3
    },
    {
      "id": 2912958,
      "postDate": "2024-07-09T06:38:59.943Z",
      "content": "<p>0.309!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13344803%2F971aaeb81832abd19847ed05d4b86ada%2Fbest_private.PNG?generation=1720507067496050&amp;alt=media\"></p>\n<p>I read some paper that described the 1DCNN method but the authors had removed the last dropout layer.</p>",
      "rawMarkdown": "0.309!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13344803%2F971aaeb81832abd19847ed05d4b86ada%2Fbest_private.PNG?generation=1720507067496050&alt=media)\n\nI read some paper that described the 1DCNN method but the authors had removed the last dropout layer.",
      "votes": 1
    },
    {
      "id": 2912753,
      "postDate": "2024-07-09T04:07:05.080Z",
      "content": "<p>My best 0.310 lol! <br>\nMore details of the solution in here! <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/518957\" target=\"_blank\">0.310 private lb</a></p>",
      "rawMarkdown": "My best 0.310 lol! \nMore details of the solution in here! [0.310 private lb](https://www.kaggle.com/competitions/leash-BELKA/discussion/518957)",
      "votes": 1,
      "replies": [
        {
          "id": 2913054,
          "postDate": "2024-07-09T07:48:00.567Z",
          "content": "<p>Who can image that 1dCNN with SMILES is the key to success😂</p>",
          "rawMarkdown": "Who can image that 1dCNN with SMILES is the key to success😂",
          "votes": 1,
          "replies": [
            {
              "id": 2913140,
              "postDate": "2024-07-09T09:28:21.990Z",
              "content": "<p>All my ensembles and advanced models are crying in corner with low private lbs! 😂</p>",
              "rawMarkdown": "All my ensembles and advanced models are crying in corner with low private lbs! 😂"
            }
          ]
        }
      ]
    },
    {
      "id": 2912660,
      "postDate": "2024-07-09T02:33:34.627Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168426%2F7f146caae9eef7ef93c4ab9cb3c52e05%2FCapture.JPG?generation=1720492033750519&amp;alt=media\"></p>\n<p>Transformer:<br>\nStep 1: pretrain - Masked input prediction on test data<br>\nStep 2: Morgan Fingerprint prediction on test data<br>\nStep 3: Single epoch on full train data</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168426%2F7f146caae9eef7ef93c4ab9cb3c52e05%2FCapture.JPG?generation=1720492033750519&alt=media)\n\nTransformer:\nStep 1: pretrain - Masked input prediction on test data\nStep 2: Morgan Fingerprint prediction on test data\nStep 3: Single epoch on full train data",
      "votes": 1
    },
    {
      "id": 2912654,
      "postDate": "2024-07-09T02:30:40.083Z",
      "content": "<p>My best private is 0.311 ! ony one fold cnn, it's just for test…. 3.89 in public ,so…. I   discarded it，crazy!</p>",
      "rawMarkdown": "My best private is 0.311 ! ony one fold cnn, it's just for test.... 3.89 in public ,so.... I   discarded it，crazy!",
      "votes": 1
    },
    {
      "id": 2912792,
      "postDate": "2024-07-09T04:23:25.267Z",
      "content": "<p>i think the host/kaggle have all the backend results. you can request the host/kaggle staff to reveal submiited score, not private scores; eg, just scores without names?</p>",
      "rawMarkdown": "i think the host/kaggle have all the backend results. you can request the host/kaggle staff to reveal submiited score, not private scores; eg, just scores without names?",
      "votes": 2
    },
    {
      "id": 2912536,
      "postDate": "2024-07-09T00:59:41.023Z",
      "content": "<p>0.261 was my best… it was my own 1dCNN before the 15 fold public notebook came out. And did no better than the 15fold 20 epoch public notebook! Crazy enough, I guess no one submitted the highest scoring public single model as is! Only eight people at 261 on the LB, the clearly best run of the clearly best model would've gotten in the 70s place (71-77 range) and silver.</p>",
      "rawMarkdown": "0.261 was my best... it was my own 1dCNN before the 15 fold public notebook came out. And did no better than the 15fold 20 epoch public notebook! Crazy enough, I guess no one submitted the highest scoring public single model as is! Only eight people at 261 on the LB, the clearly best run of the clearly best model would've gotten in the 70s place (71-77 range) and silver.",
      "votes": 2,
      "replies": [
        {
          "id": 2912537,
          "postDate": "2024-07-09T01:00:12.343Z",
          "content": "<p>Reference: <a href=\"https://www.kaggle.com/code/hugowjd/belka-1dcnn-with-all-data-15-folds-20-epoch?scriptVersionId=186125192\" target=\"_blank\">https://www.kaggle.com/code/hugowjd/belka-1dcnn-with-all-data-15-folds-20-epoch?scriptVersionId=186125192</a></p>",
          "rawMarkdown": "Reference: https://www.kaggle.com/code/hugowjd/belka-1dcnn-with-all-data-15-folds-20-epoch?scriptVersionId=186125192",
          "replies": [
            {
              "id": 2912546,
              "postDate": "2024-07-09T01:05:07.600Z",
              "content": "<p>Great! thank you. I think probably the model using only SMILES would achieve a better results in private LB. The GNN model might perform well in public LB but worse in private LB.</p>",
              "rawMarkdown": "Great! thank you. I think probably the model using only SMILES would achieve a better results in private LB. The GNN model might perform well in public LB but worse in private LB."
            }
          ]
        },
        {
          "id": 2912570,
          "postDate": "2024-07-09T01:20:32.440Z",
          "content": "<p>I also got private 0.261 (public 0.409) for my own 1dcnn model (12 folds for 25 epochs). But my far best was 0.272 (public 0.406, only 6 folds of those 12) would be silver 18. place. </p>\n<p>Unfortunately I didn't have the courage to select submission with only a single type of model.<br>\nI think that the big success of this 1dcnn trained on TPU was that it was the only possibility for a participant without a powerful HW to seriously train on all data. </p>",
          "rawMarkdown": "I also got private 0.261 (public 0.409) for my own 1dcnn model (12 folds for 25 epochs). But my far best was 0.272 (public 0.406, only 6 folds of those 12) would be silver 18. place. \n\nUnfortunately I didn't have the courage to select submission with only a single type of model.\nI think that the big success of this 1dcnn trained on TPU was that it was the only possibility for a participant without a powerful HW to seriously train on all data. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2912978,
      "postDate": "2024-07-09T06:48:00.563Z",
      "content": "<p>Smiles 1DCNN 10 Folds With 10M Sample Data<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2690746%2F612c1b9d20fa102adcceda1a5cc24834%2Fsmile.jpg?generation=1720507548241297&amp;alt=media\"></p>",
      "rawMarkdown": "Smiles 1DCNN 10 Folds With 10M Sample Data\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2690746%2F612c1b9d20fa102adcceda1a5cc24834%2Fsmile.jpg?generation=1720507548241297&alt=media)",
      "replies": [
        {
          "id": 2913050,
          "postDate": "2024-07-09T07:44:49.600Z",
          "content": "<p>I guess if you trained on all data, you'll get score near to 0.3.</p>",
          "rawMarkdown": "I guess if you trained on all data, you'll get score near to 0.3.",
          "replies": [
            {
              "id": 2913083,
              "postDate": "2024-07-09T08:31:15.490Z",
              "content": "<p>jajaja, I trained with all data with worst results.</p>",
              "rawMarkdown": "jajaja, I trained with all data with worst results."
            }
          ]
        }
      ]
    },
    {
      "id": 2912826,
      "postDate": "2024-07-09T05:14:01.450Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3678839%2Ffccb411de493f8306fcfc3e05349727d%2F2024-07-09%2008_13_41-NeurIPS%202024%20-%20Predict%20New%20Medicines%20with%20BELKA%20_%20Kaggle.png?generation=1720502039249388&amp;alt=media\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3678839%2Ffccb411de493f8306fcfc3e05349727d%2F2024-07-09%2008_13_41-NeurIPS%202024%20-%20Predict%20New%20Medicines%20with%20BELKA%20_%20Kaggle.png?generation=1720502039249388&alt=media)"
    },
    {
      "id": 2912670,
      "postDate": "2024-07-09T02:45:48.587Z",
      "content": "<p>emmmm, so weird…<br>\ncnn1d 0.291  single fold single model   <br>\nmpnn  0.284 single fold single model <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3318301%2Fb38d635e14375fa2b9cc2cd12fb2ccf5%2Fcnn1d-large.png?generation=1720492917744976&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3318301%2F77efd895a92b0b63da5c96f9e126a067%2Fmpnn_0.284.png?generation=1720493127393418&amp;alt=media\"></p>",
      "rawMarkdown": "emmmm, so weird...\ncnn1d 0.291  single fold single model   \nmpnn  0.284 single fold single model \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3318301%2Fb38d635e14375fa2b9cc2cd12fb2ccf5%2Fcnn1d-large.png?generation=1720492917744976&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3318301%2F77efd895a92b0b63da5c96f9e126a067%2Fmpnn_0.284.png?generation=1720493127393418&alt=media)\n",
      "replies": [
        {
          "id": 2912695,
          "postDate": "2024-07-09T03:20:21.850Z",
          "content": "<p>maybe one fold and less trained does not lead to overfitting and leads to more generalized models.. </p>",
          "rawMarkdown": "maybe one fold and less trained does not lead to overfitting and leads to more generalized models.. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2912667,
      "postDate": "2024-07-09T02:42:17.160Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6148721%2F5d159dd7d7e3ef1851ffdb0cc3852415%2F2024-07-09%20%2011.12.55.png?generation=1720492580675212&amp;alt=media\"><br>\nbasic mpnn result (with slight modifications from shared code by hengck23). I focused on developing 2d gnn (ex GAT, GIN..) but got worse results than the simplest mpnn in private LB. One interesting finding is that SSL (graph contrastive learning) with only building block graphs (only 1145 graphs) gives slight improvements (GAT: 0.246 -&gt; 0.263, GIN: 0.250 -&gt; 0.264). For now, sequence models seem to be the best, but I think approaches like MPNN + SSL + data augmentation also have potential..</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6148721%2F5d159dd7d7e3ef1851ffdb0cc3852415%2F2024-07-09%20%2011.12.55.png?generation=1720492580675212&alt=media)\nbasic mpnn result (with slight modifications from shared code by hengck23). I focused on developing 2d gnn (ex GAT, GIN..) but got worse results than the simplest mpnn in private LB. One interesting finding is that SSL (graph contrastive learning) with only building block graphs (only 1145 graphs) gives slight improvements (GAT: 0.246 -> 0.263, GIN: 0.250 -> 0.264). For now, sequence models seem to be the best, but I think approaches like MPNN + SSL + data augmentation also have potential.."
    },
    {
      "id": 2912665,
      "postDate": "2024-07-09T02:40:56.757Z",
      "content": "<p>In my case, it is a single-fold MolFormer</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1364892%2F7894a3533a94c1719e86428c94dc2724%2Fmolformer_best.png?generation=1720492848749301&amp;alt=media\"></p>",
      "rawMarkdown": "In my case, it is a single-fold MolFormer\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1364892%2F7894a3533a94c1719e86428c94dc2724%2Fmolformer_best.png?generation=1720492848749301&alt=media)"
    },
    {
      "id": 2912627,
      "postDate": "2024-07-09T02:10:06.217Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1206274%2Fb34aaf24834b2883643faa53d8dc599c%2F20240709-100937.jpg?generation=1720490988441362&amp;alt=media\"></p>\n<p>transformers + data-aug…</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1206274%2Fb34aaf24834b2883643faa53d8dc599c%2F20240709-100937.jpg?generation=1720490988441362&alt=media)\n\ntransformers + data-aug...",
      "replies": [
        {
          "id": 2912671,
          "postDate": "2024-07-09T02:50:54.407Z",
          "content": "<p>Wow! Expecting your solutions!</p>",
          "rawMarkdown": "Wow! Expecting your solutions!"
        }
      ]
    },
    {
      "id": 2912553,
      "postDate": "2024-07-09T01:15:26.397Z",
      "content": "<p>O.229 private, 0.398 public with 20 GNN conv layers.  My 24 layer model gives 0.477 Pub, but 0.212 private.  I have not submitted any model below 20 layers.  24 layers performed better than 20 at my local machine.  I might get better scores with 6 or less conv layers. <br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1415161%2Fbe9ad117a57b66806518df5da8efda38%2FScreenshot%202024-07-08%20at%209.24.51PM.png?generation=1720488362090317&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1415161%2Fee1c8c06c60ca28bd16d48992ddb3d72%2FScreenshot%202024-07-08%20at%209.27.02PM.png?generation=1720488482862233&amp;alt=media\"></p>",
      "rawMarkdown": "O.229 private, 0.398 public with 20 GNN conv layers.  My 24 layer model gives 0.477 Pub, but 0.212 private.  I have not submitted any model below 20 layers.  24 layers performed better than 20 at my local machine.  I might get better scores with 6 or less conv layers. \n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1415161%2Fbe9ad117a57b66806518df5da8efda38%2FScreenshot%202024-07-08%20at%209.24.51PM.png?generation=1720488362090317&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1415161%2Fee1c8c06c60ca28bd16d48992ddb3d72%2FScreenshot%202024-07-08%20at%209.27.02PM.png?generation=1720488482862233&alt=media)",
      "replies": [
        {
          "id": 2912576,
          "postDate": "2024-07-09T01:22:54.070Z",
          "content": "<p>I think the model architecture might not be the reason. Maybe the input or used features matter. It seems that the model using only SMILES will work the best. </p>",
          "rawMarkdown": "I think the model architecture might not be the reason. Maybe the input or used features matter. It seems that the model using only SMILES will work the best. ",
          "votes": 1,
          "replies": [
            {
              "id": 2912593,
              "postDate": "2024-07-09T01:31:57.583Z",
              "content": "<p>Actually I believe some people already posted leading up to this point that sequence was doing better than GNN on non-shared data.</p>\n<p>Another thing I expected to see was models tuned too much towards shared BB. Best I think is to explicitly tune against holdout non-share, (if you have the time and GPU compute), especially early stopping (sooner) when it stops generalizing. But also less complex model, fewer layers is probably going to do better.</p>",
              "rawMarkdown": "Actually I believe some people already posted leading up to this point that sequence was doing better than GNN on non-shared data.\n\nAnother thing I expected to see was models tuned too much towards shared BB. Best I think is to explicitly tune against holdout non-share, (if you have the time and GPU compute), especially early stopping (sooner) when it stops generalizing. But also less complex model, fewer layers is probably going to do better.",
              "votes": 1
            },
            {
              "id": 2912595,
              "postDate": "2024-07-09T01:34:10.290Z",
              "content": "<p>I only used SMILES features. The edge features was also from SMILES, that could be the source of the overfitting.  </p>",
              "rawMarkdown": "I only used SMILES features. The edge features was also from SMILES, that could be the source of the overfitting.  "
            },
            {
              "id": 2912602,
              "postDate": "2024-07-09T01:41:44.917Z",
              "content": "<p>I mean view the SMILES as normal sequence and using NLP methods is the best. Get the atom and bond features and using GNN might not work well in this private LB. But this is just random, depends on what private LB focus on. If the molecules in private LB are more graph-similar to the public LB, then GNN might work the best. </p>",
              "rawMarkdown": "I mean view the SMILES as normal sequence and using NLP methods is the best. Get the atom and bond features and using GNN might not work well in this private LB. But this is just random, depends on what private LB focus on. If the molecules in private LB are more graph-similar to the public LB, then GNN might work the best. ",
              "votes": 1
            },
            {
              "id": 2912608,
              "postDate": "2024-07-09T01:51:56.347Z",
              "content": "<p>I used atom features from SMILES (dim 79 vector), edge feature(dim 7 vector).  I didn't use NLP embedding.  </p>",
              "rawMarkdown": "I used atom features from SMILES (dim 79 vector), edge feature(dim 7 vector).  I didn't use NLP embedding.  "
            },
            {
              "id": 2912624,
              "postDate": "2024-07-09T02:06:01.833Z",
              "content": "<p>I also tried similar model. like using the deepchem.graph_feature to get 75d atom feature and also use 53d word2vec features. then based on adjacent matrix to design the GNN. similar to your private LB</p>",
              "rawMarkdown": "I also tried similar model. like using the deepchem.graph_feature to get 75d atom feature and also use 53d word2vec features. then based on adjacent matrix to design the GNN. similar to your private LB",
              "votes": 1
            },
            {
              "id": 2912666,
              "postDate": "2024-07-09T02:41:03.840Z",
              "content": "<p>Thanks.  Retrospectively, the regular transformer should be the better choice.  It's basically a GNN with every nodes connecting to every other nodes.  The strength of connections will be determined by att_matrix.  The model wouldn't overfit particular bond connections in the graph.  </p>",
              "rawMarkdown": "Thanks.  Retrospectively, the regular transformer should be the better choice.  It's basically a GNN with every nodes connecting to every other nodes.  The strength of connections will be determined by att_matrix.  The model wouldn't overfit particular bond connections in the graph.  "
            }
          ]
        },
        {
          "id": 2912587,
          "postDate": "2024-07-09T01:27:24.623Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2912544,
      "postDate": "2024-07-09T01:03:52.537Z",
      "content": "<p>My best submission was 0.267 in private LB (mamba single fold), crazy game…</p>",
      "rawMarkdown": "My best submission was 0.267 in private LB (mamba single fold), crazy game..."
    },
    {
      "id": 2914146,
      "postDate": "2024-07-09T18:29:53.593Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2912548,
      "postDate": "2024-07-09T01:11:22.857Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2913022,
      "author_name": "Ogurtsov",
      "author_url": "",
      "post_date": "2024-07-09T07:26:21.470000",
      "content": "<p>All this scores mean nothing until we will see separate metrics for each of 9 parts averaged in total LB score. Just small random (!) variation in scores for less predictable (or almost unpredictable) non-triazine part can lead to diferences like 0.03 or larger in total scores.</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 2912552,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-07-09T01:14:39.230000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F75c4c2c30cf8343ee1ef0e09127d1855%2FSelection_024.png?generation=1720487665043888&amp;alt=media\"></p>\n<p>tx= transformer too (fa is flash attention)</p>\n<p>transformer is the king … actually</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2912567,
          "author_name": "Yifan Wu",
          "author_url": "",
          "post_date": "2024-07-09T01:19:06.213000",
          "content": "<p>What's the input to your transformer？ Is it only SMILES?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2912571,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-07-09T01:20:42.720000",
              "content": "<p>input charcters. it is based on the code i share</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2912586,
              "author_name": "Yifan Wu",
              "author_url": "",
              "post_date": "2024-07-09T01:27:20.777000",
              "content": "<p>input characters/SMILES is the king!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2912569,
          "author_name": "Yifan Wu",
          "author_url": "",
          "post_date": "2024-07-09T01:20:31.917000",
          "content": "<p>0.315! you're the champion！</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2912573,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-07-09T01:21:16.590000",
          "content": "<p>GNN and mamba and cnn1d score<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5121bf2f956d9d50ec163f74dea12300%2FSelection_025.png?generation=1720488206406902&amp;alt=media\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2912583,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2024-07-09T01:26:32.163000",
          "content": "<p>Congrats on 5th <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> !<br>\nWhat was your transformer's valid(share) and valid(non-share) scores?</p>\n<p>How did you pick your two submissions?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2912601,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-07-09T01:40:15.807000",
              "content": "<p>the private 0.315/public0.482 solution is actually public -solution.<br>\ni will private more details (on share nonshare scores) later.<br>\ni need some sleep first.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3c5aaf52bf89941d8beafe1e30adb2de%2FSelection_026.png?generation=1720489124135543&amp;alt=media\"></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2912606,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-07-09T01:46:10.540000",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb491118540076973b5def4528aa4e4c7%2FSelection_027.png?generation=1720489561920264&amp;alt=media\"></p>\n<p>top solution.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2912639,
              "author_name": "Swikwislkdjc",
              "author_url": "",
              "post_date": "2024-07-09T02:20:43.387000",
              "content": "<p>Go sleep🤣, Looking forward to your solution!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2913130,
          "author_name": "AC",
          "author_url": "",
          "post_date": "2024-07-09T09:24:10.780000",
          "content": "<p>Congratulations! <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 🥳</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2915601,
          "author_name": "Andrius Bernatavicius",
          "author_url": "",
          "post_date": "2024-07-10T15:24:21.797000",
          "content": "<p>Could you give details on the exact transformer architecture of your top-scoring model? <br>\nThe #1 winning solution used a seemingly under-parameterized transformer encoder, which is why (in my opinion) it physically could not overfit the training dataset and ended up generalizing well.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2913328,
      "author_name": "Maciej Sypetkowski",
      "author_url": "",
      "post_date": "2024-07-09T12:21:29.897000",
      "content": "<p>I have a model (not ensemble!) that got <strong>0.347</strong> on private and 0.352 on public. Basically it was a MLP on concatenation of Morgan fingerprints, MACCS and embeddings from a pretrained ChemBERT_chEMBL. It was also pretty bad on a local evaluation, especially on splits with unseen blocks (compared to my other models). That competition was so noisy…<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2631834%2F3243f56a961c037b54c6477ddf9e10ae%2Fscreenshot.png?generation=1720527753507162&amp;alt=media\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 2913498,
          "author_name": "Andrew D. Blevins",
          "author_url": "",
          "post_date": "2024-07-09T13:52:38.697000",
          "content": "<p>We were watching this submission wondering what it was. It was head and shoulders above everything else we saw, we were hoping it was the beginning of a breakthrough, but alas it probably just got lucky on a couple of the of the building blocks. Most of the performance gains come from an unusually high performance on non-share for sEH and HSA. Kin0 is what we call the non-triazine, and the performance for this model is 0 :(<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F57970%2Fd39acf25211362a20daef30e2b460829%2FScreen%20Shot%202024-07-09%20at%207.50.55%20AM.png?generation=1720532964733655&amp;alt=media\"></p>\n<p>I will work on generating more of these plots for other submissions, and release a notebook so that people can do this themselves.</p>",
          "votes": 14,
          "replies": [
            {
              "id": 2913581,
              "author_name": "Swikwislkdjc",
              "author_url": "",
              "post_date": "2024-07-09T15:02:35.757000",
              "content": "<p>Looking forward for the notebook!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2913638,
              "author_name": "Maciej Sypetkowski",
              "author_url": "",
              "post_date": "2024-07-09T15:32:35.573000",
              "content": "<p>How ironic that a competition designed to test generalization could be won by a solution that does not generalize at all 😅</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2913655,
              "author_name": "Andrew D. Blevins",
              "author_url": "",
              "post_date": "2024-07-09T15:36:25.333000",
              "content": "<p>generalizing to new building blocks is still something we usually do not see. But yeah, we were hoping for better</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2914566,
              "author_name": "joejeo1",
              "author_url": "",
              "post_date": "2024-07-10T03:32:56.953000",
              "content": "<p>You probably should put some non-triazine samples in the public LB set.   It's already hard enough for them not being in the training set.  It would be almost impossible if their scores are also hidden. Eventually, a fuzzy winning model with large inferencing data space due to model yet to define its inferencing space via under training (like one epoch) or other ways.  If you train the model a few more epochs, the model quickly shrinking its inferencing space close to the real training data space.  Then, the score starts to drop.  Honestly, there is almost no practical values in these fuzzy models. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2955899,
              "author_name": "AC",
              "author_url": "",
              "post_date": "2024-08-11T15:13:03.127000",
              "content": "<p>Created a notebook to visualize these plots for everybody to use it. You can find it <a href=\"https://www.kaggle.com/code/ahsuna123/visualizing-the-protein-group-wise-results?scriptVersionId=192147258\" target=\"_blank\">here</a>. Hope, it's helpful! :)<br>\n<a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> <a href=\"https://www.kaggle.com/lililycai\" target=\"_blank\">@lililycai</a> </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3382839,
              "author_name": "Yijie Xu",
              "author_url": "",
              "post_date": "2025-12-28T17:16:31.370000",
              "content": "<p><a href=\"https://www.kaggle.com/ahsuna123\" target=\"_blank\">@ahsuna123</a> is this notebook still available somewhere?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2914558,
          "author_name": "Yifan Wu",
          "author_url": "",
          "post_date": "2024-07-10T03:10:09.557000",
          "content": "<p>cool! I also have a model which try to use MLP with 7 types fingerprint and 3types pretrained molecular representation model. In public LB, I found the fingerprint is very usefull (the case is training model on whole dataset for some epochs). Finally the public LB is about 0.430 but private LB is only 0.220. So I'm curious about how many epochs or how many data do you use? Just like \"all you need is frequent checkpoint\", I think the training steps is very important to control the model's performance on unseen blocks.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2914569,
              "author_name": "joejeo1",
              "author_url": "",
              "post_date": "2024-07-10T03:38:28.397000",
              "content": "<p>I think you already input too much information with \"7 types fingerprint and 3types pretrained molecular representation model\".  </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2914595,
              "author_name": "Swikwislkdjc",
              "author_url": "",
              "post_date": "2024-07-10T04:15:43.067000",
              "content": "<p>agree, too much information can easily lead to overfitting</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2914697,
              "author_name": "Yifan Wu",
              "author_url": "",
              "post_date": "2024-07-10T06:02:45.297000",
              "content": "<p>You're right. Too much information will introduce many redundance to the model, but I still think the model can choose which part are chosen to use by the gradient descent. When we are talking about \"overfitting\", I think there are two aspect of meanning. One is to overfit the training dataset. When we are training the model, more useful features will lead to a faster convergence speed, what we need to do is to stop the training at an appropriate step. That's why we need the validation set. Another one is to generalize to other totally new data (new distribution). This part I think mainly depends on the features or how we represent the input data. In this competition, the case is the distribution of data in private LB is totally different from the data in public LB. But if we select a subset from the training set, or stop at some specific steps, at this monment, the model will have the best performance in the distribution of data in private LB. <br>\nAnother thought I have is that I don't think the AI model can really generalize to some totally new distribution of data, in CV or NLP, the training/validation/test set of most of the tasks belongs to a similar distribution (near to the distribution in native). However, in biological field, the difference of distribution between different target/molecules are even larger than different tasks in CV/NLP. Maybe some zero-shot/few-shot learning techniques will work better in biological field. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2913021,
      "author_name": "Mikhail Pershin",
      "author_url": "",
      "post_date": "2024-07-09T07:26:05.880000",
      "content": "<p>Had one model that scored <strong>0.330</strong> on private, though only 0.355 on public - so it wasn't selected as a final submission 🥲!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F7a868e57aace0fb83d5eb38c62fab31a%2Finbox_283587_cacabaea5900e2c2d1800eb40257db38_Screenshot%20from%202024-07-09%2008-23-53.png?generation=1720510160991417&amp;alt=media\" alt=\"330\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 2913046,
          "author_name": "Yifan Wu",
          "author_url": "",
          "post_date": "2024-07-09T07:41:59.037000",
          "content": "<p>Great!!!! You are the best now. So what special tricks or features did you use in this model?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2913076,
              "author_name": "Mikhail Pershin",
              "author_url": "",
              "post_date": "2024-07-09T08:24:17.030000",
              "content": "<p>Honestly, nothing really interesting about this one: the model itself is just an MLP classifier on top of ECFP6 fingerprints and GIN GNN embeddings (from <a href=\"https://molfeat.datamol.io/featurizers/gin_supervised_masking\" target=\"_blank\">here</a>), trained on ~1/4 of the training set (24M train rows in the reduced version).</p>\n<p>Generally I was trying to find a point where models start to overfit on non-share - this submission was just probing into how well my validation setup is set and calibrated. I had pushed this model even further without any degradation on the non-shared part - but it seems this particular submission might've been a sweet spot for the non-triazine maybe?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2913184,
              "author_name": "Antonina Dolgorukova",
              "author_url": "",
              "post_date": "2024-07-09T10:03:54.290000",
              "content": "<p>really interesting! could yo umake a submit to check it? (mask all triazines with 0 and see th escores)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2913290,
              "author_name": "Mikhail Pershin",
              "author_url": "",
              "post_date": "2024-07-09T11:40:27.707000",
              "content": "<p>Sure thing, did a check - and it doesn't seem like it scores anything interesting on non-triazine either, giving only 0.015 private LB score<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F8c9ba36b33a29b1d5ba12039626528f5%2FScreenshot%20from%202024-07-09%2012-37-58.png?generation=1720525193036236&amp;alt=media\" alt=\"0.015\"></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2913300,
              "author_name": "Mikhail Pershin",
              "author_url": "",
              "post_date": "2024-07-09T11:48:59.163000",
              "content": "<p>Essentially it means that non-triazine had absolutely no contribution to these scores (my hypothesis above is wrong).<br>\nI even tried now a submission with non-triazine masked - and the scores are exactly the same: 0.330 and 0.355    </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2913309,
              "author_name": "Antonina Dolgorukova",
              "author_url": "",
              "post_date": "2024-07-09T11:55:52.703000",
              "content": "<p>yes, I have scored a lot of models, it's always 0.015 (and once 0.016) 🧐, and i am getting same score for each separate protein</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2913359,
              "author_name": "Mikhail Pershin",
              "author_url": "",
              "post_date": "2024-07-09T12:49:58.607000",
              "content": "<p>Well submitting zeros is not exacly the same as masking / filtering - as you'd still get an AUC-PR for that part. In fact you'd get AUC-PR ~= positive rate, so 0.015 might be just an average positive rate in the test set</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2913383,
              "author_name": "Mikhail Pershin",
              "author_url": "",
              "post_date": "2024-07-09T12:59:54.960000",
              "content": "<p>Yes, in fact submitting all zeros gives exactly same - 0.015 private 0.023 public.</p>\n<p>Significant difference from 0.005 positive rate in train!!</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2912692,
      "author_name": "Hideaki Ogasawara",
      "author_url": "",
      "post_date": "2024-07-09T03:18:37.743000",
      "content": "<p>I submitted a total of 65 entries. But after all, I got my best score of 0.275 (public score 0.403) when I just ran the shared code <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">BELKA 1DCNN Starter with all data</a> using a different random seed. I’m left with mixed feelings.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2912958,
      "author_name": "Kaiser",
      "author_url": "",
      "post_date": "2024-07-09T06:38:59.943000",
      "content": "<p>0.309!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13344803%2F971aaeb81832abd19847ed05d4b86ada%2Fbest_private.PNG?generation=1720507067496050&amp;alt=media\"></p>\n<p>I read some paper that described the 1DCNN method but the authors had removed the last dropout layer.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2912753,
      "author_name": "AC",
      "author_url": "",
      "post_date": "2024-07-09T04:07:05.080000",
      "content": "<p>My best 0.310 lol! <br>\nMore details of the solution in here! <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/518957\" target=\"_blank\">0.310 private lb</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2913054,
          "author_name": "Yifan Wu",
          "author_url": "",
          "post_date": "2024-07-09T07:48:00.567000",
          "content": "<p>Who can image that 1dCNN with SMILES is the key to success😂</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2913140,
              "author_name": "AC",
              "author_url": "",
              "post_date": "2024-07-09T09:28:21.990000",
              "content": "<p>All my ensembles and advanced models are crying in corner with low private lbs! 😂</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2912660,
      "author_name": "Elahi",
      "author_url": "",
      "post_date": "2024-07-09T02:33:34.627000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168426%2F7f146caae9eef7ef93c4ab9cb3c52e05%2FCapture.JPG?generation=1720492033750519&amp;alt=media\"></p>\n<p>Transformer:<br>\nStep 1: pretrain - Masked input prediction on test data<br>\nStep 2: Morgan Fingerprint prediction on test data<br>\nStep 3: Single epoch on full train data</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2912654,
      "author_name": "nickyxuc",
      "author_url": "",
      "post_date": "2024-07-09T02:30:40.083000",
      "content": "<p>My best private is 0.311 ! ony one fold cnn, it's just for test…. 3.89 in public ,so…. I   discarded it，crazy!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2912792,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-07-09T04:23:25.267000",
      "content": "<p>i think the host/kaggle have all the backend results. you can request the host/kaggle staff to reveal submiited score, not private scores; eg, just scores without names?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2912536,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2024-07-09T00:59:41.023000",
      "content": "<p>0.261 was my best… it was my own 1dCNN before the 15 fold public notebook came out. And did no better than the 15fold 20 epoch public notebook! Crazy enough, I guess no one submitted the highest scoring public single model as is! Only eight people at 261 on the LB, the clearly best run of the clearly best model would've gotten in the 70s place (71-77 range) and silver.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2912537,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2024-07-09T01:00:12.343000",
          "content": "<p>Reference: <a href=\"https://www.kaggle.com/code/hugowjd/belka-1dcnn-with-all-data-15-folds-20-epoch?scriptVersionId=186125192\" target=\"_blank\">https://www.kaggle.com/code/hugowjd/belka-1dcnn-with-all-data-15-folds-20-epoch?scriptVersionId=186125192</a></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2912546,
              "author_name": "Yifan Wu",
              "author_url": "",
              "post_date": "2024-07-09T01:05:07.600000",
              "content": "<p>Great! thank you. I think probably the model using only SMILES would achieve a better results in private LB. The GNN model might perform well in public LB but worse in private LB.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2912570,
          "author_name": "Allie K.",
          "author_url": "",
          "post_date": "2024-07-09T01:20:32.440000",
          "content": "<p>I also got private 0.261 (public 0.409) for my own 1dcnn model (12 folds for 25 epochs). But my far best was 0.272 (public 0.406, only 6 folds of those 12) would be silver 18. place. </p>\n<p>Unfortunately I didn't have the courage to select submission with only a single type of model.<br>\nI think that the big success of this 1dcnn trained on TPU was that it was the only possibility for a participant without a powerful HW to seriously train on all data. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2912978,
      "author_name": "Ricardo Colomer",
      "author_url": "",
      "post_date": "2024-07-09T06:48:00.563000",
      "content": "<p>Smiles 1DCNN 10 Folds With 10M Sample Data<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2690746%2F612c1b9d20fa102adcceda1a5cc24834%2Fsmile.jpg?generation=1720507548241297&amp;alt=media\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2913050,
          "author_name": "Yifan Wu",
          "author_url": "",
          "post_date": "2024-07-09T07:44:49.600000",
          "content": "<p>I guess if you trained on all data, you'll get score near to 0.3.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2913083,
              "author_name": "Ricardo Colomer",
              "author_url": "",
              "post_date": "2024-07-09T08:31:15.490000",
              "content": "<p>jajaja, I trained with all data with worst results.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2912826,
      "author_name": "Mahmoud Elshahed",
      "author_url": "",
      "post_date": "2024-07-09T05:14:01.450000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3678839%2Ffccb411de493f8306fcfc3e05349727d%2F2024-07-09%2008_13_41-NeurIPS%202024%20-%20Predict%20New%20Medicines%20with%20BELKA%20_%20Kaggle.png?generation=1720502039249388&amp;alt=media\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2912670,
      "author_name": "Magichuang",
      "author_url": "",
      "post_date": "2024-07-09T02:45:48.587000",
      "content": "<p>emmmm, so weird…<br>\ncnn1d 0.291  single fold single model   <br>\nmpnn  0.284 single fold single model <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3318301%2Fb38d635e14375fa2b9cc2cd12fb2ccf5%2Fcnn1d-large.png?generation=1720492917744976&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3318301%2F77efd895a92b0b63da5c96f9e126a067%2Fmpnn_0.284.png?generation=1720493127393418&amp;alt=media\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2912695,
          "author_name": "Swikwislkdjc",
          "author_url": "",
          "post_date": "2024-07-09T03:20:21.850000",
          "content": "<p>maybe one fold and less trained does not lead to overfitting and leads to more generalized models.. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2912667,
      "author_name": "Taegoo",
      "author_url": "",
      "post_date": "2024-07-09T02:42:17.160000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6148721%2F5d159dd7d7e3ef1851ffdb0cc3852415%2F2024-07-09%20%2011.12.55.png?generation=1720492580675212&amp;alt=media\"><br>\nbasic mpnn result (with slight modifications from shared code by hengck23). I focused on developing 2d gnn (ex GAT, GIN..) but got worse results than the simplest mpnn in private LB. One interesting finding is that SSL (graph contrastive learning) with only building block graphs (only 1145 graphs) gives slight improvements (GAT: 0.246 -&gt; 0.263, GIN: 0.250 -&gt; 0.264). For now, sequence models seem to be the best, but I think approaches like MPNN + SSL + data augmentation also have potential..</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2912665,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2024-07-09T02:40:56.757000",
      "content": "<p>In my case, it is a single-fold MolFormer</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1364892%2F7894a3533a94c1719e86428c94dc2724%2Fmolformer_best.png?generation=1720492848749301&amp;alt=media\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2912627,
      "author_name": "Angle",
      "author_url": "",
      "post_date": "2024-07-09T02:10:06.217000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1206274%2Fb34aaf24834b2883643faa53d8dc599c%2F20240709-100937.jpg?generation=1720490988441362&amp;alt=media\"></p>\n<p>transformers + data-aug…</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2912671,
          "author_name": "Swikwislkdjc",
          "author_url": "",
          "post_date": "2024-07-09T02:50:54.407000",
          "content": "<p>Wow! Expecting your solutions!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2912553,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "2024-07-09T01:15:26.397000",
      "content": "<p>O.229 private, 0.398 public with 20 GNN conv layers.  My 24 layer model gives 0.477 Pub, but 0.212 private.  I have not submitted any model below 20 layers.  24 layers performed better than 20 at my local machine.  I might get better scores with 6 or less conv layers. <br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1415161%2Fbe9ad117a57b66806518df5da8efda38%2FScreenshot%202024-07-08%20at%209.24.51PM.png?generation=1720488362090317&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1415161%2Fee1c8c06c60ca28bd16d48992ddb3d72%2FScreenshot%202024-07-08%20at%209.27.02PM.png?generation=1720488482862233&amp;alt=media\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2912576,
          "author_name": "Yifan Wu",
          "author_url": "",
          "post_date": "2024-07-09T01:22:54.070000",
          "content": "<p>I think the model architecture might not be the reason. Maybe the input or used features matter. It seems that the model using only SMILES will work the best. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2912593,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2024-07-09T01:31:57.583000",
              "content": "<p>Actually I believe some people already posted leading up to this point that sequence was doing better than GNN on non-shared data.</p>\n<p>Another thing I expected to see was models tuned too much towards shared BB. Best I think is to explicitly tune against holdout non-share, (if you have the time and GPU compute), especially early stopping (sooner) when it stops generalizing. But also less complex model, fewer layers is probably going to do better.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2912595,
              "author_name": "joejeo1",
              "author_url": "",
              "post_date": "2024-07-09T01:34:10.290000",
              "content": "<p>I only used SMILES features. The edge features was also from SMILES, that could be the source of the overfitting.  </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2912602,
              "author_name": "Yifan Wu",
              "author_url": "",
              "post_date": "2024-07-09T01:41:44.917000",
              "content": "<p>I mean view the SMILES as normal sequence and using NLP methods is the best. Get the atom and bond features and using GNN might not work well in this private LB. But this is just random, depends on what private LB focus on. If the molecules in private LB are more graph-similar to the public LB, then GNN might work the best. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2912608,
              "author_name": "joejeo1",
              "author_url": "",
              "post_date": "2024-07-09T01:51:56.347000",
              "content": "<p>I used atom features from SMILES (dim 79 vector), edge feature(dim 7 vector).  I didn't use NLP embedding.  </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2912624,
              "author_name": "Yifan Wu",
              "author_url": "",
              "post_date": "2024-07-09T02:06:01.833000",
              "content": "<p>I also tried similar model. like using the deepchem.graph_feature to get 75d atom feature and also use 53d word2vec features. then based on adjacent matrix to design the GNN. similar to your private LB</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2912666,
              "author_name": "joejeo1",
              "author_url": "",
              "post_date": "2024-07-09T02:41:03.840000",
              "content": "<p>Thanks.  Retrospectively, the regular transformer should be the better choice.  It's basically a GNN with every nodes connecting to every other nodes.  The strength of connections will be determined by att_matrix.  The model wouldn't overfit particular bond connections in the graph.  </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2912587,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-07-09T01:27:24.623000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2912544,
      "author_name": "Leo Lu",
      "author_url": "",
      "post_date": "2024-07-09T01:03:52.537000",
      "content": "<p>My best submission was 0.267 in private LB (mamba single fold), crazy game…</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2914146,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-09T18:29:53.593000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2912548,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-09T01:11:22.857000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2912532": "This is really a fucking crazy shuffle in private LB. And we can check all of our private LB results now. My best private LB is from my customed pretrained LM using only SMILES. In public LB, it is only 0.391 but in private LB, it's 0.294!. (SHIT, I should trust my dream! Why not trust my own pretrained model!!! )\n \n![fk](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1029068%2F9227f421039af5fdc45d8c33e8bb5893%2F1720486391365.jpg?generation=1720486410872878&alt=media)",
    "2913022": "All this scores mean nothing until we will see separate metrics for each of 9 parts averaged in total LB score. Just small random (!) variation in scores for less predictable (or almost unpredictable) non-triazine part can lead to diferences like 0.03 or larger in total scores.",
    "2912552": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F75c4c2c30cf8343ee1ef0e09127d1855%2FSelection_024.png?generation=1720487665043888&alt=media)\n\ntx= transformer too (fa is flash attention)\n\ntransformer is the king ... actually",
    "2913328": "I have a model (not ensemble!) that got **0.347** on private and 0.352 on public. Basically it was a MLP on concatenation of Morgan fingerprints, MACCS and embeddings from a pretrained ChemBERT_chEMBL. It was also pretty bad on a local evaluation, especially on splits with unseen blocks (compared to my other models). That competition was so noisy...\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2631834%2F3243f56a961c037b54c6477ddf9e10ae%2Fscreenshot.png?generation=1720527753507162&alt=media)",
    "2913021": "Had one model that scored **0.330** on private, though only 0.355 on public - so it wasn't selected as a final submission 🥲!\n\n![330](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F283587%2F7a868e57aace0fb83d5eb38c62fab31a%2Finbox_283587_cacabaea5900e2c2d1800eb40257db38_Screenshot%20from%202024-07-09%2008-23-53.png?generation=1720510160991417&alt=media)",
    "2912692": "I submitted a total of 65 entries. But after all, I got my best score of 0.275 (public score 0.403) when I just ran the shared code [BELKA 1DCNN Starter with all data](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data) using a different random seed. I’m left with mixed feelings.",
    "2912958": "0.309!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13344803%2F971aaeb81832abd19847ed05d4b86ada%2Fbest_private.PNG?generation=1720507067496050&alt=media)\n\nI read some paper that described the 1DCNN method but the authors had removed the last dropout layer.",
    "2912753": "My best 0.310 lol! \nMore details of the solution in here! [0.310 private lb](https://www.kaggle.com/competitions/leash-BELKA/discussion/518957)",
    "2912660": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168426%2F7f146caae9eef7ef93c4ab9cb3c52e05%2FCapture.JPG?generation=1720492033750519&alt=media)\n\nTransformer:\nStep 1: pretrain - Masked input prediction on test data\nStep 2: Morgan Fingerprint prediction on test data\nStep 3: Single epoch on full train data",
    "2912654": "My best private is 0.311 ! ony one fold cnn, it's just for test.... 3.89 in public ,so.... I   discarded it，crazy!",
    "2912792": "i think the host/kaggle have all the backend results. you can request the host/kaggle staff to reveal submiited score, not private scores; eg, just scores without names?",
    "2912536": "0.261 was my best... it was my own 1dCNN before the 15 fold public notebook came out. And did no better than the 15fold 20 epoch public notebook! Crazy enough, I guess no one submitted the highest scoring public single model as is! Only eight people at 261 on the LB, the clearly best run of the clearly best model would've gotten in the 70s place (71-77 range) and silver.",
    "2912978": "Smiles 1DCNN 10 Folds With 10M Sample Data\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2690746%2F612c1b9d20fa102adcceda1a5cc24834%2Fsmile.jpg?generation=1720507548241297&alt=media)",
    "2912826": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3678839%2Ffccb411de493f8306fcfc3e05349727d%2F2024-07-09%2008_13_41-NeurIPS%202024%20-%20Predict%20New%20Medicines%20with%20BELKA%20_%20Kaggle.png?generation=1720502039249388&alt=media)",
    "2912670": "emmmm, so weird...\ncnn1d 0.291  single fold single model   \nmpnn  0.284 single fold single model \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3318301%2Fb38d635e14375fa2b9cc2cd12fb2ccf5%2Fcnn1d-large.png?generation=1720492917744976&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3318301%2F77efd895a92b0b63da5c96f9e126a067%2Fmpnn_0.284.png?generation=1720493127393418&alt=media)\n",
    "2912667": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6148721%2F5d159dd7d7e3ef1851ffdb0cc3852415%2F2024-07-09%20%2011.12.55.png?generation=1720492580675212&alt=media)\nbasic mpnn result (with slight modifications from shared code by hengck23). I focused on developing 2d gnn (ex GAT, GIN..) but got worse results than the simplest mpnn in private LB. One interesting finding is that SSL (graph contrastive learning) with only building block graphs (only 1145 graphs) gives slight improvements (GAT: 0.246 -> 0.263, GIN: 0.250 -> 0.264). For now, sequence models seem to be the best, but I think approaches like MPNN + SSL + data augmentation also have potential..",
    "2912665": "In my case, it is a single-fold MolFormer\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1364892%2F7894a3533a94c1719e86428c94dc2724%2Fmolformer_best.png?generation=1720492848749301&alt=media)",
    "2912627": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1206274%2Fb34aaf24834b2883643faa53d8dc599c%2F20240709-100937.jpg?generation=1720490988441362&alt=media)\n\ntransformers + data-aug...",
    "2912553": "O.229 private, 0.398 public with 20 GNN conv layers.  My 24 layer model gives 0.477 Pub, but 0.212 private.  I have not submitted any model below 20 layers.  24 layers performed better than 20 at my local machine.  I might get better scores with 6 or less conv layers. \n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1415161%2Fbe9ad117a57b66806518df5da8efda38%2FScreenshot%202024-07-08%20at%209.24.51PM.png?generation=1720488362090317&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1415161%2Fee1c8c06c60ca28bd16d48992ddb3d72%2FScreenshot%202024-07-08%20at%209.27.02PM.png?generation=1720488482862233&alt=media)",
    "2912544": "My best submission was 0.267 in private LB (mamba single fold), crazy game...",
    "2914146": "",
    "2912548": ""
  }
}