{
  "id": 518939,
  "title": "What do we learn from this competition? Is the result meaningful?",
  "url": "/competitions/leash-BELKA/discussion/518939",
  "author_name": "Steven_Y",
  "post_date": "2024-07-09T00:26:06.504000",
  "votes": 11,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Although a huge shakeup was expected, but it is still so hard to believe that this is the result. From a Kaggler's perspective, I learned a lot throughout the journey. But after all, is this result helpful for the task at all? Or is it more of a proof that there's still a long way to go for ML in the field of small molecules…</p>",
  "messages": [
    {
      "id": 2912592,
      "postDate": "2024-07-09T01:31:13.620Z",
      "content": "<p>i believe such imabalance or domain shift competition (hence also the shakeup) is getting common in kaggle.<br>\ntakeaways from this:</p>\n<ul>\n<li>try many different models</li>\n<li>do less processing that would lead to overfitting</li>\n</ul>\n<p>these increase your chance of hitting the jackpot.  but … it still a jackpot</p>",
      "rawMarkdown": "i believe such imabalance or domain shift competition (hence also the shakeup) is getting common in kaggle.\ntakeaways from this:\n- try many different models\n- do less processing that would lead to overfitting\n\nthese increase your chance of hitting the jackpot.  but ... it still a jackpot",
      "votes": 11,
      "replies": [
        {
          "id": 2912609,
          "postDate": "2024-07-09T01:55:43.087Z",
          "content": "<p>Hahah yeah…. we indeed tried many different models, even KAN at the end lol. Congrats on another solo Gold! Also thank you so much for your contributions in this competition!!</p>",
          "rawMarkdown": "Hahah yeah.... we indeed tried many different models, even KAN at the end lol. Congrats on another solo Gold! Also thank you so much for your contributions in this competition!!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2912508,
      "postDate": "2024-07-09T00:26:06.503Z",
      "content": "<p>Although a huge shakeup was expected, but it is still so hard to believe that this is the result. From a Kaggler's perspective, I learned a lot throughout the journey. But after all, is this result helpful for the task at all? Or is it more of a proof that there's still a long way to go for ML in the field of small molecules…</p>",
      "rawMarkdown": "Although a huge shakeup was expected, but it is still so hard to believe that this is the result. From a Kaggler's perspective, I learned a lot throughout the journey. But after all, is this result helpful for the task at all? Or is it more of a proof that there's still a long way to go for ML in the field of small molecules...",
      "votes": 11
    },
    {
      "id": 2913575,
      "postDate": "2024-07-09T14:58:11.477Z",
      "content": "<p>Machine learning is built on the assumption that data is IID (independent and identically distributed).</p>\n<p>The test set for the public LB is not drawn from the same distribution as the examples we used to train our models, which violates this core assumption, and explains why the scores are so wildly different. This falls under the umbrella I call \"data science malpractice\". When we try to apply ML to inappropriate classes of problems, the results are unreliable/unpredictable.</p>",
      "rawMarkdown": "Machine learning is built on the assumption that data is IID (independent and identically distributed).\n\nThe test set for the public LB is not drawn from the same distribution as the examples we used to train our models, which violates this core assumption, and explains why the scores are so wildly different. This falls under the umbrella I call \"data science malpractice\". When we try to apply ML to inappropriate classes of problems, the results are unreliable/unpredictable.",
      "votes": 7,
      "replies": [
        {
          "id": 2913620,
          "postDate": "2024-07-09T15:25:48.023Z",
          "content": "<p>but IID depends on what space the data is distributed over. At the end of the day the same physics applies to all molecules and proteins, so maybe there is a representation that allows for generalization. not yet it seems. For ML to be useful in this field you need to be able to train on some data and then evaluate on catalog molecules you haven't seen before. We thought that the way we split the data was a good approximation of the real world usecase.</p>\n<p>We didn't expect this problem to be solved with this amount of data. but many academic groups and companies have been claiming to have usefully solved this problem with less data. We wanted to call their bluff. We think this problem can be solved, but we think it will take datasets with many millions of molecules and thousands of proteins.</p>",
          "rawMarkdown": "but IID depends on what space the data is distributed over. At the end of the day the same physics applies to all molecules and proteins, so maybe there is a representation that allows for generalization. not yet it seems. For ML to be useful in this field you need to be able to train on some data and then evaluate on catalog molecules you haven't seen before. We thought that the way we split the data was a good approximation of the real world usecase.\n\nWe didn't expect this problem to be solved with this amount of data. but many academic groups and companies have been claiming to have usefully solved this problem with less data. We wanted to call their bluff. We think this problem can be solved, but we think it will take datasets with many millions of molecules and thousands of proteins.",
          "votes": 3,
          "replies": [
            {
              "id": 2914276,
              "postDate": "2024-07-09T20:02:11.923Z",
              "content": "<p>I believe features that could lead to better generalization include physical properties such as 3D information (XYZ coordinates) of a conformer's overall shape and atomic-level details, and its interactions with the proteins…..Maybe Alphafold3 could do this job super excellently.</p>",
              "rawMarkdown": "I believe features that could lead to better generalization include physical properties such as 3D information (XYZ coordinates) of a conformer's overall shape and atomic-level details, and its interactions with the proteins.....Maybe Alphafold3 could do this job super excellently."
            },
            {
              "id": 2914296,
              "postDate": "2024-07-09T20:40:27.867Z",
              "content": "<p>Yeah, we were hoping to see some co-folding/blind-docking models, or models pre-trained on PDBBind, but that is probably too compute expensive.</p>",
              "rawMarkdown": "Yeah, we were hoping to see some co-folding/blind-docking models, or models pre-trained on PDBBind, but that is probably too compute expensive."
            },
            {
              "id": 2914303,
              "postDate": "2024-07-09T20:53:29.550Z",
              "content": "<p>The barriers to entry might be the bigger concern. It's extremely non trivial to get a single docker setup working. And then if you want to do a second different docker there's another big one time cost, since no two are likely to be similar in the ramp up / setup needed. </p>\n<p>And easy to decide not to do it when one person publicly posts their try and doesn't even see any real predictive power on a tiny sample of train data. </p>",
              "rawMarkdown": "The barriers to entry might be the bigger concern. It's extremely non trivial to get a single docker setup working. And then if you want to do a second different docker there's another big one time cost, since no two are likely to be similar in the ramp up / setup needed. \n\nAnd easy to decide not to do it when one person publicly posts their try and doesn't even see any real predictive power on a tiny sample of train data. ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2912517,
      "postDate": "2024-07-09T00:35:34.177Z",
      "content": "<p>I dropped more than 1000.  I have never dropped more than 50 before.  </p>",
      "rawMarkdown": "I dropped more than 1000.  I have never dropped more than 50 before.  ",
      "votes": 1,
      "replies": [
        {
          "id": 2912551,
          "postDate": "2024-07-09T01:13:38.387Z",
          "content": "<p>The final lb feels indeed like a lottery…</p>",
          "rawMarkdown": "The final lb feels indeed like a lottery...",
          "votes": 4
        }
      ]
    },
    {
      "id": 2912545,
      "postDate": "2024-07-09T01:04:32.203Z",
      "content": "<p>It's the same problem as Novozymes, the millions of interaction data is a little helpful, but really you are predicting how 'good' each BB is, and that's only 1000 datapoints, and predicting against another 1000 unseen datapoints!</p>\n<p>Too many unknowns</p>\n<p>But it doesn't make the result meaningless. You look for the success story, who was comparatively unaffected, and probably they have good models and good innovation that can advance the field another step or two.</p>",
      "rawMarkdown": "It's the same problem as Novozymes, the millions of interaction data is a little helpful, but really you are predicting how 'good' each BB is, and that's only 1000 datapoints, and predicting against another 1000 unseen datapoints!\n\nToo many unknowns\n\nBut it doesn't make the result meaningless. You look for the success story, who was comparatively unaffected, and probably they have good models and good innovation that can advance the field another step or two.",
      "votes": 2,
      "replies": [
        {
          "id": 2912550,
          "postDate": "2024-07-09T01:11:47.677Z",
          "content": "<p>Yeah I can see that it's pretty much the case with Novozymes… We were the only team who stayed in the gold zone, but looking at our past submissions, there weren't much (probably not even any) significant improvements in the private lb scores… we will post our solution soon. </p>",
          "rawMarkdown": "Yeah I can see that it's pretty much the case with Novozymes... We were the only team who stayed in the gold zone, but looking at our past submissions, there weren't much (probably not even any) significant improvements in the private lb scores... we will post our solution soon. ",
          "votes": 3,
          "replies": [
            {
              "id": 2912564,
              "postDate": "2024-07-09T01:17:55.600Z",
              "content": "<p>Congrats on Gold! I look forward to reading your solution!</p>",
              "rawMarkdown": "Congrats on Gold! I look forward to reading your solution!",
              "votes": 1
            }
          ]
        },
        {
          "id": 2912625,
          "postDate": "2024-07-09T02:08:46.220Z",
          "content": "<p>HAHAHA yep it's Novozymes 2 as I suspected. More than half of gold range seems like random submits (old and total&lt;10). It's not easy to construct a good bio competition ~sigh</p>",
          "rawMarkdown": "HAHAHA yep it's Novozymes 2 as I suspected. More than half of gold range seems like random submits (old and total<10). It's not easy to construct a good bio competition ~sigh",
          "votes": 3,
          "replies": [
            {
              "id": 2912729,
              "postDate": "2024-07-09T03:49:46.023Z",
              "content": "<p>It's true! We struggled with choosing targets, chemical libraries, splits, scoring metrics, and so on. We're hopeful that the full data release will encourage folks to try their own and make progress on the problem. </p>",
              "rawMarkdown": "It's true! We struggled with choosing targets, chemical libraries, splits, scoring metrics, and so on. We're hopeful that the full data release will encourage folks to try their own and make progress on the problem. ",
              "votes": 5
            },
            {
              "id": 2912738,
              "postDate": "2024-07-09T03:54:57.537Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2912739,
              "postDate": "2024-07-09T03:55:13.937Z",
              "content": "<p>Yeah that'd be great, thanks!</p>",
              "rawMarkdown": "Yeah that'd be great, thanks!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2912688,
      "postDate": "2024-07-09T03:15:37.950Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2912509,
      "postDate": "2024-07-09T00:27:18.580Z",
      "rawMarkdown": "",
      "votes": 4,
      "isDeleted": true,
      "replies": [
        {
          "id": 2912514,
          "postDate": "2024-07-09T00:33:36.447Z",
          "content": "<p>Indeed… not only one but two did it…</p>",
          "rawMarkdown": "Indeed... not only one but two did it...",
          "votes": 1
        },
        {
          "id": 2912525,
          "postDate": "2024-07-09T00:42:36.750Z",
          "rawMarkdown": "",
          "isDeleted": true,
          "replies": [
            {
              "id": 2914299,
              "postDate": "2024-07-09T20:51:38.280Z",
              "content": "<p><a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">BELKA 1DCNN Starter with all data</a></p>",
              "rawMarkdown": "[BELKA 1DCNN Starter with all data](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data)",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2912592,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-07-09T01:31:13.620000",
      "content": "<p>i believe such imabalance or domain shift competition (hence also the shakeup) is getting common in kaggle.<br>\ntakeaways from this:</p>\n<ul>\n<li>try many different models</li>\n<li>do less processing that would lead to overfitting</li>\n</ul>\n<p>these increase your chance of hitting the jackpot.  but … it still a jackpot</p>",
      "votes": 11,
      "replies": [
        {
          "id": 2912609,
          "author_name": "Steven_Y",
          "author_url": "",
          "post_date": "2024-07-09T01:55:43.087000",
          "content": "<p>Hahah yeah…. we indeed tried many different models, even KAN at the end lol. Congrats on another solo Gold! Also thank you so much for your contributions in this competition!!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2913575,
      "author_name": "Steven Hewitt",
      "author_url": "",
      "post_date": "2024-07-09T14:58:11.477000",
      "content": "<p>Machine learning is built on the assumption that data is IID (independent and identically distributed).</p>\n<p>The test set for the public LB is not drawn from the same distribution as the examples we used to train our models, which violates this core assumption, and explains why the scores are so wildly different. This falls under the umbrella I call \"data science malpractice\". When we try to apply ML to inappropriate classes of problems, the results are unreliable/unpredictable.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2913620,
          "author_name": "Andrew D. Blevins",
          "author_url": "",
          "post_date": "2024-07-09T15:25:48.023000",
          "content": "<p>but IID depends on what space the data is distributed over. At the end of the day the same physics applies to all molecules and proteins, so maybe there is a representation that allows for generalization. not yet it seems. For ML to be useful in this field you need to be able to train on some data and then evaluate on catalog molecules you haven't seen before. We thought that the way we split the data was a good approximation of the real world usecase.</p>\n<p>We didn't expect this problem to be solved with this amount of data. but many academic groups and companies have been claiming to have usefully solved this problem with less data. We wanted to call their bluff. We think this problem can be solved, but we think it will take datasets with many millions of molecules and thousands of proteins.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2914276,
              "author_name": "Swikwislkdjc",
              "author_url": "",
              "post_date": "2024-07-09T20:02:11.923000",
              "content": "<p>I believe features that could lead to better generalization include physical properties such as 3D information (XYZ coordinates) of a conformer's overall shape and atomic-level details, and its interactions with the proteins…..Maybe Alphafold3 could do this job super excellently.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2914296,
              "author_name": "Andrew D. Blevins",
              "author_url": "",
              "post_date": "2024-07-09T20:40:27.867000",
              "content": "<p>Yeah, we were hoping to see some co-folding/blind-docking models, or models pre-trained on PDBBind, but that is probably too compute expensive.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2914303,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2024-07-09T20:53:29.550000",
              "content": "<p>The barriers to entry might be the bigger concern. It's extremely non trivial to get a single docker setup working. And then if you want to do a second different docker there's another big one time cost, since no two are likely to be similar in the ramp up / setup needed. </p>\n<p>And easy to decide not to do it when one person publicly posts their try and doesn't even see any real predictive power on a tiny sample of train data. </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2912517,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "2024-07-09T00:35:34.177000",
      "content": "<p>I dropped more than 1000.  I have never dropped more than 50 before.  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2912551,
          "author_name": "Steven_Y",
          "author_url": "",
          "post_date": "2024-07-09T01:13:38.387000",
          "content": "<p>The final lb feels indeed like a lottery…</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2912545,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2024-07-09T01:04:32.203000",
      "content": "<p>It's the same problem as Novozymes, the millions of interaction data is a little helpful, but really you are predicting how 'good' each BB is, and that's only 1000 datapoints, and predicting against another 1000 unseen datapoints!</p>\n<p>Too many unknowns</p>\n<p>But it doesn't make the result meaningless. You look for the success story, who was comparatively unaffected, and probably they have good models and good innovation that can advance the field another step or two.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2912550,
          "author_name": "Steven_Y",
          "author_url": "",
          "post_date": "2024-07-09T01:11:47.677000",
          "content": "<p>Yeah I can see that it's pretty much the case with Novozymes… We were the only team who stayed in the gold zone, but looking at our past submissions, there weren't much (probably not even any) significant improvements in the private lb scores… we will post our solution soon. </p>",
          "votes": 3,
          "replies": [
            {
              "id": 2912564,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2024-07-09T01:17:55.600000",
              "content": "<p>Congrats on Gold! I look forward to reading your solution!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2912625,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-07-09T02:08:46.220000",
          "content": "<p>HAHAHA yep it's Novozymes 2 as I suspected. More than half of gold range seems like random submits (old and total&lt;10). It's not easy to construct a good bio competition ~sigh</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2912729,
              "author_name": "Ian Quigley",
              "author_url": "",
              "post_date": "2024-07-09T03:49:46.023000",
              "content": "<p>It's true! We struggled with choosing targets, chemical libraries, splits, scoring metrics, and so on. We're hopeful that the full data release will encourage folks to try their own and make progress on the problem. </p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2912738,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-07-09T03:54:57.537000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2912739,
              "author_name": "Steven_Y",
              "author_url": "",
              "post_date": "2024-07-09T03:55:13.937000",
              "content": "<p>Yeah that'd be great, thanks!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2912688,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-09T03:15:37.950000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2912509,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-09T00:27:18.580000",
      "content": "",
      "votes": 4,
      "replies": [
        {
          "id": 2912514,
          "author_name": "Steven_Y",
          "author_url": "",
          "post_date": "2024-07-09T00:33:36.447000",
          "content": "<p>Indeed… not only one but two did it…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2912525,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-07-09T00:42:36.750000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2914299,
              "author_name": "Andrew D. Blevins",
              "author_url": "",
              "post_date": "2024-07-09T20:51:38.280000",
              "content": "<p><a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">BELKA 1DCNN Starter with all data</a></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2912592": "i believe such imabalance or domain shift competition (hence also the shakeup) is getting common in kaggle.\ntakeaways from this:\n- try many different models\n- do less processing that would lead to overfitting\n\nthese increase your chance of hitting the jackpot.  but ... it still a jackpot",
    "2912508": "Although a huge shakeup was expected, but it is still so hard to believe that this is the result. From a Kaggler's perspective, I learned a lot throughout the journey. But after all, is this result helpful for the task at all? Or is it more of a proof that there's still a long way to go for ML in the field of small molecules...",
    "2913575": "Machine learning is built on the assumption that data is IID (independent and identically distributed).\n\nThe test set for the public LB is not drawn from the same distribution as the examples we used to train our models, which violates this core assumption, and explains why the scores are so wildly different. This falls under the umbrella I call \"data science malpractice\". When we try to apply ML to inappropriate classes of problems, the results are unreliable/unpredictable.",
    "2912517": "I dropped more than 1000.  I have never dropped more than 50 before.  ",
    "2912545": "It's the same problem as Novozymes, the millions of interaction data is a little helpful, but really you are predicting how 'good' each BB is, and that's only 1000 datapoints, and predicting against another 1000 unseen datapoints!\n\nToo many unknowns\n\nBut it doesn't make the result meaningless. You look for the success story, who was comparatively unaffected, and probably they have good models and good innovation that can advance the field another step or two.",
    "2912688": "",
    "2912509": ""
  }
}