{
  "id": 506212,
  "title": "Related question about building blocks CV, for industry use",
  "url": "/competitions/leash-BELKA/discussion/506212",
  "author_name": "",
  "post_date": "2024-05-21T00:53:42.743134800Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi everyone, I hope it's ok to post this here since it's somewhat related and it might help others working in industry.</p>\n<p>I'm working on a material science project with a somewhat similar setup to this competition: we're using building blocks (BBs) to construct molecules, and our goal is property prediction (regression).</p>\n<p>I'm wondering about the best way to train models that generalize well to new, unseen BBs (similar to the non-shared BBs here). Would it be a good strategy to use cross-validation, where each fold intentionally hides a portion of the BBs?</p>\n<p>I know this would likely lower the cross-validation score compared to random splitting, but I think the predictions might be more realistic and less prone to overfitting.</p>\n<p>What are your thoughts on this approach? What about for shared BBs (interpolation within the chemical space)? Should I use the same models trained with BBs hidden in each fold, or would random splitting be sufficient in this case?</p>\n<p>Thanks in advance for your insights!</p>",
  "messages": [
    {
      "id": "2826484",
      "postDate": "05/21/2024 00:53:42",
      "content": "<p>Hi everyone, I hope it's ok to post this here since it's somewhat related and it might help others working in industry.</p>\n<p>I'm working on a material science project with a somewhat similar setup to this competition: we're using building blocks (BBs) to construct molecules, and our goal is property prediction (regression).</p>\n<p>I'm wondering about the best way to train models that generalize well to new, unseen BBs (similar to the non-shared BBs here). Would it be a good strategy to use cross-validation, where each fold intentionally hides a portion of the BBs?</p>\n<p>I know this would likely lower the cross-validation score compared to random splitting, but I think the predictions might be more realistic and less prone to overfitting.</p>\n<p>What are your thoughts on this approach? What about for shared BBs (interpolation within the chemical space)? Should I use the same models trained with BBs hidden in each fold, or would random splitting be sufficient in this case?</p>\n<p>Thanks in advance for your insights!</p>",
      "rawMarkdown": "Hi everyone, I hope it's ok to post this here since it's somewhat related and it might help others working in industry.\n\nI'm working on a material science project with a somewhat similar setup to this competition: we're using building blocks (BBs) to construct molecules, and our goal is property prediction (regression).\n\nI'm wondering about the best way to train models that generalize well to new, unseen BBs (similar to the non-shared BBs here). Would it be a good strategy to use cross-validation, where each fold intentionally hides a portion of the BBs?\n\nI know this would likely lower the cross-validation score compared to random splitting, but I think the predictions might be more realistic and less prone to overfitting.\n\nWhat are your thoughts on this approach? What about for shared BBs (interpolation within the chemical space)? Should I use the same models trained with BBs hidden in each fold, or would random splitting be sufficient in this case?\n\nThanks in advance for your insights!",
      "votes": null
    },
    {
      "id": "2826674",
      "postDate": "05/21/2024 04:34:14",
      "content": "<p>Judging from the leader board scores not many here have yet arrived at the \"good strategy\" for unseen BB's.  </p>\n<p>Been playing here on kaggle for a number of years - the only strategy that I have had success with for the \"unseen world\" is a ton of different models and ensemble a prediction</p>\n<p>I find the notion that models (at this level of data) can generalize to be a pretty large pile of BS.</p>",
      "rawMarkdown": "Judging from the leader board scores not many here have yet arrived at the \"good strategy\" for unseen BB's.  \n\nBeen playing here on kaggle for a number of years - the only strategy that I have had success with for the \"unseen world\" is a ton of different models and ensemble a prediction\n\nI find the notion that models (at this level of data) can generalize to be a pretty large pile of BS.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2826674,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "05/21/2024 04:34:14",
      "content": "<p>Judging from the leader board scores not many here have yet arrived at the \"good strategy\" for unseen BB's.  </p>\n<p>Been playing here on kaggle for a number of years - the only strategy that I have had success with for the \"unseen world\" is a ton of different models and ensemble a prediction</p>\n<p>I find the notion that models (at this level of data) can generalize to be a pretty large pile of BS.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2826484": "Hi everyone, I hope it's ok to post this here since it's somewhat related and it might help others working in industry.\n\nI'm working on a material science project with a somewhat similar setup to this competition: we're using building blocks (BBs) to construct molecules, and our goal is property prediction (regression).\n\nI'm wondering about the best way to train models that generalize well to new, unseen BBs (similar to the non-shared BBs here). Would it be a good strategy to use cross-validation, where each fold intentionally hides a portion of the BBs?\n\nI know this would likely lower the cross-validation score compared to random splitting, but I think the predictions might be more realistic and less prone to overfitting.\n\nWhat are your thoughts on this approach? What about for shared BBs (interpolation within the chemical space)? Should I use the same models trained with BBs hidden in each fold, or would random splitting be sufficient in this case?\n\nThanks in advance for your insights!",
    "2826674": "Judging from the leader board scores not many here have yet arrived at the \"good strategy\" for unseen BB's.  \n\nBeen playing here on kaggle for a number of years - the only strategy that I have had success with for the \"unseen world\" is a ton of different models and ensemble a prediction\n\nI find the notion that models (at this level of data) can generalize to be a pretty large pile of BS."
  },
  "source": "meta"
}