{
  "id": 230745,
  "title": "What approch should take if I want to use small version of Dataset for Training?",
  "url": "/competitions/bms-molecular-translation/discussion/230745",
  "author_name": "",
  "post_date": "2021-04-05T13:32:01.819192100Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>As a noob, I'm confused about what approchs I should use for Compete with Small Dataset of Competition Dataset.<br>\nAs I assume , If I use some portion of the Dataset It may not include some of the atoms. Then how can I trust in the validation!<br>\nI need expert helps. Thanks in an advance. </p>",
  "messages": [
    {
      "id": "1263512",
      "postDate": "04/05/2021 13:32:01",
      "content": "<p>As a noob, I'm confused about what approchs I should use for Compete with Small Dataset of Competition Dataset.<br>\nAs I assume , If I use some portion of the Dataset It may not include some of the atoms. Then how can I trust in the validation!<br>\nI need expert helps. Thanks in an advance. </p>",
      "rawMarkdown": "As a noob, I'm confused about what approchs I should use for Compete with Small Dataset of Competition Dataset.\nAs I assume , If I use some portion of the Dataset It may not include some of the atoms. Then how can I trust in the validation!\nI need expert helps. Thanks in an advance.",
      "votes": null
    },
    {
      "id": "1263972",
      "postDate": "04/05/2021 19:31:32",
      "content": "<p>You might stratify your sample based on some rules like str length or something you create as feature…</p>",
      "rawMarkdown": "You might stratify your sample based on some rules like str length or something you create as feature...",
      "votes": null
    },
    {
      "id": "1266306",
      "postDate": "04/07/2021 16:43:58",
      "content": "<p>I am a real fan of 5 fold training with stratified samples.  But training time for 5 folds of the full data set is near the length my life expectancy.</p>\n<p>So I am stratifying for 50 folds right now, but only using a single fold.</p>\n<p>To understand how my model is performing I am checking to see if the predicted molecule is valid using RDKit.  Just eye balling the \"valid molecules\" it's pretty clear that short InChi's have pretty good chance of being valid molecules with my current efforts.</p>\n<p>So as suggested by <a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">Ertugrul</a> I am changing from stratification \"number of carbons\" to \"length of the InChi\" in my next revision.  At some point when I think I am getting better model I will increase the sample subset by lowing the number of folds for stratification or use a few more of the folds.</p>",
      "rawMarkdown": "I am a real fan of 5 fold training with stratified samples.  But training time for 5 folds of the full data set is near the length my life expectancy.\n\nSo I am stratifying for 50 folds right now, but only using a single fold.\n\nTo understand how my model is performing I am checking to see if the predicted molecule is valid using RDKit.  Just eye balling the \"valid molecules\" it's pretty clear that short InChi's have pretty good chance of being valid molecules with my current efforts.\n\nSo as suggested by [Ertugrul](https://www.kaggle.com/datafan07) I am changing from stratification \"number of carbons\" to \"length of the InChi\" in my next revision.  At some point when I think I am getting better model I will increase the sample subset by lowing the number of folds for stratification or use a few more of the folds.",
      "votes": null
    },
    {
      "id": "1266358",
      "postDate": "04/07/2021 17:32:00",
      "content": "<p>Thanks a lot for your detailed share. It really big help. Just one thing to know is it possible to train with TPU Pytorch? I actually not found the recources!</p>",
      "rawMarkdown": "Thanks a lot for your detailed share. It really big help. Just one thing to know is it possible to train with TPU Pytorch? I actually not found the recources!",
      "votes": null
    },
    {
      "id": "1266394",
      "postDate": "04/07/2021 17:59:04",
      "content": "<p>use this perhaps .. <a href=\"https://www.kaggle.com/tj0612/cv-strategy-by-murcko-scaffold\" target=\"_blank\">https://www.kaggle.com/tj0612/cv-strategy-by-murcko-scaffold</a></p>",
      "rawMarkdown": "use this perhaps .. https://www.kaggle.com/tj0612/cv-strategy-by-murcko-scaffold",
      "votes": null
    },
    {
      "id": "1266411",
      "postDate": "04/07/2021 18:11:34",
      "content": "<p>Thanks a lot, Sir! Really cool kernel. </p>",
      "rawMarkdown": "Thanks a lot, Sir! Really cool kernel.",
      "votes": null
    },
    {
      "id": "1266469",
      "postDate": "04/07/2021 19:24:33",
      "content": "<p>Not a Pytorch user but believe there is a shared kernel using TPU on Pytorch.  </p>\n<p><a href=\"https://www.kaggle.com/mohamed3abdelrazik/ranzcr-resnet200d-3-stage-train-step-1f4\" target=\"_blank\">https://www.kaggle.com/mohamed3abdelrazik/ranzcr-resnet200d-3-stage-train-step-1f4</a></p>",
      "rawMarkdown": "Not a Pytorch user but believe there is a shared kernel using TPU on Pytorch.  \n\nhttps://www.kaggle.com/mohamed3abdelrazik/ranzcr-resnet200d-3-stage-train-step-1f4",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1263972,
      "author_name": "datafan07",
      "author_url": "",
      "post_date": "04/05/2021 19:31:32",
      "content": "<p>You might stratify your sample based on some rules like str length or something you create as feature…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1266306,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "04/07/2021 16:43:58",
      "content": "<p>I am a real fan of 5 fold training with stratified samples.  But training time for 5 folds of the full data set is near the length my life expectancy.</p>\n<p>So I am stratifying for 50 folds right now, but only using a single fold.</p>\n<p>To understand how my model is performing I am checking to see if the predicted molecule is valid using RDKit.  Just eye balling the \"valid molecules\" it's pretty clear that short InChi's have pretty good chance of being valid molecules with my current efforts.</p>\n<p>So as suggested by <a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">Ertugrul</a> I am changing from stratification \"number of carbons\" to \"length of the InChi\" in my next revision.  At some point when I think I am getting better model I will increase the sample subset by lowing the number of folds for stratification or use a few more of the folds.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1266358,
          "author_name": "aifahim",
          "author_url": "",
          "post_date": "04/07/2021 17:32:00",
          "content": "<p>Thanks a lot for your detailed share. It really big help. Just one thing to know is it possible to train with TPU Pytorch? I actually not found the recources!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266469,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "04/07/2021 19:24:33",
          "content": "<p>Not a Pytorch user but believe there is a shared kernel using TPU on Pytorch.  </p>\n<p><a href=\"https://www.kaggle.com/mohamed3abdelrazik/ranzcr-resnet200d-3-stage-train-step-1f4\" target=\"_blank\">https://www.kaggle.com/mohamed3abdelrazik/ranzcr-resnet200d-3-stage-train-step-1f4</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1266394,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "04/07/2021 17:59:04",
      "content": "<p>use this perhaps .. <a href=\"https://www.kaggle.com/tj0612/cv-strategy-by-murcko-scaffold\" target=\"_blank\">https://www.kaggle.com/tj0612/cv-strategy-by-murcko-scaffold</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1266411,
          "author_name": "aifahim",
          "author_url": "",
          "post_date": "04/07/2021 18:11:34",
          "content": "<p>Thanks a lot, Sir! Really cool kernel. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1263512": "As a noob, I'm confused about what approchs I should use for Compete with Small Dataset of Competition Dataset.\nAs I assume , If I use some portion of the Dataset It may not include some of the atoms. Then how can I trust in the validation!\nI need expert helps. Thanks in an advance.",
    "1263972": "You might stratify your sample based on some rules like str length or something you create as feature...",
    "1266306": "I am a real fan of 5 fold training with stratified samples.  But training time for 5 folds of the full data set is near the length my life expectancy.\n\nSo I am stratifying for 50 folds right now, but only using a single fold.\n\nTo understand how my model is performing I am checking to see if the predicted molecule is valid using RDKit.  Just eye balling the \"valid molecules\" it's pretty clear that short InChi's have pretty good chance of being valid molecules with my current efforts.\n\nSo as suggested by [Ertugrul](https://www.kaggle.com/datafan07) I am changing from stratification \"number of carbons\" to \"length of the InChi\" in my next revision.  At some point when I think I am getting better model I will increase the sample subset by lowing the number of folds for stratification or use a few more of the folds.",
    "1266358": "Thanks a lot for your detailed share. It really big help. Just one thing to know is it possible to train with TPU Pytorch? I actually not found the recources!",
    "1266394": "use this perhaps .. https://www.kaggle.com/tj0612/cv-strategy-by-murcko-scaffold",
    "1266411": "Thanks a lot, Sir! Really cool kernel.",
    "1266469": "Not a Pytorch user but believe there is a shared kernel using TPU on Pytorch.  \n\nhttps://www.kaggle.com/mohamed3abdelrazik/ranzcr-resnet200d-3-stage-train-step-1f4"
  },
  "source": "meta"
}