{
  "id": 497510,
  "title": "Sharing my CV split",
  "url": "/competitions/leash-BELKA/discussion/497510",
  "author_name": "",
  "post_date": "2024-04-24T22:12:55.635858Z",
  "votes": 31,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I created a CV split that intends to mirror the test set split created by the host.</p>\n<p>Specifically, I create CV split based on:</p>\n<ul>\n<li>Scaffold split</li>\n<li>Building block split</li>\n<li>Random split</li>\n</ul>\n<p>This is based on discussion shared by host <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103\" target=\"_blank\">here</a>.</p>\n<p>DATASET WITH CV SPLIT: <a href=\"https://www.kaggle.com/datasets/thedrcat/belka-cv-split\" target=\"_blank\">https://www.kaggle.com/datasets/thedrcat/belka-cv-split</a><br>\nCODE: <a href=\"https://www.kaggle.com/code/thedrcat/belka-split-cv-like-the-host/\" target=\"_blank\">https://www.kaggle.com/code/thedrcat/belka-split-cv-like-the-host/</a></p>\n<p>I'm learning chemistry on the competition, so if I made any mistakes - apologies and would really appreciate feedback!</p>\n<p>EDIT: V2 updated based on the discussion below - removed Murcko scaffold, corrected BB split. </p>",
  "messages": [
    {
      "id": "2773779",
      "postDate": "04/24/2024 22:12:55",
      "content": "<p>I created a CV split that intends to mirror the test set split created by the host.</p>\n<p>Specifically, I create CV split based on:</p>\n<ul>\n<li>Scaffold split</li>\n<li>Building block split</li>\n<li>Random split</li>\n</ul>\n<p>This is based on discussion shared by host <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103\" target=\"_blank\">here</a>.</p>\n<p>DATASET WITH CV SPLIT: <a href=\"https://www.kaggle.com/datasets/thedrcat/belka-cv-split\" target=\"_blank\">https://www.kaggle.com/datasets/thedrcat/belka-cv-split</a><br>\nCODE: <a href=\"https://www.kaggle.com/code/thedrcat/belka-split-cv-like-the-host/\" target=\"_blank\">https://www.kaggle.com/code/thedrcat/belka-split-cv-like-the-host/</a></p>\n<p>I'm learning chemistry on the competition, so if I made any mistakes - apologies and would really appreciate feedback!</p>\n<p>EDIT: V2 updated based on the discussion below - removed Murcko scaffold, corrected BB split. </p>",
      "rawMarkdown": "I created a CV split that intends to mirror the test set split created by the host.\n\nSpecifically, I create CV split based on:\n- Scaffold split\n- Building block split\n- Random split\n\nThis is based on discussion shared by host [here](https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103).\n\nDATASET WITH CV SPLIT: https://www.kaggle.com/datasets/thedrcat/belka-cv-split\nCODE: https://www.kaggle.com/code/thedrcat/belka-split-cv-like-the-host/\n\nI'm learning chemistry on the competition, so if I made any mistakes - apologies and would really appreciate feedback!\n\nEDIT: V2 updated based on the discussion below - removed Murcko scaffold, corrected BB split.",
      "votes": null
    },
    {
      "id": "2773798",
      "postDate": "04/24/2024 22:57:22",
      "content": "<p>A word of caution- remember to exclude all BBs in BB split from train for it to mirror the no share BB split in test i.e. the group df.buildingblock1_smiles.isin(bbs1) | df.buildingblock2_smiles.isin(bbs2) | df.buildingblock3_smiles.isin(bbs2)<br>\n(| Instead &amp;)</p>",
      "rawMarkdown": "A word of caution- remember to exclude all BBs in BB split from train for it to mirror the no share BB split in test i.e. the group df.buildingblock1_smiles.isin(bbs1) | df.buildingblock2_smiles.isin(bbs2) | df.buildingblock3_smiles.isin(bbs2)\n(| Instead &)",
      "votes": null
    },
    {
      "id": "2773870",
      "postDate": "04/25/2024 00:07:46",
      "content": "<p>To be more clear, what we need is like this:</p>\n<pre><code>df[] = \ndf[].iloc[df.buildingblock1_smiles.isin(bbs1) &amp; df.buildingblock2_smiles.isin(bbs2) &amp; df.buildingblock3_smiles.isin(bbs2)] = \ndf[].iloc[df.buildingblock1_smiles.isin(bbs1) | df.buildingblock2_smiles.isin(bbs2) | df.buildingblock3_smiles.isin(bbs2)] = \n</code></pre>",
      "rawMarkdown": "To be more clear, what we need is like this:\n```python\ndf['bb_group'] = 'train'\ndf['bb_group'].iloc[df.buildingblock1_smiles.isin(bbs1) & df.buildingblock2_smiles.isin(bbs2) & df.buildingblock3_smiles.isin(bbs2)] = 'test'\ndf['bb_group'].iloc[df.buildingblock1_smiles.isin(bbs1) | df.buildingblock2_smiles.isin(bbs2) | df.buildingblock3_smiles.isin(bbs2)] = 'exclude! This is neither train nor test!'\n```",
      "votes": null
    },
    {
      "id": "2773879",
      "postDate": "04/25/2024 00:17:26",
      "content": "<p>This is great. Couple comments:</p>\n<ul>\n<li>You might consider using exactly the number of BBs in the <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/496576\" target=\"_blank\">test distribution</a>. 17BB1, 36BB2. (You might want to take a couple from BB3 that aren't in BB2, there's 181 of those in train).</li>\n<li>In the end, I will want two different test CV distributions. One for comparing with public LB, one for comparing with private LB. Those are going to be 2 different distributions, and mostly I'll care more about the private LB distribution. BUT, the public LB one is for running sanity checks and making sure that CV and LB stay in sync during feature engineering etc.</li>\n</ul>",
      "rawMarkdown": "This is great. Couple comments:\n\n* You might consider using exactly the number of BBs in the [test distribution](https://www.kaggle.com/competitions/leash-BELKA/discussion/496576). 17BB1, 36BB2. (You might want to take a couple from BB3 that aren't in BB2, there's 181 of those in train).\n* In the end, I will want two different test CV distributions. One for comparing with public LB, one for comparing with private LB. Those are going to be 2 different distributions, and mostly I'll care more about the private LB distribution. BUT, the public LB one is for running sanity checks and making sure that CV and LB stay in sync during feature engineering etc.",
      "votes": null
    },
    {
      "id": "2773889",
      "postDate": "04/25/2024 00:23:38",
      "content": "<p>Also, I'm not sure exactly what you're doing in the MurckoScaffold section, but I'm highly doubtful that there's anything in train that's at all like the non-triazines in (private) test. </p>\n<p>Instead, I think holding out an <a href=\"https://www.kaggle.com/code/chemdatafarmer/additional-seh-data\" target=\"_blank\">entire external dataset</a> would be the best we can do to test on unseen data? Keep in mind, though, that only sEH is represented in that dataset, and is much easier to bind than the other two.</p>",
      "rawMarkdown": "Also, I'm not sure exactly what you're doing in the MurckoScaffold section, but I'm highly doubtful that there's anything in train that's at all like the non-triazines in (private) test. \n\nInstead, I think holding out an [entire external dataset](https://www.kaggle.com/code/chemdatafarmer/additional-seh-data) would be the best we can do to test on unseen data? Keep in mind, though, that only sEH is represented in that dataset, and is much easier to bind than the other two.",
      "votes": null
    },
    {
      "id": "2773895",
      "postDate": "04/25/2024 00:25:49",
      "content": "<p>Another reasonable option is to hold out a LOT more BBs. Triazine core or not, there may not be a ton of difference between unseen \"everything BUT triazine core\" and unseen \"everything\".</p>",
      "rawMarkdown": "Another reasonable option is to hold out a LOT more BBs. Triazine core or not, there may not be a ton of difference between unseen \"everything BUT triazine core\" and unseen \"everything\".",
      "votes": null
    },
    {
      "id": "2773914",
      "postDate": "04/25/2024 00:53:56",
      "content": "<p>i haven't tried this yet , but i think:<br>\nif we get better results for share blocks (local CV), the \"share blocks LB\" should also get better, but the \"nonshare blocks LB\" may get worst.<br>\nto verify, submit separately for test share and nonshare blocks.</p>\n<p>in fact maybe two set of models, one for share and another for nonshare blocks prediction.<br>\nThen there is actually two set of data split</p>",
      "rawMarkdown": "i haven't tried this yet , but i think:\nif we get better results for share blocks (local CV), the \"share blocks LB\" should also get better, but the \"nonshare blocks LB\" may get worst.\nto verify, submit separately for test share and nonshare blocks.\n\nin fact maybe two set of models, one for share and another for nonshare blocks prediction.\nThen there is actually two set of data split",
      "votes": null
    },
    {
      "id": "2774472",
      "postDate": "04/25/2024 07:52:56",
      "content": "<p>Hi, thanks for the notebook and datasets!</p>\n<p>Regarding the scaffold split, just to mention that as was previously shared <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/493294\" target=\"_blank\">here</a>, all molecules in the training set contain a triazine core. I haven't checked it myself, but I would expect that the same core was used to react with the different building blocks to synthesize the final molecules (one could check that by comparing the 3 building blocks and the final molecule, taking into account some differences due to the reactions used bind the BBs and core together). I.e. the MurckoScaffold is incorporating parts of the building blocks in the detected scaffolds. So I am not sure if this is better than splitting the train dataset based on building blocks similarity.</p>",
      "rawMarkdown": "Hi, thanks for the notebook and datasets!\n\nRegarding the scaffold split, just to mention that as was previously shared [here](https://www.kaggle.com/competitions/leash-BELKA/discussion/493294), all molecules in the training set contain a triazine core. I haven't checked it myself, but I would expect that the same core was used to react with the different building blocks to synthesize the final molecules (one could check that by comparing the 3 building blocks and the final molecule, taking into account some differences due to the reactions used bind the BBs and core together). I.e. the MurckoScaffold is incorporating parts of the building blocks in the detected scaffolds. So I am not sure if this is better than splitting the train dataset based on building blocks similarity.",
      "votes": null
    },
    {
      "id": "2775732",
      "postDate": "04/25/2024 19:02:18",
      "content": "<p>Do you mean all BBs in BB split and used in the train set should be excluded from the test set?</p>",
      "rawMarkdown": "Do you mean all BBs in BB split and used in the train set should be excluded from the test set?",
      "votes": null
    },
    {
      "id": "2775805",
      "postDate": "04/25/2024 20:08:36",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> &amp; <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a>, super helpful! Yep, I think the <code>&amp;&amp;</code> group should be in the <code>noshare</code> split, but the <code>II</code> group maybe we don't need to exclude, just add to the random split…. We have quite a lot, but still don't like wasting data 😅</p>",
      "rawMarkdown": "Thanks @shlomoron & @roberthatch, super helpful! Yep, I think the `&&` group should be in the `noshare` split, but the `II` group maybe we don't need to exclude, just add to the random split.... We have quite a lot, but still don't like wasting data 😅",
      "votes": null
    },
    {
      "id": "2775816",
      "postDate": "04/25/2024 20:20:29",
      "content": "<p>Yeah, that makes sense. One option is a bit of both. You set it all aside as test data, but when you predict you look at all predictions AND look at the subset predictions (excluding the partial overlap cases) to see your score against a distribution more closely matching LB. </p>",
      "rawMarkdown": "Yeah, that makes sense. One option is a bit of both. You set it all aside as test data, but when you predict you look at all predictions AND look at the subset predictions (excluding the partial overlap cases) to see your score against a distribution more closely matching LB.",
      "votes": null
    },
    {
      "id": "2775822",
      "postDate": "04/25/2024 20:26:56",
      "content": "<p>Also, I don't recall if you do this as CV? If so, at the least you have to modify so that some data is held out of train for more than one fold. For instance If you reserve 1-30 for fold 1, and 300-330 for fold 2, then at least you can't allow the intersection of those BBs to be used in train data except when training folds 3,4,5. </p>",
      "rawMarkdown": "Also, I don't recall if you do this as CV? If so, at the least you have to modify so that some data is held out of train for more than one fold. For instance If you reserve 1-30 for fold 1, and 300-330 for fold 2, then at least you can't allow the intersection of those BBs to be used in train data except when training folds 3,4,5.",
      "votes": null
    },
    {
      "id": "2775904",
      "postDate": "04/25/2024 21:45:53",
      "content": "<p>Yep, mainly for CV - for now I only want to train single fold and find correlation between CV/LB (and ideally in a way that will transfer to private). I'll figure out how to use held out data for training once I have some good results. </p>",
      "rawMarkdown": "Yep, mainly for CV - for now I only want to train single fold and find correlation between CV/LB (and ideally in a way that will transfer to private). I'll figure out how to use held out data for training once I have some good results.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2773798,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "04/24/2024 22:57:22",
      "content": "<p>A word of caution- remember to exclude all BBs in BB split from train for it to mirror the no share BB split in test i.e. the group df.buildingblock1_smiles.isin(bbs1) | df.buildingblock2_smiles.isin(bbs2) | df.buildingblock3_smiles.isin(bbs2)<br>\n(| Instead &amp;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2773870,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "04/25/2024 00:07:46",
          "content": "<p>To be more clear, what we need is like this:</p>\n<pre><code>df[] = \ndf[].iloc[df.buildingblock1_smiles.isin(bbs1) &amp; df.buildingblock2_smiles.isin(bbs2) &amp; df.buildingblock3_smiles.isin(bbs2)] = \ndf[].iloc[df.buildingblock1_smiles.isin(bbs1) | df.buildingblock2_smiles.isin(bbs2) | df.buildingblock3_smiles.isin(bbs2)] = \n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 2775732,
              "author_name": "elbasir",
              "author_url": "",
              "post_date": "04/25/2024 19:02:18",
              "content": "<p>Do you mean all BBs in BB split and used in the train set should be excluded from the test set?</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2775805,
              "author_name": "thedrcat",
              "author_url": "",
              "post_date": "04/25/2024 20:08:36",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> &amp; <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a>, super helpful! Yep, I think the <code>&amp;&amp;</code> group should be in the <code>noshare</code> split, but the <code>II</code> group maybe we don't need to exclude, just add to the random split…. We have quite a lot, but still don't like wasting data 😅</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2775816,
                  "author_name": "roberthatch",
                  "author_url": "",
                  "post_date": "04/25/2024 20:20:29",
                  "content": "<p>Yeah, that makes sense. One option is a bit of both. You set it all aside as test data, but when you predict you look at all predictions AND look at the subset predictions (excluding the partial overlap cases) to see your score against a distribution more closely matching LB. </p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2775822,
                  "author_name": "roberthatch",
                  "author_url": "",
                  "post_date": "04/25/2024 20:26:56",
                  "content": "<p>Also, I don't recall if you do this as CV? If so, at the least you have to modify so that some data is held out of train for more than one fold. For instance If you reserve 1-30 for fold 1, and 300-330 for fold 2, then at least you can't allow the intersection of those BBs to be used in train data except when training folds 3,4,5. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2775904,
                      "author_name": "thedrcat",
                      "author_url": "",
                      "post_date": "04/25/2024 21:45:53",
                      "content": "<p>Yep, mainly for CV - for now I only want to train single fold and find correlation between CV/LB (and ideally in a way that will transfer to private). I'll figure out how to use held out data for training once I have some good results. </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2773879,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "04/25/2024 00:17:26",
      "content": "<p>This is great. Couple comments:</p>\n<ul>\n<li>You might consider using exactly the number of BBs in the <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/496576\" target=\"_blank\">test distribution</a>. 17BB1, 36BB2. (You might want to take a couple from BB3 that aren't in BB2, there's 181 of those in train).</li>\n<li>In the end, I will want two different test CV distributions. One for comparing with public LB, one for comparing with private LB. Those are going to be 2 different distributions, and mostly I'll care more about the private LB distribution. BUT, the public LB one is for running sanity checks and making sure that CV and LB stay in sync during feature engineering etc.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2773889,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "04/25/2024 00:23:38",
          "content": "<p>Also, I'm not sure exactly what you're doing in the MurckoScaffold section, but I'm highly doubtful that there's anything in train that's at all like the non-triazines in (private) test. </p>\n<p>Instead, I think holding out an <a href=\"https://www.kaggle.com/code/chemdatafarmer/additional-seh-data\" target=\"_blank\">entire external dataset</a> would be the best we can do to test on unseen data? Keep in mind, though, that only sEH is represented in that dataset, and is much easier to bind than the other two.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2773895,
              "author_name": "roberthatch",
              "author_url": "",
              "post_date": "04/25/2024 00:25:49",
              "content": "<p>Another reasonable option is to hold out a LOT more BBs. Triazine core or not, there may not be a ton of difference between unseen \"everything BUT triazine core\" and unseen \"everything\".</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2773914,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/25/2024 00:53:56",
      "content": "<p>i haven't tried this yet , but i think:<br>\nif we get better results for share blocks (local CV), the \"share blocks LB\" should also get better, but the \"nonshare blocks LB\" may get worst.<br>\nto verify, submit separately for test share and nonshare blocks.</p>\n<p>in fact maybe two set of models, one for share and another for nonshare blocks prediction.<br>\nThen there is actually two set of data split</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2774472,
      "author_name": "vicentjborrs",
      "author_url": "",
      "post_date": "04/25/2024 07:52:56",
      "content": "<p>Hi, thanks for the notebook and datasets!</p>\n<p>Regarding the scaffold split, just to mention that as was previously shared <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/493294\" target=\"_blank\">here</a>, all molecules in the training set contain a triazine core. I haven't checked it myself, but I would expect that the same core was used to react with the different building blocks to synthesize the final molecules (one could check that by comparing the 3 building blocks and the final molecule, taking into account some differences due to the reactions used bind the BBs and core together). I.e. the MurckoScaffold is incorporating parts of the building blocks in the detected scaffolds. So I am not sure if this is better than splitting the train dataset based on building blocks similarity.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2773779": "I created a CV split that intends to mirror the test set split created by the host.\n\nSpecifically, I create CV split based on:\n- Scaffold split\n- Building block split\n- Random split\n\nThis is based on discussion shared by host [here](https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103).\n\nDATASET WITH CV SPLIT: https://www.kaggle.com/datasets/thedrcat/belka-cv-split\nCODE: https://www.kaggle.com/code/thedrcat/belka-split-cv-like-the-host/\n\nI'm learning chemistry on the competition, so if I made any mistakes - apologies and would really appreciate feedback!\n\nEDIT: V2 updated based on the discussion below - removed Murcko scaffold, corrected BB split.",
    "2773798": "A word of caution- remember to exclude all BBs in BB split from train for it to mirror the no share BB split in test i.e. the group df.buildingblock1_smiles.isin(bbs1) | df.buildingblock2_smiles.isin(bbs2) | df.buildingblock3_smiles.isin(bbs2)\n(| Instead &)",
    "2773870": "To be more clear, what we need is like this:\n```python\ndf['bb_group'] = 'train'\ndf['bb_group'].iloc[df.buildingblock1_smiles.isin(bbs1) & df.buildingblock2_smiles.isin(bbs2) & df.buildingblock3_smiles.isin(bbs2)] = 'test'\ndf['bb_group'].iloc[df.buildingblock1_smiles.isin(bbs1) | df.buildingblock2_smiles.isin(bbs2) | df.buildingblock3_smiles.isin(bbs2)] = 'exclude! This is neither train nor test!'\n```",
    "2773879": "This is great. Couple comments:\n\n* You might consider using exactly the number of BBs in the [test distribution](https://www.kaggle.com/competitions/leash-BELKA/discussion/496576). 17BB1, 36BB2. (You might want to take a couple from BB3 that aren't in BB2, there's 181 of those in train).\n* In the end, I will want two different test CV distributions. One for comparing with public LB, one for comparing with private LB. Those are going to be 2 different distributions, and mostly I'll care more about the private LB distribution. BUT, the public LB one is for running sanity checks and making sure that CV and LB stay in sync during feature engineering etc.",
    "2773889": "Also, I'm not sure exactly what you're doing in the MurckoScaffold section, but I'm highly doubtful that there's anything in train that's at all like the non-triazines in (private) test. \n\nInstead, I think holding out an [entire external dataset](https://www.kaggle.com/code/chemdatafarmer/additional-seh-data) would be the best we can do to test on unseen data? Keep in mind, though, that only sEH is represented in that dataset, and is much easier to bind than the other two.",
    "2773895": "Another reasonable option is to hold out a LOT more BBs. Triazine core or not, there may not be a ton of difference between unseen \"everything BUT triazine core\" and unseen \"everything\".",
    "2773914": "i haven't tried this yet , but i think:\nif we get better results for share blocks (local CV), the \"share blocks LB\" should also get better, but the \"nonshare blocks LB\" may get worst.\nto verify, submit separately for test share and nonshare blocks.\n\nin fact maybe two set of models, one for share and another for nonshare blocks prediction.\nThen there is actually two set of data split",
    "2774472": "Hi, thanks for the notebook and datasets!\n\nRegarding the scaffold split, just to mention that as was previously shared [here](https://www.kaggle.com/competitions/leash-BELKA/discussion/493294), all molecules in the training set contain a triazine core. I haven't checked it myself, but I would expect that the same core was used to react with the different building blocks to synthesize the final molecules (one could check that by comparing the 3 building blocks and the final molecule, taking into account some differences due to the reactions used bind the BBs and core together). I.e. the MurckoScaffold is incorporating parts of the building blocks in the detected scaffolds. So I am not sure if this is better than splitting the train dataset based on building blocks similarity.",
    "2775732": "Do you mean all BBs in BB split and used in the train set should be excluded from the test set?",
    "2775805": "Thanks @shlomoron & @roberthatch, super helpful! Yep, I think the `&&` group should be in the `noshare` split, but the `II` group maybe we don't need to exclude, just add to the random split.... We have quite a lot, but still don't like wasting data 😅",
    "2775816": "Yeah, that makes sense. One option is a bit of both. You set it all aside as test data, but when you predict you look at all predictions AND look at the subset predictions (excluding the partial overlap cases) to see your score against a distribution more closely matching LB.",
    "2775822": "Also, I don't recall if you do this as CV? If so, at the least you have to modify so that some data is held out of train for more than one fold. For instance If you reserve 1-30 for fold 1, and 300-330 for fold 2, then at least you can't allow the intersection of those BBs to be used in train data except when training folds 3,4,5.",
    "2775904": "Yep, mainly for CV - for now I only want to train single fold and find correlation between CV/LB (and ideally in a way that will transfer to private). I'll figure out how to use held out data for training once I have some good results."
  },
  "source": "meta"
}