{
  "id": 358860,
  "title": "6-fold cross-validation scheme with TWO \"tests\" one is like public LB, another is like private LB",
  "url": "/competitions/open-problems-multimodal/discussion/358860",
  "author_name": "Alexander Chervov",
  "post_date": "2022-10-09T20:17:37.834000",
  "votes": 23,
  "comment_count": 4,
  "views": 0,
  "content": "<h3>Context</h3>\n<p>The private and public LB are essentially different at that competition. <br>\nThus it might lead to a shake-up - see discussion:  <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202</a></p>\n<p>So we propose some CV-scheme which might be helpful for that particular problem, see notebook: <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes</a></p>\n<p>The description is the following: </p>\n<h4>CV with TWO tests.</h4>\n<p>The main difficult is that private LB - made on additional DAY(!) (and donor), while public LB ONLY on additional donor. </p>\n<p>So we propose to make a cross-validiation which will take that into account.</p>\n<p>Thus we propose to have TWO test - 1) similar to public LB 2) similar to private LB</p>\n<p>We are lucky that there seems to be natural way to organize it.<br>\nFor CITE-seq it goes as follows:</p>\n<h4>The scheme contains 6-\"folds\" (but warning - it is not completely usual CV scheme ).</h4>\n<pre><code>Fold 0: Train: excludes Day 2 and Donor 32606 Sizes: train: 32536 Test Like Priv 21942 Test Like Publ 16510\nFold 1: Train: excludes Day 2 and Donor 31800 Sizes: train: 32638 Test Like Priv 21942 Test Like Publ 16408\nFold 2: Train: excludes Day 3 and Donor 32606 Sizes: train: 33100 Test Like Priv 20901 Test Like Publ 16987\nFold 3: Train: excludes Day 3 and Donor 31800 Sizes: train: 31543 Test Like Priv 20901 Test Like Publ 18544\nFold 4: Train: excludes Day 4 and Donor 32606 Sizes: train: 28368 Test Like Priv 28145 Test Like Publ 14475\nFold 5: Train: excludes Day 4 and Donor 31800 Sizes: train: 28189 Test Like Priv 28145 Test Like Publ 14654\n</code></pre>\n<p>Each \"Fold\" is parametrized by excluded from train day and donor. And folds are like that: </p>\n<pre><code>train_index = np.where( (df_meta['day']  != day2exclude) &amp; ( df_meta['donor']  != donor2exclude  ) )  [0]\ntest_index1_like_private_lb = np.where( df_meta['day']  == day2exclude)[0]\ntest_index2_like_public_lb = np.where( (df_meta['day']  != day2exclude) &amp;  (df_meta['donor']  == donor2exclude ) ) [0]\n</code></pre>\n<h4>Example:  assume day2exclude = 2, donor2exlude = 32606,</h4>\n<pre><code> We train on days 3,4 and without donor 32606,\n as a private test we consider day 2 , INCLUDING donor 32606 - thus it is quite similar to real private LB - where we have unseen day and unseen donor\n as a public test we consider days 3,4 donor is ONLY 32606 - thus it is similar to real public LB - the days are the same, but donor is unseen\n</code></pre>\n<h4>Male/Female are taken into account:</h4>\n<pre><code>We train always on donors where we have male and female. \nAnd predict only for MALE.\nThat is exactly as on real LB. \nThus we have 6 folds - we only exclude from train male donors 32606,  31800 - thus MALE is always in test - like on real LB.\n</code></pre>\n<p>PS</p>\n<p>The proposal as described is for CITE-seq part. Similar can be done for Multiome</p>",
  "messages": [
    {
      "id": 1979927,
      "postDate": "2022-10-09T20:17:37.833Z",
      "content": "<h3>Context</h3>\n<p>The private and public LB are essentially different at that competition. <br>\nThus it might lead to a shake-up - see discussion:  <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202</a></p>\n<p>So we propose some CV-scheme which might be helpful for that particular problem, see notebook: <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes</a></p>\n<p>The description is the following: </p>\n<h4>CV with TWO tests.</h4>\n<p>The main difficult is that private LB - made on additional DAY(!) (and donor), while public LB ONLY on additional donor. </p>\n<p>So we propose to make a cross-validiation which will take that into account.</p>\n<p>Thus we propose to have TWO test - 1) similar to public LB 2) similar to private LB</p>\n<p>We are lucky that there seems to be natural way to organize it.<br>\nFor CITE-seq it goes as follows:</p>\n<h4>The scheme contains 6-\"folds\" (but warning - it is not completely usual CV scheme ).</h4>\n<pre><code>Fold 0: Train: excludes Day 2 and Donor 32606 Sizes: train: 32536 Test Like Priv 21942 Test Like Publ 16510\nFold 1: Train: excludes Day 2 and Donor 31800 Sizes: train: 32638 Test Like Priv 21942 Test Like Publ 16408\nFold 2: Train: excludes Day 3 and Donor 32606 Sizes: train: 33100 Test Like Priv 20901 Test Like Publ 16987\nFold 3: Train: excludes Day 3 and Donor 31800 Sizes: train: 31543 Test Like Priv 20901 Test Like Publ 18544\nFold 4: Train: excludes Day 4 and Donor 32606 Sizes: train: 28368 Test Like Priv 28145 Test Like Publ 14475\nFold 5: Train: excludes Day 4 and Donor 31800 Sizes: train: 28189 Test Like Priv 28145 Test Like Publ 14654\n</code></pre>\n<p>Each \"Fold\" is parametrized by excluded from train day and donor. And folds are like that: </p>\n<pre><code>train_index = np.where( (df_meta['day']  != day2exclude) &amp; ( df_meta['donor']  != donor2exclude  ) )  [0]\ntest_index1_like_private_lb = np.where( df_meta['day']  == day2exclude)[0]\ntest_index2_like_public_lb = np.where( (df_meta['day']  != day2exclude) &amp;  (df_meta['donor']  == donor2exclude ) ) [0]\n</code></pre>\n<h4>Example:  assume day2exclude = 2, donor2exlude = 32606,</h4>\n<pre><code> We train on days 3,4 and without donor 32606,\n as a private test we consider day 2 , INCLUDING donor 32606 - thus it is quite similar to real private LB - where we have unseen day and unseen donor\n as a public test we consider days 3,4 donor is ONLY 32606 - thus it is similar to real public LB - the days are the same, but donor is unseen\n</code></pre>\n<h4>Male/Female are taken into account:</h4>\n<pre><code>We train always on donors where we have male and female. \nAnd predict only for MALE.\nThat is exactly as on real LB. \nThus we have 6 folds - we only exclude from train male donors 32606,  31800 - thus MALE is always in test - like on real LB.\n</code></pre>\n<p>PS</p>\n<p>The proposal as described is for CITE-seq part. Similar can be done for Multiome</p>",
      "rawMarkdown": "### Context\n\n\nThe private and public LB are essentially different at that competition. \nThus it might lead to a shake-up - see discussion:  https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202\n\nSo we propose some CV-scheme which might be helpful for that particular problem, see notebook: \nhttps://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\n\nThe description is the following: \n\n\n#### CV with TWO tests. \n\nThe main difficult is that private LB - made on additional DAY(!) (and donor), while public LB ONLY on additional donor. \n\nSo we propose to make a cross-validiation which will take that into account.\n\nThus we propose to have TWO test - 1) similar to public LB 2) similar to private LB\n\nWe are lucky that there seems to be natural way to organize it.\nFor CITE-seq it goes as follows:\n\n#### The scheme contains 6-\"folds\" (but warning - it is not completely usual CV scheme ).  \n\n    Fold 0: Train: excludes Day 2 and Donor 32606 Sizes: train: 32536 Test Like Priv 21942 Test Like Publ 16510\n    Fold 1: Train: excludes Day 2 and Donor 31800 Sizes: train: 32638 Test Like Priv 21942 Test Like Publ 16408\n    Fold 2: Train: excludes Day 3 and Donor 32606 Sizes: train: 33100 Test Like Priv 20901 Test Like Publ 16987\n    Fold 3: Train: excludes Day 3 and Donor 31800 Sizes: train: 31543 Test Like Priv 20901 Test Like Publ 18544\n    Fold 4: Train: excludes Day 4 and Donor 32606 Sizes: train: 28368 Test Like Priv 28145 Test Like Publ 14475\n    Fold 5: Train: excludes Day 4 and Donor 31800 Sizes: train: 28189 Test Like Priv 28145 Test Like Publ 14654\n\nEach \"Fold\" is parametrized by excluded from train day and donor. And folds are like that: \n\n    train_index = np.where( (df_meta['day']  != day2exclude) & ( df_meta['donor']  != donor2exclude  ) )  [0]\n    test_index1_like_private_lb = np.where( df_meta['day']  == day2exclude)[0]\n    test_index2_like_public_lb = np.where( (df_meta['day']  != day2exclude) &  (df_meta['donor']  == donor2exclude ) ) [0]\n\n#### Example:  assume day2exclude = 2, donor2exlude = 32606,\n     We train on days 3,4 and without donor 32606,\n     as a private test we consider day 2 , INCLUDING donor 32606 - thus it is quite similar to real private LB - where we have unseen day and unseen donor\n     as a public test we consider days 3,4 donor is ONLY 32606 - thus it is similar to real public LB - the days are the same, but donor is unseen\n\n#### Male/Female are taken into account: \n    We train always on donors where we have male and female. \n    And predict only for MALE.\n    That is exactly as on real LB. \n    Thus we have 6 folds - we only exclude from train male donors 32606,  31800 - thus MALE is always in test - like on real LB.\n    \nPS\n\nThe proposal as described is for CITE-seq part. Similar can be done for Multiome",
      "votes": 23
    },
    {
      "id": 1983180,
      "postDate": "2022-10-11T20:34:16.530Z",
      "content": "<p>The only adjustment that I might make is that test data should be in the future. So for cite: train[days 2 and 3] test[day4]. But maybe your solution will work since you are using data from different days and the sequential time difference is not as important. </p>\n<p>I think the reality is that we cannot have a good cv, because we don't have any data from day7 for cite and day10 for multiome. Therefore we can't figure out how the feature-target relationship changes by these days. For cite, we could use day 2 to day 4 change as an indicator, but we don't know what happens from day 4 to day 7. </p>\n<p>Although maybe using previous competition data and other outside day can give us a clearer picture. </p>",
      "rawMarkdown": "The only adjustment that I might make is that test data should be in the future. So for cite: train[days 2 and 3] test[day4]. But maybe your solution will work since you are using data from different days and the sequential time difference is not as important. \n\nI think the reality is that we cannot have a good cv, because we don't have any data from day7 for cite and day10 for multiome. Therefore we can't figure out how the feature-target relationship changes by these days. For cite, we could use day 2 to day 4 change as an indicator, but we don't know what happens from day 4 to day 7. \n\nAlthough maybe using previous competition data and other outside day can give us a clearer picture. ",
      "votes": 2
    },
    {
      "id": 1981449,
      "postDate": "2022-10-10T19:38:28.470Z",
      "content": "<h3>Strange conclusion - the model score on private-like-test is better, than on  public-like-test</h3>\n<p>Versions:</p>\n<p>Version 3:  simple modeling (Ridge) example with TruncatedSVD is added</p>\n<pre><code>Strange conclusion - the model on private-like-test is better, than on  public-like-test\n</code></pre>",
      "rawMarkdown": "###   Strange conclusion - the model score on private-like-test is better, than on  public-like-test\n\n Versions:\n\n Version 3:  simple modeling (Ridge) example with TruncatedSVD is added\n\n    Strange conclusion - the model on private-like-test is better, than on  public-like-test"
    },
    {
      "id": 2011494,
      "postDate": "2022-10-31T15:53:20.910Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true,
      "replies": [
        {
          "id": 2011506,
          "postDate": "2022-10-31T16:00:45.593Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1983180,
      "author_name": "Chris Miles",
      "author_url": "",
      "post_date": "2022-10-11T20:34:16.530000",
      "content": "<p>The only adjustment that I might make is that test data should be in the future. So for cite: train[days 2 and 3] test[day4]. But maybe your solution will work since you are using data from different days and the sequential time difference is not as important. </p>\n<p>I think the reality is that we cannot have a good cv, because we don't have any data from day7 for cite and day10 for multiome. Therefore we can't figure out how the feature-target relationship changes by these days. For cite, we could use day 2 to day 4 change as an indicator, but we don't know what happens from day 4 to day 7. </p>\n<p>Although maybe using previous competition data and other outside day can give us a clearer picture. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1981449,
      "author_name": "Alexander Chervov",
      "author_url": "",
      "post_date": "2022-10-10T19:38:28.470000",
      "content": "<h3>Strange conclusion - the model score on private-like-test is better, than on  public-like-test</h3>\n<p>Versions:</p>\n<p>Version 3:  simple modeling (Ridge) example with TruncatedSVD is added</p>\n<pre><code>Strange conclusion - the model on private-like-test is better, than on  public-like-test\n</code></pre>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2011494,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-10-31T15:53:20.910000",
      "content": "",
      "votes": 3,
      "replies": [
        {
          "id": 2011506,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-10-31T16:00:45.593000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1979927": "### Context\n\n\nThe private and public LB are essentially different at that competition. \nThus it might lead to a shake-up - see discussion:  https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202\n\nSo we propose some CV-scheme which might be helpful for that particular problem, see notebook: \nhttps://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\n\nThe description is the following: \n\n\n#### CV with TWO tests. \n\nThe main difficult is that private LB - made on additional DAY(!) (and donor), while public LB ONLY on additional donor. \n\nSo we propose to make a cross-validiation which will take that into account.\n\nThus we propose to have TWO test - 1) similar to public LB 2) similar to private LB\n\nWe are lucky that there seems to be natural way to organize it.\nFor CITE-seq it goes as follows:\n\n#### The scheme contains 6-\"folds\" (but warning - it is not completely usual CV scheme ).  \n\n    Fold 0: Train: excludes Day 2 and Donor 32606 Sizes: train: 32536 Test Like Priv 21942 Test Like Publ 16510\n    Fold 1: Train: excludes Day 2 and Donor 31800 Sizes: train: 32638 Test Like Priv 21942 Test Like Publ 16408\n    Fold 2: Train: excludes Day 3 and Donor 32606 Sizes: train: 33100 Test Like Priv 20901 Test Like Publ 16987\n    Fold 3: Train: excludes Day 3 and Donor 31800 Sizes: train: 31543 Test Like Priv 20901 Test Like Publ 18544\n    Fold 4: Train: excludes Day 4 and Donor 32606 Sizes: train: 28368 Test Like Priv 28145 Test Like Publ 14475\n    Fold 5: Train: excludes Day 4 and Donor 31800 Sizes: train: 28189 Test Like Priv 28145 Test Like Publ 14654\n\nEach \"Fold\" is parametrized by excluded from train day and donor. And folds are like that: \n\n    train_index = np.where( (df_meta['day']  != day2exclude) & ( df_meta['donor']  != donor2exclude  ) )  [0]\n    test_index1_like_private_lb = np.where( df_meta['day']  == day2exclude)[0]\n    test_index2_like_public_lb = np.where( (df_meta['day']  != day2exclude) &  (df_meta['donor']  == donor2exclude ) ) [0]\n\n#### Example:  assume day2exclude = 2, donor2exlude = 32606,\n     We train on days 3,4 and without donor 32606,\n     as a private test we consider day 2 , INCLUDING donor 32606 - thus it is quite similar to real private LB - where we have unseen day and unseen donor\n     as a public test we consider days 3,4 donor is ONLY 32606 - thus it is similar to real public LB - the days are the same, but donor is unseen\n\n#### Male/Female are taken into account: \n    We train always on donors where we have male and female. \n    And predict only for MALE.\n    That is exactly as on real LB. \n    Thus we have 6 folds - we only exclude from train male donors 32606,  31800 - thus MALE is always in test - like on real LB.\n    \nPS\n\nThe proposal as described is for CITE-seq part. Similar can be done for Multiome",
    "1983180": "The only adjustment that I might make is that test data should be in the future. So for cite: train[days 2 and 3] test[day4]. But maybe your solution will work since you are using data from different days and the sequential time difference is not as important. \n\nI think the reality is that we cannot have a good cv, because we don't have any data from day7 for cite and day10 for multiome. Therefore we can't figure out how the feature-target relationship changes by these days. For cite, we could use day 2 to day 4 change as an indicator, but we don't know what happens from day 4 to day 7. \n\nAlthough maybe using previous competition data and other outside day can give us a clearer picture. ",
    "1981449": "###   Strange conclusion - the model score on private-like-test is better, than on  public-like-test\n\n Versions:\n\n Version 3:  simple modeling (Ridge) example with TruncatedSVD is added\n\n    Strange conclusion - the model on private-like-test is better, than on  public-like-test",
    "2011494": ""
  }
}