{
  "id": 349591,
  "title": "CV vs LB Classical Discussion",
  "url": "/competitions/open-problems-multimodal/discussion/349591",
  "author_name": "Martin Kovacevic Buvinic",
  "post_date": "2022-09-01T22:45:24.715000",
  "votes": 32,
  "comment_count": 22,
  "views": 0,
  "content": "<p>Cite has 48663 test rows, and multi has 55935 rows (rows = amount of cells). Multi just use 30% of the rows, meaning it actually has 16780 cells.</p>\n<p>Therefore the weight for cite will be 0.743 and multi 0.257</p>\n<p>In my case my out of folds CV for cite is 0.8882 and for multi is 0.6601. If I use the weights to sum my out of folds score I actually score 0.829 which is actually the same score of my public lb</p>\n<p>This is using 5KFold for cite and multi:<br>\nCV: 0.829<br>\nLB: 0.829</p>\n<p>This is using 3 GroupKFold for cite, 4 GroupKFold for multi using day column.<br>\nCV: 0.8273<br>\nLB: 0.829</p>\n<p>Please share your results here (If you want 😄)</p>",
  "messages": [
    {
      "id": 1922997,
      "postDate": "2022-09-01T22:45:24.717Z",
      "content": "<p>Cite has 48663 test rows, and multi has 55935 rows (rows = amount of cells). Multi just use 30% of the rows, meaning it actually has 16780 cells.</p>\n<p>Therefore the weight for cite will be 0.743 and multi 0.257</p>\n<p>In my case my out of folds CV for cite is 0.8882 and for multi is 0.6601. If I use the weights to sum my out of folds score I actually score 0.829 which is actually the same score of my public lb</p>\n<p>This is using 5KFold for cite and multi:<br>\nCV: 0.829<br>\nLB: 0.829</p>\n<p>This is using 3 GroupKFold for cite, 4 GroupKFold for multi using day column.<br>\nCV: 0.8273<br>\nLB: 0.829</p>\n<p>Please share your results here (If you want 😄)</p>",
      "rawMarkdown": "Cite has 48663 test rows, and multi has 55935 rows (rows = amount of cells). Multi just use 30% of the rows, meaning it actually has 16780 cells.\n\nTherefore the weight for cite will be 0.743 and multi 0.257\n\nIn my case my out of folds CV for cite is 0.8882 and for multi is 0.6601. If I use the weights to sum my out of folds score I actually score 0.829 which is actually the same score of my public lb\n\nThis is using 5KFold for cite and multi:\nCV: 0.829\nLB: 0.829\n\nThis is using 3 GroupKFold for cite, 4 GroupKFold for multi using day column.\nCV: 0.8273\nLB: 0.829\n\nPlease share your results here (If you want 😄)",
      "votes": 31
    },
    {
      "id": 1932748,
      "postDate": "2022-09-09T21:59:55.870Z",
      "content": "<p>Just worked on cite for now .. <em>will keep updating here</em></p>\n<p><strong>Multinome</strong> 0.67~ Ensemble  </p>\n<ul>\n<li>Approach 1 NN: 0.6682</li>\n</ul>\n<p><strong>Cite</strong> : ensemble 0.905</p>\n<ul>\n<li>Approach 1 NN : 0.88992</li>\n<li>Approach 2 NN: 0.8926066667</li>\n<li>Approach 3 LGBM: 0.8844 </li>\n</ul>",
      "rawMarkdown": "Just worked on cite for now .. *will keep updating here*\n\n**Multinome** 0.67~ Ensemble  \n- Approach 1 NN: 0.6682\n\n**Cite** : ensemble 0.905\n- Approach 1 NN : 0.88992\n- Approach 2 NN: 0.8926066667\n- Approach 3 LGBM: 0.8844 \n",
      "votes": 3,
      "replies": [
        {
          "id": 1988863,
          "postDate": "2022-10-15T15:33:25.273Z",
          "content": "<p>may I ask what validation strategy(fold split) are you using?</p>",
          "rawMarkdown": "may I ask what validation strategy(fold split) are you using?",
          "votes": 1
        },
        {
          "id": 1988994,
          "postDate": "2022-10-15T17:00:57.270Z",
          "content": "<p>Hi.. By donor for now but thinking of trying few more ideas so the networks can learn but more as scores bit stuck at one point 😄</p>",
          "rawMarkdown": "Hi.. By donor for now but thinking of trying few more ideas so the networks can learn but more as scores bit stuck at one point 😄",
          "votes": 2
        }
      ]
    },
    {
      "id": 1924699,
      "postDate": "2022-09-03T10:01:42.127Z",
      "content": "<p>CV: 0.905<br>\nLB: 0.848</p>\n<p>I think I am overfitting :(</p>",
      "rawMarkdown": "CV: 0.905\nLB: 0.848\n\nI think I am overfitting :(",
      "votes": 3,
      "replies": [
        {
          "id": 1924864,
          "postDate": "2022-09-03T13:23:35.560Z",
          "content": "<p>Are you using day / patient as a feature ?</p>",
          "rawMarkdown": "Are you using day / patient as a feature ?"
        },
        {
          "id": 1925219,
          "postDate": "2022-09-03T18:18:02.880Z",
          "content": "<p>No, I don't use the metadata for anything.</p>",
          "rawMarkdown": "No, I don't use the metadata for anything.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1971305,
      "postDate": "2022-10-04T14:17:14.347Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chouchouchen\" target=\"_blank\">@chouchouchen</a> , As AmbrosM explained me, after <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933\" target=\"_blank\">The data update of 2022-09-10</a> the weights changed a bit.<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/353523\" target=\"_blank\"> They are now 0.7105 and 0.2895 </a>. There is still a difference between the CV and LB though.</p>",
      "rawMarkdown": "Hi @chouchouchen , As AmbrosM explained me, after [The data update of 2022-09-10](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933) the weights changed a bit.[ They are now 0.7105 and 0.2895 ](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/353523). There is still a difference between the CV and LB though.",
      "votes": 1
    },
    {
      "id": 1926845,
      "postDate": "2022-09-05T07:18:58.880Z",
      "content": "<p>Greetings, Martin!</p>\n<p>If we look <a href=\"http://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart#Submission\" target=\"_blank\">at this notebook (Submission part)</a> we will see the text:<br>\nThe CITEseq test predictions have 48663 rows (i.e., cells) and 140 columns (i.e. proteins). 48663 * 140 = 6812820. <strong>The final submission will have 65,744,180 rows, of which the first 6,812,820 are for the CITEseq predictions and the remaining 58,931,360 for the Multiome predictions.</strong></p>\n<p>And if we divide number of CITESEQ and MULTIOME rows in submission file by total submission length we will get 0.104 for CITESEQ and 0.896 for MULTIOME.</p>\n<p>Are You sure about 0.743 for SITE and 0.257 for MULTI?</p>",
      "rawMarkdown": "Greetings, Martin!\n\nIf we look [at this notebook (Submission part)](http://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart#Submission) we will see the text:\nThe CITEseq test predictions have 48663 rows (i.e., cells) and 140 columns (i.e. proteins). 48663 * 140 = 6812820. **The final submission will have 65,744,180 rows, of which the first 6,812,820 are for the CITEseq predictions and the remaining 58,931,360 for the Multiome predictions.**\n\nAnd if we divide number of CITESEQ and MULTIOME rows in submission file by total submission length we will get 0.104 for CITESEQ and 0.896 for MULTIOME.\n\nAre You sure about 0.743 for SITE and 0.257 for MULTI?",
      "votes": 2,
      "replies": [
        {
          "id": 1927280,
          "postDate": "2022-09-05T14:01:48.923Z",
          "content": "<p>The 6,812,820 rows of cite has more cell ids compared to the 58,931,360 rows for multi. Correlation score is computed per cell_id, the correct way to compute is to compare the amount of cell ids not rows.</p>",
          "rawMarkdown": "The 6,812,820 rows of cite has more cell ids compared to the 58,931,360 rows for multi. Correlation score is computed per cell_id, the correct way to compute is to compare the amount of cell ids not rows.",
          "votes": 5
        },
        {
          "id": 1927318,
          "postDate": "2022-09-05T14:46:57.067Z",
          "content": "<p>thank you!</p>",
          "rawMarkdown": "thank you!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1989173,
      "postDate": "2022-10-15T18:48:36.357Z",
      "content": "<p>i am not sure Kfold is doing data leakage for me how exactly are you applying kfold?</p>",
      "rawMarkdown": "i am not sure Kfold is doing data leakage for me how exactly are you applying kfold?\n"
    },
    {
      "id": 1972361,
      "postDate": "2022-10-05T05:48:53.670Z",
      "content": "<p>CITEseq:0.899(group by donor)<br>\nMultiome:0.672(Kfold) 0.670(group by donor)</p>\n<p>And I'm submitting one technique as a sample submission to check the score.<br>\nCITEseq:0.250<br>\nMultiome:-0.437<br>\nPublic:0.812</p>\n<p>if it's not my calculation error, I think the real weight of the techniques in the public is slightly off.</p>",
      "rawMarkdown": "CITEseq:0.899(group by donor)\nMultiome:0.672(Kfold) 0.670(group by donor)\n\nAnd I'm submitting one technique as a sample submission to check the score.\nCITEseq:0.250\nMultiome:-0.437\nPublic:0.812\n\nif it's not my calculation error, I think the real weight of the techniques in the public is slightly off.",
      "replies": [
        {
          "id": 1973234,
          "postDate": "2022-10-05T14:36:54.910Z",
          "content": "<p>May I ask is ur cite an ensemble or SIngle Model MLP</p>",
          "rawMarkdown": "May I ask is ur cite an ensemble or SIngle Model MLP"
        },
        {
          "id": 1974074,
          "postDate": "2022-10-06T04:13:25.583Z",
          "content": "<p>All these are for single models, but publicLB0.250/-0.437/0.812 are best even with ensembles.</p>",
          "rawMarkdown": "All these are for single models, but publicLB0.250/-0.437/0.812 are best even with ensembles."
        }
      ]
    },
    {
      "id": 1930943,
      "postDate": "2022-09-08T11:29:49.487Z",
      "content": "<p>What's the updated CV/LB scores after the re-scoring?</p>\n<p>And wouldn't donorwise split make more sense to guage the correlation between Public Test Score and CV. Since the daywise split makes sense for the Private test, which we do not get. </p>\n<p>I'm planning to make a thorough investigation of different splits and CV correlations once I finalized my training and submission pipeline. So far not getting good enough validation scores to even consider submitting 😔</p>",
      "rawMarkdown": "What's the updated CV/LB scores after the re-scoring?\n\nAnd wouldn't donorwise split make more sense to guage the correlation between Public Test Score and CV. Since the daywise split makes sense for the Private test, which we do not get. \n\nI'm planning to make a thorough investigation of different splits and CV correlations once I finalized my training and submission pipeline. So far not getting good enough validation scores to even consider submitting 😔"
    },
    {
      "id": 1949929,
      "postDate": "2022-09-22T02:06:48.813Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1930866,
      "postDate": "2022-09-08T09:24:04.307Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1930904,
          "postDate": "2022-09-08T10:04:00.420Z",
          "content": "<p>Previously, there were 65443 test rows (48663 for cite, 16780 for multi). Now, 7016 test rows from cite were removed from scoring (they are still present in the submission file but their score doesn't count). So now we have 41647 rows for cite and still 16780 rows for multi for a total of 58427 rows.<br>\nSo the new weights are 0.712 for cite and 0.288 for multi.<br>\nHope this helps :)</p>",
          "rawMarkdown": "Previously, there were 65443 test rows (48663 for cite, 16780 for multi). Now, 7016 test rows from cite were removed from scoring (they are still present in the submission file but their score doesn't count). So now we have 41647 rows for cite and still 16780 rows for multi for a total of 58427 rows.\nSo the new weights are 0.712 for cite and 0.288 for multi.\nHope this helps :)",
          "votes": 5
        },
        {
          "id": 1930939,
          "postDate": "2022-09-08T11:20:52.550Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 1931572,
          "postDate": "2022-09-08T19:36:45.450Z",
          "content": "<p>Correct, but those new weights are just for public leaderboard, for private they are the same as they where before</p>",
          "rawMarkdown": "Correct, but those new weights are just for public leaderboard, for private they are the same as they where before",
          "votes": 1
        },
        {
          "id": 1931599,
          "postDate": "2022-09-08T20:47:19.910Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1931646,
          "postDate": "2022-09-08T22:19:05.930Z",
          "content": "<p>Well, the private test set is hidden so we don't know the weights. But, in general, it is almost the same distribution as the public test set. But that's not by any mean a rule.</p>",
          "rawMarkdown": "Well, the private test set is hidden so we don't know the weights. But, in general, it is almost the same distribution as the public test set. But that's not by any mean a rule.",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1932748,
      "author_name": "Gaurav Rawat",
      "author_url": "",
      "post_date": "2022-09-09T21:59:55.870000",
      "content": "<p>Just worked on cite for now .. <em>will keep updating here</em></p>\n<p><strong>Multinome</strong> 0.67~ Ensemble  </p>\n<ul>\n<li>Approach 1 NN: 0.6682</li>\n</ul>\n<p><strong>Cite</strong> : ensemble 0.905</p>\n<ul>\n<li>Approach 1 NN : 0.88992</li>\n<li>Approach 2 NN: 0.8926066667</li>\n<li>Approach 3 LGBM: 0.8844 </li>\n</ul>",
      "votes": 3,
      "replies": [
        {
          "id": 1988863,
          "author_name": "Priyanshu Chaudhary",
          "author_url": "",
          "post_date": "2022-10-15T15:33:25.273000",
          "content": "<p>may I ask what validation strategy(fold split) are you using?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1988994,
          "author_name": "Gaurav Rawat",
          "author_url": "",
          "post_date": "2022-10-15T17:00:57.270000",
          "content": "<p>Hi.. By donor for now but thinking of trying few more ideas so the networks can learn but more as scores bit stuck at one point 😄</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1924699,
      "author_name": "Karim Fadi",
      "author_url": "",
      "post_date": "2022-09-03T10:01:42.127000",
      "content": "<p>CV: 0.905<br>\nLB: 0.848</p>\n<p>I think I am overfitting :(</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1924864,
          "author_name": "Lucas Morin",
          "author_url": "",
          "post_date": "2022-09-03T13:23:35.560000",
          "content": "<p>Are you using day / patient as a feature ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1925219,
          "author_name": "Karim Fadi",
          "author_url": "",
          "post_date": "2022-09-03T18:18:02.880000",
          "content": "<p>No, I don't use the metadata for anything.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1971305,
      "author_name": "Elias",
      "author_url": "",
      "post_date": "2022-10-04T14:17:14.347000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chouchouchen\" target=\"_blank\">@chouchouchen</a> , As AmbrosM explained me, after <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933\" target=\"_blank\">The data update of 2022-09-10</a> the weights changed a bit.<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/353523\" target=\"_blank\"> They are now 0.7105 and 0.2895 </a>. There is still a difference between the CV and LB though.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1926845,
      "author_name": "Oleg Khudyakov",
      "author_url": "",
      "post_date": "2022-09-05T07:18:58.880000",
      "content": "<p>Greetings, Martin!</p>\n<p>If we look <a href=\"http://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart#Submission\" target=\"_blank\">at this notebook (Submission part)</a> we will see the text:<br>\nThe CITEseq test predictions have 48663 rows (i.e., cells) and 140 columns (i.e. proteins). 48663 * 140 = 6812820. <strong>The final submission will have 65,744,180 rows, of which the first 6,812,820 are for the CITEseq predictions and the remaining 58,931,360 for the Multiome predictions.</strong></p>\n<p>And if we divide number of CITESEQ and MULTIOME rows in submission file by total submission length we will get 0.104 for CITESEQ and 0.896 for MULTIOME.</p>\n<p>Are You sure about 0.743 for SITE and 0.257 for MULTI?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1927280,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-09-05T14:01:48.923000",
          "content": "<p>The 6,812,820 rows of cite has more cell ids compared to the 58,931,360 rows for multi. Correlation score is computed per cell_id, the correct way to compute is to compare the amount of cell ids not rows.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1927318,
          "author_name": "Oleg Khudyakov",
          "author_url": "",
          "post_date": "2022-09-05T14:46:57.067000",
          "content": "<p>thank you!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1989173,
      "author_name": "Priyanshu Chaudhary",
      "author_url": "",
      "post_date": "2022-10-15T18:48:36.357000",
      "content": "<p>i am not sure Kfold is doing data leakage for me how exactly are you applying kfold?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1972361,
      "author_name": "museas",
      "author_url": "",
      "post_date": "2022-10-05T05:48:53.670000",
      "content": "<p>CITEseq:0.899(group by donor)<br>\nMultiome:0.672(Kfold) 0.670(group by donor)</p>\n<p>And I'm submitting one technique as a sample submission to check the score.<br>\nCITEseq:0.250<br>\nMultiome:-0.437<br>\nPublic:0.812</p>\n<p>if it's not my calculation error, I think the real weight of the techniques in the public is slightly off.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1973234,
          "author_name": "Gaurav Rawat",
          "author_url": "",
          "post_date": "2022-10-05T14:36:54.910000",
          "content": "<p>May I ask is ur cite an ensemble or SIngle Model MLP</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1974074,
          "author_name": "museas",
          "author_url": "",
          "post_date": "2022-10-06T04:13:25.583000",
          "content": "<p>All these are for single models, but publicLB0.250/-0.437/0.812 are best even with ensembles.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1930943,
      "author_name": "AbdurRafae",
      "author_url": "",
      "post_date": "2022-09-08T11:29:49.487000",
      "content": "<p>What's the updated CV/LB scores after the re-scoring?</p>\n<p>And wouldn't donorwise split make more sense to guage the correlation between Public Test Score and CV. Since the daywise split makes sense for the Private test, which we do not get. </p>\n<p>I'm planning to make a thorough investigation of different splits and CV correlations once I finalized my training and submission pipeline. So far not getting good enough validation scores to even consider submitting 😔</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1949929,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-22T02:06:48.813000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1930866,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-08T09:24:04.307000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1930904,
          "author_name": "Karim Fadi",
          "author_url": "",
          "post_date": "2022-09-08T10:04:00.420000",
          "content": "<p>Previously, there were 65443 test rows (48663 for cite, 16780 for multi). Now, 7016 test rows from cite were removed from scoring (they are still present in the submission file but their score doesn't count). So now we have 41647 rows for cite and still 16780 rows for multi for a total of 58427 rows.<br>\nSo the new weights are 0.712 for cite and 0.288 for multi.<br>\nHope this helps :)</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1930939,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-09-08T11:20:52.550000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1931572,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-09-08T19:36:45.450000",
          "content": "<p>Correct, but those new weights are just for public leaderboard, for private they are the same as they where before</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1931599,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-09-08T20:47:19.910000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1931646,
          "author_name": "Karim Fadi",
          "author_url": "",
          "post_date": "2022-09-08T22:19:05.930000",
          "content": "<p>Well, the private test set is hidden so we don't know the weights. But, in general, it is almost the same distribution as the public test set. But that's not by any mean a rule.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1922997": "Cite has 48663 test rows, and multi has 55935 rows (rows = amount of cells). Multi just use 30% of the rows, meaning it actually has 16780 cells.\n\nTherefore the weight for cite will be 0.743 and multi 0.257\n\nIn my case my out of folds CV for cite is 0.8882 and for multi is 0.6601. If I use the weights to sum my out of folds score I actually score 0.829 which is actually the same score of my public lb\n\nThis is using 5KFold for cite and multi:\nCV: 0.829\nLB: 0.829\n\nThis is using 3 GroupKFold for cite, 4 GroupKFold for multi using day column.\nCV: 0.8273\nLB: 0.829\n\nPlease share your results here (If you want 😄)",
    "1932748": "Just worked on cite for now .. *will keep updating here*\n\n**Multinome** 0.67~ Ensemble  \n- Approach 1 NN: 0.6682\n\n**Cite** : ensemble 0.905\n- Approach 1 NN : 0.88992\n- Approach 2 NN: 0.8926066667\n- Approach 3 LGBM: 0.8844 \n",
    "1924699": "CV: 0.905\nLB: 0.848\n\nI think I am overfitting :(",
    "1971305": "Hi @chouchouchen , As AmbrosM explained me, after [The data update of 2022-09-10](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933) the weights changed a bit.[ They are now 0.7105 and 0.2895 ](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/353523). There is still a difference between the CV and LB though.",
    "1926845": "Greetings, Martin!\n\nIf we look [at this notebook (Submission part)](http://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart#Submission) we will see the text:\nThe CITEseq test predictions have 48663 rows (i.e., cells) and 140 columns (i.e. proteins). 48663 * 140 = 6812820. **The final submission will have 65,744,180 rows, of which the first 6,812,820 are for the CITEseq predictions and the remaining 58,931,360 for the Multiome predictions.**\n\nAnd if we divide number of CITESEQ and MULTIOME rows in submission file by total submission length we will get 0.104 for CITESEQ and 0.896 for MULTIOME.\n\nAre You sure about 0.743 for SITE and 0.257 for MULTI?",
    "1989173": "i am not sure Kfold is doing data leakage for me how exactly are you applying kfold?\n",
    "1972361": "CITEseq:0.899(group by donor)\nMultiome:0.672(Kfold) 0.670(group by donor)\n\nAnd I'm submitting one technique as a sample submission to check the score.\nCITEseq:0.250\nMultiome:-0.437\nPublic:0.812\n\nif it's not my calculation error, I think the real weight of the techniques in the public is slightly off.",
    "1930943": "What's the updated CV/LB scores after the re-scoring?\n\nAnd wouldn't donorwise split make more sense to guage the correlation between Public Test Score and CV. Since the daywise split makes sense for the Private test, which we do not get. \n\nI'm planning to make a thorough investigation of different splits and CV correlations once I finalized my training and submission pipeline. So far not getting good enough validation scores to even consider submitting 😔",
    "1949929": "",
    "1930866": ""
  }
}