{
  "id": 364408,
  "title": "Methods to 0.813 without ensemble",
  "url": "/competitions/open-problems-multimodal/discussion/364408",
  "author_name": "Joseph Zhou",
  "post_date": "2022-11-06T10:04:47.962000",
  "votes": 54,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Few days ago I managed to get 0.813 by single model for both citeseq part and multiome part. I'm glad to share my methods. Since it's close to the end time I won't publish the code, but I have made a simple notebook of preprocessig for the multiome part <a href=\"https://www.kaggle.com/code/takanashihumbert/binarized-data-for-multiome-part\" target=\"_blank\">here</a>.</p>\n<p>For citeseq part:</p>\n<ul>\n<li>group kfold by 3 donors (3 folds)</li>\n<li>stay some important features, and reduce dimension for the rest features</li>\n<li>four layers DNN with dropout</li>\n</ul>\n<p>For multiome part:</p>\n<ul>\n<li>group kfold by 12 days x donors (4 folds)</li>\n<li></li>\n<li>standardize Y before dimension reduction</li>\n<li>four layers DNN with batch-normalization</li>\n</ul>\n<p>Hope it's helpful for you, good luck(Sadly I have no idea how to score 0.814☹️).</p>",
  "messages": [
    {
      "id": 2019062,
      "postDate": "2022-11-06T10:04:47.963Z",
      "content": "<p>Few days ago I managed to get 0.813 by single model for both citeseq part and multiome part. I'm glad to share my methods. Since it's close to the end time I won't publish the code, but I have made a simple notebook of preprocessig for the multiome part <a href=\"https://www.kaggle.com/code/takanashihumbert/binarized-data-for-multiome-part\" target=\"_blank\">here</a>.</p>\n<p>For citeseq part:</p>\n<ul>\n<li>group kfold by 3 donors (3 folds)</li>\n<li>stay some important features, and reduce dimension for the rest features</li>\n<li>four layers DNN with dropout</li>\n</ul>\n<p>For multiome part:</p>\n<ul>\n<li>group kfold by 12 days x donors (4 folds)</li>\n<li></li>\n<li>standardize Y before dimension reduction</li>\n<li>four layers DNN with batch-normalization</li>\n</ul>\n<p>Hope it's helpful for you, good luck(Sadly I have no idea how to score 0.814☹️).</p>",
      "rawMarkdown": "Few days ago I managed to get 0.813 by single model for both citeseq part and multiome part. I'm glad to share my methods. Since it's close to the end time I won't publish the code, but I have made a simple notebook of preprocessig for the multiome part [here](https://www.kaggle.com/code/takanashihumbert/binarized-data-for-multiome-part).\n\nFor citeseq part:\n- group kfold by 3 donors (3 folds)\n- stay some important features, and reduce dimension for the rest features\n- four layers DNN with dropout\n\nFor multiome part:\n- group kfold by 12 days x donors (4 folds)\n- ~~**binarize input values**(why does this could work? Maybe because of the high sparsity of inputs, and binarize could help to improve generalization)~~\n- standardize Y before dimension reduction\n- four layers DNN with batch-normalization\n\nHope it's helpful for you, good luck(Sadly I have no idea how to score 0.814☹️).",
      "votes": 54
    },
    {
      "id": 2023843,
      "postDate": "2022-11-10T04:46:45.070Z",
      "content": "<p>Thank you for posting an interesting topic. I have a question about kfold. Have you trained the number of kfolds models and ensemble them together or just used for evaluating models?</p>",
      "rawMarkdown": "Thank you for posting an interesting topic. I have a question about kfold. Have you trained the number of kfolds models and ensemble them together or just used for evaluating models?",
      "votes": 1,
      "replies": [
        {
          "id": 2023858,
          "postDate": "2022-11-10T05:11:21.480Z",
          "content": "<p>What has worked best for me is k-fold by donor, make predictions on the test set at each fold, and ensemble them.</p>",
          "rawMarkdown": "What has worked best for me is k-fold by donor, make predictions on the test set at each fold, and ensemble them."
        },
        {
          "id": 2025259,
          "postDate": "2022-11-11T05:41:01.590Z",
          "content": "<p>Thank you for teaching!<br>\nI will try it.</p>",
          "rawMarkdown": "Thank you for teaching!\nI will try it."
        }
      ]
    },
    {
      "id": 2020706,
      "postDate": "2022-11-07T18:45:08.987Z",
      "content": "<p>One more qq I had .. does improving Multinome part effects LB more or cite part .. For public LB I mean . Have observed when my cite model improves it effects Public LB more compared to Multinome !</p>",
      "rawMarkdown": "One more qq I had .. does improving Multinome part effects LB more or cite part .. For public LB I mean . Have observed when my cite model improves it effects Public LB more compared to Multinome !",
      "votes": 1,
      "replies": [
        {
          "id": 2022391,
          "postDate": "2022-11-09T02:47:15.870Z",
          "content": "<p>Here is a discussion <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359222\" target=\"_blank\">here</a> about the differential weighting of the two sets, with CITE having higher weight (0.661) than Multiome (0.339).  </p>",
          "rawMarkdown": "Here is a discussion [here](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359222) about the differential weighting of the two sets, with CITE having higher weight (0.661) than Multiome (0.339).  ",
          "votes": 2
        },
        {
          "id": 2023531,
          "postDate": "2022-11-09T20:11:41.307Z",
          "content": "<p>Thanks for the same , But Multinome is harder 😄 . They should keep separate LB haha </p>",
          "rawMarkdown": "Thanks for the same , But Multinome is harder 😄 . They should keep separate LB haha ",
          "votes": 1
        },
        {
          "id": 2023543,
          "postDate": "2022-11-09T20:27:46.743Z",
          "content": "<p>Agreed!  💯</p>",
          "rawMarkdown": "Agreed!  💯"
        }
      ]
    },
    {
      "id": 2019397,
      "postDate": "2022-11-06T15:01:27.983Z",
      "content": "<p>Interesting idea to binarize input values. Did you dimensionality reduced the targets also in your DNN ?</p>",
      "rawMarkdown": "Interesting idea to binarize input values. Did you dimensionality reduced the targets also in your DNN ?",
      "votes": 1,
      "replies": [
        {
          "id": 2019464,
          "postDate": "2022-11-06T16:02:25.713Z",
          "content": "<p>For multiome part, yes.</p>",
          "rawMarkdown": "For multiome part, yes.",
          "votes": 1
        },
        {
          "id": 2019725,
          "postDate": "2022-11-06T22:33:57.913Z",
          "content": "<p>Thanks for the good ideas most we did similar but some worth trying 😃.. binarize we also tried long back but somehow didn't work well for us.also saw people using it in last year comp</p>",
          "rawMarkdown": "Thanks for the good ideas most we did similar but some worth trying 😃.. binarize we also tried long back but somehow didn't work well for us.also saw people using it in last year comp",
          "votes": 1
        }
      ]
    },
    {
      "id": 2019773,
      "postDate": "2022-11-07T00:59:22.363Z",
      "content": "<p>I think this is one of the competition where one could have used a partner lol</p>",
      "rawMarkdown": "I think this is one of the competition where one could have used a partner lol",
      "votes": -1
    },
    {
      "id": 2028486,
      "postDate": "2022-11-14T01:51:41.430Z",
      "content": "<p>Thanks for sharing. It's too helpful for me!</p>",
      "rawMarkdown": "Thanks for sharing. It's too helpful for me!"
    },
    {
      "id": 2025264,
      "postDate": "2022-11-11T05:44:50.973Z",
      "content": "<p>Thank you for posting! <br>\nI have one thing that I do not quite understand. Did you train three different models for each donor and kfold through four days for the multiome part ?</p>",
      "rawMarkdown": "Thank you for posting! \nI have one thing that I do not quite understand. Did you train three different models for each donor and kfold through four days for the multiome part ?",
      "replies": [
        {
          "id": 2025655,
          "postDate": "2022-11-11T11:26:20.847Z",
          "content": "<p>I'm not certain but didn't he make 12 split by 3 donors × 4 days and use 3 of them to make 4 folds?</p>",
          "rawMarkdown": "I'm not certain but didn't he make 12 split by 3 donors × 4 days and use 3 of them to make 4 folds?"
        }
      ]
    },
    {
      "id": 2025014,
      "postDate": "2022-11-10T23:47:48.240Z",
      "content": "<p>Why did you standardize multiome output? Is that for using mean-square error for multiome part?</p>",
      "rawMarkdown": "Why did you standardize multiome output? Is that for using mean-square error for multiome part?"
    },
    {
      "id": 2021423,
      "postDate": "2022-11-08T08:27:51.857Z",
      "content": "<p>interestion idea. Thank you for posting this.</p>",
      "rawMarkdown": "interestion idea. Thank you for posting this."
    },
    {
      "id": 2021365,
      "postDate": "2022-11-08T07:37:24.617Z",
      "content": "<p>could u pls share ur CV scores?</p>",
      "rawMarkdown": "could u pls share ur CV scores?",
      "replies": [
        {
          "id": 2021440,
          "postDate": "2022-11-08T08:45:59.693Z",
          "content": "<p>Citeseq: 0.89407<br>\nMultiome: 0.66867</p>",
          "rawMarkdown": "Citeseq: 0.89407\nMultiome: 0.66867",
          "votes": 1
        },
        {
          "id": 2022387,
          "postDate": "2022-11-09T02:36:36.837Z",
          "content": "<p>fine, thanks👍</p>",
          "rawMarkdown": "fine, thanks👍"
        }
      ]
    },
    {
      "id": 2019408,
      "postDate": "2022-11-06T15:19:12.473Z",
      "content": "<p>Thank you for posting this.  I'm very interested to see what the top finishers have done once the competition is over.</p>\n<p>Could you explain this, please?  </p>\n<blockquote>\n  <p>group kfold by 12 days x donors (4 folds)</p>\n</blockquote>\n<p>There are 3 donors across 4 days.  Did you do group k-fold by day to get 4 fold?</p>",
      "rawMarkdown": "Thank you for posting this.  I'm very interested to see what the top finishers have done once the competition is over.\n\nCould you explain this, please?  \n\n>group kfold by 12 days x donors (4 folds)\n\nThere are 3 donors across 4 days.  Did you do group k-fold by day to get 4 fold?\n",
      "replies": [
        {
          "id": 2019471,
          "postDate": "2022-11-06T16:12:38.540Z",
          "content": "<blockquote>\n  <p>There are 3 donors across 4 days.</p>\n</blockquote>\n<p>it's 12 combinations. Each group data is one donor in one day.</p>\n<pre><code>kfold_1 = GroupKfold(n_splits=12)\nkfold_2 = GroupKfold(n_splits=4)\n</code></pre>\n<p>I use 'kfold_2'. 'kfold_1' will train 12 times and it's time-consuming.</p>",
          "rawMarkdown": "> There are 3 donors across 4 days.\n\nit's 12 combinations. Each group data is one donor in one day.\n```\nkfold_1 = GroupKfold(n_splits=12)\nkfold_2 = GroupKfold(n_splits=4)\n```\nI use 'kfold_2'. 'kfold_1' will train 12 times and it's time-consuming.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2026817,
      "postDate": "2022-11-12T11:04:24.503Z",
      "rawMarkdown": "",
      "votes": 6,
      "isDeleted": true
    },
    {
      "id": 2019078,
      "postDate": "2022-11-06T10:31:52.297Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2019101,
          "postDate": "2022-11-06T10:48:23.440Z",
          "content": "<p>Sorry for my lack of biology knowledge, I try to handle batch effect by days normalization.<br>\nBut it's totally useless for improving CV or LB. Maybe 'harmony' or 'bbknn' in scanpy could work? I'm not sure because I don't know how to implement.</p>",
          "rawMarkdown": "Sorry for my lack of biology knowledge, I try to handle batch effect by days normalization.\nBut it's totally useless for improving CV or LB. Maybe 'harmony' or 'bbknn' in scanpy could work? I'm not sure because I don't know how to implement."
        },
        {
          "id": 2019444,
          "postDate": "2022-11-06T15:41:25.047Z",
          "content": "<p>Actually I do not think dealing with batch effect is a good idea, because your output will always contain batch effect (either for protein or rna expression). If you removed the batch effect in the input side, the output side will also be affected.</p>",
          "rawMarkdown": "Actually I do not think dealing with batch effect is a good idea, because your output will always contain batch effect (either for protein or rna expression). If you removed the batch effect in the input side, the output side will also be affected."
        }
      ]
    },
    {
      "id": 2024941,
      "postDate": "2022-11-10T22:22:27.950Z",
      "content": "<p>Awesome idea….thanks</p>",
      "rawMarkdown": "Awesome idea....thanks",
      "votes": 1
    },
    {
      "id": 2030144,
      "postDate": "2022-11-15T08:16:32.833Z",
      "content": "<p>Thank U for your sharing</p>",
      "rawMarkdown": "Thank U for your sharing"
    },
    {
      "id": 2029732,
      "postDate": "2022-11-14T21:03:01.227Z",
      "content": "<p>thanks for your idea, it helped</p>",
      "rawMarkdown": "thanks for your idea, it helped"
    }
  ],
  "comments": [
    {
      "id": 2023843,
      "author_name": "GarudA_Kai",
      "author_url": "",
      "post_date": "2022-11-10T04:46:45.070000",
      "content": "<p>Thank you for posting an interesting topic. I have a question about kfold. Have you trained the number of kfolds models and ensemble them together or just used for evaluating models?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2023858,
          "author_name": "KirkDCO",
          "author_url": "",
          "post_date": "2022-11-10T05:11:21.480000",
          "content": "<p>What has worked best for me is k-fold by donor, make predictions on the test set at each fold, and ensemble them.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2025259,
          "author_name": "GarudA_Kai",
          "author_url": "",
          "post_date": "2022-11-11T05:41:01.590000",
          "content": "<p>Thank you for teaching!<br>\nI will try it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2020706,
      "author_name": "Gaurav Rawat",
      "author_url": "",
      "post_date": "2022-11-07T18:45:08.987000",
      "content": "<p>One more qq I had .. does improving Multinome part effects LB more or cite part .. For public LB I mean . Have observed when my cite model improves it effects Public LB more compared to Multinome !</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2022391,
          "author_name": "KirkDCO",
          "author_url": "",
          "post_date": "2022-11-09T02:47:15.870000",
          "content": "<p>Here is a discussion <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359222\" target=\"_blank\">here</a> about the differential weighting of the two sets, with CITE having higher weight (0.661) than Multiome (0.339).  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2023531,
          "author_name": "Gaurav Rawat",
          "author_url": "",
          "post_date": "2022-11-09T20:11:41.307000",
          "content": "<p>Thanks for the same , But Multinome is harder 😄 . They should keep separate LB haha </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2023543,
          "author_name": "KirkDCO",
          "author_url": "",
          "post_date": "2022-11-09T20:27:46.743000",
          "content": "<p>Agreed!  💯</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2019397,
      "author_name": "Vlad Vaduva",
      "author_url": "",
      "post_date": "2022-11-06T15:01:27.983000",
      "content": "<p>Interesting idea to binarize input values. Did you dimensionality reduced the targets also in your DNN ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2019464,
          "author_name": "Joseph Zhou",
          "author_url": "",
          "post_date": "2022-11-06T16:02:25.713000",
          "content": "<p>For multiome part, yes.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2019725,
          "author_name": "Gaurav Rawat",
          "author_url": "",
          "post_date": "2022-11-06T22:33:57.913000",
          "content": "<p>Thanks for the good ideas most we did similar but some worth trying 😃.. binarize we also tried long back but somehow didn't work well for us.also saw people using it in last year comp</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2019773,
      "author_name": "SergioMiguelM",
      "author_url": "",
      "post_date": "2022-11-07T00:59:22.363000",
      "content": "<p>I think this is one of the competition where one could have used a partner lol</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2028486,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-14T01:51:41.430000",
      "content": "<p>Thanks for sharing. It's too helpful for me!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2025264,
      "author_name": "GarudA_Kai",
      "author_url": "",
      "post_date": "2022-11-11T05:44:50.973000",
      "content": "<p>Thank you for posting! <br>\nI have one thing that I do not quite understand. Did you train three different models for each donor and kfold through four days for the multiome part ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2025655,
          "author_name": "junseonglee11",
          "author_url": "",
          "post_date": "2022-11-11T11:26:20.847000",
          "content": "<p>I'm not certain but didn't he make 12 split by 3 donors × 4 days and use 3 of them to make 4 folds?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2025014,
      "author_name": "junseonglee11",
      "author_url": "",
      "post_date": "2022-11-10T23:47:48.240000",
      "content": "<p>Why did you standardize multiome output? Is that for using mean-square error for multiome part?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2021423,
      "author_name": "lookey",
      "author_url": "",
      "post_date": "2022-11-08T08:27:51.857000",
      "content": "<p>interestion idea. Thank you for posting this.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2021365,
      "author_name": "sing4it Luo",
      "author_url": "",
      "post_date": "2022-11-08T07:37:24.617000",
      "content": "<p>could u pls share ur CV scores?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2021440,
          "author_name": "Joseph Zhou",
          "author_url": "",
          "post_date": "2022-11-08T08:45:59.693000",
          "content": "<p>Citeseq: 0.89407<br>\nMultiome: 0.66867</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2022387,
          "author_name": "sing4it Luo",
          "author_url": "",
          "post_date": "2022-11-09T02:36:36.837000",
          "content": "<p>fine, thanks👍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2019408,
      "author_name": "KirkDCO",
      "author_url": "",
      "post_date": "2022-11-06T15:19:12.473000",
      "content": "<p>Thank you for posting this.  I'm very interested to see what the top finishers have done once the competition is over.</p>\n<p>Could you explain this, please?  </p>\n<blockquote>\n  <p>group kfold by 12 days x donors (4 folds)</p>\n</blockquote>\n<p>There are 3 donors across 4 days.  Did you do group k-fold by day to get 4 fold?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2019471,
          "author_name": "Joseph Zhou",
          "author_url": "",
          "post_date": "2022-11-06T16:12:38.540000",
          "content": "<blockquote>\n  <p>There are 3 donors across 4 days.</p>\n</blockquote>\n<p>it's 12 combinations. Each group data is one donor in one day.</p>\n<pre><code>kfold_1 = GroupKfold(n_splits=12)\nkfold_2 = GroupKfold(n_splits=4)\n</code></pre>\n<p>I use 'kfold_2'. 'kfold_1' will train 12 times and it's time-consuming.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2026817,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-12T11:04:24.503000",
      "content": "",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2019078,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-06T10:31:52.297000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2019101,
          "author_name": "Joseph Zhou",
          "author_url": "",
          "post_date": "2022-11-06T10:48:23.440000",
          "content": "<p>Sorry for my lack of biology knowledge, I try to handle batch effect by days normalization.<br>\nBut it's totally useless for improving CV or LB. Maybe 'harmony' or 'bbknn' in scanpy could work? I'm not sure because I don't know how to implement.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2019444,
          "author_name": "TESUZI",
          "author_url": "",
          "post_date": "2022-11-06T15:41:25.047000",
          "content": "<p>Actually I do not think dealing with batch effect is a good idea, because your output will always contain batch effect (either for protein or rna expression). If you removed the batch effect in the input side, the output side will also be affected.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2024941,
      "author_name": "Raj Saha",
      "author_url": "",
      "post_date": "2022-11-10T22:22:27.950000",
      "content": "<p>Awesome idea….thanks</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2030144,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-15T08:16:32.833000",
      "content": "<p>Thank U for your sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2029732,
      "author_name": "Goat89",
      "author_url": "",
      "post_date": "2022-11-14T21:03:01.227000",
      "content": "<p>thanks for your idea, it helped</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2019062": "Few days ago I managed to get 0.813 by single model for both citeseq part and multiome part. I'm glad to share my methods. Since it's close to the end time I won't publish the code, but I have made a simple notebook of preprocessig for the multiome part [here](https://www.kaggle.com/code/takanashihumbert/binarized-data-for-multiome-part).\n\nFor citeseq part:\n- group kfold by 3 donors (3 folds)\n- stay some important features, and reduce dimension for the rest features\n- four layers DNN with dropout\n\nFor multiome part:\n- group kfold by 12 days x donors (4 folds)\n- ~~**binarize input values**(why does this could work? Maybe because of the high sparsity of inputs, and binarize could help to improve generalization)~~\n- standardize Y before dimension reduction\n- four layers DNN with batch-normalization\n\nHope it's helpful for you, good luck(Sadly I have no idea how to score 0.814☹️).",
    "2023843": "Thank you for posting an interesting topic. I have a question about kfold. Have you trained the number of kfolds models and ensemble them together or just used for evaluating models?",
    "2020706": "One more qq I had .. does improving Multinome part effects LB more or cite part .. For public LB I mean . Have observed when my cite model improves it effects Public LB more compared to Multinome !",
    "2019397": "Interesting idea to binarize input values. Did you dimensionality reduced the targets also in your DNN ?",
    "2019773": "I think this is one of the competition where one could have used a partner lol",
    "2028486": "Thanks for sharing. It's too helpful for me!",
    "2025264": "Thank you for posting! \nI have one thing that I do not quite understand. Did you train three different models for each donor and kfold through four days for the multiome part ?",
    "2025014": "Why did you standardize multiome output? Is that for using mean-square error for multiome part?",
    "2021423": "interestion idea. Thank you for posting this.",
    "2021365": "could u pls share ur CV scores?",
    "2019408": "Thank you for posting this.  I'm very interested to see what the top finishers have done once the competition is over.\n\nCould you explain this, please?  \n\n>group kfold by 12 days x donors (4 folds)\n\nThere are 3 donors across 4 days.  Did you do group k-fold by day to get 4 fold?\n",
    "2026817": "",
    "2019078": "",
    "2024941": "Awesome idea....thanks",
    "2030144": "Thank U for your sharing",
    "2029732": "thanks for your idea, it helped"
  }
}