{
  "id": 366476,
  "title": "2nd Place Solution  -  tmp's part ( code updated )",
  "url": "/competitions/open-problems-multimodal/discussion/366476",
  "author_name": "",
  "post_date": "2022-11-16T09:57:02.988705Z",
  "votes": 42,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First of all, thanks to organizers for hosting this interesting biological competition, and thanks to my teammates <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a>.   Senkin and I used quite different methods, which enables us to obtain a better blending result.</p>\n<p>I will mainly introduce some important parts in my solution, and some of these simple tricks might be part of the reasons why we remain stable in the leaderboard ( Public 1st 😄; Private 2nd 😂).  Congratulations on winning 1st place <a href=\"https://www.kaggle.com/shujisuzuki65\" target=\"_blank\">@shujisuzuki65</a>  !</p>\n<hr>\n<h1>CITEseq</h1>\n<h2>Preprocessing</h2>\n<p>This pp pipeline is different from the method commonly used in single-cell omics.  It is more like a combination. </p>\n<ol>\n<li>using raw count: <br>\n there are many ways in pp,  so we start with the original one.</li>\n<li>normalization: <br>\nsample normalization by mean values over features</li>\n<li>transformation:  <br>\nsqrt transformation</li>\n<li>standardization: <br>\nsample z-scor<br>\nfeature z-score</li>\n<li>batch-effect correction: <br>\ntake \"day\" as batch,  for each batch, we calculate the column-wise median to get a \"median-sample\" representing the batch, and then subtract this sample from each sample in this batch.  This method may not bring much improvement, but it is simple enough to avoid risks. </li>\n</ol>\n<h2>Feature engineering</h2>\n<ol>\n<li>decomposition<br>\npca (64)<br>\nipca (128)<br>\nfactor analysis (64) <br>\nIt would be strange if only pca could work. Using more decomposition features boost cv &amp; lb.</li>\n<li>features selection <br>\nCompared with selecting features that are highly correlated to the target in the whole dataset, I prefer to select features that show stable correlation with the target in each group (donor-day) even if the pcc is slightly lower.  Consistent associations are more likely to be genuine.<br>\nIn addition, some features with the same name as the target are included by relaxing the pcc threshold. </li>\n<li>cell-type (one-hot)</li>\n</ol>\n<h2>Modelling</h2>\n<ol>\n<li><p>mlp<br>\n(single model with 1 seed - public 0.815;  private 0.772)</p></li>\n<li><p>lgb<br>\ngbdt / dart</p></li>\n</ol>\n<h2>Local CV</h2>\n<ol>\n<li>We used a simple random 5-fold for both cite and multi tasks, and cv shows good consistency with lb.</li>\n<li>We are also worried about the impact of batch effect mainly caused by the new day (the impact of donor can be tested by pub lb), therefore, we use cv split by day to verify the above parts and got the consistent conclusion.  But this is only used for proof of concept,  the final submission is the bleding of random kfold results. </li>\n</ol>\n<h1>Multiome</h1>\n<p>There is nothing special in preprocessing and modeling. Only feature engineering part might be different.  <br>\n(although this is not included in the final selected submission) </p>\n<h2>Feature engineering</h2>\n<ol>\n<li>pca (64) </li>\n<li>chromosome feature<br>\n2.1 features are grouped according to their chromosomes, and only chromosomes containing &gt; 100 features are retained. <br>\n2.2 features on each chromosome are divided into 3 groups according to their position on chromosome, then calculate mean, proportion of non-zero feautres and mean of non-zero value in each group.<br>\n2.3 features on each chromosome are divided into 2 groups according to their position on chromosome, implement pca (2-dims) in each group.</li>\n<li>binary pca and chromosome feature<br>\nafter binarization, use the same method as above to obtain features of the binarized version.</li>\n</ol>\n<h1>Code</h1>\n<p>code uploaded:<br>\n<a href=\"https://github.com/baosenguo/Kaggle-Open-Problems-Multimodal-Single-Cell-Integration-2nd-Place-Solution\" target=\"_blank\">https://github.com/baosenguo/Kaggle-Open-Problems-Multimodal-Single-Cell-Integration-2nd-Place-Solution</a></p>",
  "messages": [
    {
      "id": "2031908",
      "postDate": "11/16/2022 09:57:02",
      "content": "<p>First of all, thanks to organizers for hosting this interesting biological competition, and thanks to my teammates <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a>.   Senkin and I used quite different methods, which enables us to obtain a better blending result.</p>\n<p>I will mainly introduce some important parts in my solution, and some of these simple tricks might be part of the reasons why we remain stable in the leaderboard ( Public 1st 😄; Private 2nd 😂).  Congratulations on winning 1st place <a href=\"https://www.kaggle.com/shujisuzuki65\" target=\"_blank\">@shujisuzuki65</a>  !</p>\n<hr>\n<h1>CITEseq</h1>\n<h2>Preprocessing</h2>\n<p>This pp pipeline is different from the method commonly used in single-cell omics.  It is more like a combination. </p>\n<ol>\n<li>using raw count: <br>\n there are many ways in pp,  so we start with the original one.</li>\n<li>normalization: <br>\nsample normalization by mean values over features</li>\n<li>transformation:  <br>\nsqrt transformation</li>\n<li>standardization: <br>\nsample z-scor<br>\nfeature z-score</li>\n<li>batch-effect correction: <br>\ntake \"day\" as batch,  for each batch, we calculate the column-wise median to get a \"median-sample\" representing the batch, and then subtract this sample from each sample in this batch.  This method may not bring much improvement, but it is simple enough to avoid risks. </li>\n</ol>\n<h2>Feature engineering</h2>\n<ol>\n<li>decomposition<br>\npca (64)<br>\nipca (128)<br>\nfactor analysis (64) <br>\nIt would be strange if only pca could work. Using more decomposition features boost cv &amp; lb.</li>\n<li>features selection <br>\nCompared with selecting features that are highly correlated to the target in the whole dataset, I prefer to select features that show stable correlation with the target in each group (donor-day) even if the pcc is slightly lower.  Consistent associations are more likely to be genuine.<br>\nIn addition, some features with the same name as the target are included by relaxing the pcc threshold. </li>\n<li>cell-type (one-hot)</li>\n</ol>\n<h2>Modelling</h2>\n<ol>\n<li><p>mlp<br>\n(single model with 1 seed - public 0.815;  private 0.772)</p></li>\n<li><p>lgb<br>\ngbdt / dart</p></li>\n</ol>\n<h2>Local CV</h2>\n<ol>\n<li>We used a simple random 5-fold for both cite and multi tasks, and cv shows good consistency with lb.</li>\n<li>We are also worried about the impact of batch effect mainly caused by the new day (the impact of donor can be tested by pub lb), therefore, we use cv split by day to verify the above parts and got the consistent conclusion.  But this is only used for proof of concept,  the final submission is the bleding of random kfold results. </li>\n</ol>\n<h1>Multiome</h1>\n<p>There is nothing special in preprocessing and modeling. Only feature engineering part might be different.  <br>\n(although this is not included in the final selected submission) </p>\n<h2>Feature engineering</h2>\n<ol>\n<li>pca (64) </li>\n<li>chromosome feature<br>\n2.1 features are grouped according to their chromosomes, and only chromosomes containing &gt; 100 features are retained. <br>\n2.2 features on each chromosome are divided into 3 groups according to their position on chromosome, then calculate mean, proportion of non-zero feautres and mean of non-zero value in each group.<br>\n2.3 features on each chromosome are divided into 2 groups according to their position on chromosome, implement pca (2-dims) in each group.</li>\n<li>binary pca and chromosome feature<br>\nafter binarization, use the same method as above to obtain features of the binarized version.</li>\n</ol>\n<h1>Code</h1>\n<p>code uploaded:<br>\n<a href=\"https://github.com/baosenguo/Kaggle-Open-Problems-Multimodal-Single-Cell-Integration-2nd-Place-Solution\" target=\"_blank\">https://github.com/baosenguo/Kaggle-Open-Problems-Multimodal-Single-Cell-Integration-2nd-Place-Solution</a></p>",
      "rawMarkdown": "First of all, thanks to organizers for hosting this interesting biological competition, and thanks to my teammates @senkin13.   Senkin and I used quite different methods, which enables us to obtain a better blending result.\n\nI will mainly introduce some important parts in my solution, and some of these simple tricks might be part of the reasons why we remain stable in the leaderboard ( Public 1st 😄; Private 2nd 😂).  Congratulations on winning 1st place @shujisuzuki65  !\n\n---\n\n# CITEseq\n\n## Preprocessing\n\nThis pp pipeline is different from the method commonly used in single-cell omics.  It is more like a combination. \n1. using raw count: \n     there are many ways in pp,  so we start with the original one.\n2. normalization: \nsample normalization by mean values over features\n3. transformation:  \nsqrt transformation\n4. standardization: \n    sample z-scor\n    feature z-score\n5. batch-effect correction: \ntake \"day\" as batch,  for each batch, we calculate the column-wise median to get a \"median-sample\" representing the batch, and then subtract this sample from each sample in this batch.  This method may not bring much improvement, but it is simple enough to avoid risks. \n\n## Feature engineering\n\n1. decomposition\npca (64)\nipca (128)\nfactor analysis (64) \nIt would be strange if only pca could work. Using more decomposition features boost cv & lb.\n2. features selection \nCompared with selecting features that are highly correlated to the target in the whole dataset, I prefer to select features that show stable correlation with the target in each group (donor-day) even if the pcc is slightly lower.  Consistent associations are more likely to be genuine.\nIn addition, some features with the same name as the target are included by relaxing the pcc threshold. \n3. cell-type (one-hot)\n\n## Modelling\n1. mlp\n (single model with 1 seed - public 0.815;  private 0.772)\n\n2. lgb\ngbdt / dart\n\n## Local CV\n1. We used a simple random 5-fold for both cite and multi tasks, and cv shows good consistency with lb.\n2. We are also worried about the impact of batch effect mainly caused by the new day (the impact of donor can be tested by pub lb), therefore, we use cv split by day to verify the above parts and got the consistent conclusion.  But this is only used for proof of concept,  the final submission is the bleding of random kfold results. \n\n# Multiome\n There is nothing special in preprocessing and modeling. Only feature engineering part might be different.  \n(although this is not included in the final selected submission) \n\n## Feature engineering\n1. pca (64) \n2. chromosome feature\n2.1 features are grouped according to their chromosomes, and only chromosomes containing > 100 features are retained. \n2.2 features on each chromosome are divided into 3 groups according to their position on chromosome, then calculate mean, proportion of non-zero feautres and mean of non-zero value in each group.\n2.3 features on each chromosome are divided into 2 groups according to their position on chromosome, implement pca (2-dims) in each group.\n3. binary pca and chromosome feature\nafter binarization, use the same method as above to obtain features of the binarized version.\n\n# Code\ncode uploaded:\nhttps://github.com/baosenguo/Kaggle-Open-Problems-Multimodal-Single-Cell-Integration-2nd-Place-Solution",
      "votes": null
    },
    {
      "id": "2032217",
      "postDate": "11/16/2022 13:44:35",
      "content": "<p>Congratulations for your simple solution that beat every one consistently in PB. I passionately watched you in 1st place through out competition. Congrats. </p>",
      "rawMarkdown": "Congratulations for your simple solution that beat every one consistently in PB. I passionately watched you in 1st place through out competition. Congrats.",
      "votes": null
    },
    {
      "id": "2032259",
      "postDate": "11/16/2022 14:11:36",
      "content": "<p>Thanks for sharing. The feature selection based on a stable cv startegy is veryyy nice method! Also feature preprocessing like batch effects also a good idea. Once the stable inputs are well-prepared, you used other more consistent cv strategy like random k fold without worrying too much overfitting problem. Thanks again. Learned a lot.</p>\n<p>Our team did batch effects correction also but haven't thinking much about how to selecting more robust features, which caused we shaked from 40+ to 700+ ranking.🤕</p>",
      "rawMarkdown": "Thanks for sharing. The feature selection based on a stable cv startegy is veryyy nice method! Also feature preprocessing like batch effects also a good idea. Once the stable inputs are well-prepared, you used other more consistent cv strategy like random k fold without worrying too much overfitting problem. Thanks again. Learned a lot.\n\nOur team did batch effects correction also but haven't thinking much about how to selecting more robust features, which caused we shaked from 40+ to 700+ ranking.🤕",
      "votes": null
    },
    {
      "id": "2033919",
      "postDate": "11/17/2022 16:52:53",
      "content": "<p>Can I ask the structure of your 13-layer MLP for CITE-seq? I tried to increase the number of layers but only make it worse….</p>",
      "rawMarkdown": "Can I ask the structure of your 13-layer MLP for CITE-seq? I tried to increase the number of layers but only make it worse....",
      "votes": null
    },
    {
      "id": "2035511",
      "postDate": "11/18/2022 23:58:36",
      "content": "<p>Congratulations  <a href=\"https://www.kaggle.com/baosenguo\" target=\"_blank\">@baosenguo</a> and <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> Great job staying on top of public LB during and great job building a generalizing model to stay top on private LB! </p>",
      "rawMarkdown": "Congratulations  @baosenguo and @senkin13 Great job staying on top of public LB during and great job building a generalizing model to stay top on private LB!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2032217,
      "author_name": "venkatapadavala",
      "author_url": "",
      "post_date": "11/16/2022 13:44:35",
      "content": "<p>Congratulations for your simple solution that beat every one consistently in PB. I passionately watched you in 1st place through out competition. Congrats. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2032259,
      "author_name": "alvinai9603",
      "author_url": "",
      "post_date": "11/16/2022 14:11:36",
      "content": "<p>Thanks for sharing. The feature selection based on a stable cv startegy is veryyy nice method! Also feature preprocessing like batch effects also a good idea. Once the stable inputs are well-prepared, you used other more consistent cv strategy like random k fold without worrying too much overfitting problem. Thanks again. Learned a lot.</p>\n<p>Our team did batch effects correction also but haven't thinking much about how to selecting more robust features, which caused we shaked from 40+ to 700+ ranking.🤕</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2033919,
      "author_name": "jinyang18",
      "author_url": "",
      "post_date": "11/17/2022 16:52:53",
      "content": "<p>Can I ask the structure of your 13-layer MLP for CITE-seq? I tried to increase the number of layers but only make it worse….</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2035511,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/18/2022 23:58:36",
      "content": "<p>Congratulations  <a href=\"https://www.kaggle.com/baosenguo\" target=\"_blank\">@baosenguo</a> and <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> Great job staying on top of public LB during and great job building a generalizing model to stay top on private LB! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2031908": "First of all, thanks to organizers for hosting this interesting biological competition, and thanks to my teammates @senkin13.   Senkin and I used quite different methods, which enables us to obtain a better blending result.\n\nI will mainly introduce some important parts in my solution, and some of these simple tricks might be part of the reasons why we remain stable in the leaderboard ( Public 1st 😄; Private 2nd 😂).  Congratulations on winning 1st place @shujisuzuki65  !\n\n---\n\n# CITEseq\n\n## Preprocessing\n\nThis pp pipeline is different from the method commonly used in single-cell omics.  It is more like a combination. \n1. using raw count: \n     there are many ways in pp,  so we start with the original one.\n2. normalization: \nsample normalization by mean values over features\n3. transformation:  \nsqrt transformation\n4. standardization: \n    sample z-scor\n    feature z-score\n5. batch-effect correction: \ntake \"day\" as batch,  for each batch, we calculate the column-wise median to get a \"median-sample\" representing the batch, and then subtract this sample from each sample in this batch.  This method may not bring much improvement, but it is simple enough to avoid risks. \n\n## Feature engineering\n\n1. decomposition\npca (64)\nipca (128)\nfactor analysis (64) \nIt would be strange if only pca could work. Using more decomposition features boost cv & lb.\n2. features selection \nCompared with selecting features that are highly correlated to the target in the whole dataset, I prefer to select features that show stable correlation with the target in each group (donor-day) even if the pcc is slightly lower.  Consistent associations are more likely to be genuine.\nIn addition, some features with the same name as the target are included by relaxing the pcc threshold. \n3. cell-type (one-hot)\n\n## Modelling\n1. mlp\n (single model with 1 seed - public 0.815;  private 0.772)\n\n2. lgb\ngbdt / dart\n\n## Local CV\n1. We used a simple random 5-fold for both cite and multi tasks, and cv shows good consistency with lb.\n2. We are also worried about the impact of batch effect mainly caused by the new day (the impact of donor can be tested by pub lb), therefore, we use cv split by day to verify the above parts and got the consistent conclusion.  But this is only used for proof of concept,  the final submission is the bleding of random kfold results. \n\n# Multiome\n There is nothing special in preprocessing and modeling. Only feature engineering part might be different.  \n(although this is not included in the final selected submission) \n\n## Feature engineering\n1. pca (64) \n2. chromosome feature\n2.1 features are grouped according to their chromosomes, and only chromosomes containing > 100 features are retained. \n2.2 features on each chromosome are divided into 3 groups according to their position on chromosome, then calculate mean, proportion of non-zero feautres and mean of non-zero value in each group.\n2.3 features on each chromosome are divided into 2 groups according to their position on chromosome, implement pca (2-dims) in each group.\n3. binary pca and chromosome feature\nafter binarization, use the same method as above to obtain features of the binarized version.\n\n# Code\ncode uploaded:\nhttps://github.com/baosenguo/Kaggle-Open-Problems-Multimodal-Single-Cell-Integration-2nd-Place-Solution",
    "2032217": "Congratulations for your simple solution that beat every one consistently in PB. I passionately watched you in 1st place through out competition. Congrats.",
    "2032259": "Thanks for sharing. The feature selection based on a stable cv startegy is veryyy nice method! Also feature preprocessing like batch effects also a good idea. Once the stable inputs are well-prepared, you used other more consistent cv strategy like random k fold without worrying too much overfitting problem. Thanks again. Learned a lot.\n\nOur team did batch effects correction also but haven't thinking much about how to selecting more robust features, which caused we shaked from 40+ to 700+ ranking.🤕",
    "2033919": "Can I ask the structure of your 13-layer MLP for CITE-seq? I tried to increase the number of layers but only make it worse....",
    "2035511": "Congratulations  @baosenguo and @senkin13 Great job staying on top of public LB during and great job building a generalizing model to stay top on private LB!"
  },
  "source": "meta"
}