{
  "id": 366453,
  "title": "2nd place solution(senkin part with code)",
  "url": "/competitions/open-problems-multimodal/writeups/senkin-tmp-2nd-place-solution-senkin-part-with-cod",
  "author_name": "",
  "post_date": "2022-12-07T08:14:00.837Z",
  "votes": 95,
  "comment_count": 29,
  "views": 0,
  "content": "<p>Thanks to all the organizers and kaggle team hosting such a challengeable competition.Thanks my team mate <a href=\"https://www.kaggle.com/baosenguo\" target=\"_blank\">@baosenguo</a>, I have no knowledge about  bioinformatics ,learned a lot from him.I thought my team could win this competition as we were at the 1st place of LB from start to end,but the time domain shift is unpredictable,we accept this result and congratulas to <a href=\"https://www.kaggle.com/shujisuzuki65\" target=\"_blank\">@shujisuzuki65</a> ,great shakeup!</p>\n<h1>Overview</h1>\n<h2>Cite</h2>\n<p><a href=\"url\" target=\"_blank\">[<img src=\"https://i.postimg.cc/8kBQqQL2/cite.png\" alt=\"cite.png\">](https://postimg.cc/s1XNhL6m)</a></p>\n<h2>Multi</h2>\n<p><a href=\"url\" target=\"_blank\">[<img src=\"https://i.postimg.cc/tgHz0f3c/multi.png\" alt=\"multi.png\">](https://postimg.cc/sMwW7T0P)</a></p>\n<h1>preprocessing</h1>\n<p>1)  <strong>centered log ratio transformation (CLR)</strong> is the best normalization method for both of cite and multi, I found the method from nature articles. <a href=\"https://www.nature.com/articles/s41467-022-29356-8\" target=\"_blank\">https://www.nature.com/articles/s41467-022-29356-8</a></p>\n<p>2) high correlation raw features with target </p>\n<p>3) <a href=\"https://www.kaggle.com/baosenguo\" target=\"_blank\">@baosenguo</a> designed fine tuned process</p>\n<ul>\n<li>using raw count:</li>\n<li>normalization:sample normalization by mean values over features</li>\n<li>transformation:sqrt transformation</li>\n<li>standardization:feature z-score</li>\n<li>batch-effect correction:take \"day\" as batch, for each batch, we calculate the column-wise median to get a \"median-sample\" representing the batch, and then subtract this sample from each sample in this batch. This method may not bring much improvement, but it is simple enough to avoid risks.</li>\n</ul>\n<p>4) row-wise zscore transformation before input to neural network</p>\n<h1>validation</h1>\n<p>The biggest challenge in this competition is how to build a robust model for unseen donor in public test and unseen day&amp;donor in private test. At the early stage I used random kfold, cross validation and LB score matched very well, so we don't need to worry about donor domain shift.But time domain shift is unpredictable, after team merge, we check our features one by one with out-of-day validation(groupkfold by day) to make sure all the features can improve every day.</p>\n<h1>model</h1>\n<ul>\n<li><p>Lightgbm<br>\ntrain 4 lightgbm models with different input features,then transform oof predictions to tsvd as nn model's meta features<br>\n-- library-size normalized and log1p transformed counts -&gt; tsvd<br>\n-- raw counts -&gt; clr -&gt; tsvd<br>\n-- raw counts<br>\n-- raw counts with raw target<br>\none trick is input sparse matrix of raw count to lightgbm directly with small \"feature_fraction\": 0.1,it brings nn model much improvment.</p></li>\n<li><p>NN<br>\nBasiclly 3layers MLP works well,one trick is to use GRU to replace first dense layer or add GRU after final dense layer.<br>\nCite target is transformed to dsb having negative values, compared to ReLU, ELU is much better to deal with negative target values,Swish is also work well for both of cite and multi.<br>\nAt the early stage I found cosine similarity is best as loss funtion for my model, after team merge, I learned from teammate to use MSE and Huber to build more different models.</p></li>\n</ul>\n<h1>notebook</h1>\n<p>[simple cite version]<br>\n<a href=\"https://www.kaggle.com/code/senkin13/2nd-place-gru-cite\" target=\"_blank\">https://www.kaggle.com/code/senkin13/2nd-place-gru-cite</a></p>\n<h1>github</h1>\n<p><a href=\"https://github.com/senkin13/kaggle/tree/master/Open-Problems-Multimodal-Single-Cell-Integration-2nd-Place-Solution\" target=\"_blank\">https://github.com/senkin13/kaggle/tree/master/Open-Problems-Multimodal-Single-Cell-Integration-2nd-Place-Solution</a></p>",
  "messages": [
    {
      "id": "2031724",
      "postDate": "11/16/2022 08:18:50",
      "content": "<p>Thanks to all the organizers and kaggle team hosting such a challengeable competition.Thanks my team mate <a href=\"https://www.kaggle.com/baosenguo\" target=\"_blank\">@baosenguo</a>, I have no knowledge about  bioinformatics ,learned a lot from him.I thought my team could win this competition as we were at the 1st place of LB from start to end,but the time domain shift is unpredictable,we accept this result and congratulas to <a href=\"https://www.kaggle.com/shujisuzuki65\" target=\"_blank\">@shujisuzuki65</a> ,great shakeup!</p>\n<h1>Overview</h1>\n<h2>Cite</h2>\n<p><a href=\"url\" target=\"_blank\">[<img src=\"https://i.postimg.cc/8kBQqQL2/cite.png\" alt=\"cite.png\">](https://postimg.cc/s1XNhL6m)</a></p>\n<h2>Multi</h2>\n<p><a href=\"url\" target=\"_blank\">[<img src=\"https://i.postimg.cc/tgHz0f3c/multi.png\" alt=\"multi.png\">](https://postimg.cc/sMwW7T0P)</a></p>\n<h1>preprocessing</h1>\n<p>1)  <strong>centered log ratio transformation (CLR)</strong> is the best normalization method for both of cite and multi, I found the method from nature articles. <a href=\"https://www.nature.com/articles/s41467-022-29356-8\" target=\"_blank\">https://www.nature.com/articles/s41467-022-29356-8</a></p>\n<p>2) high correlation raw features with target </p>\n<p>3) <a href=\"https://www.kaggle.com/baosenguo\" target=\"_blank\">@baosenguo</a> designed fine tuned process</p>\n<ul>\n<li>using raw count:</li>\n<li>normalization:sample normalization by mean values over features</li>\n<li>transformation:sqrt transformation</li>\n<li>standardization:feature z-score</li>\n<li>batch-effect correction:take \"day\" as batch, for each batch, we calculate the column-wise median to get a \"median-sample\" representing the batch, and then subtract this sample from each sample in this batch. This method may not bring much improvement, but it is simple enough to avoid risks.</li>\n</ul>\n<p>4) row-wise zscore transformation before input to neural network</p>\n<h1>validation</h1>\n<p>The biggest challenge in this competition is how to build a robust model for unseen donor in public test and unseen day&amp;donor in private test. At the early stage I used random kfold, cross validation and LB score matched very well, so we don't need to worry about donor domain shift.But time domain shift is unpredictable, after team merge, we check our features one by one with out-of-day validation(groupkfold by day) to make sure all the features can improve every day.</p>\n<h1>model</h1>\n<ul>\n<li><p>Lightgbm<br>\ntrain 4 lightgbm models with different input features,then transform oof predictions to tsvd as nn model's meta features<br>\n-- library-size normalized and log1p transformed counts -&gt; tsvd<br>\n-- raw counts -&gt; clr -&gt; tsvd<br>\n-- raw counts<br>\n-- raw counts with raw target<br>\none trick is input sparse matrix of raw count to lightgbm directly with small \"feature_fraction\": 0.1,it brings nn model much improvment.</p></li>\n<li><p>NN<br>\nBasiclly 3layers MLP works well,one trick is to use GRU to replace first dense layer or add GRU after final dense layer.<br>\nCite target is transformed to dsb having negative values, compared to ReLU, ELU is much better to deal with negative target values,Swish is also work well for both of cite and multi.<br>\nAt the early stage I found cosine similarity is best as loss funtion for my model, after team merge, I learned from teammate to use MSE and Huber to build more different models.</p></li>\n</ul>\n<h1>notebook</h1>\n<p>[simple cite version]<br>\n<a href=\"https://www.kaggle.com/code/senkin13/2nd-place-gru-cite\" target=\"_blank\">https://www.kaggle.com/code/senkin13/2nd-place-gru-cite</a></p>\n<h1>github</h1>\n<p><a href=\"https://github.com/senkin13/kaggle/tree/master/Open-Problems-Multimodal-Single-Cell-Integration-2nd-Place-Solution\" target=\"_blank\">https://github.com/senkin13/kaggle/tree/master/Open-Problems-Multimodal-Single-Cell-Integration-2nd-Place-Solution</a></p>",
      "rawMarkdown": "Thanks to all the organizers and kaggle team hosting such a challengeable competition.Thanks my team mate @baosenguo, I have no knowledge about  bioinformatics ,learned a lot from him.I thought my team could win this competition as we were at the 1st place of LB from start to end,but the time domain shift is unpredictable,we accept this result and congratulas to @shujisuzuki65 ,great shakeup!\n\n# Overview\n## Cite\n[[![cite.png](https://i.postimg.cc/8kBQqQL2/cite.png)](https://postimg.cc/s1XNhL6m)](url)\n\n## Multi\n[[![multi.png](https://i.postimg.cc/tgHz0f3c/multi.png)](https://postimg.cc/sMwW7T0P)](url)\n \n# preprocessing\n1)  **centered log ratio transformation (CLR)** is the best normalization method for both of cite and multi, I found the method from nature articles. https://www.nature.com/articles/s41467-022-29356-8\n\n2) high correlation raw features with target \n\n3) @baosenguo designed fine tuned process\n- using raw count:\n- normalization:sample normalization by mean values over features\n- transformation:sqrt transformation\n- standardization:feature z-score\n- batch-effect correction:take \"day\" as batch, for each batch, we calculate the column-wise median to get a \"median-sample\" representing the batch, and then subtract this sample from each sample in this batch. This method may not bring much improvement, but it is simple enough to avoid risks.\n\n4) row-wise zscore transformation before input to neural network\n\n# validation\nThe biggest challenge in this competition is how to build a robust model for unseen donor in public test and unseen day&donor in private test. At the early stage I used random kfold, cross validation and LB score matched very well, so we don't need to worry about donor domain shift.But time domain shift is unpredictable, after team merge, we check our features one by one with out-of-day validation(groupkfold by day) to make sure all the features can improve every day.\n\n# model\n- Lightgbm\ntrain 4 lightgbm models with different input features,then transform oof predictions to tsvd as nn model's meta features\n-- library-size normalized and log1p transformed counts -> tsvd\n-- raw counts -> clr -> tsvd\n-- raw counts\n-- raw counts with raw target\none trick is input sparse matrix of raw count to lightgbm directly with small \"feature_fraction\": 0.1,it brings nn model much improvment.\n\n- NN\nBasiclly 3layers MLP works well,one trick is to use GRU to replace first dense layer or add GRU after final dense layer.\nCite target is transformed to dsb having negative values, compared to ReLU, ELU is much better to deal with negative target values,Swish is also work well for both of cite and multi.\nAt the early stage I found cosine similarity is best as loss funtion for my model, after team merge, I learned from teammate to use MSE and Huber to build more different models.\n\n# notebook\n[simple cite version]\nhttps://www.kaggle.com/code/senkin13/2nd-place-gru-cite\n\n# github\nhttps://github.com/senkin13/kaggle/tree/master/Open-Problems-Multimodal-Single-Cell-Integration-2nd-Place-Solution",
      "votes": null
    },
    {
      "id": "2031838",
      "postDate": "11/16/2022 09:08:40",
      "content": "<p>I also used cosine similarity on some models, but at the end I also used weighted combination of mse, cs, and corr loss, and it works well</p>",
      "rawMarkdown": "I also used cosine similarity on some models, but at the end I also used weighted combination of mse, cs, and corr loss, and it works well",
      "votes": null
    },
    {
      "id": "2031849",
      "postDate": "11/16/2022 09:13:51",
      "content": "<p>yes, we should ensemble models as many as possible, I regret I didn't ensemble lots of models.</p>",
      "rawMarkdown": "yes, we should ensemble models as many as possible, I regret I didn't ensemble lots of models.",
      "votes": null
    },
    {
      "id": "2032172",
      "postDate": "11/16/2022 13:08:26",
      "content": "<p>Interesting.  I did not have ideas about TF-IDF, it seems that this approach is always used in NLP. Glad to see it works.</p>",
      "rawMarkdown": "Interesting.  I did not have ideas about TF-IDF, it seems that this approach is always used in NLP. Glad to see it works.",
      "votes": null
    },
    {
      "id": "2032195",
      "postDate": "11/16/2022 13:27:40",
      "content": "<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> Thank you very much for your efforts on the competition ! <br>\nYou were the one who broke all the limits which seemed to be unbrokable, thus giving the others the example to follow  !<br>\nSo mainly due to your efforts we kind of can estimate how much we can extract from that data - that it is important for the research community.</p>\n<p>PS</p>\n<p>If you would have time to share your experience in a zoom webinar - it would be very great ! <br>\nPSPS and welcome to join our telegram chat: <a href=\"https://t.me/sberlogacompete\" target=\"_blank\">https://t.me/sberlogacompete</a> </p>",
      "rawMarkdown": "senkin13 Thank you very much for your efforts on the competition ! \nYou were the one who broke all the limits which seemed to be unbrokable, thus giving the others the example to follow  !\nSo mainly due to your efforts we kind of can estimate how much we can extract from that data - that it is important for the research community.\n\nPS\n\nIf you would have time to share your experience in a zoom webinar - it would be very great ! \nPSPS and welcome to join our telegram chat: https://t.me/sberlogacompete",
      "votes": null
    },
    {
      "id": "2032227",
      "postDate": "11/16/2022 13:49:11",
      "content": "<p>Thank you for invitation, I am going to share more details at NeurIPS workshop at 7/Dec, you can follow that.</p>",
      "rawMarkdown": "Thank you for invitation, I am going to share more details at NeurIPS workshop at 7/Dec, you can follow that.",
      "votes": null
    },
    {
      "id": "2032230",
      "postDate": "11/16/2022 13:51:53",
      "content": "<p>yes, it also can be used for GBDT or ridge model.By the way,TF-IDF is the old version data provided,you can check Dataset Description page.</p>",
      "rawMarkdown": "yes, it also can be used for GBDT or ridge model.By the way,TF-IDF is the old version data provided,you can check Dataset Description page.",
      "votes": null
    },
    {
      "id": "2032265",
      "postDate": "11/16/2022 14:18:27",
      "content": "<p>Super strong solution! Thank you for sharing!</p>",
      "rawMarkdown": "Super strong solution! Thank you for sharing!",
      "votes": null
    },
    {
      "id": "2032538",
      "postDate": "11/16/2022 17:02:42",
      "content": "<p>May I know how much the row-wise z-score step help you score up?</p>",
      "rawMarkdown": "May I know how much the row-wise z-score step help you score up?",
      "votes": null
    },
    {
      "id": "2032582",
      "postDate": "11/16/2022 17:31:18",
      "content": "<p>Thank you for sharing! Would you like to share the consideration why gru works in this dataset, does it mean that through forget and add new information through features, gru layer extract more information than noise compared to mlp? </p>",
      "rawMarkdown": "Thank you for sharing! Would you like to share the consideration why gru works in this dataset, does it mean that through forget and add new information through features, gru layer extract more information than noise compared to mlp?",
      "votes": null
    },
    {
      "id": "2033035",
      "postDate": "11/17/2022 02:03:28",
      "content": "<p>I don't have precise theory supported,but I assume you are right,gru layer extract some more and different information than mlp.<br>\nthere were two successful experience in the past。<br>\n<a href=\"https://www.kaggle.com/competitions/favorita-grocery-sales-forecasting/discussion/47582\" target=\"_blank\">https://www.kaggle.com/competitions/favorita-grocery-sales-forecasting/discussion/47582</a><br>\n<a href=\"https://www.kaggle.com/competitions/talkingdata-adtracking-fraud-detection/discussion/56262\" target=\"_blank\">https://www.kaggle.com/competitions/talkingdata-adtracking-fraud-detection/discussion/56262</a></p>",
      "rawMarkdown": "I don't have precise theory supported,but I assume you are right,gru layer extract some more and different information than mlp.\nthere were two successful experience in the past。\nhttps://www.kaggle.com/competitions/favorita-grocery-sales-forecasting/discussion/47582\nhttps://www.kaggle.com/competitions/talkingdata-adtracking-fraud-detection/discussion/56262",
      "votes": null
    },
    {
      "id": "2033038",
      "postDate": "11/17/2022 02:04:04",
      "content": "<p>I remember it boosted cv 0.0003</p>",
      "rawMarkdown": "I remember it boosted cv 0.0003",
      "votes": null
    },
    {
      "id": "2033935",
      "postDate": "11/17/2022 17:01:39",
      "content": "<p>awesome solution, thanks!</p>",
      "rawMarkdown": "awesome solution, thanks!",
      "votes": null
    },
    {
      "id": "2034172",
      "postDate": "11/17/2022 22:07:42",
      "content": "<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a>  Big congratulation and thanks a lot for sharing your team solutions. <br>\nMay I ask in the early stage why you typically choose the cosine similarity as your loss function? Did you try a batch of the different loss function and find cosine-similarity was the best score?</p>",
      "rawMarkdown": "senkin13  Big congratulation and thanks a lot for sharing your team solutions. \nMay I ask in the early stage why you typically choose the cosine similarity as your loss function? Did you try a batch of the different loss function and find cosine-similarity was the best score?",
      "votes": null
    },
    {
      "id": "2034637",
      "postDate": "11/18/2022 10:00:04",
      "content": "<p>Hi! Tried to upgrade your notebook, changed your GRU net on my best CNN with 1D and 2D convs. Got a little <a href=\"https://www.kaggle.com/code/bejeweled/2nd-place-cite-2d-cnn/notebook\" target=\"_blank\">improvements</a></p>",
      "rawMarkdown": "Hi! Tried to upgrade your notebook, changed your GRU net on my best CNN with 1D and 2D convs. Got a little [improvements](https://www.kaggle.com/code/bejeweled/2nd-place-cite-2d-cnn/notebook)",
      "votes": null
    },
    {
      "id": "2035112",
      "postDate": "11/18/2022 16:14:18",
      "content": "<p>Do you use the raw data for cite exclusively? and not use the kaggle preprocessed version at all? Thank you.</p>",
      "rawMarkdown": "Do you use the raw data for cite exclusively? and not use the kaggle preprocessed version at all? Thank you.",
      "votes": null
    },
    {
      "id": "2035509",
      "postDate": "11/18/2022 23:57:53",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> and <a href=\"https://www.kaggle.com/baosenguo\" target=\"_blank\">@baosenguo</a> Great job leading the public LB during the competition and great job building a generalizing model to stay top on private LB!</p>",
      "rawMarkdown": "Congratulations @senkin13 and @baosenguo Great job leading the public LB during the competition and great job building a generalizing model to stay top on private LB!",
      "votes": null
    },
    {
      "id": "2035768",
      "postDate": "11/19/2022 07:27:46",
      "content": "<p>Thank you for sharing! I still don't understand the data processing. Where can I see more detailed information about（train_cite_X.shape）→(70988, 1009)，Can you explain why it is 1009 columns？</p>",
      "rawMarkdown": "Thank you for sharing! I still don't understand the data processing. Where can I see more detailed information about（train_cite_X.shape）→(70988, 1009)，Can you explain why it is 1009 columns？",
      "votes": null
    },
    {
      "id": "2037060",
      "postDate": "11/20/2022 12:00:39",
      "content": "<p>Congratulations again  with the great solution! </p>\n<p>May I ask you - you write:<br>\n\"we check our features one by one with out-of-day validation(groupkfold by day) to make sure all the features can improve every day\"</p>\n<p>How exactly that check can be done ?<br>\nAnd how much resources it takes ? I mean you have something like 1000 features - so train new models dropping features one by one - 1000 trains - probably not the way you used ?</p>",
      "rawMarkdown": "Congratulations again  with the great solution! \n\nMay I ask you - you write:\n\"we check our features one by one with out-of-day validation(groupkfold by day) to make sure all the features can improve every day\"\n\nHow exactly that check can be done ?\nAnd how much resources it takes ? I mean you have something like 1000 features - so train new models dropping features one by one - 1000 trains - probably not the way you used ?",
      "votes": null
    },
    {
      "id": "2038802",
      "postDate": "11/21/2022 17:01:30",
      "content": "<p>Congratulation! and thanks for sharing the notebook, its super useful!</p>",
      "rawMarkdown": "Congratulation! and thanks for sharing the notebook, its super useful!",
      "votes": null
    },
    {
      "id": "2041639",
      "postDate": "11/24/2022 05:17:28",
      "content": "<p>Congratulations and thanks for posting your great solution senkin13!</p>",
      "rawMarkdown": "Congratulations and thanks for posting your great solution senkin13!",
      "votes": null
    },
    {
      "id": "2041695",
      "postDate": "11/24/2022 06:20:24",
      "content": "<p>sorry, I mean one group by one group, for example clr of raw count is one group, lightgbm oof predictions is one group.</p>",
      "rawMarkdown": "sorry, I mean one group by one group, for example clr of raw count is one group, lightgbm oof predictions is one group.",
      "votes": null
    },
    {
      "id": "2041699",
      "postDate": "11/24/2022 06:21:58",
      "content": "<p>yes, I only use the raw data for cite exclusively</p>",
      "rawMarkdown": "yes, I only use the raw data for cite exclusively",
      "votes": null
    },
    {
      "id": "2041701",
      "postDate": "11/24/2022 06:23:09",
      "content": "<p>good job, it seems lower cv,higher plb, our random kfold didn't fit private very well</p>",
      "rawMarkdown": "good job, it seems lower cv,higher plb, our random kfold didn't fit private very well",
      "votes": null
    },
    {
      "id": "2041703",
      "postDate": "11/24/2022 06:24:07",
      "content": "<p>I tried mse,mae,pearson correlation,cosine similarity, cosine similarity is best for my model</p>",
      "rawMarkdown": "I tried mse,mae,pearson correlation,cosine similarity, cosine similarity is best for my model",
      "votes": null
    },
    {
      "id": "2041708",
      "postDate": "11/24/2022 06:26:28",
      "content": "<p>it includes clr transformation(200), lgb oof(400), fine-tuned transformation(164), high correlation and important features(245)</p>",
      "rawMarkdown": "it includes clr transformation(200), lgb oof(400), fine-tuned transformation(164), high correlation and important features(245)",
      "votes": null
    },
    {
      "id": "2057578",
      "postDate": "12/07/2022 08:14:02",
      "content": "<p>source code uploaded</p>",
      "rawMarkdown": "source code uploaded",
      "votes": null
    },
    {
      "id": "2174401",
      "postDate": "03/09/2023 05:23:07",
      "content": "<p>I'm really new to those topic. I'm just confused about the purpose of LightBGM model, could you please briefly explain it to me? Thank in advance!!!</p>",
      "rawMarkdown": "I'm really new to those topic. I'm just confused about the purpose of LightBGM model, could you please briefly explain it to me? Thank in advance!!!",
      "votes": null
    },
    {
      "id": "2305320",
      "postDate": "06/16/2023 15:36:05",
      "content": "<p>I think, LightBGM use extracting features</p>",
      "rawMarkdown": "I think, LightBGM use extracting features",
      "votes": null
    },
    {
      "id": "2310250",
      "postDate": "06/20/2023 08:35:10",
      "content": "<p>I am interested in your solutions. So, I tried to run your code, but I running LightLBM on the CPU very slowly. Thus, I wonder whether your team uses GPU when training the LightLBM model. I try to config the LightLBM model to access GPU, but I meet the bug environment. </p>",
      "rawMarkdown": "I am interested in your solutions. So, I tried to run your code, but I running LightLBM on the CPU very slowly. Thus, I wonder whether your team uses GPU when training the LightLBM model. I try to config the LightLBM model to access GPU, but I meet the bug environment.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2031838,
      "author_name": "bejeweled",
      "author_url": "",
      "post_date": "11/16/2022 09:08:40",
      "content": "<p>I also used cosine similarity on some models, but at the end I also used weighted combination of mse, cs, and corr loss, and it works well</p>",
      "votes": null,
      "replies": [
        {
          "id": 2031849,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "11/16/2022 09:13:51",
          "content": "<p>yes, we should ensemble models as many as possible, I regret I didn't ensemble lots of models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2032172,
      "author_name": "llttyy",
      "author_url": "",
      "post_date": "11/16/2022 13:08:26",
      "content": "<p>Interesting.  I did not have ideas about TF-IDF, it seems that this approach is always used in NLP. Glad to see it works.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2032230,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "11/16/2022 13:51:53",
          "content": "<p>yes, it also can be used for GBDT or ridge model.By the way,TF-IDF is the old version data provided,you can check Dataset Description page.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2032195,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "11/16/2022 13:27:40",
      "content": "<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> Thank you very much for your efforts on the competition ! <br>\nYou were the one who broke all the limits which seemed to be unbrokable, thus giving the others the example to follow  !<br>\nSo mainly due to your efforts we kind of can estimate how much we can extract from that data - that it is important for the research community.</p>\n<p>PS</p>\n<p>If you would have time to share your experience in a zoom webinar - it would be very great ! <br>\nPSPS and welcome to join our telegram chat: <a href=\"https://t.me/sberlogacompete\" target=\"_blank\">https://t.me/sberlogacompete</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 2032227,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "11/16/2022 13:49:11",
          "content": "<p>Thank you for invitation, I am going to share more details at NeurIPS workshop at 7/Dec, you can follow that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2032265,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/16/2022 14:18:27",
      "content": "<p>Super strong solution! Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2032538,
      "author_name": "yjyang027",
      "author_url": "",
      "post_date": "11/16/2022 17:02:42",
      "content": "<p>May I know how much the row-wise z-score step help you score up?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2033038,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "11/17/2022 02:04:04",
          "content": "<p>I remember it boosted cv 0.0003</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2032582,
      "author_name": "learnmore1",
      "author_url": "",
      "post_date": "11/16/2022 17:31:18",
      "content": "<p>Thank you for sharing! Would you like to share the consideration why gru works in this dataset, does it mean that through forget and add new information through features, gru layer extract more information than noise compared to mlp? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2033035,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "11/17/2022 02:03:28",
          "content": "<p>I don't have precise theory supported,but I assume you are right,gru layer extract some more and different information than mlp.<br>\nthere were two successful experience in the past。<br>\n<a href=\"https://www.kaggle.com/competitions/favorita-grocery-sales-forecasting/discussion/47582\" target=\"_blank\">https://www.kaggle.com/competitions/favorita-grocery-sales-forecasting/discussion/47582</a><br>\n<a href=\"https://www.kaggle.com/competitions/talkingdata-adtracking-fraud-detection/discussion/56262\" target=\"_blank\">https://www.kaggle.com/competitions/talkingdata-adtracking-fraud-detection/discussion/56262</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2033935,
          "author_name": "learnmore1",
          "author_url": "",
          "post_date": "11/17/2022 17:01:39",
          "content": "<p>awesome solution, thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2034172,
      "author_name": "bwhale",
      "author_url": "",
      "post_date": "11/17/2022 22:07:42",
      "content": "<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a>  Big congratulation and thanks a lot for sharing your team solutions. <br>\nMay I ask in the early stage why you typically choose the cosine similarity as your loss function? Did you try a batch of the different loss function and find cosine-similarity was the best score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2041703,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "11/24/2022 06:24:07",
          "content": "<p>I tried mse,mae,pearson correlation,cosine similarity, cosine similarity is best for my model</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2034637,
      "author_name": "bejeweled",
      "author_url": "",
      "post_date": "11/18/2022 10:00:04",
      "content": "<p>Hi! Tried to upgrade your notebook, changed your GRU net on my best CNN with 1D and 2D convs. Got a little <a href=\"https://www.kaggle.com/code/bejeweled/2nd-place-cite-2d-cnn/notebook\" target=\"_blank\">improvements</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2041701,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "11/24/2022 06:23:09",
          "content": "<p>good job, it seems lower cv,higher plb, our random kfold didn't fit private very well</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2035112,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/18/2022 16:14:18",
      "content": "<p>Do you use the raw data for cite exclusively? and not use the kaggle preprocessed version at all? Thank you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2041699,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "11/24/2022 06:21:58",
          "content": "<p>yes, I only use the raw data for cite exclusively</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2035509,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/18/2022 23:57:53",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> and <a href=\"https://www.kaggle.com/baosenguo\" target=\"_blank\">@baosenguo</a> Great job leading the public LB during the competition and great job building a generalizing model to stay top on private LB!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2035768,
      "author_name": "julyhappy",
      "author_url": "",
      "post_date": "11/19/2022 07:27:46",
      "content": "<p>Thank you for sharing! I still don't understand the data processing. Where can I see more detailed information about（train_cite_X.shape）→(70988, 1009)，Can you explain why it is 1009 columns？</p>",
      "votes": null,
      "replies": [
        {
          "id": 2041708,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "11/24/2022 06:26:28",
          "content": "<p>it includes clr transformation(200), lgb oof(400), fine-tuned transformation(164), high correlation and important features(245)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2037060,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "11/20/2022 12:00:39",
      "content": "<p>Congratulations again  with the great solution! </p>\n<p>May I ask you - you write:<br>\n\"we check our features one by one with out-of-day validation(groupkfold by day) to make sure all the features can improve every day\"</p>\n<p>How exactly that check can be done ?<br>\nAnd how much resources it takes ? I mean you have something like 1000 features - so train new models dropping features one by one - 1000 trains - probably not the way you used ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2041695,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "11/24/2022 06:20:24",
          "content": "<p>sorry, I mean one group by one group, for example clr of raw count is one group, lightgbm oof predictions is one group.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2038802,
      "author_name": "nicolasfrateg",
      "author_url": "",
      "post_date": "11/21/2022 17:01:30",
      "content": "<p>Congratulation! and thanks for sharing the notebook, its super useful!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2041639,
      "author_name": "songqizhou",
      "author_url": "",
      "post_date": "11/24/2022 05:17:28",
      "content": "<p>Congratulations and thanks for posting your great solution senkin13!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2057578,
      "author_name": "senkin13",
      "author_url": "",
      "post_date": "12/07/2022 08:14:02",
      "content": "<p>source code uploaded</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2174401,
      "author_name": "notran",
      "author_url": "",
      "post_date": "03/09/2023 05:23:07",
      "content": "<p>I'm really new to those topic. I'm just confused about the purpose of LightBGM model, could you please briefly explain it to me? Thank in advance!!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2305320,
          "author_name": "pxtri2156",
          "author_url": "",
          "post_date": "06/16/2023 15:36:05",
          "content": "<p>I think, LightBGM use extracting features</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2310250,
      "author_name": "pxtri2156",
      "author_url": "",
      "post_date": "06/20/2023 08:35:10",
      "content": "<p>I am interested in your solutions. So, I tried to run your code, but I running LightLBM on the CPU very slowly. Thus, I wonder whether your team uses GPU when training the LightLBM model. I try to config the LightLBM model to access GPU, but I meet the bug environment. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2031724": "Thanks to all the organizers and kaggle team hosting such a challengeable competition.Thanks my team mate @baosenguo, I have no knowledge about  bioinformatics ,learned a lot from him.I thought my team could win this competition as we were at the 1st place of LB from start to end,but the time domain shift is unpredictable,we accept this result and congratulas to @shujisuzuki65 ,great shakeup!\n\n# Overview\n## Cite\n[[![cite.png](https://i.postimg.cc/8kBQqQL2/cite.png)](https://postimg.cc/s1XNhL6m)](url)\n\n## Multi\n[[![multi.png](https://i.postimg.cc/tgHz0f3c/multi.png)](https://postimg.cc/sMwW7T0P)](url)\n \n# preprocessing\n1)  **centered log ratio transformation (CLR)** is the best normalization method for both of cite and multi, I found the method from nature articles. https://www.nature.com/articles/s41467-022-29356-8\n\n2) high correlation raw features with target \n\n3) @baosenguo designed fine tuned process\n- using raw count:\n- normalization:sample normalization by mean values over features\n- transformation:sqrt transformation\n- standardization:feature z-score\n- batch-effect correction:take \"day\" as batch, for each batch, we calculate the column-wise median to get a \"median-sample\" representing the batch, and then subtract this sample from each sample in this batch. This method may not bring much improvement, but it is simple enough to avoid risks.\n\n4) row-wise zscore transformation before input to neural network\n\n# validation\nThe biggest challenge in this competition is how to build a robust model for unseen donor in public test and unseen day&donor in private test. At the early stage I used random kfold, cross validation and LB score matched very well, so we don't need to worry about donor domain shift.But time domain shift is unpredictable, after team merge, we check our features one by one with out-of-day validation(groupkfold by day) to make sure all the features can improve every day.\n\n# model\n- Lightgbm\ntrain 4 lightgbm models with different input features,then transform oof predictions to tsvd as nn model's meta features\n-- library-size normalized and log1p transformed counts -> tsvd\n-- raw counts -> clr -> tsvd\n-- raw counts\n-- raw counts with raw target\none trick is input sparse matrix of raw count to lightgbm directly with small \"feature_fraction\": 0.1,it brings nn model much improvment.\n\n- NN\nBasiclly 3layers MLP works well,one trick is to use GRU to replace first dense layer or add GRU after final dense layer.\nCite target is transformed to dsb having negative values, compared to ReLU, ELU is much better to deal with negative target values,Swish is also work well for both of cite and multi.\nAt the early stage I found cosine similarity is best as loss funtion for my model, after team merge, I learned from teammate to use MSE and Huber to build more different models.\n\n# notebook\n[simple cite version]\nhttps://www.kaggle.com/code/senkin13/2nd-place-gru-cite\n\n# github\nhttps://github.com/senkin13/kaggle/tree/master/Open-Problems-Multimodal-Single-Cell-Integration-2nd-Place-Solution",
    "2031838": "I also used cosine similarity on some models, but at the end I also used weighted combination of mse, cs, and corr loss, and it works well",
    "2031849": "yes, we should ensemble models as many as possible, I regret I didn't ensemble lots of models.",
    "2032172": "Interesting.  I did not have ideas about TF-IDF, it seems that this approach is always used in NLP. Glad to see it works.",
    "2032195": "senkin13 Thank you very much for your efforts on the competition ! \nYou were the one who broke all the limits which seemed to be unbrokable, thus giving the others the example to follow  !\nSo mainly due to your efforts we kind of can estimate how much we can extract from that data - that it is important for the research community.\n\nPS\n\nIf you would have time to share your experience in a zoom webinar - it would be very great ! \nPSPS and welcome to join our telegram chat: https://t.me/sberlogacompete",
    "2032227": "Thank you for invitation, I am going to share more details at NeurIPS workshop at 7/Dec, you can follow that.",
    "2032230": "yes, it also can be used for GBDT or ridge model.By the way,TF-IDF is the old version data provided,you can check Dataset Description page.",
    "2032265": "Super strong solution! Thank you for sharing!",
    "2032538": "May I know how much the row-wise z-score step help you score up?",
    "2032582": "Thank you for sharing! Would you like to share the consideration why gru works in this dataset, does it mean that through forget and add new information through features, gru layer extract more information than noise compared to mlp?",
    "2033035": "I don't have precise theory supported,but I assume you are right,gru layer extract some more and different information than mlp.\nthere were two successful experience in the past。\nhttps://www.kaggle.com/competitions/favorita-grocery-sales-forecasting/discussion/47582\nhttps://www.kaggle.com/competitions/talkingdata-adtracking-fraud-detection/discussion/56262",
    "2033038": "I remember it boosted cv 0.0003",
    "2033935": "awesome solution, thanks!",
    "2034172": "senkin13  Big congratulation and thanks a lot for sharing your team solutions. \nMay I ask in the early stage why you typically choose the cosine similarity as your loss function? Did you try a batch of the different loss function and find cosine-similarity was the best score?",
    "2034637": "Hi! Tried to upgrade your notebook, changed your GRU net on my best CNN with 1D and 2D convs. Got a little [improvements](https://www.kaggle.com/code/bejeweled/2nd-place-cite-2d-cnn/notebook)",
    "2035112": "Do you use the raw data for cite exclusively? and not use the kaggle preprocessed version at all? Thank you.",
    "2035509": "Congratulations @senkin13 and @baosenguo Great job leading the public LB during the competition and great job building a generalizing model to stay top on private LB!",
    "2035768": "Thank you for sharing! I still don't understand the data processing. Where can I see more detailed information about（train_cite_X.shape）→(70988, 1009)，Can you explain why it is 1009 columns？",
    "2037060": "Congratulations again  with the great solution! \n\nMay I ask you - you write:\n\"we check our features one by one with out-of-day validation(groupkfold by day) to make sure all the features can improve every day\"\n\nHow exactly that check can be done ?\nAnd how much resources it takes ? I mean you have something like 1000 features - so train new models dropping features one by one - 1000 trains - probably not the way you used ?",
    "2038802": "Congratulation! and thanks for sharing the notebook, its super useful!",
    "2041639": "Congratulations and thanks for posting your great solution senkin13!",
    "2041695": "sorry, I mean one group by one group, for example clr of raw count is one group, lightgbm oof predictions is one group.",
    "2041699": "yes, I only use the raw data for cite exclusively",
    "2041701": "good job, it seems lower cv,higher plb, our random kfold didn't fit private very well",
    "2041703": "I tried mse,mae,pearson correlation,cosine similarity, cosine similarity is best for my model",
    "2041708": "it includes clr transformation(200), lgb oof(400), fine-tuned transformation(164), high correlation and important features(245)",
    "2057578": "source code uploaded",
    "2174401": "I'm really new to those topic. I'm just confused about the purpose of LightBGM model, could you please briefly explain it to me? Thank in advance!!!",
    "2305320": "I think, LightBGM use extracting features",
    "2310250": "I am interested in your solutions. So, I tried to run your code, but I running LightLBM on the CPU very slowly. Thus, I wonder whether your team uses GPU when training the LightLBM model. I try to config the LightLBM model to access GPU, but I meet the bug environment."
  },
  "source": "meta"
}