{
  "id": 366460,
  "title": "4th place solution (with code)",
  "url": "/competitions/open-problems-multimodal/writeups/oliver-wang-4th-place-solution-with-code",
  "author_name": "",
  "post_date": "2022-12-06T17:46:42.947Z",
  "votes": 38,
  "comment_count": 17,
  "views": 0,
  "content": "<h2>Intro</h2>\n<p>To begin with, thanks to the Kaggle team and Open Problems team for hosting such a wonderful contest. I would also like to share my gratitude to all the competitors, especially those generous competitors who are willing to share their codes and thoughts like <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> <a href=\"https://www.kaggle.com/baosenguo\" target=\"_blank\">@baosenguo</a> <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a> and so on. I couldn't have gone so far without their help.</p>\n<p>You can find full code on Github <a href=\"https://github.com/oliverwang15/4th-Place-Solution-for-Open-Problems-Multimodal-Single-Cell\" target=\"_blank\">here</a></p>\n<h2>Cite</h2>\n<h3>Data preprocessing</h3>\n<p>At first, all of my feature engineering methods are based on the <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/data\" target=\"_blank\">original</a>  data, but my public score raised from 0.812 to 0.813 after I merely change the data source so all the feature engineering processes and based on the <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359355\" target=\"_blank\">raw</a> data. </p>\n<p>My preprocessing method is using <code>np.log1p</code> to change the raw data. I have also tried other preprocessing methods like <code>MAGIC</code> and <code>TF-IDF</code> but they can't improve my CV score.</p>\n<h3>Feature engineering</h3>\n<p>The final inputs of the models consist of mainly six parts. Three of them are dimension reduction parts including <code>Tsvd</code>, <code>UMAP</code>, and <code>Novel’s method</code>. The rest are feature selection parts including <code>name importance</code>, <code>corr importance,</code> and <code>rf importance</code>.</p>\n<ul>\n<li><p><code>Tsvd</code>: <code>TruncatedSVD(n_components=128, random_state=42)</code></p></li>\n<li><p><code>UMAP</code>: <code>UMAP(n_neighbors = 16,n_components=128, random_state=42,verbose = True)</code></p></li>\n<li><p><code>Novel’s method</code>: The original method can be found <a href=\"https://github.com/openproblems-bio/neurips2021_multimodal_topmethods/blob/dc7bd58dacbe804dcc7be047531d795b1b04741e/src/predict_modality/methods/novel/resources/helper_functions.py\" target=\"_blank\">here</a>. At first, I wanted to implement the preprocessing method to replace simple <code>log1p</code> but after I replaced the <code>Tsvd</code> results of <code>log1p</code> by  the <code>Tsvd</code> results of <code>Novel’s method</code> I found that my CV went down. But if I kept both of them, the CV score would increase a little bit. So I kept the <code>Tsvd</code> results of <code>Novel’s method</code>.</p></li>\n<li><p><code>name importance</code>: It 's mainly based on AmbrosM's <a href=\"https://www.kaggle.com/code/ambrosm/msci-eda-which-makes-sense/notebook#Name-matching\" target=\"_blank\">notebook</a>. But I added additional information from <code>mygene</code> while matching. I will release my complete preprocessing code later and specific results can be found there.</p></li>\n<li><p><code>corr importance</code>: As the name suggested, I chose the top 3 features that correlated with the targets. There was overlap and the number of selected features was about 104</p></li>\n<li><p><code>rf importance</code>: Since the feature importances of random forest may apply to NN and other models as well. So I selected 128 top feature importances of the random forest model.</p></li>\n</ul>\n<p>I have also tried other mothed including <code>PCA</code>, <code>KernelPCA</code>, <code>LocallyLinearEmbedding</code>, and <code>SpectralEmbedding</code>.<code>PCA</code> gives little help and it will cause severe overfitting when used with <code>Tsvd</code>. I could' t finish the manifold methods in 24 hours so I gave them up.</p>\n<h3>Models</h3>\n<p>I have implemented the CV strategy like the private test, but it turns out that the strategy like the public test is better. So all of the results are based on <code>GroupKFold</code> on <code>donors</code>. I have done there-layers stacking in the competition. and I have also done the ensemble on the stacking results and the results of independent models. Here are the models I used and I will also release the code later. </p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Stacking</th>\n<th>NN</th>\n<th>NN_online</th>\n<th>CNN</th>\n<th>kernel_rigde</th>\n<th>LGBM</th>\n<th>Catboost</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CV</td>\n<td>0.89677</td>\n<td>0.89596</td>\n<td>0.89580</td>\n<td>0.89530</td>\n<td>0.89326</td>\n<td>0.89270</td>\n<td>0.89100</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><p><code>NN</code>: A personal-designed NN network, trying to do something like the transformers. I used MLP to replace the dot product in the mechanism of attention. This may not be so reasonable and I am also aware of the importance of feature vectors and dot products. But I was so fascinated by attention and I also tried <code>tabnet</code> and <code>rtdl</code> but they didn't work very well. But my method seemed to work even better than simple MLP. <br>\n<a href=\"https://www.kaggle.com/oliverwang15/4th-solution-cite-nn\" target=\"_blank\">Demo notebook</a></p></li>\n<li><p><code>CNN</code>: Inspired by the tmp method <a href=\"https://www.kaggle.com/competitions/lish-moa/discussion/202256\" target=\"_blank\">here</a> and also added multidimensional convolution kernel like the Resnet. </p></li>\n<li><p><code>NN(Online)</code>: This model is mainly based on pourchot's method <a href=\"https://www.kaggle.com/code/pourchot/all-in-one-citeseq-multiome-with-keras\" target=\"_blank\">here</a> and only some tiny change was made.<br>\n<a href=\"https://www.kaggle.com/code/oliverwang15/4th-solution-cite-online-nn\" target=\"_blank\">Demo notebook</a></p></li>\n<li><p><code>Kernel Rigde</code>: This model is inspired by the best solution of last year's competition. I used <a href=\"https://docs.ray.io/en/master/tune/index.html\" target=\"_blank\">Ray Tune</a> to optimize the hypermeters<br>\n<a href=\"https://www.kaggle.com/code/oliverwang15/4th-solution-ray-tune-krr\" target=\"_blank\">Demo notebook with ray tune</a></p></li>\n<li><p><code>Catboost</code>: There are many options for <code>catboost</code> here. Using <code>MultiOutputRegressor</code> or <code>MultiRMSE</code> as <code>objective</code>.But we can't do earlystopping to prevent overfitting in the first method and the result of the second method is not good enough so I made a class <code>MultiOutputCatboostRegressor</code> personally, using <code>MSE</code> to fit the normalized targets.</p></li>\n<li><p><code>LGBM</code>: I also wrote <code>MultiOutputLGBMRegressor</code> and the results seem to be better and the training process was so slow that I had to give it up in the stacking. However, I still trained a independent LGBM model and used it in the final training.  <br>\n<a href=\"https://www.kaggle.com/code/oliverwang15/4th-solution-cite-multioutputlgbm\" target=\"_blank\">Demo notebook</a> </p></li>\n<li><p><code>stacking</code>: I used <code>KNN</code>,<code>CNN</code>,<code>ridge</code>,<code>rf</code>,<code>catboost</code>,<code>NN</code> in the first layer and only <code>CNN</code>,<code>catboost</code>,<code>NN</code> in the second and just a simple <code>MLP</code> in the last layer. To avoid overfitting, I used <code>KFold</code> and oof predictions between layers, and every stacking model are using <code>GroupKFold</code>(so there are 3 stacking models here). It seems to be a little bit to understand so you may refer to the picture. If you still have confusion please feel free to ask me.<br>\n<a href=\"https://www.kaggle.com/code/oliverwang15/4th-solution-cite-stacking-train\" target=\"_blank\">Demo notebook train</a> <br>\n<a href=\"https://www.kaggle.com/code/oliverwang15/4th-solution-cite-stacking-predict\" target=\"_blank\">Demo notebook predict</a> </p></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2641164%2F2dedb86bf1f9498fb9da9232da1e579a%2FStacking%20Training.png?generation=1668608507518581&amp;alt=media\" alt=\"\"></p>\n<table>\n<thead>\n<tr>\n<th>CV Results</th>\n<th>Model Ⅰ (vaild 32606)</th>\n<th>Model Ⅱ (vaild 13176)</th>\n<th>Model Ⅲ (vaild 31800)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Fold 1</td>\n<td>0.8989</td>\n<td>0.8967</td>\n<td>0.8947</td>\n</tr>\n<tr>\n<td>Fold 2</td>\n<td>0.8995</td>\n<td>0.8967</td>\n<td>0.8951</td>\n</tr>\n<tr>\n<td>Fold 3</td>\n<td>0.8985</td>\n<td>0.8959</td>\n<td>0.8949</td>\n</tr>\n<tr>\n<td>Fold Mean</td>\n<td>0.89897</td>\n<td>0.89643</td>\n<td>0.89490</td>\n</tr>\n<tr>\n<td>Model Mean</td>\n<td>0.89677</td>\n<td>-</td>\n<td>-</td>\n</tr>\n</tbody>\n</table>\n<h3>Ensemble</h3>\n<p><a href=\"https://www.kaggle.com/code/oliverwang15/4th-solution-ensemble/notebook\" target=\"_blank\">notebook</a> </p>\n<h2>Multi</h2>\n<p>To be honest, I put most of my efforts on cite part so there is nothing very special here and I will make a brief introduction. </p>\n<h3>Data preprocessing &amp; Feature engineering</h3>\n<h4>inputs:</h4>\n<ol>\n<li>TF-IDF normalization</li>\n<li><code>np.log1p(data * 1e4)</code></li>\n<li>Tsvd -&gt; 512</li>\n</ol>\n<h4>targets:</h4>\n<ol>\n<li>Normalization -&gt; mean = 0, std = 1</li>\n<li>Tsvd -&gt; 1024</li>\n</ol>\n<h3>Models</h3>\n<ul>\n<li><p><code>NN</code>: A personal-designed NN network as mentioned above. The output of the model is 1024 dim and make dot product with <code>tsvd.components_</code>(constant) to get the final prediction than use <code>correl_loss</code> to calculate the loss then back propagate the grads.</p></li>\n<li><p><code>Catboost</code>: The results from online <a href=\"https://www.kaggle.com/code/xiafire/lb-t15-msci-multiome-catboostregressor\" target=\"_blank\">notebook</a></p></li>\n<li><p><code>LGBM</code>: The same as the <code>MultiOutputLGBMRegressor</code> mentioned above. Using <code>MSE</code> to fit the tsvd results of normalized targets.</p></li>\n</ul>\n<h3>Ensemble</h3>\n<p>The same notebook as mentioned above.<br>\n<a href=\"https://www.kaggle.com/code/oliverwang15/4th-solution-ensemble/notebook\" target=\"_blank\">notebook</a></p>",
  "messages": [
    {
      "id": "2031785",
      "postDate": "11/16/2022 08:40:45",
      "content": "<h2>Intro</h2>\n<p>To begin with, thanks to the Kaggle team and Open Problems team for hosting such a wonderful contest. I would also like to share my gratitude to all the competitors, especially those generous competitors who are willing to share their codes and thoughts like <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> <a href=\"https://www.kaggle.com/baosenguo\" target=\"_blank\">@baosenguo</a> <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a> and so on. I couldn't have gone so far without their help.</p>\n<p>You can find full code on Github <a href=\"https://github.com/oliverwang15/4th-Place-Solution-for-Open-Problems-Multimodal-Single-Cell\" target=\"_blank\">here</a></p>\n<h2>Cite</h2>\n<h3>Data preprocessing</h3>\n<p>At first, all of my feature engineering methods are based on the <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/data\" target=\"_blank\">original</a>  data, but my public score raised from 0.812 to 0.813 after I merely change the data source so all the feature engineering processes and based on the <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359355\" target=\"_blank\">raw</a> data. </p>\n<p>My preprocessing method is using <code>np.log1p</code> to change the raw data. I have also tried other preprocessing methods like <code>MAGIC</code> and <code>TF-IDF</code> but they can't improve my CV score.</p>\n<h3>Feature engineering</h3>\n<p>The final inputs of the models consist of mainly six parts. Three of them are dimension reduction parts including <code>Tsvd</code>, <code>UMAP</code>, and <code>Novel’s method</code>. The rest are feature selection parts including <code>name importance</code>, <code>corr importance,</code> and <code>rf importance</code>.</p>\n<ul>\n<li><p><code>Tsvd</code>: <code>TruncatedSVD(n_components=128, random_state=42)</code></p></li>\n<li><p><code>UMAP</code>: <code>UMAP(n_neighbors = 16,n_components=128, random_state=42,verbose = True)</code></p></li>\n<li><p><code>Novel’s method</code>: The original method can be found <a href=\"https://github.com/openproblems-bio/neurips2021_multimodal_topmethods/blob/dc7bd58dacbe804dcc7be047531d795b1b04741e/src/predict_modality/methods/novel/resources/helper_functions.py\" target=\"_blank\">here</a>. At first, I wanted to implement the preprocessing method to replace simple <code>log1p</code> but after I replaced the <code>Tsvd</code> results of <code>log1p</code> by  the <code>Tsvd</code> results of <code>Novel’s method</code> I found that my CV went down. But if I kept both of them, the CV score would increase a little bit. So I kept the <code>Tsvd</code> results of <code>Novel’s method</code>.</p></li>\n<li><p><code>name importance</code>: It 's mainly based on AmbrosM's <a href=\"https://www.kaggle.com/code/ambrosm/msci-eda-which-makes-sense/notebook#Name-matching\" target=\"_blank\">notebook</a>. But I added additional information from <code>mygene</code> while matching. I will release my complete preprocessing code later and specific results can be found there.</p></li>\n<li><p><code>corr importance</code>: As the name suggested, I chose the top 3 features that correlated with the targets. There was overlap and the number of selected features was about 104</p></li>\n<li><p><code>rf importance</code>: Since the feature importances of random forest may apply to NN and other models as well. So I selected 128 top feature importances of the random forest model.</p></li>\n</ul>\n<p>I have also tried other mothed including <code>PCA</code>, <code>KernelPCA</code>, <code>LocallyLinearEmbedding</code>, and <code>SpectralEmbedding</code>.<code>PCA</code> gives little help and it will cause severe overfitting when used with <code>Tsvd</code>. I could' t finish the manifold methods in 24 hours so I gave them up.</p>\n<h3>Models</h3>\n<p>I have implemented the CV strategy like the private test, but it turns out that the strategy like the public test is better. So all of the results are based on <code>GroupKFold</code> on <code>donors</code>. I have done there-layers stacking in the competition. and I have also done the ensemble on the stacking results and the results of independent models. Here are the models I used and I will also release the code later. </p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Stacking</th>\n<th>NN</th>\n<th>NN_online</th>\n<th>CNN</th>\n<th>kernel_rigde</th>\n<th>LGBM</th>\n<th>Catboost</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CV</td>\n<td>0.89677</td>\n<td>0.89596</td>\n<td>0.89580</td>\n<td>0.89530</td>\n<td>0.89326</td>\n<td>0.89270</td>\n<td>0.89100</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><p><code>NN</code>: A personal-designed NN network, trying to do something like the transformers. I used MLP to replace the dot product in the mechanism of attention. This may not be so reasonable and I am also aware of the importance of feature vectors and dot products. But I was so fascinated by attention and I also tried <code>tabnet</code> and <code>rtdl</code> but they didn't work very well. But my method seemed to work even better than simple MLP. <br>\n<a href=\"https://www.kaggle.com/oliverwang15/4th-solution-cite-nn\" target=\"_blank\">Demo notebook</a></p></li>\n<li><p><code>CNN</code>: Inspired by the tmp method <a href=\"https://www.kaggle.com/competitions/lish-moa/discussion/202256\" target=\"_blank\">here</a> and also added multidimensional convolution kernel like the Resnet. </p></li>\n<li><p><code>NN(Online)</code>: This model is mainly based on pourchot's method <a href=\"https://www.kaggle.com/code/pourchot/all-in-one-citeseq-multiome-with-keras\" target=\"_blank\">here</a> and only some tiny change was made.<br>\n<a href=\"https://www.kaggle.com/code/oliverwang15/4th-solution-cite-online-nn\" target=\"_blank\">Demo notebook</a></p></li>\n<li><p><code>Kernel Rigde</code>: This model is inspired by the best solution of last year's competition. I used <a href=\"https://docs.ray.io/en/master/tune/index.html\" target=\"_blank\">Ray Tune</a> to optimize the hypermeters<br>\n<a href=\"https://www.kaggle.com/code/oliverwang15/4th-solution-ray-tune-krr\" target=\"_blank\">Demo notebook with ray tune</a></p></li>\n<li><p><code>Catboost</code>: There are many options for <code>catboost</code> here. Using <code>MultiOutputRegressor</code> or <code>MultiRMSE</code> as <code>objective</code>.But we can't do earlystopping to prevent overfitting in the first method and the result of the second method is not good enough so I made a class <code>MultiOutputCatboostRegressor</code> personally, using <code>MSE</code> to fit the normalized targets.</p></li>\n<li><p><code>LGBM</code>: I also wrote <code>MultiOutputLGBMRegressor</code> and the results seem to be better and the training process was so slow that I had to give it up in the stacking. However, I still trained a independent LGBM model and used it in the final training.  <br>\n<a href=\"https://www.kaggle.com/code/oliverwang15/4th-solution-cite-multioutputlgbm\" target=\"_blank\">Demo notebook</a> </p></li>\n<li><p><code>stacking</code>: I used <code>KNN</code>,<code>CNN</code>,<code>ridge</code>,<code>rf</code>,<code>catboost</code>,<code>NN</code> in the first layer and only <code>CNN</code>,<code>catboost</code>,<code>NN</code> in the second and just a simple <code>MLP</code> in the last layer. To avoid overfitting, I used <code>KFold</code> and oof predictions between layers, and every stacking model are using <code>GroupKFold</code>(so there are 3 stacking models here). It seems to be a little bit to understand so you may refer to the picture. If you still have confusion please feel free to ask me.<br>\n<a href=\"https://www.kaggle.com/code/oliverwang15/4th-solution-cite-stacking-train\" target=\"_blank\">Demo notebook train</a> <br>\n<a href=\"https://www.kaggle.com/code/oliverwang15/4th-solution-cite-stacking-predict\" target=\"_blank\">Demo notebook predict</a> </p></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2641164%2F2dedb86bf1f9498fb9da9232da1e579a%2FStacking%20Training.png?generation=1668608507518581&amp;alt=media\" alt=\"\"></p>\n<table>\n<thead>\n<tr>\n<th>CV Results</th>\n<th>Model Ⅰ (vaild 32606)</th>\n<th>Model Ⅱ (vaild 13176)</th>\n<th>Model Ⅲ (vaild 31800)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Fold 1</td>\n<td>0.8989</td>\n<td>0.8967</td>\n<td>0.8947</td>\n</tr>\n<tr>\n<td>Fold 2</td>\n<td>0.8995</td>\n<td>0.8967</td>\n<td>0.8951</td>\n</tr>\n<tr>\n<td>Fold 3</td>\n<td>0.8985</td>\n<td>0.8959</td>\n<td>0.8949</td>\n</tr>\n<tr>\n<td>Fold Mean</td>\n<td>0.89897</td>\n<td>0.89643</td>\n<td>0.89490</td>\n</tr>\n<tr>\n<td>Model Mean</td>\n<td>0.89677</td>\n<td>-</td>\n<td>-</td>\n</tr>\n</tbody>\n</table>\n<h3>Ensemble</h3>\n<p><a href=\"https://www.kaggle.com/code/oliverwang15/4th-solution-ensemble/notebook\" target=\"_blank\">notebook</a> </p>\n<h2>Multi</h2>\n<p>To be honest, I put most of my efforts on cite part so there is nothing very special here and I will make a brief introduction. </p>\n<h3>Data preprocessing &amp; Feature engineering</h3>\n<h4>inputs:</h4>\n<ol>\n<li>TF-IDF normalization</li>\n<li><code>np.log1p(data * 1e4)</code></li>\n<li>Tsvd -&gt; 512</li>\n</ol>\n<h4>targets:</h4>\n<ol>\n<li>Normalization -&gt; mean = 0, std = 1</li>\n<li>Tsvd -&gt; 1024</li>\n</ol>\n<h3>Models</h3>\n<ul>\n<li><p><code>NN</code>: A personal-designed NN network as mentioned above. The output of the model is 1024 dim and make dot product with <code>tsvd.components_</code>(constant) to get the final prediction than use <code>correl_loss</code> to calculate the loss then back propagate the grads.</p></li>\n<li><p><code>Catboost</code>: The results from online <a href=\"https://www.kaggle.com/code/xiafire/lb-t15-msci-multiome-catboostregressor\" target=\"_blank\">notebook</a></p></li>\n<li><p><code>LGBM</code>: The same as the <code>MultiOutputLGBMRegressor</code> mentioned above. Using <code>MSE</code> to fit the tsvd results of normalized targets.</p></li>\n</ul>\n<h3>Ensemble</h3>\n<p>The same notebook as mentioned above.<br>\n<a href=\"https://www.kaggle.com/code/oliverwang15/4th-solution-ensemble/notebook\" target=\"_blank\">notebook</a></p>",
      "rawMarkdown": "## Intro\n\nTo begin with, thanks to the Kaggle team and Open Problems team for hosting such a wonderful contest. I would also like to share my gratitude to all the competitors, especially those generous competitors who are willing to share their codes and thoughts like @ambrosm @alexandervc @baosenguo @pourchot and so on. I couldn't have gone so far without their help.\n\nYou can find full code on Github [here](https://github.com/oliverwang15/4th-Place-Solution-for-Open-Problems-Multimodal-Single-Cell)\n\n## Cite\n\n### Data preprocessing\n\nAt first, all of my feature engineering methods are based on the [original](https://www.kaggle.com/competitions/open-problems-multimodal/data)  data, but my public score raised from 0.812 to 0.813 after I merely change the data source so all the feature engineering processes and based on the [raw](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359355) data. \n\nMy preprocessing method is using `np.log1p` to change the raw data. I have also tried other preprocessing methods like `MAGIC` and `TF-IDF` but they can't improve my CV score.\n\n### Feature engineering\n\nThe final inputs of the models consist of mainly six parts. Three of them are dimension reduction parts including `Tsvd`, `UMAP`, and `Novel’s method`. The rest are feature selection parts including `name importance`, `corr importance,` and `rf importance`.\n\n- `Tsvd`: `TruncatedSVD(n_components=128, random_state=42)`\n\n- `UMAP`: `UMAP(n_neighbors = 16,n_components=128, random_state=42,verbose = True)`\n\n- `Novel’s method`: The original method can be found [here](https://github.com/openproblems-bio/neurips2021_multimodal_topmethods/blob/dc7bd58dacbe804dcc7be047531d795b1b04741e/src/predict_modality/methods/novel/resources/helper_functions.py). At first, I wanted to implement the preprocessing method to replace simple `log1p` but after I replaced the `Tsvd` results of `log1p` by  the `Tsvd` results of `Novel’s method ` I found that my CV went down. But if I kept both of them, the CV score would increase a little bit. So I kept the `Tsvd` results of `Novel’s method `.\n\n- `name importance`: It 's mainly based on AmbrosM's [notebook](https://www.kaggle.com/code/ambrosm/msci-eda-which-makes-sense/notebook#Name-matching). But I added additional information from `mygene` while matching. I will release my complete preprocessing code later and specific results can be found there.\n\n- `corr importance`: As the name suggested, I chose the top 3 features that correlated with the targets. There was overlap and the number of selected features was about 104\n\n- `rf importance`: Since the feature importances of random forest may apply to NN and other models as well. So I selected 128 top feature importances of the random forest model.\n\nI have also tried other mothed including `PCA`, `KernelPCA`, `LocallyLinearEmbedding`, and `SpectralEmbedding`.`PCA` gives little help and it will cause severe overfitting when used with `Tsvd`. I could' t finish the manifold methods in 24 hours so I gave them up.\n\n### Models\n\nI have implemented the CV strategy like the private test, but it turns out that the strategy like the public test is better. So all of the results are based on `GroupKFold` on `donors`. I have done there-layers stacking in the competition. and I have also done the ensemble on the stacking results and the results of independent models. Here are the models I used and I will also release the code later. \n| Method | Stacking |  NN       |  NN_online  | CNN | kernel_rigde | LGBM    | Catboost |\n| :------: | :------: | :------: | :------: | :------: | :------: | :------: | :------: |\n| CV | 0.89677 | 0.89596 | 0.89580 | 0.89530 | 0.89326 | 0.89270 | 0.89100 |\n\n\n\n\n- `NN`: A personal-designed NN network, trying to do something like the transformers. I used MLP to replace the dot product in the mechanism of attention. This may not be so reasonable and I am also aware of the importance of feature vectors and dot products. But I was so fascinated by attention and I also tried `tabnet` and `rtdl` but they didn't work very well. But my method seemed to work even better than simple MLP. \n  [Demo notebook](https://www.kaggle.com/oliverwang15/4th-solution-cite-nn)\n\n- `CNN`: Inspired by the tmp method [here](https://www.kaggle.com/competitions/lish-moa/discussion/202256) and also added multidimensional convolution kernel like the Resnet. \n\n- `NN(Online)`: This model is mainly based on pourchot's method [here](https://www.kaggle.com/code/pourchot/all-in-one-citeseq-multiome-with-keras) and only some tiny change was made.\n[Demo notebook](https://www.kaggle.com/code/oliverwang15/4th-solution-cite-online-nn)\n\n- `Kernel Rigde `: This model is inspired by the best solution of last year's competition. I used [Ray Tune](https://docs.ray.io/en/master/tune/index.html) to optimize the hypermeters\n  [Demo notebook with ray tune](https://www.kaggle.com/code/oliverwang15/4th-solution-ray-tune-krr)\n\n- `Catboost`: There are many options for `catboost` here. Using `MultiOutputRegressor` or `MultiRMSE` as `objective`.But we can't do earlystopping to prevent overfitting in the first method and the result of the second method is not good enough so I made a class `MultiOutputCatboostRegressor` personally, using `MSE` to fit the normalized targets.\n\n- `LGBM`: I also wrote `MultiOutputLGBMRegressor` and the results seem to be better and the training process was so slow that I had to give it up in the stacking. However, I still trained a independent LGBM model and used it in the final training.  \n [Demo notebook](https://www.kaggle.com/code/oliverwang15/4th-solution-cite-multioutputlgbm) \n\n-  `stacking`: I used `KNN`,`CNN`,`ridge`,`rf`,`catboost`,`NN` in the first layer and only `CNN`,`catboost`,`NN` in the second and just a simple `MLP` in the last layer. To avoid overfitting, I used `KFold` and oof predictions between layers, and every stacking model are using `GroupKFold`(so there are 3 stacking models here). It seems to be a little bit to understand so you may refer to the picture. If you still have confusion please feel free to ask me.\n[Demo notebook train](https://www.kaggle.com/code/oliverwang15/4th-solution-cite-stacking-train) \n[Demo notebook predict](https://www.kaggle.com/code/oliverwang15/4th-solution-cite-stacking-predict) \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2641164%2F2dedb86bf1f9498fb9da9232da1e579a%2FStacking%20Training.png?generation=1668608507518581&alt=media)\n\n| CV Results    | Model Ⅰ (vaild 32606) | Model Ⅱ (vaild 13176) | Model Ⅲ (vaild 31800) |\n| ---------- | --------------------- | --------------------- | --------------------- |\n| Fold 1     | 0.8989                | 0.8967                | 0.8947                |\n| Fold 2     | 0.8995                | 0.8967                | 0.8951                |\n| Fold 3     | 0.8985                | 0.8959                | 0.8949                |\n| Fold Mean  | 0.89897              | 0.89643              | 0.89490                |\n| Model Mean | 0.89677              | -                     | -                     |\n\n### Ensemble\n[notebook](https://www.kaggle.com/code/oliverwang15/4th-solution-ensemble/notebook) \n\n## Multi\nTo be honest, I put most of my efforts on cite part so there is nothing very special here and I will make a brief introduction. \n\n### Data preprocessing & Feature engineering\n#### inputs:\n1. TF-IDF normalization\n2. `np.log1p(data * 1e4)`\n3. Tsvd -> 512\n#### targets: \n1. Normalization -> mean = 0, std = 1\n2. Tsvd -> 1024\n### Models\n- `NN`: A personal-designed NN network as mentioned above. The output of the model is 1024 dim and make dot product with `tsvd.components_`(constant) to get the final prediction than use `correl_loss` to calculate the loss then back propagate the grads.\n\n- `Catboost`: The results from online [notebook](https://www.kaggle.com/code/xiafire/lb-t15-msci-multiome-catboostregressor)\n\n- `LGBM`: The same as the `MultiOutputLGBMRegressor` mentioned above. Using `MSE` to fit the tsvd results of normalized targets.\n\n### Ensemble\nThe same notebook as mentioned above.\n[notebook](https://www.kaggle.com/code/oliverwang15/4th-solution-ensemble/notebook)",
      "votes": null
    },
    {
      "id": "2031857",
      "postDate": "11/16/2022 09:17:37",
      "content": "<p><a href=\"https://www.kaggle.com/oliverwang15\" target=\"_blank\">@oliverwang15</a> , congrats for your work ! 👍</p>",
      "rawMarkdown": "oliverwang15 , congrats for your work ! 👍",
      "votes": null
    },
    {
      "id": "2032250",
      "postDate": "11/16/2022 14:06:32",
      "content": "<p>Thanks and your work is very profound and helpful ! 👍</p>",
      "rawMarkdown": "Thanks and your work is very profound and helpful ! 👍",
      "votes": null
    },
    {
      "id": "2033450",
      "postDate": "11/17/2022 09:17:00",
      "content": "<p>Congratulations for your solo gold, and thank you for sharing your solution.<br>\nI have a question. I also did stacking, but its score is lower than simple average or even single model's score. Did you do something special?</p>",
      "rawMarkdown": "Congratulations for your solo gold, and thank you for sharing your solution.\nI have a question. I also did stacking, but its score is lower than simple average or even single model's score. Did you do something special?",
      "votes": null
    },
    {
      "id": "2033463",
      "postDate": "11/17/2022 09:24:40",
      "content": "<p>Thanks. From my perspective, the reason why stacking is lower is often related to overfitting. The CV score is very high but the LB score is relatively low. So in order to avoid or alleviate overfitting, I used the special KFold strategy and simple MLP as the last layer, which is illustrated in the picture.</p>",
      "rawMarkdown": "Thanks. From my perspective, the reason why stacking is lower is often related to overfitting. The CV score is very high but the LB score is relatively low. So in order to avoid or alleviate overfitting, I used the special KFold strategy and simple MLP as the last layer, which is illustrated in the picture.",
      "votes": null
    },
    {
      "id": "2033471",
      "postDate": "11/17/2022 09:28:50",
      "content": "<p>I understood. Thank you!</p>",
      "rawMarkdown": "I understood. Thank you!",
      "votes": null
    },
    {
      "id": "2033490",
      "postDate": "11/17/2022 09:40:55",
      "content": "<p>Here are the codes of the folds selection process. Hope they may help you get a better understand</p>\n<pre><code> ():  \n    random.seed()\n    random.shuffle(lis)\n    num_fold = ((lis)/folds)\n     [lis[i::folds]  i  (folds)]\n\nmeta_train[] = [i  i  (meta_train.shape[])]\npeople_list = [,,]\nfold_list = []\nnum_fold = \n\n val_people  tqdm([,,]):\n    train_people = [i  i  people_list  i != val_people]\n    train_idx = meta_train[meta_train.donor.isin(train_people)]..to_list()\n    val_idx = meta_train[meta_train.donor == val_people]..to_list()\n    useless_idx = [i  i  meta_train..to_list()  i   train_idx+val_idx]\n    train_fold_1,train_fold_2,train_fold_3 = get_folds(train_idx,num_fold)\n    val_fold_1,val_fold_2,val_fold_3 = get_folds(val_idx,num_fold)\n\n    one_fold = [\n        [[train_fold_1+train_fold_2,val_fold_1+val_fold_2],train_fold_3+val_fold_3],\n        [[train_fold_1+train_fold_3,val_fold_1+val_fold_3],train_fold_2+val_fold_2],\n        [[train_fold_2+train_fold_3,val_fold_2+val_fold_3],train_fold_1+val_fold_1+useless_idx],\n    ]\n    fold_list.append(one_fold)\n</code></pre>",
      "rawMarkdown": "Here are the codes of the folds selection process. Hope they may help you get a better understand\n``` python\ndef get_folds(lis,folds):  # randomly split to n parts\n    random.seed(42)\n    random.shuffle(lis)\n    num_fold = int(len(lis)/folds)\n    return [lis[i::folds] for i in range(folds)]\n\nmeta_train[\"id\"] = [i for i in range(meta_train.shape[0])]\npeople_list = [32606,13176,31800]\nfold_list = []\nnum_fold = 3\n\nfor val_people in tqdm([32606,13176,31800]):\n    train_people = [i for i in people_list if i != val_people]\n    train_idx = meta_train[meta_train.donor.isin(train_people)].id.to_list()\n    val_idx = meta_train[meta_train.donor == val_people].id.to_list()\n    useless_idx = [i for i in meta_train.id.to_list() if i not in train_idx+val_idx]\n    train_fold_1,train_fold_2,train_fold_3 = get_folds(train_idx,num_fold)\n    val_fold_1,val_fold_2,val_fold_3 = get_folds(val_idx,num_fold)\n\n    one_fold = [\n        [[train_fold_1+train_fold_2,val_fold_1+val_fold_2],train_fold_3+val_fold_3],\n        [[train_fold_1+train_fold_3,val_fold_1+val_fold_3],train_fold_2+val_fold_2],\n        [[train_fold_2+train_fold_3,val_fold_2+val_fold_3],train_fold_1+val_fold_1+useless_idx],\n    ]\n    fold_list.append(one_fold)\n```",
      "votes": null
    },
    {
      "id": "2033882",
      "postDate": "11/17/2022 16:11:21",
      "content": "<p>Very clear visualization, and congrats on the solo gold!</p>",
      "rawMarkdown": "Very clear visualization, and congrats on the solo gold!",
      "votes": null
    },
    {
      "id": "2034216",
      "postDate": "11/17/2022 23:22:29",
      "content": "<p>Congratulations for you 4th position and thanks for sharing your solution!</p>\n<p>I understood the following about your fold strategy:</p>\n<ul>\n<li>You first divide the training data into 2 groups: train and validation, based on the donor.</li>\n<li>Then you split each group into 3 folds, e.g., train_fold_1,train_fold_2, train_fold_3, ending up with 6 folds.</li>\n<li>You train L1 with \"train_fold_1 + train_fold_2\" and validate with \"validation_fold_1 + validation_fold_2\".</li>\n<li>You make predictions for \"train_fold_3 + validation_fold_3\" (oof)</li>\n<li>You break the L1 oofs in the same manner as the original data and repeat the process for L2, and then for L3</li>\n</ul>\n<p>Did I get it right?</p>\n<p>Did you have a chance to compare it's PB performance against that of using the validation folds to generate the oof for the next level?</p>",
      "rawMarkdown": "Congratulations for you 4th position and thanks for sharing your solution!\n\nI understood the following about your fold strategy:\n- You first divide the training data into 2 groups: train and validation, based on the donor.\n- Then you split each group into 3 folds, e.g., train_fold_1,train_fold_2, train_fold_3, ending up with 6 folds.\n- You train L1 with \"train_fold_1 + train_fold_2\" and validate with \"validation_fold_1 + validation_fold_2\".\n- You make predictions for \"train_fold_3 + validation_fold_3\" (oof)\n- You break the L1 oofs in the same manner as the original data and repeat the process for L2, and then for L3\n\nDid I get it right?\n\nDid you have a chance to compare it's PB performance against that of using the validation folds to generate the oof for the next level?",
      "votes": null
    },
    {
      "id": "2034217",
      "postDate": "11/17/2022 23:25:34",
      "content": "<p>Warm Congratulations! and thanks for sharing your solution.</p>\n<p>May I ask about the stacking parts, What do you mean about the first layer and second layer? Does it mean that you first select all features to train several models (e.g KNN,CNN,ridge,rf,catboost,NN) and get the prediction results for each of them, after that connate all \"first layer model\" predicted results as input for the \"second layer\" input? </p>",
      "rawMarkdown": "Warm Congratulations! and thanks for sharing your solution.\n\nMay I ask about the stacking parts, What do you mean about the first layer and second layer? Does it mean that you first select all features to train several models (e.g KNN,CNN,ridge,rf,catboost,NN) and get the prediction results for each of them, after that connate all \"first layer model\" predicted results as input for the \"second layer\" input?",
      "votes": null
    },
    {
      "id": "2034376",
      "postDate": "11/18/2022 05:07:56",
      "content": "<p>Thanks! Hope they can help you!</p>",
      "rawMarkdown": "Thanks! Hope they can help you!",
      "votes": null
    },
    {
      "id": "2034386",
      "postDate": "11/18/2022 05:22:36",
      "content": "<p>Thanks. Yes, you are right. But in the second layer input there are not only the output of those models in the first layer but the original features. Actually I was inspired my Mu Li‘s idea <a href=\"https://www.bilibili.com/video/BV1PZ4y197CX/\" target=\"_blank\">here</a>. If you are interested you may have a look</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2641164%2F3b0bff6cf6ec5543e26f3310cdf42f2e%2F_20221118131601.png?generation=1668748770938919&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks. Yes, you are right. But in the second layer input there are not only the output of those models in the first layer but the original features. Actually I was inspired my Mu Li‘s idea [here](https://www.bilibili.com/video/BV1PZ4y197CX/). If you are interested you may have a look\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2641164%2F3b0bff6cf6ec5543e26f3310cdf42f2e%2F_20221118131601.png?generation=1668748770938919&alt=media)",
      "votes": null
    },
    {
      "id": "2034396",
      "postDate": "11/18/2022 05:30:58",
      "content": "<p>Thanks. Yes, you get it right. I'm sorry but I'm not quite sure about the last sentence. Do you mean comparing the PB performance of the final results between using oof and not using oof? Or comparing the PB performance of the oof predictions of the first layer and the final results? Or anything else？</p>",
      "rawMarkdown": "Thanks. Yes, you get it right. I'm sorry but I'm not quite sure about the last sentence. Do you mean comparing the PB performance of the final results between using oof and not using oof? Or comparing the PB performance of the oof predictions of the first layer and the final results? Or anything else？",
      "votes": null
    },
    {
      "id": "2034769",
      "postDate": "11/18/2022 12:19:19",
      "content": "<p>Thanks for the response. Let me illustrate my last question with two examples:</p>\n<ul>\n<li><p>Approach 1) train \"train_fold_1 + train_fold_2\", validate \"validation_fold_1 + validation_fold_2\", oof \"train_fold_3 + validation_fold_3\" (oof)</p></li>\n<li><p>Approach 2) train \"train_fold_1 + train_fold_2 + train_fold_3\", validate \"validation_fold_1 + validation_fold_2 + validation_fold_3\", oof \"validation_fold_1 + validation_fold_2 + validation_fold_3\"</p></li>\n</ul>\n<p>Did you have a chance to compare the CV/LB/PB performance of the two approaches? 1 worked very well and I wonder if and how much better it was in this competition than 2.</p>",
      "rawMarkdown": "Thanks for the response. Let me illustrate my last question with two examples:\n\n- Approach 1) train \"train_fold_1 + train_fold_2\", validate \"validation_fold_1 + validation_fold_2\", oof \"train_fold_3 + validation_fold_3\" (oof)\n\n- Approach 2) train \"train_fold_1 + train_fold_2 + train_fold_3\", validate \"validation_fold_1 + validation_fold_2 + validation_fold_3\", oof \"validation_fold_1 + validation_fold_2 + validation_fold_3\"\n\nDid you have a chance to compare the CV/LB/PB performance of the two approaches? 1 worked very well and I wonder if and how much better it was in this competition than 2.",
      "votes": null
    },
    {
      "id": "2034775",
      "postDate": "11/18/2022 12:30:45",
      "content": "<p>Ok. I see. I will try it. But since I made the stacking parts in a hurry because of the approaching deadline at that time. I will first finish reorganizing the code then make this experiment. If you can't wait to know the answer you may also try it by yourself after I make the stacking parts public.</p>",
      "rawMarkdown": "Ok. I see. I will try it. But since I made the stacking parts in a hurry because of the approaching deadline at that time. I will first finish reorganizing the code then make this experiment. If you can't wait to know the answer you may also try it by yourself after I make the stacking parts public.",
      "votes": null
    },
    {
      "id": "2055443",
      "postDate": "12/05/2022 04:50:51",
      "content": "<p>Thank you for your sharing.<br>\nI have a question. You used some dimension reduction method for cite part.<br>\nHow do you use them differently? Did you simply want diversity for the sake of the ensemble?</p>",
      "rawMarkdown": "Thank you for your sharing.\nI have a question. You used some dimension reduction method for cite part.\nHow do you use them differently? Did you simply want diversity for the sake of the ensemble?",
      "votes": null
    },
    {
      "id": "2055512",
      "postDate": "12/05/2022 06:05:49",
      "content": "<p>Hi, Aesop. Thanks for your question. Indeed, for the sake of diversity, we should use different data to train different models and ensemble them to get the best results. But actually, when I was doing this part since the training part was time-consuming, and trying the best coefficient or doing stacking would also be time-consuming. So I reduced the diversity in data and trained models with all the data I mentioned above, even in the stacking part. So maybe we could try to use different data to train the models and then ensemble them to see whether the results would be better.</p>",
      "rawMarkdown": "Hi, Aesop. Thanks for your question. Indeed, for the sake of diversity, we should use different data to train different models and ensemble them to get the best results. But actually, when I was doing this part since the training part was time-consuming, and trying the best coefficient or doing stacking would also be time-consuming. So I reduced the diversity in data and trained models with all the data I mentioned above, even in the stacking part. So maybe we could try to use different data to train the models and then ensemble them to see whether the results would be better.",
      "votes": null
    },
    {
      "id": "2055613",
      "postDate": "12/05/2022 08:23:50",
      "content": "<p>I understood. Thanks!!</p>",
      "rawMarkdown": "I understood. Thanks!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2031857,
      "author_name": "pourchot",
      "author_url": "",
      "post_date": "11/16/2022 09:17:37",
      "content": "<p><a href=\"https://www.kaggle.com/oliverwang15\" target=\"_blank\">@oliverwang15</a> , congrats for your work ! 👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 2032250,
          "author_name": "oliverwang15",
          "author_url": "",
          "post_date": "11/16/2022 14:06:32",
          "content": "<p>Thanks and your work is very profound and helpful ! 👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2033450,
      "author_name": "ludditep",
      "author_url": "",
      "post_date": "11/17/2022 09:17:00",
      "content": "<p>Congratulations for your solo gold, and thank you for sharing your solution.<br>\nI have a question. I also did stacking, but its score is lower than simple average or even single model's score. Did you do something special?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2033463,
          "author_name": "oliverwang15",
          "author_url": "",
          "post_date": "11/17/2022 09:24:40",
          "content": "<p>Thanks. From my perspective, the reason why stacking is lower is often related to overfitting. The CV score is very high but the LB score is relatively low. So in order to avoid or alleviate overfitting, I used the special KFold strategy and simple MLP as the last layer, which is illustrated in the picture.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2033471,
          "author_name": "ludditep",
          "author_url": "",
          "post_date": "11/17/2022 09:28:50",
          "content": "<p>I understood. Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2033490,
          "author_name": "oliverwang15",
          "author_url": "",
          "post_date": "11/17/2022 09:40:55",
          "content": "<p>Here are the codes of the folds selection process. Hope they may help you get a better understand</p>\n<pre><code> ():  \n    random.seed()\n    random.shuffle(lis)\n    num_fold = ((lis)/folds)\n     [lis[i::folds]  i  (folds)]\n\nmeta_train[] = [i  i  (meta_train.shape[])]\npeople_list = [,,]\nfold_list = []\nnum_fold = \n\n val_people  tqdm([,,]):\n    train_people = [i  i  people_list  i != val_people]\n    train_idx = meta_train[meta_train.donor.isin(train_people)]..to_list()\n    val_idx = meta_train[meta_train.donor == val_people]..to_list()\n    useless_idx = [i  i  meta_train..to_list()  i   train_idx+val_idx]\n    train_fold_1,train_fold_2,train_fold_3 = get_folds(train_idx,num_fold)\n    val_fold_1,val_fold_2,val_fold_3 = get_folds(val_idx,num_fold)\n\n    one_fold = [\n        [[train_fold_1+train_fold_2,val_fold_1+val_fold_2],train_fold_3+val_fold_3],\n        [[train_fold_1+train_fold_3,val_fold_1+val_fold_3],train_fold_2+val_fold_2],\n        [[train_fold_2+train_fold_3,val_fold_2+val_fold_3],train_fold_1+val_fold_1+useless_idx],\n    ]\n    fold_list.append(one_fold)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2034216,
          "author_name": "vialactea",
          "author_url": "",
          "post_date": "11/17/2022 23:22:29",
          "content": "<p>Congratulations for you 4th position and thanks for sharing your solution!</p>\n<p>I understood the following about your fold strategy:</p>\n<ul>\n<li>You first divide the training data into 2 groups: train and validation, based on the donor.</li>\n<li>Then you split each group into 3 folds, e.g., train_fold_1,train_fold_2, train_fold_3, ending up with 6 folds.</li>\n<li>You train L1 with \"train_fold_1 + train_fold_2\" and validate with \"validation_fold_1 + validation_fold_2\".</li>\n<li>You make predictions for \"train_fold_3 + validation_fold_3\" (oof)</li>\n<li>You break the L1 oofs in the same manner as the original data and repeat the process for L2, and then for L3</li>\n</ul>\n<p>Did I get it right?</p>\n<p>Did you have a chance to compare it's PB performance against that of using the validation folds to generate the oof for the next level?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2034396,
          "author_name": "oliverwang15",
          "author_url": "",
          "post_date": "11/18/2022 05:30:58",
          "content": "<p>Thanks. Yes, you get it right. I'm sorry but I'm not quite sure about the last sentence. Do you mean comparing the PB performance of the final results between using oof and not using oof? Or comparing the PB performance of the oof predictions of the first layer and the final results? Or anything else？</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2034769,
          "author_name": "vialactea",
          "author_url": "",
          "post_date": "11/18/2022 12:19:19",
          "content": "<p>Thanks for the response. Let me illustrate my last question with two examples:</p>\n<ul>\n<li><p>Approach 1) train \"train_fold_1 + train_fold_2\", validate \"validation_fold_1 + validation_fold_2\", oof \"train_fold_3 + validation_fold_3\" (oof)</p></li>\n<li><p>Approach 2) train \"train_fold_1 + train_fold_2 + train_fold_3\", validate \"validation_fold_1 + validation_fold_2 + validation_fold_3\", oof \"validation_fold_1 + validation_fold_2 + validation_fold_3\"</p></li>\n</ul>\n<p>Did you have a chance to compare the CV/LB/PB performance of the two approaches? 1 worked very well and I wonder if and how much better it was in this competition than 2.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2034775,
          "author_name": "oliverwang15",
          "author_url": "",
          "post_date": "11/18/2022 12:30:45",
          "content": "<p>Ok. I see. I will try it. But since I made the stacking parts in a hurry because of the approaching deadline at that time. I will first finish reorganizing the code then make this experiment. If you can't wait to know the answer you may also try it by yourself after I make the stacking parts public.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2033882,
      "author_name": "jcerpentier",
      "author_url": "",
      "post_date": "11/17/2022 16:11:21",
      "content": "<p>Very clear visualization, and congrats on the solo gold!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2034376,
          "author_name": "oliverwang15",
          "author_url": "",
          "post_date": "11/18/2022 05:07:56",
          "content": "<p>Thanks! Hope they can help you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2034217,
      "author_name": "bwhale",
      "author_url": "",
      "post_date": "11/17/2022 23:25:34",
      "content": "<p>Warm Congratulations! and thanks for sharing your solution.</p>\n<p>May I ask about the stacking parts, What do you mean about the first layer and second layer? Does it mean that you first select all features to train several models (e.g KNN,CNN,ridge,rf,catboost,NN) and get the prediction results for each of them, after that connate all \"first layer model\" predicted results as input for the \"second layer\" input? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2034386,
          "author_name": "oliverwang15",
          "author_url": "",
          "post_date": "11/18/2022 05:22:36",
          "content": "<p>Thanks. Yes, you are right. But in the second layer input there are not only the output of those models in the first layer but the original features. Actually I was inspired my Mu Li‘s idea <a href=\"https://www.bilibili.com/video/BV1PZ4y197CX/\" target=\"_blank\">here</a>. If you are interested you may have a look</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2641164%2F3b0bff6cf6ec5543e26f3310cdf42f2e%2F_20221118131601.png?generation=1668748770938919&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2055443,
      "author_name": "aesoptacit",
      "author_url": "",
      "post_date": "12/05/2022 04:50:51",
      "content": "<p>Thank you for your sharing.<br>\nI have a question. You used some dimension reduction method for cite part.<br>\nHow do you use them differently? Did you simply want diversity for the sake of the ensemble?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2055512,
          "author_name": "oliverwang15",
          "author_url": "",
          "post_date": "12/05/2022 06:05:49",
          "content": "<p>Hi, Aesop. Thanks for your question. Indeed, for the sake of diversity, we should use different data to train different models and ensemble them to get the best results. But actually, when I was doing this part since the training part was time-consuming, and trying the best coefficient or doing stacking would also be time-consuming. So I reduced the diversity in data and trained models with all the data I mentioned above, even in the stacking part. So maybe we could try to use different data to train the models and then ensemble them to see whether the results would be better.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2055613,
          "author_name": "aesoptacit",
          "author_url": "",
          "post_date": "12/05/2022 08:23:50",
          "content": "<p>I understood. Thanks!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2031785": "## Intro\n\nTo begin with, thanks to the Kaggle team and Open Problems team for hosting such a wonderful contest. I would also like to share my gratitude to all the competitors, especially those generous competitors who are willing to share their codes and thoughts like @ambrosm @alexandervc @baosenguo @pourchot and so on. I couldn't have gone so far without their help.\n\nYou can find full code on Github [here](https://github.com/oliverwang15/4th-Place-Solution-for-Open-Problems-Multimodal-Single-Cell)\n\n## Cite\n\n### Data preprocessing\n\nAt first, all of my feature engineering methods are based on the [original](https://www.kaggle.com/competitions/open-problems-multimodal/data)  data, but my public score raised from 0.812 to 0.813 after I merely change the data source so all the feature engineering processes and based on the [raw](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359355) data. \n\nMy preprocessing method is using `np.log1p` to change the raw data. I have also tried other preprocessing methods like `MAGIC` and `TF-IDF` but they can't improve my CV score.\n\n### Feature engineering\n\nThe final inputs of the models consist of mainly six parts. Three of them are dimension reduction parts including `Tsvd`, `UMAP`, and `Novel’s method`. The rest are feature selection parts including `name importance`, `corr importance,` and `rf importance`.\n\n- `Tsvd`: `TruncatedSVD(n_components=128, random_state=42)`\n\n- `UMAP`: `UMAP(n_neighbors = 16,n_components=128, random_state=42,verbose = True)`\n\n- `Novel’s method`: The original method can be found [here](https://github.com/openproblems-bio/neurips2021_multimodal_topmethods/blob/dc7bd58dacbe804dcc7be047531d795b1b04741e/src/predict_modality/methods/novel/resources/helper_functions.py). At first, I wanted to implement the preprocessing method to replace simple `log1p` but after I replaced the `Tsvd` results of `log1p` by  the `Tsvd` results of `Novel’s method ` I found that my CV went down. But if I kept both of them, the CV score would increase a little bit. So I kept the `Tsvd` results of `Novel’s method `.\n\n- `name importance`: It 's mainly based on AmbrosM's [notebook](https://www.kaggle.com/code/ambrosm/msci-eda-which-makes-sense/notebook#Name-matching). But I added additional information from `mygene` while matching. I will release my complete preprocessing code later and specific results can be found there.\n\n- `corr importance`: As the name suggested, I chose the top 3 features that correlated with the targets. There was overlap and the number of selected features was about 104\n\n- `rf importance`: Since the feature importances of random forest may apply to NN and other models as well. So I selected 128 top feature importances of the random forest model.\n\nI have also tried other mothed including `PCA`, `KernelPCA`, `LocallyLinearEmbedding`, and `SpectralEmbedding`.`PCA` gives little help and it will cause severe overfitting when used with `Tsvd`. I could' t finish the manifold methods in 24 hours so I gave them up.\n\n### Models\n\nI have implemented the CV strategy like the private test, but it turns out that the strategy like the public test is better. So all of the results are based on `GroupKFold` on `donors`. I have done there-layers stacking in the competition. and I have also done the ensemble on the stacking results and the results of independent models. Here are the models I used and I will also release the code later. \n| Method | Stacking |  NN       |  NN_online  | CNN | kernel_rigde | LGBM    | Catboost |\n| :------: | :------: | :------: | :------: | :------: | :------: | :------: | :------: |\n| CV | 0.89677 | 0.89596 | 0.89580 | 0.89530 | 0.89326 | 0.89270 | 0.89100 |\n\n\n\n\n- `NN`: A personal-designed NN network, trying to do something like the transformers. I used MLP to replace the dot product in the mechanism of attention. This may not be so reasonable and I am also aware of the importance of feature vectors and dot products. But I was so fascinated by attention and I also tried `tabnet` and `rtdl` but they didn't work very well. But my method seemed to work even better than simple MLP. \n  [Demo notebook](https://www.kaggle.com/oliverwang15/4th-solution-cite-nn)\n\n- `CNN`: Inspired by the tmp method [here](https://www.kaggle.com/competitions/lish-moa/discussion/202256) and also added multidimensional convolution kernel like the Resnet. \n\n- `NN(Online)`: This model is mainly based on pourchot's method [here](https://www.kaggle.com/code/pourchot/all-in-one-citeseq-multiome-with-keras) and only some tiny change was made.\n[Demo notebook](https://www.kaggle.com/code/oliverwang15/4th-solution-cite-online-nn)\n\n- `Kernel Rigde `: This model is inspired by the best solution of last year's competition. I used [Ray Tune](https://docs.ray.io/en/master/tune/index.html) to optimize the hypermeters\n  [Demo notebook with ray tune](https://www.kaggle.com/code/oliverwang15/4th-solution-ray-tune-krr)\n\n- `Catboost`: There are many options for `catboost` here. Using `MultiOutputRegressor` or `MultiRMSE` as `objective`.But we can't do earlystopping to prevent overfitting in the first method and the result of the second method is not good enough so I made a class `MultiOutputCatboostRegressor` personally, using `MSE` to fit the normalized targets.\n\n- `LGBM`: I also wrote `MultiOutputLGBMRegressor` and the results seem to be better and the training process was so slow that I had to give it up in the stacking. However, I still trained a independent LGBM model and used it in the final training.  \n [Demo notebook](https://www.kaggle.com/code/oliverwang15/4th-solution-cite-multioutputlgbm) \n\n-  `stacking`: I used `KNN`,`CNN`,`ridge`,`rf`,`catboost`,`NN` in the first layer and only `CNN`,`catboost`,`NN` in the second and just a simple `MLP` in the last layer. To avoid overfitting, I used `KFold` and oof predictions between layers, and every stacking model are using `GroupKFold`(so there are 3 stacking models here). It seems to be a little bit to understand so you may refer to the picture. If you still have confusion please feel free to ask me.\n[Demo notebook train](https://www.kaggle.com/code/oliverwang15/4th-solution-cite-stacking-train) \n[Demo notebook predict](https://www.kaggle.com/code/oliverwang15/4th-solution-cite-stacking-predict) \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2641164%2F2dedb86bf1f9498fb9da9232da1e579a%2FStacking%20Training.png?generation=1668608507518581&alt=media)\n\n| CV Results    | Model Ⅰ (vaild 32606) | Model Ⅱ (vaild 13176) | Model Ⅲ (vaild 31800) |\n| ---------- | --------------------- | --------------------- | --------------------- |\n| Fold 1     | 0.8989                | 0.8967                | 0.8947                |\n| Fold 2     | 0.8995                | 0.8967                | 0.8951                |\n| Fold 3     | 0.8985                | 0.8959                | 0.8949                |\n| Fold Mean  | 0.89897              | 0.89643              | 0.89490                |\n| Model Mean | 0.89677              | -                     | -                     |\n\n### Ensemble\n[notebook](https://www.kaggle.com/code/oliverwang15/4th-solution-ensemble/notebook) \n\n## Multi\nTo be honest, I put most of my efforts on cite part so there is nothing very special here and I will make a brief introduction. \n\n### Data preprocessing & Feature engineering\n#### inputs:\n1. TF-IDF normalization\n2. `np.log1p(data * 1e4)`\n3. Tsvd -> 512\n#### targets: \n1. Normalization -> mean = 0, std = 1\n2. Tsvd -> 1024\n### Models\n- `NN`: A personal-designed NN network as mentioned above. The output of the model is 1024 dim and make dot product with `tsvd.components_`(constant) to get the final prediction than use `correl_loss` to calculate the loss then back propagate the grads.\n\n- `Catboost`: The results from online [notebook](https://www.kaggle.com/code/xiafire/lb-t15-msci-multiome-catboostregressor)\n\n- `LGBM`: The same as the `MultiOutputLGBMRegressor` mentioned above. Using `MSE` to fit the tsvd results of normalized targets.\n\n### Ensemble\nThe same notebook as mentioned above.\n[notebook](https://www.kaggle.com/code/oliverwang15/4th-solution-ensemble/notebook)",
    "2031857": "oliverwang15 , congrats for your work ! 👍",
    "2032250": "Thanks and your work is very profound and helpful ! 👍",
    "2033450": "Congratulations for your solo gold, and thank you for sharing your solution.\nI have a question. I also did stacking, but its score is lower than simple average or even single model's score. Did you do something special?",
    "2033463": "Thanks. From my perspective, the reason why stacking is lower is often related to overfitting. The CV score is very high but the LB score is relatively low. So in order to avoid or alleviate overfitting, I used the special KFold strategy and simple MLP as the last layer, which is illustrated in the picture.",
    "2033471": "I understood. Thank you!",
    "2033490": "Here are the codes of the folds selection process. Hope they may help you get a better understand\n``` python\ndef get_folds(lis,folds):  # randomly split to n parts\n    random.seed(42)\n    random.shuffle(lis)\n    num_fold = int(len(lis)/folds)\n    return [lis[i::folds] for i in range(folds)]\n\nmeta_train[\"id\"] = [i for i in range(meta_train.shape[0])]\npeople_list = [32606,13176,31800]\nfold_list = []\nnum_fold = 3\n\nfor val_people in tqdm([32606,13176,31800]):\n    train_people = [i for i in people_list if i != val_people]\n    train_idx = meta_train[meta_train.donor.isin(train_people)].id.to_list()\n    val_idx = meta_train[meta_train.donor == val_people].id.to_list()\n    useless_idx = [i for i in meta_train.id.to_list() if i not in train_idx+val_idx]\n    train_fold_1,train_fold_2,train_fold_3 = get_folds(train_idx,num_fold)\n    val_fold_1,val_fold_2,val_fold_3 = get_folds(val_idx,num_fold)\n\n    one_fold = [\n        [[train_fold_1+train_fold_2,val_fold_1+val_fold_2],train_fold_3+val_fold_3],\n        [[train_fold_1+train_fold_3,val_fold_1+val_fold_3],train_fold_2+val_fold_2],\n        [[train_fold_2+train_fold_3,val_fold_2+val_fold_3],train_fold_1+val_fold_1+useless_idx],\n    ]\n    fold_list.append(one_fold)\n```",
    "2033882": "Very clear visualization, and congrats on the solo gold!",
    "2034216": "Congratulations for you 4th position and thanks for sharing your solution!\n\nI understood the following about your fold strategy:\n- You first divide the training data into 2 groups: train and validation, based on the donor.\n- Then you split each group into 3 folds, e.g., train_fold_1,train_fold_2, train_fold_3, ending up with 6 folds.\n- You train L1 with \"train_fold_1 + train_fold_2\" and validate with \"validation_fold_1 + validation_fold_2\".\n- You make predictions for \"train_fold_3 + validation_fold_3\" (oof)\n- You break the L1 oofs in the same manner as the original data and repeat the process for L2, and then for L3\n\nDid I get it right?\n\nDid you have a chance to compare it's PB performance against that of using the validation folds to generate the oof for the next level?",
    "2034217": "Warm Congratulations! and thanks for sharing your solution.\n\nMay I ask about the stacking parts, What do you mean about the first layer and second layer? Does it mean that you first select all features to train several models (e.g KNN,CNN,ridge,rf,catboost,NN) and get the prediction results for each of them, after that connate all \"first layer model\" predicted results as input for the \"second layer\" input?",
    "2034376": "Thanks! Hope they can help you!",
    "2034386": "Thanks. Yes, you are right. But in the second layer input there are not only the output of those models in the first layer but the original features. Actually I was inspired my Mu Li‘s idea [here](https://www.bilibili.com/video/BV1PZ4y197CX/). If you are interested you may have a look\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2641164%2F3b0bff6cf6ec5543e26f3310cdf42f2e%2F_20221118131601.png?generation=1668748770938919&alt=media)",
    "2034396": "Thanks. Yes, you get it right. I'm sorry but I'm not quite sure about the last sentence. Do you mean comparing the PB performance of the final results between using oof and not using oof? Or comparing the PB performance of the oof predictions of the first layer and the final results? Or anything else？",
    "2034769": "Thanks for the response. Let me illustrate my last question with two examples:\n\n- Approach 1) train \"train_fold_1 + train_fold_2\", validate \"validation_fold_1 + validation_fold_2\", oof \"train_fold_3 + validation_fold_3\" (oof)\n\n- Approach 2) train \"train_fold_1 + train_fold_2 + train_fold_3\", validate \"validation_fold_1 + validation_fold_2 + validation_fold_3\", oof \"validation_fold_1 + validation_fold_2 + validation_fold_3\"\n\nDid you have a chance to compare the CV/LB/PB performance of the two approaches? 1 worked very well and I wonder if and how much better it was in this competition than 2.",
    "2034775": "Ok. I see. I will try it. But since I made the stacking parts in a hurry because of the approaching deadline at that time. I will first finish reorganizing the code then make this experiment. If you can't wait to know the answer you may also try it by yourself after I make the stacking parts public.",
    "2055443": "Thank you for your sharing.\nI have a question. You used some dimension reduction method for cite part.\nHow do you use them differently? Did you simply want diversity for the sake of the ensemble?",
    "2055512": "Hi, Aesop. Thanks for your question. Indeed, for the sake of diversity, we should use different data to train different models and ensemble them to get the best results. But actually, when I was doing this part since the training part was time-consuming, and trying the best coefficient or doing stacking would also be time-consuming. So I reduced the diversity in data and trained models with all the data I mentioned above, even in the stacking part. So maybe we could try to use different data to train the models and then ensemble them to see whether the results would be better.",
    "2055613": "I understood. Thanks!!"
  },
  "source": "meta"
}