{
  "id": 458750,
  "title": "3rd Place Solution for the Open Problems – Single-Cell Perturbations",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/458750",
  "author_name": "Rafał Pawłowski",
  "post_date": "2023-12-01T10:30:52.212000",
  "votes": 33,
  "comment_count": 10,
  "views": 0,
  "content": "<h1>1. Integration of Biological Knowledge</h1>\n<p>Generally, I treated this problem as a regression with 2 feature columns and 18211 targets, but I had tried to utilize SMILES sequences in neural network with LSTM unit. Both the sm_name and SMILES columns can be encoded exactly in the same way, so the sm_name column can be replaced with the SMILES column. Moreover, SMILES column is more informative, because every single character of the sequence can be encoded (not only a single value like in sm_name)  and the order of these characters provides extra information. Theoretically, In the worst case, the performance of the neural network using SMILES column instead of sn_name should be not worse than using original columns. Unfortunate, I have reached an unsatisfactory public score with this neural network and I stopped further research. </p>\n<h1>2. Exploration of the problem</h1>\n<p>For simplicity, the analysis is done for the single best model without pseudolabeling. Since mrrmse metric is sensitive to outliers, a distribution of ranges for columns should be checked.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F523b491c0054eae38b031b945b7eb168%2FScreenshot%20from%202023-12-05%2009-36-00.png?generation=1701849069283853&amp;alt=media\" alt=\"\"><br>\nThe majority of columns have a range of values in an interval (4, 50). Naturally, the columns with high range lead to high mae or mse, so in order to determine genes which are easy and hard to predict, the standardized (divided by std) colwise mse is applied. The table below shows the hardest and easiest genes for prediction.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2Fd8bf3ee03e6d6154593170a5db43cab3%2FScreenshot%20from%202023-12-05%2009-45-05.png?generation=1701849594906441&amp;alt=media\" alt=\"\"><br>\nThe scheme of the applied cross validation.  Every fold contains one cell type chosen from NK cells, T cells CD4+, T cells CD8+, T regulatory cells and only sm_names being in public and private test was involved. The lowest value of this validation split corresponds to the lowest value on public and private dataset. In my opinion this is a reliable split and the perfect split depends on the model architecture, so every model can have its perfect validation split. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2Fb4a580c8349343ab853a40cf6ee8b43e%2FScreenshot%20from%202023-12-05%2009-49-01.png?generation=1701849988973830&amp;alt=media\" alt=\"\"><br>\nThe easiness of learning per cell types is shown below. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F64d7f90af54659c0e8477d4faa3a3c51%2FScreenshot%20from%202023-12-05%2009-56-05.png?generation=1701850210915961&amp;alt=media\" alt=\"\"><br>\nThe values of loss are different, because they are calculated in truncated space. T cells CD4+ and NK cells are learning well. T regulatory cells are harder for training. The T cells CD8 are weakly improving on validation dataset. This split uses about 25% of the dataset for validation, so I believe more reliable splits exist. My new proposition of split is: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F718b580e65a5654a92880f8a6bcf6d1b%2FScreenshot%20from%202023-12-05%2010-08-16.png?generation=1701850452329548&amp;alt=media\" alt=\"\"><br>\nThis split is similar to the previous one, but for each validation fold, randomly selected examples are moved to training. Let's check an impact of the new split for training. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F8a68fc6a4b3025b48d37dabc67a50c05%2FScreenshot%20from%202023-12-05%2010-12-19.png?generation=1701850788086187&amp;alt=media\" alt=\"\"><br>\nFor each cell type, more data improved performance on this same, constant and small validation set. Metric is calculated on full dimension this time. The first value at 480 training examples corresponds to the previous split. Theoretically, decreasing the size of validation set to one example can lead to the best performance. Since it will be very similar to training on whole dataset, what is done finally.</p>\n<h1>3. Model design</h1>\n<h2>Solution</h2>\n<p>The prediction system is two staged, so I publish two versions of the notebook. <br>\nThe first stage predicts pseudolabels. To be honest, if I stopped on this version, I would not be the third. <br>\nThe predicted pseudolabels on all test data (255 rows) are added to training in the second stage.</p>\n<h3>Stage 1 preparing pseudolabels</h3>\n<p>The main part of this system is a neural network. Every neural network and its environment was optimized by optuna. Hyperparameters that have been optimized:<br>\na dropout value, a number of neurons in particular layers, an output dimension of an embedding layer, a number of epochs, a learning rate, a batch size, a number of dimension of truncated singular value decomposition.<br>\nThe optimization was done on custom 4-folds cross validation.  In order to avoid overfitting to cross validation by optuna I applied 2 repeats for every fold and took an average. Generally, the more, the better. The optuna's criterion was MRRMSE. <br>\nFinally, 7 models were ensembled. Optuna was applied again to determine best weights of linear combination. The prediction of test set is the pseudolabels now and will be used in second stage.</p>\n<h3>Stage 2 retraining with pseudolabels</h3>\n<p>The pseudolabels (255 rows) were added to the training dataset. I applied 20 models with optimized parameters in different experiments for a model diversity.<br>\nOptuna selected optimal weights for the linear combination of the prediction again.<br>\nModels had high variance, so every model was trained 10 times on all dataset and the median of prediction is taken as a final prediction.  The prediction was additionally clipped to colwise min and max. </p>\n<p><strong>History of improvements:</strong></p>\n<ol>\n<li>a replacing onehot encoding with an embedding layer</li>\n<li>a replacing MAE loss with MRRMSE loss</li>\n<li>an ensembing of models with mean</li>\n<li>a dimension reduction with truncated singular value decomposition</li>\n<li>an ensembling of models with weighted mean</li>\n<li>using pseudolabeling</li>\n<li>using pseudolabeling and ensembling of 20 models and weighted mean. </li>\n</ol>\n<p><strong>What did not work for me</strong>:</p>\n<ul>\n<li>a label normalization, standardization</li>\n<li>a chained regression</li>\n<li>a denoising dataset</li>\n<li>a removal of outliers</li>\n<li>an adding noise to labels</li>\n<li>a training on selected easy / hard to predict columns</li>\n<li>a huber loss.</li>\n</ul>\n<h1>4. Robustness</h1>\n<p>I have tested 3 types of the robustness: increasing dataset size, adding noise to labels and inputs. Adding the noise to inputs failed totally, it is logical for me, because the nominal and categorical values are changed and are behaving like the continues values, what is not beneficial.<br>\nLet's see the performance on 40%, 50%,…, 100% of training dataset. It started from 40%, because singular values decomposition is limited by number of examples.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2Fb684812f2feb70c2b9f28f2df5714cd7%2FScreenshot%20from%202023-12-05%2013-11-10.png?generation=1701852018611279&amp;alt=media\" alt=\"\"><br>\nThe experiment was 5 times repeated, so the interval of uncertainty is visible. More data improves the performance significantly. <br>\nThe last test of robustness is adding noise to the labels. The random gaussian (a distribution with 0 mean and scale * std) noise was added to the labels.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F8625a5cb1af874f4986277c98e76ab8e%2FScreenshot%20from%202023-12-05%2011-45-22.png?generation=1701852528532795&amp;alt=media\" alt=\"\"><br>\nAdding some noise (0.01 * std) can even improve the model's performance.  Generally, the model is robust to the noise. </p>\n<h1>5. Documentation &amp; code style</h1>\n<p>The code on GitHub is documented. </p>\n<h1>6. Reproducibility</h1>\n<p>GitHub code:<br>\n<a href=\"https://github.com/okon2000/single_cell_perturbations\" target=\"_blank\">repo</a><br>\nNotebook. The version 264 is first a stage and 266 the second one:<br>\n<a href=\"https://www.kaggle.com/code/jankowalski2000/3rd-place-solution\" target=\"_blank\">notebook</a>.</p>\n<p>The code runs in approximately 1 hour using CPU Intel(R) Core(TM) i5-9300H CPU @ 2.40GHz and 8GB RAM.</p>",
  "messages": [
    {
      "id": 2545160,
      "postDate": "2023-12-01T10:30:52.213Z",
      "content": "<h1>1. Integration of Biological Knowledge</h1>\n<p>Generally, I treated this problem as a regression with 2 feature columns and 18211 targets, but I had tried to utilize SMILES sequences in neural network with LSTM unit. Both the sm_name and SMILES columns can be encoded exactly in the same way, so the sm_name column can be replaced with the SMILES column. Moreover, SMILES column is more informative, because every single character of the sequence can be encoded (not only a single value like in sm_name)  and the order of these characters provides extra information. Theoretically, In the worst case, the performance of the neural network using SMILES column instead of sn_name should be not worse than using original columns. Unfortunate, I have reached an unsatisfactory public score with this neural network and I stopped further research. </p>\n<h1>2. Exploration of the problem</h1>\n<p>For simplicity, the analysis is done for the single best model without pseudolabeling. Since mrrmse metric is sensitive to outliers, a distribution of ranges for columns should be checked.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F523b491c0054eae38b031b945b7eb168%2FScreenshot%20from%202023-12-05%2009-36-00.png?generation=1701849069283853&amp;alt=media\" alt=\"\"><br>\nThe majority of columns have a range of values in an interval (4, 50). Naturally, the columns with high range lead to high mae or mse, so in order to determine genes which are easy and hard to predict, the standardized (divided by std) colwise mse is applied. The table below shows the hardest and easiest genes for prediction.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2Fd8bf3ee03e6d6154593170a5db43cab3%2FScreenshot%20from%202023-12-05%2009-45-05.png?generation=1701849594906441&amp;alt=media\" alt=\"\"><br>\nThe scheme of the applied cross validation.  Every fold contains one cell type chosen from NK cells, T cells CD4+, T cells CD8+, T regulatory cells and only sm_names being in public and private test was involved. The lowest value of this validation split corresponds to the lowest value on public and private dataset. In my opinion this is a reliable split and the perfect split depends on the model architecture, so every model can have its perfect validation split. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2Fb4a580c8349343ab853a40cf6ee8b43e%2FScreenshot%20from%202023-12-05%2009-49-01.png?generation=1701849988973830&amp;alt=media\" alt=\"\"><br>\nThe easiness of learning per cell types is shown below. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F64d7f90af54659c0e8477d4faa3a3c51%2FScreenshot%20from%202023-12-05%2009-56-05.png?generation=1701850210915961&amp;alt=media\" alt=\"\"><br>\nThe values of loss are different, because they are calculated in truncated space. T cells CD4+ and NK cells are learning well. T regulatory cells are harder for training. The T cells CD8 are weakly improving on validation dataset. This split uses about 25% of the dataset for validation, so I believe more reliable splits exist. My new proposition of split is: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F718b580e65a5654a92880f8a6bcf6d1b%2FScreenshot%20from%202023-12-05%2010-08-16.png?generation=1701850452329548&amp;alt=media\" alt=\"\"><br>\nThis split is similar to the previous one, but for each validation fold, randomly selected examples are moved to training. Let's check an impact of the new split for training. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F8a68fc6a4b3025b48d37dabc67a50c05%2FScreenshot%20from%202023-12-05%2010-12-19.png?generation=1701850788086187&amp;alt=media\" alt=\"\"><br>\nFor each cell type, more data improved performance on this same, constant and small validation set. Metric is calculated on full dimension this time. The first value at 480 training examples corresponds to the previous split. Theoretically, decreasing the size of validation set to one example can lead to the best performance. Since it will be very similar to training on whole dataset, what is done finally.</p>\n<h1>3. Model design</h1>\n<h2>Solution</h2>\n<p>The prediction system is two staged, so I publish two versions of the notebook. <br>\nThe first stage predicts pseudolabels. To be honest, if I stopped on this version, I would not be the third. <br>\nThe predicted pseudolabels on all test data (255 rows) are added to training in the second stage.</p>\n<h3>Stage 1 preparing pseudolabels</h3>\n<p>The main part of this system is a neural network. Every neural network and its environment was optimized by optuna. Hyperparameters that have been optimized:<br>\na dropout value, a number of neurons in particular layers, an output dimension of an embedding layer, a number of epochs, a learning rate, a batch size, a number of dimension of truncated singular value decomposition.<br>\nThe optimization was done on custom 4-folds cross validation.  In order to avoid overfitting to cross validation by optuna I applied 2 repeats for every fold and took an average. Generally, the more, the better. The optuna's criterion was MRRMSE. <br>\nFinally, 7 models were ensembled. Optuna was applied again to determine best weights of linear combination. The prediction of test set is the pseudolabels now and will be used in second stage.</p>\n<h3>Stage 2 retraining with pseudolabels</h3>\n<p>The pseudolabels (255 rows) were added to the training dataset. I applied 20 models with optimized parameters in different experiments for a model diversity.<br>\nOptuna selected optimal weights for the linear combination of the prediction again.<br>\nModels had high variance, so every model was trained 10 times on all dataset and the median of prediction is taken as a final prediction.  The prediction was additionally clipped to colwise min and max. </p>\n<p><strong>History of improvements:</strong></p>\n<ol>\n<li>a replacing onehot encoding with an embedding layer</li>\n<li>a replacing MAE loss with MRRMSE loss</li>\n<li>an ensembing of models with mean</li>\n<li>a dimension reduction with truncated singular value decomposition</li>\n<li>an ensembling of models with weighted mean</li>\n<li>using pseudolabeling</li>\n<li>using pseudolabeling and ensembling of 20 models and weighted mean. </li>\n</ol>\n<p><strong>What did not work for me</strong>:</p>\n<ul>\n<li>a label normalization, standardization</li>\n<li>a chained regression</li>\n<li>a denoising dataset</li>\n<li>a removal of outliers</li>\n<li>an adding noise to labels</li>\n<li>a training on selected easy / hard to predict columns</li>\n<li>a huber loss.</li>\n</ul>\n<h1>4. Robustness</h1>\n<p>I have tested 3 types of the robustness: increasing dataset size, adding noise to labels and inputs. Adding the noise to inputs failed totally, it is logical for me, because the nominal and categorical values are changed and are behaving like the continues values, what is not beneficial.<br>\nLet's see the performance on 40%, 50%,…, 100% of training dataset. It started from 40%, because singular values decomposition is limited by number of examples.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2Fb684812f2feb70c2b9f28f2df5714cd7%2FScreenshot%20from%202023-12-05%2013-11-10.png?generation=1701852018611279&amp;alt=media\" alt=\"\"><br>\nThe experiment was 5 times repeated, so the interval of uncertainty is visible. More data improves the performance significantly. <br>\nThe last test of robustness is adding noise to the labels. The random gaussian (a distribution with 0 mean and scale * std) noise was added to the labels.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F8625a5cb1af874f4986277c98e76ab8e%2FScreenshot%20from%202023-12-05%2011-45-22.png?generation=1701852528532795&amp;alt=media\" alt=\"\"><br>\nAdding some noise (0.01 * std) can even improve the model's performance.  Generally, the model is robust to the noise. </p>\n<h1>5. Documentation &amp; code style</h1>\n<p>The code on GitHub is documented. </p>\n<h1>6. Reproducibility</h1>\n<p>GitHub code:<br>\n<a href=\"https://github.com/okon2000/single_cell_perturbations\" target=\"_blank\">repo</a><br>\nNotebook. The version 264 is first a stage and 266 the second one:<br>\n<a href=\"https://www.kaggle.com/code/jankowalski2000/3rd-place-solution\" target=\"_blank\">notebook</a>.</p>\n<p>The code runs in approximately 1 hour using CPU Intel(R) Core(TM) i5-9300H CPU @ 2.40GHz and 8GB RAM.</p>",
      "rawMarkdown": "# 1. Integration of Biological Knowledge\nGenerally, I treated this problem as a regression with 2 feature columns and 18211 targets, but I had tried to utilize SMILES sequences in neural network with LSTM unit. Both the sm_name and SMILES columns can be encoded exactly in the same way, so the sm_name column can be replaced with the SMILES column. Moreover, SMILES column is more informative, because every single character of the sequence can be encoded (not only a single value like in sm_name)  and the order of these characters provides extra information. Theoretically, In the worst case, the performance of the neural network using SMILES column instead of sn_name should be not worse than using original columns. Unfortunate, I have reached an unsatisfactory public score with this neural network and I stopped further research. \n\n# 2. Exploration of the problem\nFor simplicity, the analysis is done for the single best model without pseudolabeling. Since mrrmse metric is sensitive to outliers, a distribution of ranges for columns should be checked.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F523b491c0054eae38b031b945b7eb168%2FScreenshot%20from%202023-12-05%2009-36-00.png?generation=1701849069283853&alt=media)\nThe majority of columns have a range of values in an interval (4, 50). Naturally, the columns with high range lead to high mae or mse, so in order to determine genes which are easy and hard to predict, the standardized (divided by std) colwise mse is applied. The table below shows the hardest and easiest genes for prediction.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2Fd8bf3ee03e6d6154593170a5db43cab3%2FScreenshot%20from%202023-12-05%2009-45-05.png?generation=1701849594906441&alt=media)\nThe scheme of the applied cross validation.  Every fold contains one cell type chosen from NK cells, T cells CD4+, T cells CD8+, T regulatory cells and only sm_names being in public and private test was involved. The lowest value of this validation split corresponds to the lowest value on public and private dataset. In my opinion this is a reliable split and the perfect split depends on the model architecture, so every model can have its perfect validation split. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2Fb4a580c8349343ab853a40cf6ee8b43e%2FScreenshot%20from%202023-12-05%2009-49-01.png?generation=1701849988973830&alt=media)\nThe easiness of learning per cell types is shown below. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F64d7f90af54659c0e8477d4faa3a3c51%2FScreenshot%20from%202023-12-05%2009-56-05.png?generation=1701850210915961&alt=media)\nThe values of loss are different, because they are calculated in truncated space. T cells CD4+ and NK cells are learning well. T regulatory cells are harder for training. The T cells CD8 are weakly improving on validation dataset. This split uses about 25% of the dataset for validation, so I believe more reliable splits exist. My new proposition of split is: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F718b580e65a5654a92880f8a6bcf6d1b%2FScreenshot%20from%202023-12-05%2010-08-16.png?generation=1701850452329548&alt=media)\nThis split is similar to the previous one, but for each validation fold, randomly selected examples are moved to training. Let's check an impact of the new split for training. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F8a68fc6a4b3025b48d37dabc67a50c05%2FScreenshot%20from%202023-12-05%2010-12-19.png?generation=1701850788086187&alt=media)\nFor each cell type, more data improved performance on this same, constant and small validation set. Metric is calculated on full dimension this time. The first value at 480 training examples corresponds to the previous split. Theoretically, decreasing the size of validation set to one example can lead to the best performance. Since it will be very similar to training on whole dataset, what is done finally.\n\n#3. Model design\n## Solution\nThe prediction system is two staged, so I publish two versions of the notebook. \nThe first stage predicts pseudolabels. To be honest, if I stopped on this version, I would not be the third. \nThe predicted pseudolabels on all test data (255 rows) are added to training in the second stage.\n### Stage 1 preparing pseudolabels\nThe main part of this system is a neural network. Every neural network and its environment was optimized by optuna. Hyperparameters that have been optimized:\na dropout value, a number of neurons in particular layers, an output dimension of an embedding layer, a number of epochs, a learning rate, a batch size, a number of dimension of truncated singular value decomposition.\nThe optimization was done on custom 4-folds cross validation.  In order to avoid overfitting to cross validation by optuna I applied 2 repeats for every fold and took an average. Generally, the more, the better. The optuna's criterion was MRRMSE. \nFinally, 7 models were ensembled. Optuna was applied again to determine best weights of linear combination. The prediction of test set is the pseudolabels now and will be used in second stage.\n### Stage 2 retraining with pseudolabels\nThe pseudolabels (255 rows) were added to the training dataset. I applied 20 models with optimized parameters in different experiments for a model diversity.\nOptuna selected optimal weights for the linear combination of the prediction again.\nModels had high variance, so every model was trained 10 times on all dataset and the median of prediction is taken as a final prediction.  The prediction was additionally clipped to colwise min and max. \n\n**History of improvements:**\n1. a replacing onehot encoding with an embedding layer\n2. a replacing MAE loss with MRRMSE loss\n3. an ensembing of models with mean\n4. a dimension reduction with truncated singular value decomposition\n5. an ensembling of models with weighted mean\n6. using pseudolabeling\n7. using pseudolabeling and ensembling of 20 models and weighted mean. \n\n**What did not work for me**:\n- a label normalization, standardization\n- a chained regression\n- a denoising dataset\n- a removal of outliers\n- an adding noise to labels\n- a training on selected easy / hard to predict columns\n- a huber loss.\n\n# 4. Robustness\nI have tested 3 types of the robustness: increasing dataset size, adding noise to labels and inputs. Adding the noise to inputs failed totally, it is logical for me, because the nominal and categorical values are changed and are behaving like the continues values, what is not beneficial.\nLet's see the performance on 40%, 50%,..., 100% of training dataset. It started from 40%, because singular values decomposition is limited by number of examples.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2Fb684812f2feb70c2b9f28f2df5714cd7%2FScreenshot%20from%202023-12-05%2013-11-10.png?generation=1701852018611279&alt=media)\nThe experiment was 5 times repeated, so the interval of uncertainty is visible. More data improves the performance significantly. \nThe last test of robustness is adding noise to the labels. The random gaussian (a distribution with 0 mean and scale * std) noise was added to the labels.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F8625a5cb1af874f4986277c98e76ab8e%2FScreenshot%20from%202023-12-05%2011-45-22.png?generation=1701852528532795&alt=media)\nAdding some noise (0.01 * std) can even improve the model's performance.  Generally, the model is robust to the noise. \n\n# 5. Documentation & code style\nThe code on GitHub is documented. \n\n# 6. Reproducibility\nGitHub code:\n[repo](https://github.com/okon2000/single_cell_perturbations)\nNotebook. The version 264 is first a stage and 266 the second one:\n[notebook](https://www.kaggle.com/code/jankowalski2000/3rd-place-solution).\n\nThe code runs in approximately 1 hour using CPU Intel(R) Core(TM) i5-9300H CPU @ 2.40GHz and 8GB RAM.",
      "votes": 33
    },
    {
      "id": 2545958,
      "postDate": "2023-12-02T01:03:13.280Z",
      "content": "<p>Congratulations🎉🎉🎉 <a href=\"https://www.kaggle.com/jankowalski2000\" target=\"_blank\">@jankowalski2000</a> ! Thank you very much for sharing! Your code is as beautiful as a poem, and reading it is a pleasure. From the updates on your notebook, I can see your perseverance in solving problems, and this honor is the best reward for your efforts. You are so young and talented, respect!</p>",
      "rawMarkdown": "Congratulations🎉🎉🎉 @jankowalski2000 ! Thank you very much for sharing! Your code is as beautiful as a poem, and reading it is a pleasure. From the updates on your notebook, I can see your perseverance in solving problems, and this honor is the best reward for your efforts. You are so young and talented, respect!",
      "votes": 6
    },
    {
      "id": 2554858,
      "postDate": "2023-12-09T13:21:10.383Z",
      "content": "<p>Great work 👏</p>",
      "rawMarkdown": "Great work 👏",
      "votes": 1
    },
    {
      "id": 2553850,
      "postDate": "2023-12-08T14:52:24.483Z",
      "content": "<p>congratulations <a href=\"https://www.kaggle.com/jankowalski2000\" target=\"_blank\">@jankowalski2000</a> ! ,  will your model be presented (by you or someone of your team) in the NeurIPS workshop next week ?</p>",
      "rawMarkdown": "congratulations @jankowalski2000 ! ,  will your model be presented (by you or someone of your team) in the NeurIPS workshop next week ?",
      "votes": 1,
      "replies": [
        {
          "id": 2553893,
          "postDate": "2023-12-08T15:33:15.643Z",
          "content": "<p>I am not sure, I have to provide a lot of documents and refactor my code to fulfill all requirements for a leaderboard award. </p>",
          "rawMarkdown": "I am not sure, I have to provide a lot of documents and refactor my code to fulfill all requirements for a leaderboard award. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2550810,
      "postDate": "2023-12-06T09:40:27.353Z",
      "content": "<p>The analysis and github code are added. </p>",
      "rawMarkdown": "The analysis and github code are added. ",
      "votes": 1
    },
    {
      "id": 2547310,
      "postDate": "2023-12-03T12:10:11.610Z",
      "content": "<p>Congratulations on 3rd place in this competition. Thanks for sharing the details of your solution. </p>",
      "rawMarkdown": "Congratulations on 3rd place in this competition. Thanks for sharing the details of your solution. ",
      "votes": 1
    },
    {
      "id": 2545687,
      "postDate": "2023-12-01T17:00:23.080Z",
      "content": "<p>Great Work, I've gone through all.</p>",
      "rawMarkdown": "Great Work, I've gone through all.",
      "votes": 1
    },
    {
      "id": 2629015,
      "postDate": "2024-01-31T15:27:52.797Z",
      "content": "<p>It seems like pseudolabelling was a big part of the success of your solution. For classification, there is some explanation that pseudolabelling can help as a regularizer similar to entropy regularization (Lee 2013), but I've not seen it used before or studied in the regression setting. Why do you think it helps here, or have you seen it used in regression before? Thanks for the writeup!</p>",
      "rawMarkdown": "It seems like pseudolabelling was a big part of the success of your solution. For classification, there is some explanation that pseudolabelling can help as a regularizer similar to entropy regularization (Lee 2013), but I've not seen it used before or studied in the regression setting. Why do you think it helps here, or have you seen it used in regression before? Thanks for the writeup!"
    },
    {
      "id": 2559124,
      "postDate": "2023-12-12T16:08:36.207Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2923607,
      "postDate": "2024-07-16T02:38:38.687Z",
      "content": "<p>Helps Me A lot, thanks!</p>",
      "rawMarkdown": "Helps Me A lot, thanks!"
    }
  ],
  "comments": [
    {
      "id": 2545958,
      "author_name": "BarryZhou",
      "author_url": "",
      "post_date": "2023-12-02T01:03:13.280000",
      "content": "<p>Congratulations🎉🎉🎉 <a href=\"https://www.kaggle.com/jankowalski2000\" target=\"_blank\">@jankowalski2000</a> ! Thank you very much for sharing! Your code is as beautiful as a poem, and reading it is a pleasure. From the updates on your notebook, I can see your perseverance in solving problems, and this honor is the best reward for your efforts. You are so young and talented, respect!</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2554858,
      "author_name": "Ankita Nain",
      "author_url": "",
      "post_date": "2023-12-09T13:21:10.383000",
      "content": "<p>Great work 👏</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2553850,
      "author_name": "johanna-galvis",
      "author_url": "",
      "post_date": "2023-12-08T14:52:24.483000",
      "content": "<p>congratulations <a href=\"https://www.kaggle.com/jankowalski2000\" target=\"_blank\">@jankowalski2000</a> ! ,  will your model be presented (by you or someone of your team) in the NeurIPS workshop next week ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2553893,
          "author_name": "Rafał Pawłowski",
          "author_url": "",
          "post_date": "2023-12-08T15:33:15.643000",
          "content": "<p>I am not sure, I have to provide a lot of documents and refactor my code to fulfill all requirements for a leaderboard award. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2550810,
      "author_name": "Rafał Pawłowski",
      "author_url": "",
      "post_date": "2023-12-06T09:40:27.353000",
      "content": "<p>The analysis and github code are added. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2547310,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-12-03T12:10:11.610000",
      "content": "<p>Congratulations on 3rd place in this competition. Thanks for sharing the details of your solution. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2545687,
      "author_name": "VIKRAM MISHRA",
      "author_url": "",
      "post_date": "2023-12-01T17:00:23.080000",
      "content": "<p>Great Work, I've gone through all.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2629015,
      "author_name": "Maximus Power",
      "author_url": "",
      "post_date": "2024-01-31T15:27:52.797000",
      "content": "<p>It seems like pseudolabelling was a big part of the success of your solution. For classification, there is some explanation that pseudolabelling can help as a regularizer similar to entropy regularization (Lee 2013), but I've not seen it used before or studied in the regression setting. Why do you think it helps here, or have you seen it used in regression before? Thanks for the writeup!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2559124,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-12T16:08:36.207000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2923607,
      "author_name": "Sixian Hsu",
      "author_url": "",
      "post_date": "2024-07-16T02:38:38.687000",
      "content": "<p>Helps Me A lot, thanks!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2545160": "# 1. Integration of Biological Knowledge\nGenerally, I treated this problem as a regression with 2 feature columns and 18211 targets, but I had tried to utilize SMILES sequences in neural network with LSTM unit. Both the sm_name and SMILES columns can be encoded exactly in the same way, so the sm_name column can be replaced with the SMILES column. Moreover, SMILES column is more informative, because every single character of the sequence can be encoded (not only a single value like in sm_name)  and the order of these characters provides extra information. Theoretically, In the worst case, the performance of the neural network using SMILES column instead of sn_name should be not worse than using original columns. Unfortunate, I have reached an unsatisfactory public score with this neural network and I stopped further research. \n\n# 2. Exploration of the problem\nFor simplicity, the analysis is done for the single best model without pseudolabeling. Since mrrmse metric is sensitive to outliers, a distribution of ranges for columns should be checked.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F523b491c0054eae38b031b945b7eb168%2FScreenshot%20from%202023-12-05%2009-36-00.png?generation=1701849069283853&alt=media)\nThe majority of columns have a range of values in an interval (4, 50). Naturally, the columns with high range lead to high mae or mse, so in order to determine genes which are easy and hard to predict, the standardized (divided by std) colwise mse is applied. The table below shows the hardest and easiest genes for prediction.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2Fd8bf3ee03e6d6154593170a5db43cab3%2FScreenshot%20from%202023-12-05%2009-45-05.png?generation=1701849594906441&alt=media)\nThe scheme of the applied cross validation.  Every fold contains one cell type chosen from NK cells, T cells CD4+, T cells CD8+, T regulatory cells and only sm_names being in public and private test was involved. The lowest value of this validation split corresponds to the lowest value on public and private dataset. In my opinion this is a reliable split and the perfect split depends on the model architecture, so every model can have its perfect validation split. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2Fb4a580c8349343ab853a40cf6ee8b43e%2FScreenshot%20from%202023-12-05%2009-49-01.png?generation=1701849988973830&alt=media)\nThe easiness of learning per cell types is shown below. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F64d7f90af54659c0e8477d4faa3a3c51%2FScreenshot%20from%202023-12-05%2009-56-05.png?generation=1701850210915961&alt=media)\nThe values of loss are different, because they are calculated in truncated space. T cells CD4+ and NK cells are learning well. T regulatory cells are harder for training. The T cells CD8 are weakly improving on validation dataset. This split uses about 25% of the dataset for validation, so I believe more reliable splits exist. My new proposition of split is: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F718b580e65a5654a92880f8a6bcf6d1b%2FScreenshot%20from%202023-12-05%2010-08-16.png?generation=1701850452329548&alt=media)\nThis split is similar to the previous one, but for each validation fold, randomly selected examples are moved to training. Let's check an impact of the new split for training. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F8a68fc6a4b3025b48d37dabc67a50c05%2FScreenshot%20from%202023-12-05%2010-12-19.png?generation=1701850788086187&alt=media)\nFor each cell type, more data improved performance on this same, constant and small validation set. Metric is calculated on full dimension this time. The first value at 480 training examples corresponds to the previous split. Theoretically, decreasing the size of validation set to one example can lead to the best performance. Since it will be very similar to training on whole dataset, what is done finally.\n\n#3. Model design\n## Solution\nThe prediction system is two staged, so I publish two versions of the notebook. \nThe first stage predicts pseudolabels. To be honest, if I stopped on this version, I would not be the third. \nThe predicted pseudolabels on all test data (255 rows) are added to training in the second stage.\n### Stage 1 preparing pseudolabels\nThe main part of this system is a neural network. Every neural network and its environment was optimized by optuna. Hyperparameters that have been optimized:\na dropout value, a number of neurons in particular layers, an output dimension of an embedding layer, a number of epochs, a learning rate, a batch size, a number of dimension of truncated singular value decomposition.\nThe optimization was done on custom 4-folds cross validation.  In order to avoid overfitting to cross validation by optuna I applied 2 repeats for every fold and took an average. Generally, the more, the better. The optuna's criterion was MRRMSE. \nFinally, 7 models were ensembled. Optuna was applied again to determine best weights of linear combination. The prediction of test set is the pseudolabels now and will be used in second stage.\n### Stage 2 retraining with pseudolabels\nThe pseudolabels (255 rows) were added to the training dataset. I applied 20 models with optimized parameters in different experiments for a model diversity.\nOptuna selected optimal weights for the linear combination of the prediction again.\nModels had high variance, so every model was trained 10 times on all dataset and the median of prediction is taken as a final prediction.  The prediction was additionally clipped to colwise min and max. \n\n**History of improvements:**\n1. a replacing onehot encoding with an embedding layer\n2. a replacing MAE loss with MRRMSE loss\n3. an ensembing of models with mean\n4. a dimension reduction with truncated singular value decomposition\n5. an ensembling of models with weighted mean\n6. using pseudolabeling\n7. using pseudolabeling and ensembling of 20 models and weighted mean. \n\n**What did not work for me**:\n- a label normalization, standardization\n- a chained regression\n- a denoising dataset\n- a removal of outliers\n- an adding noise to labels\n- a training on selected easy / hard to predict columns\n- a huber loss.\n\n# 4. Robustness\nI have tested 3 types of the robustness: increasing dataset size, adding noise to labels and inputs. Adding the noise to inputs failed totally, it is logical for me, because the nominal and categorical values are changed and are behaving like the continues values, what is not beneficial.\nLet's see the performance on 40%, 50%,..., 100% of training dataset. It started from 40%, because singular values decomposition is limited by number of examples.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2Fb684812f2feb70c2b9f28f2df5714cd7%2FScreenshot%20from%202023-12-05%2013-11-10.png?generation=1701852018611279&alt=media)\nThe experiment was 5 times repeated, so the interval of uncertainty is visible. More data improves the performance significantly. \nThe last test of robustness is adding noise to the labels. The random gaussian (a distribution with 0 mean and scale * std) noise was added to the labels.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6882624%2F8625a5cb1af874f4986277c98e76ab8e%2FScreenshot%20from%202023-12-05%2011-45-22.png?generation=1701852528532795&alt=media)\nAdding some noise (0.01 * std) can even improve the model's performance.  Generally, the model is robust to the noise. \n\n# 5. Documentation & code style\nThe code on GitHub is documented. \n\n# 6. Reproducibility\nGitHub code:\n[repo](https://github.com/okon2000/single_cell_perturbations)\nNotebook. The version 264 is first a stage and 266 the second one:\n[notebook](https://www.kaggle.com/code/jankowalski2000/3rd-place-solution).\n\nThe code runs in approximately 1 hour using CPU Intel(R) Core(TM) i5-9300H CPU @ 2.40GHz and 8GB RAM.",
    "2545958": "Congratulations🎉🎉🎉 @jankowalski2000 ! Thank you very much for sharing! Your code is as beautiful as a poem, and reading it is a pleasure. From the updates on your notebook, I can see your perseverance in solving problems, and this honor is the best reward for your efforts. You are so young and talented, respect!",
    "2554858": "Great work 👏",
    "2553850": "congratulations @jankowalski2000 ! ,  will your model be presented (by you or someone of your team) in the NeurIPS workshop next week ?",
    "2550810": "The analysis and github code are added. ",
    "2547310": "Congratulations on 3rd place in this competition. Thanks for sharing the details of your solution. ",
    "2545687": "Great Work, I've gone through all.",
    "2629015": "It seems like pseudolabelling was a big part of the success of your solution. For classification, there is some explanation that pseudolabelling can help as a regularizer similar to entropy regularization (Lee 2013), but I've not seen it used before or studied in the regression setting. Why do you think it helps here, or have you seen it used in regression before? Thanks for the writeup!",
    "2559124": "",
    "2923607": "Helps Me A lot, thanks!"
  }
}