{
  "id": 459258,
  "title": "1st Place Solution Writeup for Open Problems – Single-Cell Perturbations",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/459258",
  "author_name": "JK-Piece",
  "post_date": "2023-12-04T12:40:52.913000",
  "votes": 50,
  "comment_count": 22,
  "views": 0,
  "content": "<p>I would like to thank the organizers and Kaggle for hosting this exciting competition. I am also grateful to the participants who shared starter notebooks, datasets, and insightful ideas. Below is a more detailed writeup of my solution, including late findings.</p>\n<h1>Competition Page</h1>\n<p><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview</a></p>\n<p><a href=\"https://openproblems.bio/\" target=\"_blank\">https://openproblems.bio/</a></p>\n<h1>1. Integration of Biological Knowledge</h1>\n<p>Since the input features only consisted of pairs of short keywords, that is, cell types and small molecule names, and given the large size of the target variable, I was quickly convinced  that I needed to somehow enrich the input feature space. I therefore dedicated my first days of the competition to this task. First, I searched for biological word/term embeddings in the literature, and found the paper  ''BioWordVec, improving biomedical word embeddings with subword information and MeSH’’ by Zhang et al [1]. The paper directed me to the code on Github where I could find pretrained embeddings for biological terms. This was motivated by the fact that 1) I would be able to find most cell types, and small molecule names in such embeddings, and 2) The embeddings would encode rich information about the general meaning of each term. With these embeddings, I created larger input features and trained a regression model. This achieved 0.767 on the public leaderboard. With a better hyperparameter search and feature engineering, I improved the score to 0.614. As this seemed to be a good direction to go, I decided to further enrich the input features. This time, I searched for the definition of each cell type and small molecule name on wikipedia. For this, I used the python library <code>wikipedia</code> <a href=\"https://pypi.org/project/wikipedia/\" target=\"_blank\">https://pypi.org/project/wikipedia/</a>. I then represented each cell type and small molecule name by a few sentences describing it, then I bootstrapped an embedding from the descriptions. For example, Nk cells were described by: ''Natural killer cells, also known as NK cells or large granular lymphocytes (LGL), are a type of cytotoxic lymphocyte critical to the innate immune system that belong to the rapidly expanding family of known innate lymphoid cells (ILC) and represent 5–20% of all circulating lymphocytes in humans. The role of NK cells is analogous to that of cytotoxic T cells in the vertebrate adaptive immune response. NK cells provide rapid responses to virus-infected cell and other intracellular pathogens acting at around 3 days after infection, and respond to tumor formation.’’ I also explored different numbers of sentences to describe each cell type and small molecule.<br>\nWhile this is interesting from the biological point of view, it did not improve the leaderboard score. In fact, the score became worse (0.656 vs. 0.614 previously). This can be explained by the fact that such natural language descriptions came with some noise, and pretrained embeddings were probably not computed to deal with this. Fine-tuning the embeddings on natural language descriptions of biological terms also fell short.</p>\n<p>Because my initial idea about input feature enrichment did not meet my expectations, I decided to look for alternatives. Thanks to the discussions in the forum, I came across a notebook proposing to use SMILES to encode chemical structures of small molecules. I immediately decided to use ChemBERTa embeddings of SMILES encodings and observed a significant improvement in the evaluation metric MRRMSE on the validation data splits (I used a 5 fold cross-validation setting throughout the competition). With this, I developed additional data augmentation techniques, including the mean, standard deviation, and (25%, 50%, 75%) percentiles of differential expressions per cell type and small molecule in the training data.</p>\n<h1>2. Exploration of the Problem</h1>\n<p>As mentioned in the previous section, I started the competition by trying to build rich features for the input pairs (cell_type, sm_name). Ultimately, the use of ChemBERTa features of small molecules’ SMILES appeared to be an important step towards this goal.  Combined with the mean, standard deviation, and (25%, 50%, 75%) percentiles per cell type and small molecule, I achieved an optimal input feature representation.</p>\n<p>In my experiments, I used a 5-fold cross-validation setting with a fixed seed (42). It was hard to achieve a good score on the validation sets of the 2nd and 4th folds. On these folds, the MRRMSE on the validation set was approximately 1.19, and 1.15 on average, respectively. On the 1st, 3rd and 5th folds the average scores were 0.86, 0.86, and 0.90, respectively. The scores are the average across different model architectures (LSTM, 1d-CNN, GRU) and different input feature combinations (''initial’’, ''light’’, ''heavy’’). The three different input feature representations are as follows:<br>\n“initial”: ChemBERTa embeddings, 1 hot encoding of cell_type/sm_name pairs, mean, std, percentiles of targets per cell_type and sm_name<br>\n“light”: ChemBERTa embeddings, 1 hot encoding of cell_type/sm_name pairs, mean targets per cell_type and sm_name<br>\n“heavy”: ChemBERTa embeddings, 1 hot encoding of cell_type/sm_name pairs, mean, 25%, 50%, 75% percentiles of targets per cell_type and sm_name<br>\nThe figure below shows the training curves (MRRMSE) per fold averaged over all three model architectures.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3937164%2F8e7a5107e722900173e206496e4e8ee6%2FMRRMSE.png?generation=1702286999246540&amp;alt=media\" alt=\"\"></p>\n<p>The differences in the validation MRRMSE in the above figure motivated me to take a closer look into the validation sets, where I found different distributions of cell types. The figure below shows the predominant cell types per fold and the corresponding average (across models and different input feature representations) validation MRRMSE. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3937164%2F8e816183ed88332f09fb95731adcf775%2FHardCelltypes.png?generation=1702287123989655&amp;alt=media\" alt=\"\"></p>\n<p>In the validation sets of the 1st, 3rd, and 5th folds, the predominant cell types (in terms of percentage) are ''T regulatory cells’’, ''B cells’’, and ''Nk cells’’, respectively. On the 2nd and 4th folds, ''T cells CD8+’’ and ''Myeloid cells’’ are the most represented cell types in the validation sets, respectively. The percentage is computed as the number of occurrences of a cell type in the validation set divided by the number of occurrences of that cell type in the training set.<br>\nFrom the bar-plots above, the cell types that are easier to predict are ''T regulatory cells’’, ''B cells’’, and ''Nk cells’’, while ''T cells CD8+’’ and ''Myeloid cells’’ are the hardest to predict. Based on this observation, an ideal training set should include more  ''T cells CD8+’’ and ''Myeloid cells’’ than the rest of the cell types. In this way, trained ML models would be able to generalize to other cell types.</p>\n<h1>3. Model Design</h1>\n<h2>Model Architecture</h2>\n<p>I tried different model architectures, including gradient boosting models, MLP, and 2D CNN which did not work so well. I finally selected LSTM, GRU, and 1-d CNN architectures as they performed better on the validation sets. Below I show a rough implementation of the GRU model.</p>\n<pre><code>dims_dict = {: {: , : , : },\n                       : {: {: , : , : },\n                                  : {: [,], : [,], : [,]}\n                     }}\n (nn.Module):\n     ():\n        (GRU, self).__init__()\n        self.name = \n        self.scheme = scheme\n        self.gru = nn.GRU(dims_dict[][][self.scheme][], , num_layers=, batch_first=)\n        self.linear = nn.Sequential(\n            nn.Linear(dims_dict[][][self.scheme], ),\n            nn.Dropout(),\n            nn.ReLU(),\n            nn.Linear(, ),\n            nn.Dropout(),\n            nn.ReLU())\n        self.head = nn.Linear(, )\n\n        self.loss1 = nn.MSELoss()\n        self.loss2 = LogCoshLoss()\n        self.loss3 = nn.L1Loss()\n        self.loss4 = nn.BCELoss()\n\n     ():\n        shape1, shape2 = dims_dict[][][self.scheme]\n        x = x.reshape(x.shape[],shape1,shape2)\n         y  :\n            out, hn = self.gru(x)\n            out = out.reshape(out.shape[],-)\n            out = torch.cat([out, hn.reshape(hn.shape[], -)], dim=)\n            out = self.head(self.linear(out))\n             out\n        :\n            out, hn = self.gru(x)\n            out = out.reshape(out.shape[],-)\n            out = torch.cat([out, hn.reshape(hn.shape[], -)], dim=)\n            out = self.head(self.linear(out))\n            loss1 = *self.loss1(out, y) + *self.loss2(out, y) + *self.loss3(out, y)\n            yhat = torch.sigmoid(out)\n            yy = torch.sigmoid(y)\n            loss2 = self.loss4(yhat, yy)\n             *loss1 + *loss2\n</code></pre>\n<p>In my late experiments, I realized that 1d-CNN and GRU are actually the best architectures as they achieve the best scores alone (0.733  for GRU and 0.745 for 1d-CNN on Private LB). LSTM alone achieves 0.839 on Private LB. With 0.25xLSTM + 0.65xCNN the Private LB is 0.725, and with 0.25xLSTM + 0.65xGRU the Private LB is 0.723.</p>\n<h2>Loss Functions and Optimizer</h2>\n<p>I simultaneously optimized 4 loss functions via weighted averaging: MSE, MAE, LogCosh, and BCE. The weights are 0.32, 0.24, 0.24, and 0.2, respectively.<br>\nThis was found to enhance the predictive performance of models. The Adam optimizer with learning rate 0.001 for LSTM and CNN, and 0.0003 for GRU was used to train the models. LogCosh is defined as:</p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n\n     ():\n        ey_t = (y_t - y_prime_t)/ \n         torch.mean(torch.log(torch.cosh(ey_t + )))\n</code></pre>\n<p>LogCosh is similar to MAE with the difference being that it is a softer version that can allow smoother convergence. It was adapted from <a href=\"https://github.com/tuantle/regression-losses-pytorch\" target=\"_blank\">https://github.com/tuantle/regression-losses-pytorch</a>.</p>\n<p>The BCE loss is indeed special as it is often used for classification tasks. However, I argue that it sends better signals to the models and optimizers when the target values are close to zero. To demonstrate this, consider the following two pieces of code:</p>\n<pre><code>m1 = nn.Sigmoid()\nloss = nn.BCELoss()\n = torch.tensor([], requires_grad=).unsqueeze()\ntarget = torch.sigmoid(torch.tensor([-], requires_grad=).unsqueeze())\noutput1 = loss(m1(), target)\n(output1.item()) \n\nm2 = nn.Identity()\nloss = nn.MSELoss()\n = torch.tensor([], requires_grad=).unsqueeze()\ntarget = torch.tensor([-], requires_grad=).unsqueeze()\noutput2 = loss(m2(), target)\n(output2.item())\n</code></pre>\n<p>With this example, one can observe that the MSELoss tells the model and optimizer that \"it is ok, there is no mistake here\". Obviously, there is a mistake, and BCELoss can see it as it returns a high loss value (0.694 compared to 0.010 for MSELoss). My choice of the BCELoss in this competition is motivated by the fact that most target values are from a Gaussian distribution with mean 0 as can be seen in the figure below.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3937164%2F7d9566cc821cac514ad9ebc4fd4e658d%2FGaussian.png?generation=1702287278847099&amp;alt=media\" alt=\"\"></p>\n<h2>Hyperparameters</h2>\n<ul>\n<li>250 epochs of training</li>\n<li>Learning rate 0.001 for LSTM and CNN, and 0.0003 for GRU</li>\n<li>Gradient norm clip value: [5.0, 1.0, 1.0] for the three schemes ''initial'', ''light'', and ''heavy''</li>\n</ul>\n<h1>4. Robustness</h1>\n<p>I conducted 4 experiments using different subsets of the training data, and monitored the private leaderboard score. I considered subsets of the initial training data (de_train) with sizes 25%, 50%, 75%, and 100%. Below 25%, we cannot cover all small molecules in the test set (id_map) even with a stratified split on sm_name, and hence the one hot encoding algorithm cannot run. With 25%, I achieved 0.946. With 50%, I achieved 0.815. With 75% of the training data, it is 0.769, and with the full data the private leaderboard is 0.719 (which is better than my winning submission because I removed padding in the ChemBERTa model). The figure below shows the robustness of my approach as a decreasing curve, i.e., improvement of the MRRMSE with increasing training data amount.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3937164%2F1eeb2e54c1c23a317402cb5c8ede967d%2Frobustness.png?generation=1702287304367920&amp;alt=media\" alt=\"\"></p>\n<p>My second data augmentation technique can be regarded as noise addition. I randomly replace 30% of the input features’ entries with zeros, and add the resulting input feature together with the correct target as a new training datapoint. This has proven to improve the predictive performance of my models. In this sense, my models are robust to the noise as their performance is not hindered but rather improved. The biological motivation here is that we might not need to know the complete chemical structure of a molecule (assuming the dropped input features are from sm_name) to know its impact on a cell. Similarly, there might be a biological disorder in a cell, and we would still expect that cell to respond to a molecule (drug) in the same way as a normal cell.</p>\n<p>Below is the data augmentation function</p>\n<pre><code> ():\n    copy_x = x_.copy()\n    new_x = []\n    new_y = y_.copy()\n    dim = x_.shape[]\n    k = (*dim)\n     i  (x_.shape[]):\n        idx = random.sample((dim), k=k)\n        copy_x[i,:,idx] = \n        new_x.append(copy_x[i])\n     np.stack(new_x, axis=), new_y\n</code></pre>\n<h1>5. Documentation and Code Style</h1>\n<p>The documentation and software dependencies are available on Github at  <a href=\"https://github.com/Jean-KOUAGOU/1st-place-solution-single-cell-pbs\" target=\"_blank\">https://github.com/Jean-KOUAGOU/1st-place-solution-single-cell-pbs</a></p>\n<h1>6. Reproducibility</h1>\n<p>Code is available and well documented on Github at <a href=\"https://github.com/Jean-KOUAGOU/1st-place-solution-single-cell-pbs\" target=\"_blank\">https://github.com/Jean-KOUAGOU/1st-place-solution-single-cell-pbs</a>. Reproduction scripts are added.</p>\n<h1>Sources</h1>\n<p><a href=\"https://www.nature.com/articles/s41597-019-0055-0\" target=\"_blank\">[1] BioWordVec, improving biomedical word embeddings with subword information and MeSH</a><br>\n<a href=\"https://github.com/tuantle/regression-losses-pytorch\" target=\"_blank\">Pytorch Regression Loss Functions</a><br>\n[ChemBERTa](<a href=\"https://huggingface.co/DeepChem/ChemBERTa-77M-MTR\" target=\"_blank\">https://huggingface.co/DeepChem/ChemBERTa-77M-MTR</a></p>",
  "messages": [
    {
      "id": 2548444,
      "postDate": "2023-12-04T12:40:52.913Z",
      "content": "<p>I would like to thank the organizers and Kaggle for hosting this exciting competition. I am also grateful to the participants who shared starter notebooks, datasets, and insightful ideas. Below is a more detailed writeup of my solution, including late findings.</p>\n<h1>Competition Page</h1>\n<p><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview</a></p>\n<p><a href=\"https://openproblems.bio/\" target=\"_blank\">https://openproblems.bio/</a></p>\n<h1>1. Integration of Biological Knowledge</h1>\n<p>Since the input features only consisted of pairs of short keywords, that is, cell types and small molecule names, and given the large size of the target variable, I was quickly convinced  that I needed to somehow enrich the input feature space. I therefore dedicated my first days of the competition to this task. First, I searched for biological word/term embeddings in the literature, and found the paper  ''BioWordVec, improving biomedical word embeddings with subword information and MeSH’’ by Zhang et al [1]. The paper directed me to the code on Github where I could find pretrained embeddings for biological terms. This was motivated by the fact that 1) I would be able to find most cell types, and small molecule names in such embeddings, and 2) The embeddings would encode rich information about the general meaning of each term. With these embeddings, I created larger input features and trained a regression model. This achieved 0.767 on the public leaderboard. With a better hyperparameter search and feature engineering, I improved the score to 0.614. As this seemed to be a good direction to go, I decided to further enrich the input features. This time, I searched for the definition of each cell type and small molecule name on wikipedia. For this, I used the python library <code>wikipedia</code> <a href=\"https://pypi.org/project/wikipedia/\" target=\"_blank\">https://pypi.org/project/wikipedia/</a>. I then represented each cell type and small molecule name by a few sentences describing it, then I bootstrapped an embedding from the descriptions. For example, Nk cells were described by: ''Natural killer cells, also known as NK cells or large granular lymphocytes (LGL), are a type of cytotoxic lymphocyte critical to the innate immune system that belong to the rapidly expanding family of known innate lymphoid cells (ILC) and represent 5–20% of all circulating lymphocytes in humans. The role of NK cells is analogous to that of cytotoxic T cells in the vertebrate adaptive immune response. NK cells provide rapid responses to virus-infected cell and other intracellular pathogens acting at around 3 days after infection, and respond to tumor formation.’’ I also explored different numbers of sentences to describe each cell type and small molecule.<br>\nWhile this is interesting from the biological point of view, it did not improve the leaderboard score. In fact, the score became worse (0.656 vs. 0.614 previously). This can be explained by the fact that such natural language descriptions came with some noise, and pretrained embeddings were probably not computed to deal with this. Fine-tuning the embeddings on natural language descriptions of biological terms also fell short.</p>\n<p>Because my initial idea about input feature enrichment did not meet my expectations, I decided to look for alternatives. Thanks to the discussions in the forum, I came across a notebook proposing to use SMILES to encode chemical structures of small molecules. I immediately decided to use ChemBERTa embeddings of SMILES encodings and observed a significant improvement in the evaluation metric MRRMSE on the validation data splits (I used a 5 fold cross-validation setting throughout the competition). With this, I developed additional data augmentation techniques, including the mean, standard deviation, and (25%, 50%, 75%) percentiles of differential expressions per cell type and small molecule in the training data.</p>\n<h1>2. Exploration of the Problem</h1>\n<p>As mentioned in the previous section, I started the competition by trying to build rich features for the input pairs (cell_type, sm_name). Ultimately, the use of ChemBERTa features of small molecules’ SMILES appeared to be an important step towards this goal.  Combined with the mean, standard deviation, and (25%, 50%, 75%) percentiles per cell type and small molecule, I achieved an optimal input feature representation.</p>\n<p>In my experiments, I used a 5-fold cross-validation setting with a fixed seed (42). It was hard to achieve a good score on the validation sets of the 2nd and 4th folds. On these folds, the MRRMSE on the validation set was approximately 1.19, and 1.15 on average, respectively. On the 1st, 3rd and 5th folds the average scores were 0.86, 0.86, and 0.90, respectively. The scores are the average across different model architectures (LSTM, 1d-CNN, GRU) and different input feature combinations (''initial’’, ''light’’, ''heavy’’). The three different input feature representations are as follows:<br>\n“initial”: ChemBERTa embeddings, 1 hot encoding of cell_type/sm_name pairs, mean, std, percentiles of targets per cell_type and sm_name<br>\n“light”: ChemBERTa embeddings, 1 hot encoding of cell_type/sm_name pairs, mean targets per cell_type and sm_name<br>\n“heavy”: ChemBERTa embeddings, 1 hot encoding of cell_type/sm_name pairs, mean, 25%, 50%, 75% percentiles of targets per cell_type and sm_name<br>\nThe figure below shows the training curves (MRRMSE) per fold averaged over all three model architectures.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3937164%2F8e7a5107e722900173e206496e4e8ee6%2FMRRMSE.png?generation=1702286999246540&amp;alt=media\" alt=\"\"></p>\n<p>The differences in the validation MRRMSE in the above figure motivated me to take a closer look into the validation sets, where I found different distributions of cell types. The figure below shows the predominant cell types per fold and the corresponding average (across models and different input feature representations) validation MRRMSE. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3937164%2F8e816183ed88332f09fb95731adcf775%2FHardCelltypes.png?generation=1702287123989655&amp;alt=media\" alt=\"\"></p>\n<p>In the validation sets of the 1st, 3rd, and 5th folds, the predominant cell types (in terms of percentage) are ''T regulatory cells’’, ''B cells’’, and ''Nk cells’’, respectively. On the 2nd and 4th folds, ''T cells CD8+’’ and ''Myeloid cells’’ are the most represented cell types in the validation sets, respectively. The percentage is computed as the number of occurrences of a cell type in the validation set divided by the number of occurrences of that cell type in the training set.<br>\nFrom the bar-plots above, the cell types that are easier to predict are ''T regulatory cells’’, ''B cells’’, and ''Nk cells’’, while ''T cells CD8+’’ and ''Myeloid cells’’ are the hardest to predict. Based on this observation, an ideal training set should include more  ''T cells CD8+’’ and ''Myeloid cells’’ than the rest of the cell types. In this way, trained ML models would be able to generalize to other cell types.</p>\n<h1>3. Model Design</h1>\n<h2>Model Architecture</h2>\n<p>I tried different model architectures, including gradient boosting models, MLP, and 2D CNN which did not work so well. I finally selected LSTM, GRU, and 1-d CNN architectures as they performed better on the validation sets. Below I show a rough implementation of the GRU model.</p>\n<pre><code>dims_dict = {: {: , : , : },\n                       : {: {: , : , : },\n                                  : {: [,], : [,], : [,]}\n                     }}\n (nn.Module):\n     ():\n        (GRU, self).__init__()\n        self.name = \n        self.scheme = scheme\n        self.gru = nn.GRU(dims_dict[][][self.scheme][], , num_layers=, batch_first=)\n        self.linear = nn.Sequential(\n            nn.Linear(dims_dict[][][self.scheme], ),\n            nn.Dropout(),\n            nn.ReLU(),\n            nn.Linear(, ),\n            nn.Dropout(),\n            nn.ReLU())\n        self.head = nn.Linear(, )\n\n        self.loss1 = nn.MSELoss()\n        self.loss2 = LogCoshLoss()\n        self.loss3 = nn.L1Loss()\n        self.loss4 = nn.BCELoss()\n\n     ():\n        shape1, shape2 = dims_dict[][][self.scheme]\n        x = x.reshape(x.shape[],shape1,shape2)\n         y  :\n            out, hn = self.gru(x)\n            out = out.reshape(out.shape[],-)\n            out = torch.cat([out, hn.reshape(hn.shape[], -)], dim=)\n            out = self.head(self.linear(out))\n             out\n        :\n            out, hn = self.gru(x)\n            out = out.reshape(out.shape[],-)\n            out = torch.cat([out, hn.reshape(hn.shape[], -)], dim=)\n            out = self.head(self.linear(out))\n            loss1 = *self.loss1(out, y) + *self.loss2(out, y) + *self.loss3(out, y)\n            yhat = torch.sigmoid(out)\n            yy = torch.sigmoid(y)\n            loss2 = self.loss4(yhat, yy)\n             *loss1 + *loss2\n</code></pre>\n<p>In my late experiments, I realized that 1d-CNN and GRU are actually the best architectures as they achieve the best scores alone (0.733  for GRU and 0.745 for 1d-CNN on Private LB). LSTM alone achieves 0.839 on Private LB. With 0.25xLSTM + 0.65xCNN the Private LB is 0.725, and with 0.25xLSTM + 0.65xGRU the Private LB is 0.723.</p>\n<h2>Loss Functions and Optimizer</h2>\n<p>I simultaneously optimized 4 loss functions via weighted averaging: MSE, MAE, LogCosh, and BCE. The weights are 0.32, 0.24, 0.24, and 0.2, respectively.<br>\nThis was found to enhance the predictive performance of models. The Adam optimizer with learning rate 0.001 for LSTM and CNN, and 0.0003 for GRU was used to train the models. LogCosh is defined as:</p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n\n     ():\n        ey_t = (y_t - y_prime_t)/ \n         torch.mean(torch.log(torch.cosh(ey_t + )))\n</code></pre>\n<p>LogCosh is similar to MAE with the difference being that it is a softer version that can allow smoother convergence. It was adapted from <a href=\"https://github.com/tuantle/regression-losses-pytorch\" target=\"_blank\">https://github.com/tuantle/regression-losses-pytorch</a>.</p>\n<p>The BCE loss is indeed special as it is often used for classification tasks. However, I argue that it sends better signals to the models and optimizers when the target values are close to zero. To demonstrate this, consider the following two pieces of code:</p>\n<pre><code>m1 = nn.Sigmoid()\nloss = nn.BCELoss()\n = torch.tensor([], requires_grad=).unsqueeze()\ntarget = torch.sigmoid(torch.tensor([-], requires_grad=).unsqueeze())\noutput1 = loss(m1(), target)\n(output1.item()) \n\nm2 = nn.Identity()\nloss = nn.MSELoss()\n = torch.tensor([], requires_grad=).unsqueeze()\ntarget = torch.tensor([-], requires_grad=).unsqueeze()\noutput2 = loss(m2(), target)\n(output2.item())\n</code></pre>\n<p>With this example, one can observe that the MSELoss tells the model and optimizer that \"it is ok, there is no mistake here\". Obviously, there is a mistake, and BCELoss can see it as it returns a high loss value (0.694 compared to 0.010 for MSELoss). My choice of the BCELoss in this competition is motivated by the fact that most target values are from a Gaussian distribution with mean 0 as can be seen in the figure below.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3937164%2F7d9566cc821cac514ad9ebc4fd4e658d%2FGaussian.png?generation=1702287278847099&amp;alt=media\" alt=\"\"></p>\n<h2>Hyperparameters</h2>\n<ul>\n<li>250 epochs of training</li>\n<li>Learning rate 0.001 for LSTM and CNN, and 0.0003 for GRU</li>\n<li>Gradient norm clip value: [5.0, 1.0, 1.0] for the three schemes ''initial'', ''light'', and ''heavy''</li>\n</ul>\n<h1>4. Robustness</h1>\n<p>I conducted 4 experiments using different subsets of the training data, and monitored the private leaderboard score. I considered subsets of the initial training data (de_train) with sizes 25%, 50%, 75%, and 100%. Below 25%, we cannot cover all small molecules in the test set (id_map) even with a stratified split on sm_name, and hence the one hot encoding algorithm cannot run. With 25%, I achieved 0.946. With 50%, I achieved 0.815. With 75% of the training data, it is 0.769, and with the full data the private leaderboard is 0.719 (which is better than my winning submission because I removed padding in the ChemBERTa model). The figure below shows the robustness of my approach as a decreasing curve, i.e., improvement of the MRRMSE with increasing training data amount.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3937164%2F1eeb2e54c1c23a317402cb5c8ede967d%2Frobustness.png?generation=1702287304367920&amp;alt=media\" alt=\"\"></p>\n<p>My second data augmentation technique can be regarded as noise addition. I randomly replace 30% of the input features’ entries with zeros, and add the resulting input feature together with the correct target as a new training datapoint. This has proven to improve the predictive performance of my models. In this sense, my models are robust to the noise as their performance is not hindered but rather improved. The biological motivation here is that we might not need to know the complete chemical structure of a molecule (assuming the dropped input features are from sm_name) to know its impact on a cell. Similarly, there might be a biological disorder in a cell, and we would still expect that cell to respond to a molecule (drug) in the same way as a normal cell.</p>\n<p>Below is the data augmentation function</p>\n<pre><code> ():\n    copy_x = x_.copy()\n    new_x = []\n    new_y = y_.copy()\n    dim = x_.shape[]\n    k = (*dim)\n     i  (x_.shape[]):\n        idx = random.sample((dim), k=k)\n        copy_x[i,:,idx] = \n        new_x.append(copy_x[i])\n     np.stack(new_x, axis=), new_y\n</code></pre>\n<h1>5. Documentation and Code Style</h1>\n<p>The documentation and software dependencies are available on Github at  <a href=\"https://github.com/Jean-KOUAGOU/1st-place-solution-single-cell-pbs\" target=\"_blank\">https://github.com/Jean-KOUAGOU/1st-place-solution-single-cell-pbs</a></p>\n<h1>6. Reproducibility</h1>\n<p>Code is available and well documented on Github at <a href=\"https://github.com/Jean-KOUAGOU/1st-place-solution-single-cell-pbs\" target=\"_blank\">https://github.com/Jean-KOUAGOU/1st-place-solution-single-cell-pbs</a>. Reproduction scripts are added.</p>\n<h1>Sources</h1>\n<p><a href=\"https://www.nature.com/articles/s41597-019-0055-0\" target=\"_blank\">[1] BioWordVec, improving biomedical word embeddings with subword information and MeSH</a><br>\n<a href=\"https://github.com/tuantle/regression-losses-pytorch\" target=\"_blank\">Pytorch Regression Loss Functions</a><br>\n[ChemBERTa](<a href=\"https://huggingface.co/DeepChem/ChemBERTa-77M-MTR\" target=\"_blank\">https://huggingface.co/DeepChem/ChemBERTa-77M-MTR</a></p>",
      "rawMarkdown": "I would like to thank the organizers and Kaggle for hosting this exciting competition. I am also grateful to the participants who shared starter notebooks, datasets, and insightful ideas. Below is a more detailed writeup of my solution, including late findings.\n\n# Competition Page\nhttps://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\n\nhttps://openproblems.bio/\n\n# 1. Integration of Biological Knowledge\nSince the input features only consisted of pairs of short keywords, that is, cell types and small molecule names, and given the large size of the target variable, I was quickly convinced  that I needed to somehow enrich the input feature space. I therefore dedicated my first days of the competition to this task. First, I searched for biological word/term embeddings in the literature, and found the paper  ''BioWordVec, improving biomedical word embeddings with subword information and MeSH’’ by Zhang et al [1]. The paper directed me to the code on Github where I could find pretrained embeddings for biological terms. This was motivated by the fact that 1) I would be able to find most cell types, and small molecule names in such embeddings, and 2) The embeddings would encode rich information about the general meaning of each term. With these embeddings, I created larger input features and trained a regression model. This achieved 0.767 on the public leaderboard. With a better hyperparameter search and feature engineering, I improved the score to 0.614. As this seemed to be a good direction to go, I decided to further enrich the input features. This time, I searched for the definition of each cell type and small molecule name on wikipedia. For this, I used the python library ``wikipedia`` https://pypi.org/project/wikipedia/. I then represented each cell type and small molecule name by a few sentences describing it, then I bootstrapped an embedding from the descriptions. For example, Nk cells were described by: ''Natural killer cells, also known as NK cells or large granular lymphocytes (LGL), are a type of cytotoxic lymphocyte critical to the innate immune system that belong to the rapidly expanding family of known innate lymphoid cells (ILC) and represent 5–20% of all circulating lymphocytes in humans. The role of NK cells is analogous to that of cytotoxic T cells in the vertebrate adaptive immune response. NK cells provide rapid responses to virus-infected cell and other intracellular pathogens acting at around 3 days after infection, and respond to tumor formation.’’ I also explored different numbers of sentences to describe each cell type and small molecule.\nWhile this is interesting from the biological point of view, it did not improve the leaderboard score. In fact, the score became worse (0.656 vs. 0.614 previously). This can be explained by the fact that such natural language descriptions came with some noise, and pretrained embeddings were probably not computed to deal with this. Fine-tuning the embeddings on natural language descriptions of biological terms also fell short.\n\nBecause my initial idea about input feature enrichment did not meet my expectations, I decided to look for alternatives. Thanks to the discussions in the forum, I came across a notebook proposing to use SMILES to encode chemical structures of small molecules. I immediately decided to use ChemBERTa embeddings of SMILES encodings and observed a significant improvement in the evaluation metric MRRMSE on the validation data splits (I used a 5 fold cross-validation setting throughout the competition). With this, I developed additional data augmentation techniques, including the mean, standard deviation, and (25%, 50%, 75%) percentiles of differential expressions per cell type and small molecule in the training data.\n\n# 2. Exploration of the Problem\nAs mentioned in the previous section, I started the competition by trying to build rich features for the input pairs (cell_type, sm_name). Ultimately, the use of ChemBERTa features of small molecules’ SMILES appeared to be an important step towards this goal.  Combined with the mean, standard deviation, and (25%, 50%, 75%) percentiles per cell type and small molecule, I achieved an optimal input feature representation.\n\nIn my experiments, I used a 5-fold cross-validation setting with a fixed seed (42). It was hard to achieve a good score on the validation sets of the 2nd and 4th folds. On these folds, the MRRMSE on the validation set was approximately 1.19, and 1.15 on average, respectively. On the 1st, 3rd and 5th folds the average scores were 0.86, 0.86, and 0.90, respectively. The scores are the average across different model architectures (LSTM, 1d-CNN, GRU) and different input feature combinations (''initial’’, ''light’’, ''heavy’’). The three different input feature representations are as follows:\n“initial”: ChemBERTa embeddings, 1 hot encoding of cell_type/sm_name pairs, mean, std, percentiles of targets per cell_type and sm_name\n“light”: ChemBERTa embeddings, 1 hot encoding of cell_type/sm_name pairs, mean targets per cell_type and sm_name\n“heavy”: ChemBERTa embeddings, 1 hot encoding of cell_type/sm_name pairs, mean, 25%, 50%, 75% percentiles of targets per cell_type and sm_name\nThe figure below shows the training curves (MRRMSE) per fold averaged over all three model architectures.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3937164%2F8e7a5107e722900173e206496e4e8ee6%2FMRRMSE.png?generation=1702286999246540&alt=media)\n\nThe differences in the validation MRRMSE in the above figure motivated me to take a closer look into the validation sets, where I found different distributions of cell types. The figure below shows the predominant cell types per fold and the corresponding average (across models and different input feature representations) validation MRRMSE. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3937164%2F8e816183ed88332f09fb95731adcf775%2FHardCelltypes.png?generation=1702287123989655&alt=media)\n\nIn the validation sets of the 1st, 3rd, and 5th folds, the predominant cell types (in terms of percentage) are ''T regulatory cells’’, ''B cells’’, and ''Nk cells’’, respectively. On the 2nd and 4th folds, ''T cells CD8+’’ and ''Myeloid cells’’ are the most represented cell types in the validation sets, respectively. The percentage is computed as the number of occurrences of a cell type in the validation set divided by the number of occurrences of that cell type in the training set.\nFrom the bar-plots above, the cell types that are easier to predict are ''T regulatory cells’’, ''B cells’’, and ''Nk cells’’, while ''T cells CD8+’’ and ''Myeloid cells’’ are the hardest to predict. Based on this observation, an ideal training set should include more  ''T cells CD8+’’ and ''Myeloid cells’’ than the rest of the cell types. In this way, trained ML models would be able to generalize to other cell types.\n \n# 3. Model Design\n## Model Architecture\nI tried different model architectures, including gradient boosting models, MLP, and 2D CNN which did not work so well. I finally selected LSTM, GRU, and 1-d CNN architectures as they performed better on the validation sets. Below I show a rough implementation of the GRU model.\n\n```python\ndims_dict = {'conv': {'heavy': 13400, 'light': 4576, 'initial': 8992},\n                       'rnn': {'linear': {'heavy': 99968, 'light': 24192, 'initial': 29568},\n                                  'input_shape': {'heavy': [779,142], 'light': [187,202], 'initial': [229,324]}\n                     }}\nclass GRU(nn.Module):\n    def __init__(self, scheme):\n        super(GRU, self).__init__()\n        self.name = 'GRU'\n        self.scheme = scheme\n        self.gru = nn.GRU(dims_dict['rnn']['input_shape'][self.scheme][1], 128, num_layers=2, batch_first=True)\n        self.linear = nn.Sequential(\n            nn.Linear(dims_dict['rnn']['linear'][self.scheme], 1024),\n            nn.Dropout(0.3),\n            nn.ReLU(),\n            nn.Linear(1024, 512),\n            nn.Dropout(0.3),\n            nn.ReLU())\n        self.head = nn.Linear(512, 18211)\n        \n        self.loss1 = nn.MSELoss()\n        self.loss2 = LogCoshLoss()\n        self.loss3 = nn.L1Loss()\n        self.loss4 = nn.BCELoss()\n        \n    def forward(self, x, y=None):\n        shape1, shape2 = dims_dict['rnn']['input_shape'][self.scheme]\n        x = x.reshape(x.shape[0],shape1,shape2)\n        if y is None:\n            out, hn = self.gru(x)\n            out = out.reshape(out.shape[0],-1)\n            out = torch.cat([out, hn.reshape(hn.shape[1], -1)], dim=1)\n            out = self.head(self.linear(out))\n            return out\n        else:\n            out, hn = self.gru(x)\n            out = out.reshape(out.shape[0],-1)\n            out = torch.cat([out, hn.reshape(hn.shape[1], -1)], dim=1)\n            out = self.head(self.linear(out))\n            loss1 = 0.4*self.loss1(out, y) + 0.3*self.loss2(out, y) + 0.3*self.loss3(out, y)\n            yhat = torch.sigmoid(out)\n            yy = torch.sigmoid(y)\n            loss2 = self.loss4(yhat, yy)\n            return 0.8*loss1 + 0.2*loss2\n```\nIn my late experiments, I realized that 1d-CNN and GRU are actually the best architectures as they achieve the best scores alone (0.733  for GRU and 0.745 for 1d-CNN on Private LB). LSTM alone achieves 0.839 on Private LB. With 0.25xLSTM + 0.65xCNN the Private LB is 0.725, and with 0.25xLSTM + 0.65xGRU the Private LB is 0.723.\n\n## Loss Functions and Optimizer\nI simultaneously optimized 4 loss functions via weighted averaging: MSE, MAE, LogCosh, and BCE. The weights are 0.32, 0.24, 0.24, and 0.2, respectively.\nThis was found to enhance the predictive performance of models. The Adam optimizer with learning rate 0.001 for LSTM and CNN, and 0.0003 for GRU was used to train the models. LogCosh is defined as:\n\n```python\nclass LogCoshLoss(nn.Module):\n    def __init__(self):\n        super().__init__()\n\n    def forward(self, y_prime_t, y_t):\n        ey_t = (y_t - y_prime_t)/3 # divide by 3 to avoid numerical overflow in cosh\n        return torch.mean(torch.log(torch.cosh(ey_t + 1e-12)))\n```\nLogCosh is similar to MAE with the difference being that it is a softer version that can allow smoother convergence. It was adapted from https://github.com/tuantle/regression-losses-pytorch.\n\nThe BCE loss is indeed special as it is often used for classification tasks. However, I argue that it sends better signals to the models and optimizers when the target values are close to zero. To demonstrate this, consider the following two pieces of code:\n```python\nm1 = nn.Sigmoid()\nloss = nn.BCELoss()\ninput = torch.tensor([0.05], requires_grad=True).unsqueeze(0)\ntarget = torch.sigmoid(torch.tensor([-0.05], requires_grad=False).unsqueeze(0))\noutput1 = loss(m1(input), target)\nprint(output1.item()) # 0.694\n\nm2 = nn.Identity()\nloss = nn.MSELoss()\ninput = torch.tensor([0.05], requires_grad=True).unsqueeze(0)\ntarget = torch.tensor([-0.05], requires_grad=False).unsqueeze(0)\noutput2 = loss(m2(input), target)\nprint(output2.item())# 0.010\n```\nWith this example, one can observe that the MSELoss tells the model and optimizer that \"it is ok, there is no mistake here\". Obviously, there is a mistake, and BCELoss can see it as it returns a high loss value (0.694 compared to 0.010 for MSELoss). My choice of the BCELoss in this competition is motivated by the fact that most target values are from a Gaussian distribution with mean 0 as can be seen in the figure below.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3937164%2F7d9566cc821cac514ad9ebc4fd4e658d%2FGaussian.png?generation=1702287278847099&alt=media)\n## Hyperparameters\n- 250 epochs of training\n- Learning rate 0.001 for LSTM and CNN, and 0.0003 for GRU\n- Gradient norm clip value: [5.0, 1.0, 1.0] for the three schemes ''initial'', ''light'', and ''heavy''\n# 4. Robustness\nI conducted 4 experiments using different subsets of the training data, and monitored the private leaderboard score. I considered subsets of the initial training data (de_train) with sizes 25%, 50%, 75%, and 100%. Below 25%, we cannot cover all small molecules in the test set (id_map) even with a stratified split on sm_name, and hence the one hot encoding algorithm cannot run. With 25%, I achieved 0.946. With 50%, I achieved 0.815. With 75% of the training data, it is 0.769, and with the full data the private leaderboard is 0.719 (which is better than my winning submission because I removed padding in the ChemBERTa model). The figure below shows the robustness of my approach as a decreasing curve, i.e., improvement of the MRRMSE with increasing training data amount.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3937164%2F1eeb2e54c1c23a317402cb5c8ede967d%2Frobustness.png?generation=1702287304367920&alt=media)\n\nMy second data augmentation technique can be regarded as noise addition. I randomly replace 30% of the input features’ entries with zeros, and add the resulting input feature together with the correct target as a new training datapoint. This has proven to improve the predictive performance of my models. In this sense, my models are robust to the noise as their performance is not hindered but rather improved. The biological motivation here is that we might not need to know the complete chemical structure of a molecule (assuming the dropped input features are from sm_name) to know its impact on a cell. Similarly, there might be a biological disorder in a cell, and we would still expect that cell to respond to a molecule (drug) in the same way as a normal cell.\n\nBelow is the data augmentation function\n```python\ndef augment_data(x_, y_):\n    copy_x = x_.copy()\n    new_x = []\n    new_y = y_.copy()\n    dim = x_.shape[2]\n    k = int(0.3*dim)\n    for i in range(x_.shape[0]):\n        idx = random.sample(range(dim), k=k)\n        copy_x[i,:,idx] = 0\n        new_x.append(copy_x[i])\n    return np.stack(new_x, axis=0), new_y\n```\n# 5. Documentation and Code Style\nThe documentation and software dependencies are available on Github at  https://github.com/Jean-KOUAGOU/1st-place-solution-single-cell-pbs\n\n# 6. Reproducibility\nCode is available and well documented on Github at https://github.com/Jean-KOUAGOU/1st-place-solution-single-cell-pbs. Reproduction scripts are added.\n\n# Sources\n[[1] BioWordVec, improving biomedical word embeddings with subword information and MeSH](https://www.nature.com/articles/s41597-019-0055-0)\n[Pytorch Regression Loss Functions](https://github.com/tuantle/regression-losses-pytorch)\n[ChemBERTa](https://huggingface.co/DeepChem/ChemBERTa-77M-MTR",
      "votes": 50
    },
    {
      "id": 2559350,
      "postDate": "2023-12-12T18:53:26.113Z",
      "content": "<p>Reproduce 1st Place Private  Leaderboard Notebook: <a href=\"https://www.kaggle.com/code/jeannkouagou/1st-place-solution\" target=\"_blank\">https://www.kaggle.com/code/jeannkouagou/1st-place-solution</a><br>\nNote: Run ChemBERTa Model with random LM head (600-d) and no padding to obtain 0.719 on Private LB</p>",
      "rawMarkdown": "Reproduce 1st Place Private  Leaderboard Notebook: https://www.kaggle.com/code/jeannkouagou/1st-place-solution\nNote: Run ChemBERTa Model with random LM head (600-d) and no padding to obtain 0.719 on Private LB",
      "votes": 3
    },
    {
      "id": 2548712,
      "postDate": "2023-12-04T16:19:08.157Z",
      "content": "<p>Congratulations on the 1st place!  Thank you very much for sharing your detailed solution and what worked and what did not.  I see an exciting intersection of yours and 2nd one - using mean along with std of targets per cell_type and sm_name as features. </p>",
      "rawMarkdown": "Congratulations on the 1st place!  Thank you very much for sharing your detailed solution and what worked and what did not.  I see an exciting intersection of yours and 2nd one - using mean along with std of targets per cell_type and sm_name as features. ",
      "votes": 4,
      "replies": [
        {
          "id": 2548764,
          "postDate": "2023-12-04T17:04:04.113Z",
          "content": "<p>Thank you Makio. Indeed, there is an intersection, which is interesting!</p>",
          "rawMarkdown": "Thank you Makio. Indeed, there is an intersection, which is interesting!",
          "votes": 3,
          "replies": [
            {
              "id": 2548991,
              "postDate": "2023-12-04T23:15:42.863Z",
              "content": "<p>I wish you success in future competitions as well!</p>",
              "rawMarkdown": "I wish you success in future competitions as well!",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2570783,
      "postDate": "2023-12-22T15:24:00.693Z",
      "content": "<p>Clear code in github and clear explanation in the post! Congrates!</p>",
      "rawMarkdown": "Clear code in github and clear explanation in the post! Congrates!",
      "votes": 1,
      "replies": [
        {
          "id": 2571151,
          "postDate": "2023-12-23T00:01:53.473Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/zechengyin\" target=\"_blank\">@zechengyin</a> </p>",
          "rawMarkdown": "Thank you @zechengyin "
        }
      ]
    },
    {
      "id": 2559208,
      "postDate": "2023-12-12T16:50:01.090Z",
      "content": "<p>Congratulations with the victory ! <br>\nWould you be so kind to share your submission files for further research purposes , please.</p>",
      "rawMarkdown": "Congratulations with the victory ! \nWould you be so kind to share your submission files for further research purposes , please.",
      "votes": 1,
      "replies": [
        {
          "id": 2559270,
          "postDate": "2023-12-12T17:29:44.087Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a>. Thanks. Here are the submissions: <a href=\"https://www.kaggle.com/code/jeannkouagou/submit-best\" target=\"_blank\">https://www.kaggle.com/code/jeannkouagou/submit-best</a></p>",
          "rawMarkdown": "Hi @alexandervc. Thanks. Here are the submissions: https://www.kaggle.com/code/jeannkouagou/submit-best",
          "votes": 2,
          "replies": [
            {
              "id": 2559345,
              "postDate": "2023-12-12T18:50:21.610Z",
              "content": "<p>Thanks a lot ! </p>",
              "rawMarkdown": "Thanks a lot ! ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2551860,
      "postDate": "2023-12-07T02:55:57.800Z",
      "content": "<p>class LogCoshLoss(nn.Module):<br>\n    def <strong>init</strong>(self):<br>\n        super().<strong>init</strong>()</p>\n<pre><code> ( - )/ \n     torch.(torch.(torch.( + )))\n</code></pre>\n<p>Is 1e-12 here misplaced?</p>",
      "rawMarkdown": "class LogCoshLoss(nn.Module):\n    def __init__(self):\n        super().__init__()\n\n    def forward(self, y_prime_t, y_t):\n        ey_t = (y_t - y_prime_t)/3 # divide by 3 to avoid numerical overflow in cosh\n        return torch.mean(torch.log(torch.cosh(ey_t + 1e-12)))\nIs 1e-12 here misplaced?",
      "votes": 1,
      "replies": [
        {
          "id": 2552161,
          "postDate": "2023-12-07T08:55:34.383Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jingfengou\" target=\"_blank\">@jingfengou</a>. Adding 1e-12 is not necessary. But it can act as a small regularizer</p>",
          "rawMarkdown": "Hi @jingfengou. Adding 1e-12 is not necessary. But it can act as a small regularizer",
          "votes": 1
        }
      ]
    },
    {
      "id": 2549049,
      "postDate": "2023-12-05T01:01:15.500Z",
      "content": "<p>Congrats, what's the rationale behind the sigmoid output head?</p>",
      "rawMarkdown": "Congrats, what's the rationale behind the sigmoid output head?",
      "votes": 1,
      "replies": [
        {
          "id": 2549480,
          "postDate": "2023-12-05T09:42:37.513Z",
          "content": "<p>Thanks for the question. Maybe I was not precise enough.<br>\nIn Section \"Loss Functions and Optimizers\", I wrote <code>The BCE loss is indeed special as it is often used for classification tasks. However, I argue that it helps models better learn the signs of the predicted values as they are mostly close to zero.</code></p>\n<p>I actually meant that the BCE loss applied on the sigmoid of predictions and targets returns a higher loss (i.e., sends a better signal to the optimizer) when the values involved are close to zero; indeed, most values are close to zero, see plot above. To demonstrate this, let's consider the two pieces of code below:</p>",
          "rawMarkdown": "Thanks for the question. Maybe I was not precise enough.\nIn Section \"Loss Functions and Optimizers\", I wrote `The BCE loss is indeed special as it is often used for classification tasks. However, I argue that it helps models better learn the signs of the predicted values as they are mostly close to zero.`\n\nI actually meant that the BCE loss applied on the sigmoid of predictions and targets returns a higher loss (i.e., sends a better signal to the optimizer) when the values involved are close to zero; indeed, most values are close to zero, see plot above. To demonstrate this, let's consider the two pieces of code below:",
          "votes": 2,
          "replies": [
            {
              "id": 2549489,
              "postDate": "2023-12-05T09:46:40.230Z",
              "content": "<pre><code>m1 = nn.Sigmoid()\nloss = nn.BCELoss()\n = torch.tensor([], requires_grad=).unsqueeze()\ntarget = torch.sigmoid(torch.tensor([-], requires_grad=).unsqueeze())\noutput1 = loss(m1(), target)\n(output1.item()) \n\nm2 = nn.Identity()\nloss = nn.MSELoss()\n = torch.tensor([], requires_grad=).unsqueeze()\ntarget = torch.tensor([-], requires_grad=).unsqueeze()\noutput2 = loss(m2(), target)\n(output2.item())\n</code></pre>\n<p>With this, one can observe that the MSELoss tells the model and optimizer that \"it is ok, there is no mistake here\"<br>\nOppositely, the BCELoss considers this to be a huge mistake.</p>",
              "rawMarkdown": "``` python\nm1 = nn.Sigmoid()\nloss = nn.BCELoss()\ninput = torch.tensor([0.05], requires_grad=True).unsqueeze(0)\ntarget = torch.sigmoid(torch.tensor([-0.05], requires_grad=False).unsqueeze(0))\noutput1 = loss(m1(input), target)\nprint(output1.item()) # 0.694\n\nm2 = nn.Identity()\nloss = nn.MSELoss()\ninput = torch.tensor([0.05], requires_grad=True).unsqueeze(0)\ntarget = torch.tensor([-0.05], requires_grad=False).unsqueeze(0)\noutput2 = loss(m2(input), target)\nprint(output2.item())# 0.010\n```\n\nWith this, one can observe that the MSELoss tells the model and optimizer that \"it is ok, there is no mistake here\"\nOppositely, the BCELoss considers this to be a huge mistake.",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 2548818,
      "postDate": "2023-12-04T18:07:29.473Z",
      "content": "<p>Congratulations! Excellent work</p>",
      "rawMarkdown": "Congratulations! Excellent work",
      "votes": 2,
      "replies": [
        {
          "id": 2548872,
          "postDate": "2023-12-04T19:45:00.590Z",
          "content": "<p>Thank you Miraj. I appreciate </p>",
          "rawMarkdown": "Thank you Miraj. I appreciate "
        }
      ]
    },
    {
      "id": 3024339,
      "postDate": "2024-10-21T14:33:25.280Z",
      "content": "<p>Why do you think CNN was a good choice for this task</p>",
      "rawMarkdown": "Why do you think CNN was a good choice for this task"
    },
    {
      "id": 2711824,
      "postDate": "2024-03-23T05:34:19.617Z",
      "content": "<p>Good work, congrats!</p>",
      "rawMarkdown": "Good work, congrats!"
    },
    {
      "id": 2581202,
      "postDate": "2023-12-31T14:53:17.403Z",
      "content": "<p>Congrates sir</p>",
      "rawMarkdown": "Congrates sir",
      "replies": [
        {
          "id": 2595912,
          "postDate": "2024-01-10T18:07:29.550Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/abhishekgupta18895\" target=\"_blank\">@abhishekgupta18895</a> </p>",
          "rawMarkdown": "Thank you @abhishekgupta18895 "
        }
      ]
    },
    {
      "id": 2559109,
      "postDate": "2023-12-12T16:04:22.293Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 2559184,
          "postDate": "2023-12-12T16:36:51.283Z",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/manavtrivedi\" target=\"_blank\">@manavtrivedi</a> </p>",
          "rawMarkdown": "Thanks a lot @manavtrivedi ",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2559350,
      "author_name": "JK-Piece",
      "author_url": "",
      "post_date": "2023-12-12T18:53:26.113000",
      "content": "<p>Reproduce 1st Place Private  Leaderboard Notebook: <a href=\"https://www.kaggle.com/code/jeannkouagou/1st-place-solution\" target=\"_blank\">https://www.kaggle.com/code/jeannkouagou/1st-place-solution</a><br>\nNote: Run ChemBERTa Model with random LM head (600-d) and no padding to obtain 0.719 on Private LB</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2548712,
      "author_name": "makio323",
      "author_url": "",
      "post_date": "2023-12-04T16:19:08.157000",
      "content": "<p>Congratulations on the 1st place!  Thank you very much for sharing your detailed solution and what worked and what did not.  I see an exciting intersection of yours and 2nd one - using mean along with std of targets per cell_type and sm_name as features. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 2548764,
          "author_name": "JK-Piece",
          "author_url": "",
          "post_date": "2023-12-04T17:04:04.113000",
          "content": "<p>Thank you Makio. Indeed, there is an intersection, which is interesting!</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2548991,
              "author_name": "makio323",
              "author_url": "",
              "post_date": "2023-12-04T23:15:42.863000",
              "content": "<p>I wish you success in future competitions as well!</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2570783,
      "author_name": "Yonggie",
      "author_url": "",
      "post_date": "2023-12-22T15:24:00.693000",
      "content": "<p>Clear code in github and clear explanation in the post! Congrates!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2571151,
          "author_name": "JK-Piece",
          "author_url": "",
          "post_date": "2023-12-23T00:01:53.473000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/zechengyin\" target=\"_blank\">@zechengyin</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2559208,
      "author_name": "Alexander Chervov",
      "author_url": "",
      "post_date": "2023-12-12T16:50:01.090000",
      "content": "<p>Congratulations with the victory ! <br>\nWould you be so kind to share your submission files for further research purposes , please.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2559270,
          "author_name": "JK-Piece",
          "author_url": "",
          "post_date": "2023-12-12T17:29:44.087000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a>. Thanks. Here are the submissions: <a href=\"https://www.kaggle.com/code/jeannkouagou/submit-best\" target=\"_blank\">https://www.kaggle.com/code/jeannkouagou/submit-best</a></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2559345,
              "author_name": "Alexander Chervov",
              "author_url": "",
              "post_date": "2023-12-12T18:50:21.610000",
              "content": "<p>Thanks a lot ! </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2551860,
      "author_name": "jingfeng ou",
      "author_url": "",
      "post_date": "2023-12-07T02:55:57.800000",
      "content": "<p>class LogCoshLoss(nn.Module):<br>\n    def <strong>init</strong>(self):<br>\n        super().<strong>init</strong>()</p>\n<pre><code> ( - )/ \n     torch.(torch.(torch.( + )))\n</code></pre>\n<p>Is 1e-12 here misplaced?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2552161,
          "author_name": "JK-Piece",
          "author_url": "",
          "post_date": "2023-12-07T08:55:34.383000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jingfengou\" target=\"_blank\">@jingfengou</a>. Adding 1e-12 is not necessary. But it can act as a small regularizer</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2549049,
      "author_name": "maxleverage",
      "author_url": "",
      "post_date": "2023-12-05T01:01:15.500000",
      "content": "<p>Congrats, what's the rationale behind the sigmoid output head?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2549480,
          "author_name": "JK-Piece",
          "author_url": "",
          "post_date": "2023-12-05T09:42:37.513000",
          "content": "<p>Thanks for the question. Maybe I was not precise enough.<br>\nIn Section \"Loss Functions and Optimizers\", I wrote <code>The BCE loss is indeed special as it is often used for classification tasks. However, I argue that it helps models better learn the signs of the predicted values as they are mostly close to zero.</code></p>\n<p>I actually meant that the BCE loss applied on the sigmoid of predictions and targets returns a higher loss (i.e., sends a better signal to the optimizer) when the values involved are close to zero; indeed, most values are close to zero, see plot above. To demonstrate this, let's consider the two pieces of code below:</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2549489,
              "author_name": "JK-Piece",
              "author_url": "",
              "post_date": "2023-12-05T09:46:40.230000",
              "content": "<pre><code>m1 = nn.Sigmoid()\nloss = nn.BCELoss()\n = torch.tensor([], requires_grad=).unsqueeze()\ntarget = torch.sigmoid(torch.tensor([-], requires_grad=).unsqueeze())\noutput1 = loss(m1(), target)\n(output1.item()) \n\nm2 = nn.Identity()\nloss = nn.MSELoss()\n = torch.tensor([], requires_grad=).unsqueeze()\ntarget = torch.tensor([-], requires_grad=).unsqueeze()\noutput2 = loss(m2(), target)\n(output2.item())\n</code></pre>\n<p>With this, one can observe that the MSELoss tells the model and optimizer that \"it is ok, there is no mistake here\"<br>\nOppositely, the BCELoss considers this to be a huge mistake.</p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2548818,
      "author_name": "Meraj Hasan",
      "author_url": "",
      "post_date": "2023-12-04T18:07:29.473000",
      "content": "<p>Congratulations! Excellent work</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2548872,
          "author_name": "JK-Piece",
          "author_url": "",
          "post_date": "2023-12-04T19:45:00.590000",
          "content": "<p>Thank you Miraj. I appreciate </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3024339,
      "author_name": "Ankush H V",
      "author_url": "",
      "post_date": "2024-10-21T14:33:25.280000",
      "content": "<p>Why do you think CNN was a good choice for this task</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2711824,
      "author_name": "Vinit Shah",
      "author_url": "",
      "post_date": "2024-03-23T05:34:19.617000",
      "content": "<p>Good work, congrats!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2581202,
      "author_name": "Abhishek Gupta",
      "author_url": "",
      "post_date": "2023-12-31T14:53:17.403000",
      "content": "<p>Congrates sir</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2595912,
          "author_name": "JK-Piece",
          "author_url": "",
          "post_date": "2024-01-10T18:07:29.550000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/abhishekgupta18895\" target=\"_blank\">@abhishekgupta18895</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2559109,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-12T16:04:22.293000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2559184,
          "author_name": "JK-Piece",
          "author_url": "",
          "post_date": "2023-12-12T16:36:51.283000",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/manavtrivedi\" target=\"_blank\">@manavtrivedi</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2548444": "I would like to thank the organizers and Kaggle for hosting this exciting competition. I am also grateful to the participants who shared starter notebooks, datasets, and insightful ideas. Below is a more detailed writeup of my solution, including late findings.\n\n# Competition Page\nhttps://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\n\nhttps://openproblems.bio/\n\n# 1. Integration of Biological Knowledge\nSince the input features only consisted of pairs of short keywords, that is, cell types and small molecule names, and given the large size of the target variable, I was quickly convinced  that I needed to somehow enrich the input feature space. I therefore dedicated my first days of the competition to this task. First, I searched for biological word/term embeddings in the literature, and found the paper  ''BioWordVec, improving biomedical word embeddings with subword information and MeSH’’ by Zhang et al [1]. The paper directed me to the code on Github where I could find pretrained embeddings for biological terms. This was motivated by the fact that 1) I would be able to find most cell types, and small molecule names in such embeddings, and 2) The embeddings would encode rich information about the general meaning of each term. With these embeddings, I created larger input features and trained a regression model. This achieved 0.767 on the public leaderboard. With a better hyperparameter search and feature engineering, I improved the score to 0.614. As this seemed to be a good direction to go, I decided to further enrich the input features. This time, I searched for the definition of each cell type and small molecule name on wikipedia. For this, I used the python library ``wikipedia`` https://pypi.org/project/wikipedia/. I then represented each cell type and small molecule name by a few sentences describing it, then I bootstrapped an embedding from the descriptions. For example, Nk cells were described by: ''Natural killer cells, also known as NK cells or large granular lymphocytes (LGL), are a type of cytotoxic lymphocyte critical to the innate immune system that belong to the rapidly expanding family of known innate lymphoid cells (ILC) and represent 5–20% of all circulating lymphocytes in humans. The role of NK cells is analogous to that of cytotoxic T cells in the vertebrate adaptive immune response. NK cells provide rapid responses to virus-infected cell and other intracellular pathogens acting at around 3 days after infection, and respond to tumor formation.’’ I also explored different numbers of sentences to describe each cell type and small molecule.\nWhile this is interesting from the biological point of view, it did not improve the leaderboard score. In fact, the score became worse (0.656 vs. 0.614 previously). This can be explained by the fact that such natural language descriptions came with some noise, and pretrained embeddings were probably not computed to deal with this. Fine-tuning the embeddings on natural language descriptions of biological terms also fell short.\n\nBecause my initial idea about input feature enrichment did not meet my expectations, I decided to look for alternatives. Thanks to the discussions in the forum, I came across a notebook proposing to use SMILES to encode chemical structures of small molecules. I immediately decided to use ChemBERTa embeddings of SMILES encodings and observed a significant improvement in the evaluation metric MRRMSE on the validation data splits (I used a 5 fold cross-validation setting throughout the competition). With this, I developed additional data augmentation techniques, including the mean, standard deviation, and (25%, 50%, 75%) percentiles of differential expressions per cell type and small molecule in the training data.\n\n# 2. Exploration of the Problem\nAs mentioned in the previous section, I started the competition by trying to build rich features for the input pairs (cell_type, sm_name). Ultimately, the use of ChemBERTa features of small molecules’ SMILES appeared to be an important step towards this goal.  Combined with the mean, standard deviation, and (25%, 50%, 75%) percentiles per cell type and small molecule, I achieved an optimal input feature representation.\n\nIn my experiments, I used a 5-fold cross-validation setting with a fixed seed (42). It was hard to achieve a good score on the validation sets of the 2nd and 4th folds. On these folds, the MRRMSE on the validation set was approximately 1.19, and 1.15 on average, respectively. On the 1st, 3rd and 5th folds the average scores were 0.86, 0.86, and 0.90, respectively. The scores are the average across different model architectures (LSTM, 1d-CNN, GRU) and different input feature combinations (''initial’’, ''light’’, ''heavy’’). The three different input feature representations are as follows:\n“initial”: ChemBERTa embeddings, 1 hot encoding of cell_type/sm_name pairs, mean, std, percentiles of targets per cell_type and sm_name\n“light”: ChemBERTa embeddings, 1 hot encoding of cell_type/sm_name pairs, mean targets per cell_type and sm_name\n“heavy”: ChemBERTa embeddings, 1 hot encoding of cell_type/sm_name pairs, mean, 25%, 50%, 75% percentiles of targets per cell_type and sm_name\nThe figure below shows the training curves (MRRMSE) per fold averaged over all three model architectures.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3937164%2F8e7a5107e722900173e206496e4e8ee6%2FMRRMSE.png?generation=1702286999246540&alt=media)\n\nThe differences in the validation MRRMSE in the above figure motivated me to take a closer look into the validation sets, where I found different distributions of cell types. The figure below shows the predominant cell types per fold and the corresponding average (across models and different input feature representations) validation MRRMSE. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3937164%2F8e816183ed88332f09fb95731adcf775%2FHardCelltypes.png?generation=1702287123989655&alt=media)\n\nIn the validation sets of the 1st, 3rd, and 5th folds, the predominant cell types (in terms of percentage) are ''T regulatory cells’’, ''B cells’’, and ''Nk cells’’, respectively. On the 2nd and 4th folds, ''T cells CD8+’’ and ''Myeloid cells’’ are the most represented cell types in the validation sets, respectively. The percentage is computed as the number of occurrences of a cell type in the validation set divided by the number of occurrences of that cell type in the training set.\nFrom the bar-plots above, the cell types that are easier to predict are ''T regulatory cells’’, ''B cells’’, and ''Nk cells’’, while ''T cells CD8+’’ and ''Myeloid cells’’ are the hardest to predict. Based on this observation, an ideal training set should include more  ''T cells CD8+’’ and ''Myeloid cells’’ than the rest of the cell types. In this way, trained ML models would be able to generalize to other cell types.\n \n# 3. Model Design\n## Model Architecture\nI tried different model architectures, including gradient boosting models, MLP, and 2D CNN which did not work so well. I finally selected LSTM, GRU, and 1-d CNN architectures as they performed better on the validation sets. Below I show a rough implementation of the GRU model.\n\n```python\ndims_dict = {'conv': {'heavy': 13400, 'light': 4576, 'initial': 8992},\n                       'rnn': {'linear': {'heavy': 99968, 'light': 24192, 'initial': 29568},\n                                  'input_shape': {'heavy': [779,142], 'light': [187,202], 'initial': [229,324]}\n                     }}\nclass GRU(nn.Module):\n    def __init__(self, scheme):\n        super(GRU, self).__init__()\n        self.name = 'GRU'\n        self.scheme = scheme\n        self.gru = nn.GRU(dims_dict['rnn']['input_shape'][self.scheme][1], 128, num_layers=2, batch_first=True)\n        self.linear = nn.Sequential(\n            nn.Linear(dims_dict['rnn']['linear'][self.scheme], 1024),\n            nn.Dropout(0.3),\n            nn.ReLU(),\n            nn.Linear(1024, 512),\n            nn.Dropout(0.3),\n            nn.ReLU())\n        self.head = nn.Linear(512, 18211)\n        \n        self.loss1 = nn.MSELoss()\n        self.loss2 = LogCoshLoss()\n        self.loss3 = nn.L1Loss()\n        self.loss4 = nn.BCELoss()\n        \n    def forward(self, x, y=None):\n        shape1, shape2 = dims_dict['rnn']['input_shape'][self.scheme]\n        x = x.reshape(x.shape[0],shape1,shape2)\n        if y is None:\n            out, hn = self.gru(x)\n            out = out.reshape(out.shape[0],-1)\n            out = torch.cat([out, hn.reshape(hn.shape[1], -1)], dim=1)\n            out = self.head(self.linear(out))\n            return out\n        else:\n            out, hn = self.gru(x)\n            out = out.reshape(out.shape[0],-1)\n            out = torch.cat([out, hn.reshape(hn.shape[1], -1)], dim=1)\n            out = self.head(self.linear(out))\n            loss1 = 0.4*self.loss1(out, y) + 0.3*self.loss2(out, y) + 0.3*self.loss3(out, y)\n            yhat = torch.sigmoid(out)\n            yy = torch.sigmoid(y)\n            loss2 = self.loss4(yhat, yy)\n            return 0.8*loss1 + 0.2*loss2\n```\nIn my late experiments, I realized that 1d-CNN and GRU are actually the best architectures as they achieve the best scores alone (0.733  for GRU and 0.745 for 1d-CNN on Private LB). LSTM alone achieves 0.839 on Private LB. With 0.25xLSTM + 0.65xCNN the Private LB is 0.725, and with 0.25xLSTM + 0.65xGRU the Private LB is 0.723.\n\n## Loss Functions and Optimizer\nI simultaneously optimized 4 loss functions via weighted averaging: MSE, MAE, LogCosh, and BCE. The weights are 0.32, 0.24, 0.24, and 0.2, respectively.\nThis was found to enhance the predictive performance of models. The Adam optimizer with learning rate 0.001 for LSTM and CNN, and 0.0003 for GRU was used to train the models. LogCosh is defined as:\n\n```python\nclass LogCoshLoss(nn.Module):\n    def __init__(self):\n        super().__init__()\n\n    def forward(self, y_prime_t, y_t):\n        ey_t = (y_t - y_prime_t)/3 # divide by 3 to avoid numerical overflow in cosh\n        return torch.mean(torch.log(torch.cosh(ey_t + 1e-12)))\n```\nLogCosh is similar to MAE with the difference being that it is a softer version that can allow smoother convergence. It was adapted from https://github.com/tuantle/regression-losses-pytorch.\n\nThe BCE loss is indeed special as it is often used for classification tasks. However, I argue that it sends better signals to the models and optimizers when the target values are close to zero. To demonstrate this, consider the following two pieces of code:\n```python\nm1 = nn.Sigmoid()\nloss = nn.BCELoss()\ninput = torch.tensor([0.05], requires_grad=True).unsqueeze(0)\ntarget = torch.sigmoid(torch.tensor([-0.05], requires_grad=False).unsqueeze(0))\noutput1 = loss(m1(input), target)\nprint(output1.item()) # 0.694\n\nm2 = nn.Identity()\nloss = nn.MSELoss()\ninput = torch.tensor([0.05], requires_grad=True).unsqueeze(0)\ntarget = torch.tensor([-0.05], requires_grad=False).unsqueeze(0)\noutput2 = loss(m2(input), target)\nprint(output2.item())# 0.010\n```\nWith this example, one can observe that the MSELoss tells the model and optimizer that \"it is ok, there is no mistake here\". Obviously, there is a mistake, and BCELoss can see it as it returns a high loss value (0.694 compared to 0.010 for MSELoss). My choice of the BCELoss in this competition is motivated by the fact that most target values are from a Gaussian distribution with mean 0 as can be seen in the figure below.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3937164%2F7d9566cc821cac514ad9ebc4fd4e658d%2FGaussian.png?generation=1702287278847099&alt=media)\n## Hyperparameters\n- 250 epochs of training\n- Learning rate 0.001 for LSTM and CNN, and 0.0003 for GRU\n- Gradient norm clip value: [5.0, 1.0, 1.0] for the three schemes ''initial'', ''light'', and ''heavy''\n# 4. Robustness\nI conducted 4 experiments using different subsets of the training data, and monitored the private leaderboard score. I considered subsets of the initial training data (de_train) with sizes 25%, 50%, 75%, and 100%. Below 25%, we cannot cover all small molecules in the test set (id_map) even with a stratified split on sm_name, and hence the one hot encoding algorithm cannot run. With 25%, I achieved 0.946. With 50%, I achieved 0.815. With 75% of the training data, it is 0.769, and with the full data the private leaderboard is 0.719 (which is better than my winning submission because I removed padding in the ChemBERTa model). The figure below shows the robustness of my approach as a decreasing curve, i.e., improvement of the MRRMSE with increasing training data amount.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3937164%2F1eeb2e54c1c23a317402cb5c8ede967d%2Frobustness.png?generation=1702287304367920&alt=media)\n\nMy second data augmentation technique can be regarded as noise addition. I randomly replace 30% of the input features’ entries with zeros, and add the resulting input feature together with the correct target as a new training datapoint. This has proven to improve the predictive performance of my models. In this sense, my models are robust to the noise as their performance is not hindered but rather improved. The biological motivation here is that we might not need to know the complete chemical structure of a molecule (assuming the dropped input features are from sm_name) to know its impact on a cell. Similarly, there might be a biological disorder in a cell, and we would still expect that cell to respond to a molecule (drug) in the same way as a normal cell.\n\nBelow is the data augmentation function\n```python\ndef augment_data(x_, y_):\n    copy_x = x_.copy()\n    new_x = []\n    new_y = y_.copy()\n    dim = x_.shape[2]\n    k = int(0.3*dim)\n    for i in range(x_.shape[0]):\n        idx = random.sample(range(dim), k=k)\n        copy_x[i,:,idx] = 0\n        new_x.append(copy_x[i])\n    return np.stack(new_x, axis=0), new_y\n```\n# 5. Documentation and Code Style\nThe documentation and software dependencies are available on Github at  https://github.com/Jean-KOUAGOU/1st-place-solution-single-cell-pbs\n\n# 6. Reproducibility\nCode is available and well documented on Github at https://github.com/Jean-KOUAGOU/1st-place-solution-single-cell-pbs. Reproduction scripts are added.\n\n# Sources\n[[1] BioWordVec, improving biomedical word embeddings with subword information and MeSH](https://www.nature.com/articles/s41597-019-0055-0)\n[Pytorch Regression Loss Functions](https://github.com/tuantle/regression-losses-pytorch)\n[ChemBERTa](https://huggingface.co/DeepChem/ChemBERTa-77M-MTR",
    "2559350": "Reproduce 1st Place Private  Leaderboard Notebook: https://www.kaggle.com/code/jeannkouagou/1st-place-solution\nNote: Run ChemBERTa Model with random LM head (600-d) and no padding to obtain 0.719 on Private LB",
    "2548712": "Congratulations on the 1st place!  Thank you very much for sharing your detailed solution and what worked and what did not.  I see an exciting intersection of yours and 2nd one - using mean along with std of targets per cell_type and sm_name as features. ",
    "2570783": "Clear code in github and clear explanation in the post! Congrates!",
    "2559208": "Congratulations with the victory ! \nWould you be so kind to share your submission files for further research purposes , please.",
    "2551860": "class LogCoshLoss(nn.Module):\n    def __init__(self):\n        super().__init__()\n\n    def forward(self, y_prime_t, y_t):\n        ey_t = (y_t - y_prime_t)/3 # divide by 3 to avoid numerical overflow in cosh\n        return torch.mean(torch.log(torch.cosh(ey_t + 1e-12)))\nIs 1e-12 here misplaced?",
    "2549049": "Congrats, what's the rationale behind the sigmoid output head?",
    "2548818": "Congratulations! Excellent work",
    "3024339": "Why do you think CNN was a good choice for this task",
    "2711824": "Good work, congrats!",
    "2581202": "Congrates sir",
    "2559109": ""
  }
}