{
  "id": 461363,
  "title": "6th Place Solution for the Open Problems – Single-Cell Perturbations",
  "url": "/competitions/open-problems-single-cell-perturbations/writeups/xlearning-scu-6th-place-solution-for-the-open-prob",
  "author_name": "",
  "post_date": "2023-12-14T11:55:24.487Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<h2>1. Integration of biological knowledge</h2>\n<p>I tried integrating different biological knowledge, but unfortunately, most didn't work very well.</p>\n<p>Firstly, I came up with the idea of using the GO term to reduce genes into modules. I adopted the KEGG 2016 genes set and filtered out gene sets with a P-value less than 0.05. However, genes from the same set exhibit considerably different differential expression (DE) values. As a result, I feel the hypothesis that genes from the same GO term would have similar DE values cannot stand. Then, I resorted to finding co-expressed genes based on the gene expression matrix. I found the DE values between co-expressed genes do have a high Pearson correlation. However, I checked the predicted DE values of co-expressed genes and found they yield strong Pearson correlations as well. In other words, adding the correlation constraints to the optimization process won't have much effect.</p>\n<p>Next is about the molecule representation. I tried using the features from the pre-trained molecule models including ChemBERT and MolCLR, as well as the TFIDF-transformed element count proposed in the public notebooks. In my experiments, using the pre-trained features led to slightly inferior results, while the TFIDF-transformed element count and simply learnable molecule features give similar performance. I wonder if it is because the pre-training models capture more coarse-grained differences between molecules, or because the molecules in the contest are quite different from those used for pre-training.</p>\n<p>Finally is the integration of ATAC data. I did not have a good idea of integrating ATAC data into the model. All that I had come up with is that genes corresponding to a small ATAC count would be less affected by molecules. Anyway, I am not sure about such a hypothesis, and in practice, I simply concatenated the ATAC feature and found it slightly improves the results.</p>\n<h2>2. Exploration of the problem</h2>\n<p>According to the problem definition, I think three targets could be predicted to achieve the task. The first target is directly the DE values. The variables in DE value prediction would be cell types and compounds. A model could be trained to predict the DE value given cell type-compound pairs. The second target is the bulk expression values, namely, the input to the Limma model. The variables in bulk prediction would be only the compounds (as I will elaborate on later). I feel such a paradigm is the most promising and robust solution. The most important reason is that the results could be precisely validated. To be specific, we could send the predicted and provided bulk expressions to the Limma model, and check how close the predicted DE values are to the provided ones. I feel such a validation scheme would be much less risks in overfitting. The third target is the single-cell gene expression. Though the training data is most sufficient in such a paradigm, it is the most challenging paradigm since we do not have the paired before-perturbation and after-perturbation gene expression values.</p>\n<h2>3. Model design</h2>\n<p>In this competition, I mainly explored the first two prediction paradigms, namely the DE values and bulk expression values.</p>\n<h3>3.1 DE value prediction</h3>\n<h4>3.1.1 Preprocess</h4>\n<p>No data preprocessing is applied. Common preprocessing strategies such as normalizing and scaling even harm the performance in my experience.</p>\n<h4>3.1.2 Model architecture</h4>\n<p>The plainest model learns the cell type and compound representation, as well as makes the predictions at the same time. The model could be written within a few lines, namely,</p>\n<pre><code> (nn.Module):\n     ():\n        (Net, self).__init__()\n        self.type_num = type_num\n        self.compound_num = compound_num\n        self.gene_num = gene_num\n\n        self.type_embedding = nn.Embedding(self.type_num, )\n        self.compound_embedding = nn.Embedding(self.compound_num, )\n\n        self.predictor = nn.Sequential(\n            nn.Linear(, ),\n            nn.BatchNorm1d(),\n            nn.Dropout(),\n            nn.ReLU(),\n            nn.Linear(, ),\n            nn.BatchNorm1d(),\n            nn.Dropout(),\n            nn.ReLU(),\n            nn.Linear(, self.gene_num),\n        )\n\n     ():\n        type_embedding = self.type_embedding()\n        compound_embedding = self.compound_embedding(compound)\n        embedding = torch.cat([type_embedding, compound_embedding], dim=)\n         self.predictor(embedding)\n</code></pre>\n<p>Optionally, we could replace the learnable cell type representation with RNA or ATAC counts averaged by cell type, and replace the compound representation with features from pre-trained models such as ChemBERT and MolCLR. Training such a vanilla model gives 0.586 MRRMSE on the public test split and 0.785 MRRMSE on the private test split.</p>\n<p>I also tried Transformer architecture like the Performer used in the scBERT paper. The idea is to treat each gene as a token and replace the positional embedding with cell type and compound representation. The token of each gene is learned together with the Transformer model. Due to the space limitation, I will not attach the model configuration code here, please refer to the GitHub repository attached below. Sadly, though the Transformer architecture generally gives better performance, in my practice, I just cannot successfully train the model (i.e., the training loss stays very high), and achieved 0.611 MRRMSE on the public test split and 0.767 MRRMSE on the private test split. It was not until the competition ended that I found the Transformer architecture achieved a much better score on the private test split, despite its poor performance on the public test split. Such an interesting result is worth exploring (or it is simply due to randomness and luck).</p>\n<h4>3.1.3 Learning objective</h4>\n<p>To directly predict the DE values, I tried the MSE loss, L1 loss, and MRRMSE loss. The MRRMSE loss gives the best performance in practice. I did not use fancy tuning on the optimizer but simply used a standard Adam optimizer without any learning rate scheduler.</p>\n<h3>3.2 Bulk expression prediction</h3>\n<h4>3.2.1 Preprocess</h4>\n<p>The bulk count data is first normalized following the classic scRNA-seq preprocessing pipeline, namely, i) scaled to have 1e4 counts per bulk, ii) applied the log1p transformation, and iii) scaled to have zero mean and unit variance.</p>\n<h4>3.2.2 Model architecture</h4>\n<p>After many tries on the DE value prediction paradigm, I felt a bit frustrated since the performance gain only comes from the fancy tuning of the model architecture. Hence, I turned to try the bulk expression prediction paradigm in the later stage of the competition. According to the experiment design, I noticed that there is a negative control spot in each row on the plate. In other words, the only difference between the negative spot and the remaining spots lies in the added compound (please correct me if that is wrong). Consequently, we could predict the perturbed bulk expression values based on the baseline counts and the compound. To this end, I built a conditional autoencoder as follows,</p>\n<pre><code> (nn.Module):\n     ():\n        (Net, self).__init__()\n        self.compound_num = compound_num\n        self.gene_num = gene_num\n\n         sm_feature  :\n            self.sm_emb = nn.Embedding(self.gene_num, )\n            self.sm_enc = \n        :\n            self.sm_emb = sm_feature\n            self.sm_enc = nn.Sequential(\n                nn.Linear(self.sm_emb.shape[], ),\n                nn.BatchNorm1d(),\n                nn.ReLU(),\n                nn.Linear(, ),\n            )\n        self.type_atac = type_atac\n        self.atac_enc = nn.Sequential(\n            nn.Linear(self.type_atac.shape[], ),\n            nn.BatchNorm1d(),\n            nn.ReLU(),\n            nn.Linear(, ),\n        )\n        self.encoder = nn.Sequential(\n            nn.Linear(self.gene_num, ),\n            nn.BatchNorm1d(),\n            nn.ReLU(),\n            nn.Linear(, ),\n        )\n        self.decoder = ConditionalNet(\n            dim_feature=,\n            dim_cond_embed=,\n            dim_hidden=,\n            dim_out=self.gene_num,\n            n_blocks=,\n            skip_layers=(),\n        )\n\n     ():\n         self.sm_enc  :\n            sm = self.sm_emb(sm_name)\n        :\n            sm = self.sm_enc(self.sm_emb[sm_name])\n        encode = self.encoder(x)\n        atac = self.atac_enc(self.type_atac[])\n\n        x = sm\n        cond = torch.cat([encode, atac], dim=)\n        pred = self.decoder(x, cond)\n\n         pred\n</code></pre>\n<p>Notably, here I unnaturally chose the negative count and cell type-averaged ATAC count as conditions, while letting the compound be the input. In practice, such a configuration leads to better results than reversely setting negative count as input and compound as the condition. Such a result could probably be attributed to the over-fitting of compound representation for the later configuration. Likewise, I also tried the Transformer architecture but it does not work very well.</p>\n<h4>3.2.3 Learning objective</h4>\n<p>The model was trained to predict the normalized bulk count value change after adding the compound. The predicted normalized value is then recovered to the raw count according to the previous scaling and normalization factors. I tried the MSE loss, L1 loss, Smooth L1 loss, and MRRMSE loss, and found the Smooth L1 loss leads to the best performance. The predicted bulk value was then concatenated with the known bulk value and fed into the Limma model to compute the DE values.</p>\n<h4>3.2.4 Post preprocessing</h4>\n<p>Though the bulk expression prediction paradigm seems technically sound, its performance is not that satisfying. In my experiments, the prediction DE values differ a lot on the known compounds. Here, I would like to point out that comparing the DE values computed by Limma on the predicted bulk expression and the provided DE values is a solid validation metric, which would be less affected by the over-fitting problem. Anyway, I found the computed DE values have a larger mean compared with the provided ones. I wonder if it is because the Limma is applied on different sets of compounds (i.e., partially on the provided DE values but fully on the private DE values). Thus, I scaled the DE values computed by Limma to have the same mean on the known compounds. Such a paradigm ends up in the best MRRMSE of 0.587 on the public test split and 0.809 on the private test split.</p>\n<h3>3.3 Ensembling</h3>\n<p>Ensembling the results of the above two different paradigms led to 0.563 MRRMSE on the public test split and 0.755 on the private test split. By further ensembling the EDA results of 0.567 public score, the performance further improved to have 0.552 MRRMSE on the public test split and 0.728 on the private test split. The performance gain by ensembling results of different paradigms is surprising. Yes, my best history submission is even better than the 1st place in the leaderboard. However, it does not correspond to the best public score so I did not choose it as the final submission.</p>\n<h2>4. Robustness</h2>\n<p>For the DE value prediction paradigm, I tried removing compounds with extreme values such as \"MLN 2238\", but did not observe performance improvements. Adding noise to the inputs is less effective than adding the Dropout layers into the network. For the bulk expression paradigm, I tried removing the most dissimilar cell type \"T regulatory cells\" when training the model. Interestingly, the results were almost not influenced. The stronger robustness of the bulk expression paradigm could probably be attributed to the larger training data sample number. As for directly predicting the DE values, hundreds of data samples are too little to predict thousands of genes.</p>\n<h2>5. Documentation, code style, and reproducibility</h2>\n<p>The code and documentation can be accessed from <a href=\"https://github.com/Yunfan-Li/PerturbPrediction\" target=\"_blank\">https://github.com/Yunfan-Li/PerturbPrediction</a>.</p>",
  "messages": [
    {
      "id": "2560820",
      "postDate": "12/14/2023 03:19:51",
      "content": "<h2>1. Integration of biological knowledge</h2>\n<p>I tried integrating different biological knowledge, but unfortunately, most didn't work very well.</p>\n<p>Firstly, I came up with the idea of using the GO term to reduce genes into modules. I adopted the KEGG 2016 genes set and filtered out gene sets with a P-value less than 0.05. However, genes from the same set exhibit considerably different differential expression (DE) values. As a result, I feel the hypothesis that genes from the same GO term would have similar DE values cannot stand. Then, I resorted to finding co-expressed genes based on the gene expression matrix. I found the DE values between co-expressed genes do have a high Pearson correlation. However, I checked the predicted DE values of co-expressed genes and found they yield strong Pearson correlations as well. In other words, adding the correlation constraints to the optimization process won't have much effect.</p>\n<p>Next is about the molecule representation. I tried using the features from the pre-trained molecule models including ChemBERT and MolCLR, as well as the TFIDF-transformed element count proposed in the public notebooks. In my experiments, using the pre-trained features led to slightly inferior results, while the TFIDF-transformed element count and simply learnable molecule features give similar performance. I wonder if it is because the pre-training models capture more coarse-grained differences between molecules, or because the molecules in the contest are quite different from those used for pre-training.</p>\n<p>Finally is the integration of ATAC data. I did not have a good idea of integrating ATAC data into the model. All that I had come up with is that genes corresponding to a small ATAC count would be less affected by molecules. Anyway, I am not sure about such a hypothesis, and in practice, I simply concatenated the ATAC feature and found it slightly improves the results.</p>\n<h2>2. Exploration of the problem</h2>\n<p>According to the problem definition, I think three targets could be predicted to achieve the task. The first target is directly the DE values. The variables in DE value prediction would be cell types and compounds. A model could be trained to predict the DE value given cell type-compound pairs. The second target is the bulk expression values, namely, the input to the Limma model. The variables in bulk prediction would be only the compounds (as I will elaborate on later). I feel such a paradigm is the most promising and robust solution. The most important reason is that the results could be precisely validated. To be specific, we could send the predicted and provided bulk expressions to the Limma model, and check how close the predicted DE values are to the provided ones. I feel such a validation scheme would be much less risks in overfitting. The third target is the single-cell gene expression. Though the training data is most sufficient in such a paradigm, it is the most challenging paradigm since we do not have the paired before-perturbation and after-perturbation gene expression values.</p>\n<h2>3. Model design</h2>\n<p>In this competition, I mainly explored the first two prediction paradigms, namely the DE values and bulk expression values.</p>\n<h3>3.1 DE value prediction</h3>\n<h4>3.1.1 Preprocess</h4>\n<p>No data preprocessing is applied. Common preprocessing strategies such as normalizing and scaling even harm the performance in my experience.</p>\n<h4>3.1.2 Model architecture</h4>\n<p>The plainest model learns the cell type and compound representation, as well as makes the predictions at the same time. The model could be written within a few lines, namely,</p>\n<pre><code> (nn.Module):\n     ():\n        (Net, self).__init__()\n        self.type_num = type_num\n        self.compound_num = compound_num\n        self.gene_num = gene_num\n\n        self.type_embedding = nn.Embedding(self.type_num, )\n        self.compound_embedding = nn.Embedding(self.compound_num, )\n\n        self.predictor = nn.Sequential(\n            nn.Linear(, ),\n            nn.BatchNorm1d(),\n            nn.Dropout(),\n            nn.ReLU(),\n            nn.Linear(, ),\n            nn.BatchNorm1d(),\n            nn.Dropout(),\n            nn.ReLU(),\n            nn.Linear(, self.gene_num),\n        )\n\n     ():\n        type_embedding = self.type_embedding()\n        compound_embedding = self.compound_embedding(compound)\n        embedding = torch.cat([type_embedding, compound_embedding], dim=)\n         self.predictor(embedding)\n</code></pre>\n<p>Optionally, we could replace the learnable cell type representation with RNA or ATAC counts averaged by cell type, and replace the compound representation with features from pre-trained models such as ChemBERT and MolCLR. Training such a vanilla model gives 0.586 MRRMSE on the public test split and 0.785 MRRMSE on the private test split.</p>\n<p>I also tried Transformer architecture like the Performer used in the scBERT paper. The idea is to treat each gene as a token and replace the positional embedding with cell type and compound representation. The token of each gene is learned together with the Transformer model. Due to the space limitation, I will not attach the model configuration code here, please refer to the GitHub repository attached below. Sadly, though the Transformer architecture generally gives better performance, in my practice, I just cannot successfully train the model (i.e., the training loss stays very high), and achieved 0.611 MRRMSE on the public test split and 0.767 MRRMSE on the private test split. It was not until the competition ended that I found the Transformer architecture achieved a much better score on the private test split, despite its poor performance on the public test split. Such an interesting result is worth exploring (or it is simply due to randomness and luck).</p>\n<h4>3.1.3 Learning objective</h4>\n<p>To directly predict the DE values, I tried the MSE loss, L1 loss, and MRRMSE loss. The MRRMSE loss gives the best performance in practice. I did not use fancy tuning on the optimizer but simply used a standard Adam optimizer without any learning rate scheduler.</p>\n<h3>3.2 Bulk expression prediction</h3>\n<h4>3.2.1 Preprocess</h4>\n<p>The bulk count data is first normalized following the classic scRNA-seq preprocessing pipeline, namely, i) scaled to have 1e4 counts per bulk, ii) applied the log1p transformation, and iii) scaled to have zero mean and unit variance.</p>\n<h4>3.2.2 Model architecture</h4>\n<p>After many tries on the DE value prediction paradigm, I felt a bit frustrated since the performance gain only comes from the fancy tuning of the model architecture. Hence, I turned to try the bulk expression prediction paradigm in the later stage of the competition. According to the experiment design, I noticed that there is a negative control spot in each row on the plate. In other words, the only difference between the negative spot and the remaining spots lies in the added compound (please correct me if that is wrong). Consequently, we could predict the perturbed bulk expression values based on the baseline counts and the compound. To this end, I built a conditional autoencoder as follows,</p>\n<pre><code> (nn.Module):\n     ():\n        (Net, self).__init__()\n        self.compound_num = compound_num\n        self.gene_num = gene_num\n\n         sm_feature  :\n            self.sm_emb = nn.Embedding(self.gene_num, )\n            self.sm_enc = \n        :\n            self.sm_emb = sm_feature\n            self.sm_enc = nn.Sequential(\n                nn.Linear(self.sm_emb.shape[], ),\n                nn.BatchNorm1d(),\n                nn.ReLU(),\n                nn.Linear(, ),\n            )\n        self.type_atac = type_atac\n        self.atac_enc = nn.Sequential(\n            nn.Linear(self.type_atac.shape[], ),\n            nn.BatchNorm1d(),\n            nn.ReLU(),\n            nn.Linear(, ),\n        )\n        self.encoder = nn.Sequential(\n            nn.Linear(self.gene_num, ),\n            nn.BatchNorm1d(),\n            nn.ReLU(),\n            nn.Linear(, ),\n        )\n        self.decoder = ConditionalNet(\n            dim_feature=,\n            dim_cond_embed=,\n            dim_hidden=,\n            dim_out=self.gene_num,\n            n_blocks=,\n            skip_layers=(),\n        )\n\n     ():\n         self.sm_enc  :\n            sm = self.sm_emb(sm_name)\n        :\n            sm = self.sm_enc(self.sm_emb[sm_name])\n        encode = self.encoder(x)\n        atac = self.atac_enc(self.type_atac[])\n\n        x = sm\n        cond = torch.cat([encode, atac], dim=)\n        pred = self.decoder(x, cond)\n\n         pred\n</code></pre>\n<p>Notably, here I unnaturally chose the negative count and cell type-averaged ATAC count as conditions, while letting the compound be the input. In practice, such a configuration leads to better results than reversely setting negative count as input and compound as the condition. Such a result could probably be attributed to the over-fitting of compound representation for the later configuration. Likewise, I also tried the Transformer architecture but it does not work very well.</p>\n<h4>3.2.3 Learning objective</h4>\n<p>The model was trained to predict the normalized bulk count value change after adding the compound. The predicted normalized value is then recovered to the raw count according to the previous scaling and normalization factors. I tried the MSE loss, L1 loss, Smooth L1 loss, and MRRMSE loss, and found the Smooth L1 loss leads to the best performance. The predicted bulk value was then concatenated with the known bulk value and fed into the Limma model to compute the DE values.</p>\n<h4>3.2.4 Post preprocessing</h4>\n<p>Though the bulk expression prediction paradigm seems technically sound, its performance is not that satisfying. In my experiments, the prediction DE values differ a lot on the known compounds. Here, I would like to point out that comparing the DE values computed by Limma on the predicted bulk expression and the provided DE values is a solid validation metric, which would be less affected by the over-fitting problem. Anyway, I found the computed DE values have a larger mean compared with the provided ones. I wonder if it is because the Limma is applied on different sets of compounds (i.e., partially on the provided DE values but fully on the private DE values). Thus, I scaled the DE values computed by Limma to have the same mean on the known compounds. Such a paradigm ends up in the best MRRMSE of 0.587 on the public test split and 0.809 on the private test split.</p>\n<h3>3.3 Ensembling</h3>\n<p>Ensembling the results of the above two different paradigms led to 0.563 MRRMSE on the public test split and 0.755 on the private test split. By further ensembling the EDA results of 0.567 public score, the performance further improved to have 0.552 MRRMSE on the public test split and 0.728 on the private test split. The performance gain by ensembling results of different paradigms is surprising. Yes, my best history submission is even better than the 1st place in the leaderboard. However, it does not correspond to the best public score so I did not choose it as the final submission.</p>\n<h2>4. Robustness</h2>\n<p>For the DE value prediction paradigm, I tried removing compounds with extreme values such as \"MLN 2238\", but did not observe performance improvements. Adding noise to the inputs is less effective than adding the Dropout layers into the network. For the bulk expression paradigm, I tried removing the most dissimilar cell type \"T regulatory cells\" when training the model. Interestingly, the results were almost not influenced. The stronger robustness of the bulk expression paradigm could probably be attributed to the larger training data sample number. As for directly predicting the DE values, hundreds of data samples are too little to predict thousands of genes.</p>\n<h2>5. Documentation, code style, and reproducibility</h2>\n<p>The code and documentation can be accessed from <a href=\"https://github.com/Yunfan-Li/PerturbPrediction\" target=\"_blank\">https://github.com/Yunfan-Li/PerturbPrediction</a>.</p>",
      "rawMarkdown": "## 1. Integration of biological knowledge\nI tried integrating different biological knowledge, but unfortunately, most didn't work very well.\n\nFirstly, I came up with the idea of using the GO term to reduce genes into modules. I adopted the KEGG 2016 genes set and filtered out gene sets with a P-value less than 0.05. However, genes from the same set exhibit considerably different differential expression (DE) values. As a result, I feel the hypothesis that genes from the same GO term would have similar DE values cannot stand. Then, I resorted to finding co-expressed genes based on the gene expression matrix. I found the DE values between co-expressed genes do have a high Pearson correlation. However, I checked the predicted DE values of co-expressed genes and found they yield strong Pearson correlations as well. In other words, adding the correlation constraints to the optimization process won't have much effect.\n\nNext is about the molecule representation. I tried using the features from the pre-trained molecule models including ChemBERT and MolCLR, as well as the TFIDF-transformed element count proposed in the public notebooks. In my experiments, using the pre-trained features led to slightly inferior results, while the TFIDF-transformed element count and simply learnable molecule features give similar performance. I wonder if it is because the pre-training models capture more coarse-grained differences between molecules, or because the molecules in the contest are quite different from those used for pre-training.\n\nFinally is the integration of ATAC data. I did not have a good idea of integrating ATAC data into the model. All that I had come up with is that genes corresponding to a small ATAC count would be less affected by molecules. Anyway, I am not sure about such a hypothesis, and in practice, I simply concatenated the ATAC feature and found it slightly improves the results.\n\n## 2. Exploration of the problem\nAccording to the problem definition, I think three targets could be predicted to achieve the task. The first target is directly the DE values. The variables in DE value prediction would be cell types and compounds. A model could be trained to predict the DE value given cell type-compound pairs. The second target is the bulk expression values, namely, the input to the Limma model. The variables in bulk prediction would be only the compounds (as I will elaborate on later). I feel such a paradigm is the most promising and robust solution. The most important reason is that the results could be precisely validated. To be specific, we could send the predicted and provided bulk expressions to the Limma model, and check how close the predicted DE values are to the provided ones. I feel such a validation scheme would be much less risks in overfitting. The third target is the single-cell gene expression. Though the training data is most sufficient in such a paradigm, it is the most challenging paradigm since we do not have the paired before-perturbation and after-perturbation gene expression values.\n\n## 3. Model design\nIn this competition, I mainly explored the first two prediction paradigms, namely the DE values and bulk expression values.\n\n### 3.1 DE value prediction\n\n#### 3.1.1 Preprocess\n\nNo data preprocessing is applied. Common preprocessing strategies such as normalizing and scaling even harm the performance in my experience.\n\n#### 3.1.2 Model architecture\n\nThe plainest model learns the cell type and compound representation, as well as makes the predictions at the same time. The model could be written within a few lines, namely,\n\n```python\nclass Net(nn.Module):\n    def __init__(self, type_num, compound_num, gene_num):\n        super(Net, self).__init__()\n        self.type_num = type_num\n        self.compound_num = compound_num\n        self.gene_num = gene_num\n\n        self.type_embedding = nn.Embedding(self.type_num, 32)\n        self.compound_embedding = nn.Embedding(self.compound_num, 32)\n\n        self.predictor = nn.Sequential(\n            nn.Linear(64, 256),\n            nn.BatchNorm1d(256),\n            nn.Dropout(),\n            nn.ReLU(),\n            nn.Linear(256, 1024),\n            nn.BatchNorm1d(1024),\n            nn.Dropout(),\n            nn.ReLU(),\n            nn.Linear(1024, self.gene_num),\n        )\n\n    def forward(self, type, compound):\n        type_embedding = self.type_embedding(type)\n        compound_embedding = self.compound_embedding(compound)\n        embedding = torch.cat([type_embedding, compound_embedding], dim=1)\n        return self.predictor(embedding)\n```\n\nOptionally, we could replace the learnable cell type representation with RNA or ATAC counts averaged by cell type, and replace the compound representation with features from pre-trained models such as ChemBERT and MolCLR. Training such a vanilla model gives 0.586 MRRMSE on the public test split and 0.785 MRRMSE on the private test split.\n\nI also tried Transformer architecture like the Performer used in the scBERT paper. The idea is to treat each gene as a token and replace the positional embedding with cell type and compound representation. The token of each gene is learned together with the Transformer model. Due to the space limitation, I will not attach the model configuration code here, please refer to the GitHub repository attached below. Sadly, though the Transformer architecture generally gives better performance, in my practice, I just cannot successfully train the model (i.e., the training loss stays very high), and achieved 0.611 MRRMSE on the public test split and 0.767 MRRMSE on the private test split. It was not until the competition ended that I found the Transformer architecture achieved a much better score on the private test split, despite its poor performance on the public test split. Such an interesting result is worth exploring (or it is simply due to randomness and luck).\n\n#### 3.1.3 Learning objective\n\nTo directly predict the DE values, I tried the MSE loss, L1 loss, and MRRMSE loss. The MRRMSE loss gives the best performance in practice. I did not use fancy tuning on the optimizer but simply used a standard Adam optimizer without any learning rate scheduler.\n\n### 3.2 Bulk expression prediction\n\n#### 3.2.1 Preprocess\n\nThe bulk count data is first normalized following the classic scRNA-seq preprocessing pipeline, namely, i) scaled to have 1e4 counts per bulk, ii) applied the log1p transformation, and iii) scaled to have zero mean and unit variance.\n\n#### 3.2.2 Model architecture\n\nAfter many tries on the DE value prediction paradigm, I felt a bit frustrated since the performance gain only comes from the fancy tuning of the model architecture. Hence, I turned to try the bulk expression prediction paradigm in the later stage of the competition. According to the experiment design, I noticed that there is a negative control spot in each row on the plate. In other words, the only difference between the negative spot and the remaining spots lies in the added compound (please correct me if that is wrong). Consequently, we could predict the perturbed bulk expression values based on the baseline counts and the compound. To this end, I built a conditional autoencoder as follows,\n\n```python\nclass Net(nn.Module):\n    def __init__(\n        self,\n        gene_num,\n        compound_num,\n        sm_feature=None,\n        type_rna=None,\n        type_atac=None,\n    ):\n        super(Net, self).__init__()\n        self.compound_num = compound_num\n        self.gene_num = gene_num\n\n        if sm_feature is None:\n            self.sm_emb = nn.Embedding(self.gene_num, 64)\n            self.sm_enc = None\n        else:\n            self.sm_emb = sm_feature\n            self.sm_enc = nn.Sequential(\n                nn.Linear(self.sm_emb.shape[1], 512),\n                nn.BatchNorm1d(512),\n                nn.ReLU(),\n                nn.Linear(512, 64),\n            )\n        self.type_atac = type_atac\n        self.atac_enc = nn.Sequential(\n            nn.Linear(self.type_atac.shape[1], 512),\n            nn.BatchNorm1d(512),\n            nn.ReLU(),\n            nn.Linear(512, 64),\n        )\n        self.encoder = nn.Sequential(\n            nn.Linear(self.gene_num, 512),\n            nn.BatchNorm1d(512),\n            nn.ReLU(),\n            nn.Linear(512, 64),\n        )\n        self.decoder = ConditionalNet(\n            dim_feature=64,\n            dim_cond_embed=128,\n            dim_hidden=1024,\n            dim_out=self.gene_num,\n            n_blocks=4,\n            skip_layers=(),\n        )\n\n    def forward(self, x, type, sm_name):\n        if self.sm_enc is None:\n            sm = self.sm_emb(sm_name)\n        else:\n            sm = self.sm_enc(self.sm_emb[sm_name])\n        encode = self.encoder(x)\n        atac = self.atac_enc(self.type_atac[type])\n\n        x = sm\n        cond = torch.cat([encode, atac], dim=1)\n        pred = self.decoder(x, cond)\n\n        return pred\n```\n\nNotably, here I unnaturally chose the negative count and cell type-averaged ATAC count as conditions, while letting the compound be the input. In practice, such a configuration leads to better results than reversely setting negative count as input and compound as the condition. Such a result could probably be attributed to the over-fitting of compound representation for the later configuration. Likewise, I also tried the Transformer architecture but it does not work very well.\n\n#### 3.2.3 Learning objective\nThe model was trained to predict the normalized bulk count value change after adding the compound. The predicted normalized value is then recovered to the raw count according to the previous scaling and normalization factors. I tried the MSE loss, L1 loss, Smooth L1 loss, and MRRMSE loss, and found the Smooth L1 loss leads to the best performance. The predicted bulk value was then concatenated with the known bulk value and fed into the Limma model to compute the DE values.\n\n#### 3.2.4 Post preprocessing\nThough the bulk expression prediction paradigm seems technically sound, its performance is not that satisfying. In my experiments, the prediction DE values differ a lot on the known compounds. Here, I would like to point out that comparing the DE values computed by Limma on the predicted bulk expression and the provided DE values is a solid validation metric, which would be less affected by the over-fitting problem. Anyway, I found the computed DE values have a larger mean compared with the provided ones. I wonder if it is because the Limma is applied on different sets of compounds (i.e., partially on the provided DE values but fully on the private DE values). Thus, I scaled the DE values computed by Limma to have the same mean on the known compounds. Such a paradigm ends up in the best MRRMSE of 0.587 on the public test split and 0.809 on the private test split.\n\n### 3.3 Ensembling\nEnsembling the results of the above two different paradigms led to 0.563 MRRMSE on the public test split and 0.755 on the private test split. By further ensembling the EDA results of 0.567 public score, the performance further improved to have 0.552 MRRMSE on the public test split and 0.728 on the private test split. The performance gain by ensembling results of different paradigms is surprising. Yes, my best history submission is even better than the 1st place in the leaderboard. However, it does not correspond to the best public score so I did not choose it as the final submission.\n\n## 4. Robustness\nFor the DE value prediction paradigm, I tried removing compounds with extreme values such as \"MLN 2238\", but did not observe performance improvements. Adding noise to the inputs is less effective than adding the Dropout layers into the network. For the bulk expression paradigm, I tried removing the most dissimilar cell type \"T regulatory cells\" when training the model. Interestingly, the results were almost not influenced. The stronger robustness of the bulk expression paradigm could probably be attributed to the larger training data sample number. As for directly predicting the DE values, hundreds of data samples are too little to predict thousands of genes.\n\n## 5. Documentation, code style, and reproducibility\nThe code and documentation can be accessed from <https://github.com/Yunfan-Li/PerturbPrediction>.",
      "votes": null
    },
    {
      "id": "2564223",
      "postDate": "12/16/2023 23:11:21",
      "content": "<p>May I have access to the submission.csv  file 0.736?</p>",
      "rawMarkdown": "May I have access to the submission.csv  file 0.736?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2564223,
      "author_name": "",
      "author_url": "",
      "post_date": "12/16/2023 23:11:21",
      "content": "<p>May I have access to the submission.csv  file 0.736?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2560820": "## 1. Integration of biological knowledge\nI tried integrating different biological knowledge, but unfortunately, most didn't work very well.\n\nFirstly, I came up with the idea of using the GO term to reduce genes into modules. I adopted the KEGG 2016 genes set and filtered out gene sets with a P-value less than 0.05. However, genes from the same set exhibit considerably different differential expression (DE) values. As a result, I feel the hypothesis that genes from the same GO term would have similar DE values cannot stand. Then, I resorted to finding co-expressed genes based on the gene expression matrix. I found the DE values between co-expressed genes do have a high Pearson correlation. However, I checked the predicted DE values of co-expressed genes and found they yield strong Pearson correlations as well. In other words, adding the correlation constraints to the optimization process won't have much effect.\n\nNext is about the molecule representation. I tried using the features from the pre-trained molecule models including ChemBERT and MolCLR, as well as the TFIDF-transformed element count proposed in the public notebooks. In my experiments, using the pre-trained features led to slightly inferior results, while the TFIDF-transformed element count and simply learnable molecule features give similar performance. I wonder if it is because the pre-training models capture more coarse-grained differences between molecules, or because the molecules in the contest are quite different from those used for pre-training.\n\nFinally is the integration of ATAC data. I did not have a good idea of integrating ATAC data into the model. All that I had come up with is that genes corresponding to a small ATAC count would be less affected by molecules. Anyway, I am not sure about such a hypothesis, and in practice, I simply concatenated the ATAC feature and found it slightly improves the results.\n\n## 2. Exploration of the problem\nAccording to the problem definition, I think three targets could be predicted to achieve the task. The first target is directly the DE values. The variables in DE value prediction would be cell types and compounds. A model could be trained to predict the DE value given cell type-compound pairs. The second target is the bulk expression values, namely, the input to the Limma model. The variables in bulk prediction would be only the compounds (as I will elaborate on later). I feel such a paradigm is the most promising and robust solution. The most important reason is that the results could be precisely validated. To be specific, we could send the predicted and provided bulk expressions to the Limma model, and check how close the predicted DE values are to the provided ones. I feel such a validation scheme would be much less risks in overfitting. The third target is the single-cell gene expression. Though the training data is most sufficient in such a paradigm, it is the most challenging paradigm since we do not have the paired before-perturbation and after-perturbation gene expression values.\n\n## 3. Model design\nIn this competition, I mainly explored the first two prediction paradigms, namely the DE values and bulk expression values.\n\n### 3.1 DE value prediction\n\n#### 3.1.1 Preprocess\n\nNo data preprocessing is applied. Common preprocessing strategies such as normalizing and scaling even harm the performance in my experience.\n\n#### 3.1.2 Model architecture\n\nThe plainest model learns the cell type and compound representation, as well as makes the predictions at the same time. The model could be written within a few lines, namely,\n\n```python\nclass Net(nn.Module):\n    def __init__(self, type_num, compound_num, gene_num):\n        super(Net, self).__init__()\n        self.type_num = type_num\n        self.compound_num = compound_num\n        self.gene_num = gene_num\n\n        self.type_embedding = nn.Embedding(self.type_num, 32)\n        self.compound_embedding = nn.Embedding(self.compound_num, 32)\n\n        self.predictor = nn.Sequential(\n            nn.Linear(64, 256),\n            nn.BatchNorm1d(256),\n            nn.Dropout(),\n            nn.ReLU(),\n            nn.Linear(256, 1024),\n            nn.BatchNorm1d(1024),\n            nn.Dropout(),\n            nn.ReLU(),\n            nn.Linear(1024, self.gene_num),\n        )\n\n    def forward(self, type, compound):\n        type_embedding = self.type_embedding(type)\n        compound_embedding = self.compound_embedding(compound)\n        embedding = torch.cat([type_embedding, compound_embedding], dim=1)\n        return self.predictor(embedding)\n```\n\nOptionally, we could replace the learnable cell type representation with RNA or ATAC counts averaged by cell type, and replace the compound representation with features from pre-trained models such as ChemBERT and MolCLR. Training such a vanilla model gives 0.586 MRRMSE on the public test split and 0.785 MRRMSE on the private test split.\n\nI also tried Transformer architecture like the Performer used in the scBERT paper. The idea is to treat each gene as a token and replace the positional embedding with cell type and compound representation. The token of each gene is learned together with the Transformer model. Due to the space limitation, I will not attach the model configuration code here, please refer to the GitHub repository attached below. Sadly, though the Transformer architecture generally gives better performance, in my practice, I just cannot successfully train the model (i.e., the training loss stays very high), and achieved 0.611 MRRMSE on the public test split and 0.767 MRRMSE on the private test split. It was not until the competition ended that I found the Transformer architecture achieved a much better score on the private test split, despite its poor performance on the public test split. Such an interesting result is worth exploring (or it is simply due to randomness and luck).\n\n#### 3.1.3 Learning objective\n\nTo directly predict the DE values, I tried the MSE loss, L1 loss, and MRRMSE loss. The MRRMSE loss gives the best performance in practice. I did not use fancy tuning on the optimizer but simply used a standard Adam optimizer without any learning rate scheduler.\n\n### 3.2 Bulk expression prediction\n\n#### 3.2.1 Preprocess\n\nThe bulk count data is first normalized following the classic scRNA-seq preprocessing pipeline, namely, i) scaled to have 1e4 counts per bulk, ii) applied the log1p transformation, and iii) scaled to have zero mean and unit variance.\n\n#### 3.2.2 Model architecture\n\nAfter many tries on the DE value prediction paradigm, I felt a bit frustrated since the performance gain only comes from the fancy tuning of the model architecture. Hence, I turned to try the bulk expression prediction paradigm in the later stage of the competition. According to the experiment design, I noticed that there is a negative control spot in each row on the plate. In other words, the only difference between the negative spot and the remaining spots lies in the added compound (please correct me if that is wrong). Consequently, we could predict the perturbed bulk expression values based on the baseline counts and the compound. To this end, I built a conditional autoencoder as follows,\n\n```python\nclass Net(nn.Module):\n    def __init__(\n        self,\n        gene_num,\n        compound_num,\n        sm_feature=None,\n        type_rna=None,\n        type_atac=None,\n    ):\n        super(Net, self).__init__()\n        self.compound_num = compound_num\n        self.gene_num = gene_num\n\n        if sm_feature is None:\n            self.sm_emb = nn.Embedding(self.gene_num, 64)\n            self.sm_enc = None\n        else:\n            self.sm_emb = sm_feature\n            self.sm_enc = nn.Sequential(\n                nn.Linear(self.sm_emb.shape[1], 512),\n                nn.BatchNorm1d(512),\n                nn.ReLU(),\n                nn.Linear(512, 64),\n            )\n        self.type_atac = type_atac\n        self.atac_enc = nn.Sequential(\n            nn.Linear(self.type_atac.shape[1], 512),\n            nn.BatchNorm1d(512),\n            nn.ReLU(),\n            nn.Linear(512, 64),\n        )\n        self.encoder = nn.Sequential(\n            nn.Linear(self.gene_num, 512),\n            nn.BatchNorm1d(512),\n            nn.ReLU(),\n            nn.Linear(512, 64),\n        )\n        self.decoder = ConditionalNet(\n            dim_feature=64,\n            dim_cond_embed=128,\n            dim_hidden=1024,\n            dim_out=self.gene_num,\n            n_blocks=4,\n            skip_layers=(),\n        )\n\n    def forward(self, x, type, sm_name):\n        if self.sm_enc is None:\n            sm = self.sm_emb(sm_name)\n        else:\n            sm = self.sm_enc(self.sm_emb[sm_name])\n        encode = self.encoder(x)\n        atac = self.atac_enc(self.type_atac[type])\n\n        x = sm\n        cond = torch.cat([encode, atac], dim=1)\n        pred = self.decoder(x, cond)\n\n        return pred\n```\n\nNotably, here I unnaturally chose the negative count and cell type-averaged ATAC count as conditions, while letting the compound be the input. In practice, such a configuration leads to better results than reversely setting negative count as input and compound as the condition. Such a result could probably be attributed to the over-fitting of compound representation for the later configuration. Likewise, I also tried the Transformer architecture but it does not work very well.\n\n#### 3.2.3 Learning objective\nThe model was trained to predict the normalized bulk count value change after adding the compound. The predicted normalized value is then recovered to the raw count according to the previous scaling and normalization factors. I tried the MSE loss, L1 loss, Smooth L1 loss, and MRRMSE loss, and found the Smooth L1 loss leads to the best performance. The predicted bulk value was then concatenated with the known bulk value and fed into the Limma model to compute the DE values.\n\n#### 3.2.4 Post preprocessing\nThough the bulk expression prediction paradigm seems technically sound, its performance is not that satisfying. In my experiments, the prediction DE values differ a lot on the known compounds. Here, I would like to point out that comparing the DE values computed by Limma on the predicted bulk expression and the provided DE values is a solid validation metric, which would be less affected by the over-fitting problem. Anyway, I found the computed DE values have a larger mean compared with the provided ones. I wonder if it is because the Limma is applied on different sets of compounds (i.e., partially on the provided DE values but fully on the private DE values). Thus, I scaled the DE values computed by Limma to have the same mean on the known compounds. Such a paradigm ends up in the best MRRMSE of 0.587 on the public test split and 0.809 on the private test split.\n\n### 3.3 Ensembling\nEnsembling the results of the above two different paradigms led to 0.563 MRRMSE on the public test split and 0.755 on the private test split. By further ensembling the EDA results of 0.567 public score, the performance further improved to have 0.552 MRRMSE on the public test split and 0.728 on the private test split. The performance gain by ensembling results of different paradigms is surprising. Yes, my best history submission is even better than the 1st place in the leaderboard. However, it does not correspond to the best public score so I did not choose it as the final submission.\n\n## 4. Robustness\nFor the DE value prediction paradigm, I tried removing compounds with extreme values such as \"MLN 2238\", but did not observe performance improvements. Adding noise to the inputs is less effective than adding the Dropout layers into the network. For the bulk expression paradigm, I tried removing the most dissimilar cell type \"T regulatory cells\" when training the model. Interestingly, the results were almost not influenced. The stronger robustness of the bulk expression paradigm could probably be attributed to the larger training data sample number. As for directly predicting the DE values, hundreds of data samples are too little to predict thousands of genes.\n\n## 5. Documentation, code style, and reproducibility\nThe code and documentation can be accessed from <https://github.com/Yunfan-Li/PerturbPrediction>.",
    "2564223": "May I have access to the submission.csv  file 0.736?"
  },
  "source": "meta"
}