{
  "id": 461225,
  "title": "16th Place Solution Writeup for the Open Problems – Single-Cell Perturbations (Los Rodriguez)",
  "url": "/competitions/open-problems-single-cell-perturbations/writeups/los-rodriguez-16th-place-solution-writeup-for-the-",
  "author_name": "",
  "post_date": "2023-12-13T10:42:20.563Z",
  "votes": 7,
  "comment_count": 4,
  "views": 0,
  "content": "<p>We finally had some time to writeup our strategy for the OP2 challenge. It was a super engaging competition, and we're really thankful to both the organizers and fellow competitors for making it such a blast! The whole experience taught us a ton, and we're happy to share what we did/discover along the way. Can't wait for the next challenge!</p>\n<h2>Context</h2>\n<p>In this competition, the main objective was to predict the effect of drug perturbations on peripheral blood mononuclear cells (PBMCs) from several patient samples. For convenience, we have created a Python package with the model here <a href=\"https://github.com/scapeML/scape\" target=\"_blank\">https://github.com/scapeML/scape</a>. </p>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data</a></li>\n</ul>\n<h2>Overview of the approach</h2>\n<p>Similar to most problems in biological research via omics data, we encountered a high-dimensional feature space (~18k genes) and a low-dimensional observation space (~614 cell/drug combinations) with a low signal-to-noise ratio, where most of the genes show random fluctuations after perturbation. The main data modality to be predicted consisted of signed and log-transformed P-values from differential expression (DE) analysis. In the DE analysis, pseudo-bulk expression profiles from drug-treated cells were compared against the profiles of cells treated with Dimethyl Sulfoxide (DMSO). In addition, challenge organizers also provided the raw data from the single-cell RNA-Seq experiment and from an accompanying ATAC-Seq experiment, conducted only in basal state.</p>\n<p>At the beginning of the challenge, we tested different models using the signed log-pvalues (“de_train” data) alone, such as simple linear models, ensembles of gradient boosting with drug and cell features, conditional variational autoencoders, etc. We soon realized that a simple Neural Network using only a small subset of genes to compute drug and cell features (median of the genes grouped by drug and cell) was enough to have a competitive model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2F72933c3ebc89980311c9824a42bffde2%2Fnn-architecture.png?generation=1702418417955250&amp;alt=media\" alt=\"\"></p>\n<p>The figure above shows the final architecture used for our submission. We used a Neural Network that takes as inputs drug and cell features and produces signed log-pvalues. Features were computed as the median of the signed log-pvalues grouped by drugs and cells, calculated from the de_train.parquet file. Additionally, we also estimated log fold-changes (LFCs) from pseudobulk expression, to produce a matrix of the same shape as the de_train data but containing LFCs instead. We also computed the median per cell/drug as features.</p>\n<p>Similar to a Conditional Variational Autoencoder (CVAE), we used cell features both in the encoding part and the decoding part of the NN. Initially, the model consisted of a CVAE that was trained using the cell features as the conditional features to learn an encoding/decoding function conditioned on the particular cell type. However, after testing different ways to train the CVAE (similar to a beta-VAE with different annealing strategies for the KL term), we finally considered a non probabilistic NN since we did not find any practical advantage in this case with respect to a simpler non-probabilistic NN. </p>\n<h3>Neural Net</h3>\n<p>We created a method to parametrize the architecture of the NN and the feature extraction from different data sources. This is the code needed to create the NN through the scAPE library specifically created for this challenge: <a href=\"https://github.com/scapeML/scape/blob/222b19f47a32afb8d157aecd5e46b23d90b73e9d/scape/_model.py#L678\" target=\"_blank\">https://github.com/scapeML/scape/blob/222b19f47a32afb8d157aecd5e46b23d90b73e9d/scape/_model.py#L678</a></p>\n<p>We used <code>n_genes=64</code> (top 64 genes sorted by variance across conditions).This generates a NN with 9637475 parameters (36.76 MB). The inputs are computed from de_train and from log-fold changes calculated from pseudobulk. Cell features are duplicated both in the encoder and decoder. We did some permutations to estimate the distribution of CV errors, permuting drug and cell features, also in the encoder and decoder part. Drug features have more impact on the final error (something to expect since there are 146 datapoints per cell type + B/Myeloid drugs). For the cell features, in general we observed that when used through the encoder and decoder, the NN places more importance in the cell features on the decoder rather than the encoder. This might suggest that cell features are more important for performing a conditional decoding of the drug features which is cell-type specific.</p>\n<h3>Model selection</h3>\n<p>Using the previous NN, we did a leave-one-drug-out for NK cells, which resulted in 146 models. We used the median of the predictions from the 146 models on the cell/drugs for the submission to generate what we call <em>base predictions</em>.</p>\n<p>The idea of using this strategy is motivated by the fact that NK cells are the most similar ones to B/Myeloid cells. The other advantage of adopting a leave-one-drug-out approach for NK-cells is that it allows us to estimate how well the model generalizes to unseen drugs, on a per-drug basis per cell type. We also observed that in general, the median was much better than the mean for aggregating the results of the 146 models, since it is more robust to outliers (some models did not generalize well on some drugs, and early stopping selected bad models in those situations).</p>\n<p>We also trained a second neural network with the same hyperparameters, but this time using only the top 256 most variable genes and focusing on the 60 most variable drugs. In this second set of predictions, instead of predicting the 18211 genes, the NN predicts only the top 256 genes used as inputs. We did this because we realized the NN was learning to decide if there was an effect on a given cell type from a small set of genes (essentially, determining where to place values close to 0 in the matrix). We reasoned that training again on only a subset of the data, where most of the changes were concentrated, would help increase performance for that subset of genes. We generated 60 models and computed the median of the predictions, which we referred to as <strong>enhanced predictions</strong>.</p>\n<p>We finally replaced the base predictions with the enhanced predictions (on the subset of genes/drugs). For the final submission, to be more conservative, we mixed the predictions in 0.80/0.20 proportions (0.80 given to the enhanced predictions). <strong>We tested this strategy with several different base predictions, and it always resulted in a boost of performance, which was also the case on the private leaderboard.</strong> A reproducible notebook for the submission is available at: <a href=\"https://github.com/scapeML/scape/blob/main/docs/notebooks/solution.ipynb\" target=\"_blank\">https://github.com/scapeML/scape/blob/main/docs/notebooks/solution.ipynb</a>.</p>\n<p>One limitation of this strategy is that most of the trained models are very similar, and the blending with the median is very conservative. We also tested different CV strategies, and we found that using blendings of models trained on a 4-CV setting with handpicked drugs on both B/Myeloid cells provided better results in the private leaderboard. However, we didn't trust this strategy that much since it was not very stable and it was hard to understand how well those models were performing in particular cell/drugs combinations.</p>\n<h3>Baselines</h3>\n<p>We think that having simple baselines is important to understand 1) if the model works, and 2) how good it is. </p>\n<p>We decide to use two simple baselines: predicting zeros (as the baseline used in the competition, which achieves a 0.666 error in the public LB), and the median of the genes grouped by drugs (computed on the training data). The second baseline is more informative.</p>\n<p>We combined those baselines with our leave-on-drug-out strategy to produce plots per drug, so we could have an upper bound estimation of the generalization for each drug. Here is an example for NK and Prednisolone:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2Fba8aa85e97594aed0ab09f324528e0fa%2Fexample-nk-prednisolone.png?generation=1702418460653383&amp;alt=media\" alt=\"\"></p>\n<h2>Specific questions</h2>\n<h3>1. Use of prior knowledge</h3>\n<p>We decided against using LINCS data in our model because it primarily focuses on cancer cell data, which tends to have a molecular state quite distinct from the PBMCs we were investigating. Despite our exploration of published work and datasets related to predicting drug-induced changes in single-cell states, none of them encompassed the vast array of drug perturbations examined in the challenge. Additionally, we chose not to integrate external data into our approach due to concerns about handling batch effects caused by differences in laboratory settings, protocols, and other related factors.</p>\n<p>We've also tried to use ATAC-seq with no success. We believe that this data would be useful in the case of not having any measurement for B/Myeloid. However, more informative than ATAC-seq are the actual perturbational profiles on the small subset of drugs on those cells.</p>\n<p>Here is a summary of different features we tested:</p>\n<ol>\n<li>Dummy binary variables for cell types and drugs.</li>\n<li>Basal omics features, including average expression in DMSO and average accessibility per the ATAC-Seq data.</li>\n<li>Summary statistics of the drug response after grouping by cell type and drug, including standard deviation, mean and median.</li>\n<li>A “raw” fold-change computed over the raw counts of the single-cell RNA-Seq data (this is, without the corrections applied by limma).</li>\n<li>Centroids of the principal component space of the drug response data, using cell-type and drug as grouping variables.</li>\n</ol>\n<p>And we obtained the best results using the median of the response after grouping by cell type and drug in combination with the raw fold changes, using only a subset with the most variable genes in the dataset.</p>\n<h3>2. Exploration of the problem</h3>\n<p>We found that the error distribution for the drugs was more or less even except for the first four drugs, which accounted for 15% of the total error. As expected, we also found out that the response of drugs that were harder to predict was very different from training cell-types in comparison to the test cell-type. For instance, the drug that accounted for the maximum proportion of the error (IN1451), produced a strong response in NK cells, but seemed to have little effect in T cells CD4+, T cells CD8+ and T regulatory cells.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2Fec079268e6ac589644cf28ae1921517f%2Ffig6.png?generation=1702418626577518&amp;alt=media\" alt=\"\"></p>\n<p>Our approach was refined to better understand cell-type errors, aiming to identify the most challenging cell type for accurate prediction. We evaluated 15 drugs across all cell types, selecting 4 at random for testing. This test set was used for cell-type cross-validation, where the model was trained on data from all 15 drugs, excluding the 4 test drugs within a specific cell type. Our method facilitated evaluation of predictive performance for each cell type. Our findings, illustrated in the figure below, indicated that myeloid cells were more difficult to predict than others, corroborating RNA-Seq PCA analysis results.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2F86646a0d4e453e58e2856a74b029b11d%2Ffig7.png?generation=1702418748195660&amp;alt=media\" alt=\"\"></p>\n<p>Regarding genes, we investigated if specific biological functions were harder to predict. An enrichment analysis of the top 5% genes with the highest average error in our local CV setup was conducted using MSigDB hallmarks and <a href=\"https://www.kaggle.com/code/pablormier/op2-biologically-aware-dimensionality-reduction\" target=\"_blank\">decoupleR</a>. This revealed that certain hallmarks, such as epithelial mesenchymal transition and TNF alpha signaling, had a significant number of genes with high error rates:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2Fcc674087faeccb96578b85de39123a7a%2Ffig8.png?generation=1702418832195389&amp;alt=media\" alt=\"\"></p>\n<h3>3. Model design</h3>\n<p>We wanted to check if simpler models could perform just as well. So, we cut down the input features in our models. Considering that our architecture was simple already, we aimed to find the fewest input features that could match the local CV performance of using the top 128 genes with the highest variance.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2Ff29dbc0d67029908ca8bcf86be308426%2Ffig9.png?generation=1702418909501736&amp;alt=media\" alt=\"\"></p>\n<p>Interestingly, we found that models with 8 to 64 input features would achieve similar performance that the model that employed 128 features.</p>\n<p>Regarding explainability, even though the model is not easily interpretable, we put some extra care in understanding better how the NN behaved through the leave-on-drug-out + baselines, and by doing permutations on the input data after training a model, to asses the impact on the validation loss.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2F309451c68977bac8fa6e310a788103df%2Ffig10.png?generation=1702418986950956&amp;alt=media\" alt=\"\"></p>\n<p>We observed that while both components had a direct impact on the performance of our model, the mean of the errors after drug features permutation was higher compared to the average error after cell features permutation. This is something to expect, as we have more data points of gene values per drug (146 data points per cell type except B/Myeloid), but we only have 6 data points grouping by cell type. We used this type of permutation tests to estimate the importance that different features had in the CV error.</p>\n<h3>4. Robustness</h3>\n<p>Our model included a Gaussian Noise layer from Keras to perturb the input data. We used this to test CV errors for different noise levels. The following figure shows that a gaussian noise of std=0.01 w. This is the value we selected for training the final models:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2F21ba9f226003e9246391f7181ba834bb%2Ffig11.png?generation=1702419081567894&amp;alt=media\" alt=\"\"></p>\n<h3>5. Documentation &amp; code style</h3>\n<p>For convenience, we refactored the code and created a package called “scape” (<a href=\"https://github.com/scapeML/scape\" target=\"_blank\">https://github.com/scapeML/scape</a>) using <a href=\"https://pdm-project.org/\" target=\"_blank\">https://pdm-project.org/</a>, which contains the code that we finally used for the submission. The code is documented using numpydoc docstrings, and we included a series of notebooks using the scape package to learn how to use it and how to manually create the setup for generating our submission. We have put effort into developing a library that allows for the comfortable configuration and parameterization of neural networks, with an automatic mode for calculating diverse features from drugs and cell lines.</p>\n<h3>6. Reproducibility</h3>\n<p>In order to improve reproducibility, we show how the tool package can be installed and used directly from Google Colab <a href=\"https://colab.research.google.com/drive/1-o_lT-ttoKS-nbozj2RQusGoi-vm0-XL?usp=sharing\" target=\"_blank\">https://colab.research.google.com/drive/1-o_lT-ttoKS-nbozj2RQusGoi-vm0-XL?usp=sharing</a>. </p>\n<p>We also included an environment.yml file to exactly recreate the environment we used for testing using conda.</p>\n<h2>Sources</h2>\n<ul>\n<li><a href=\"https://github.com/scapeML/scape\" target=\"_blank\">https://github.com/scapeML/scape</a></li>\n<li><a href=\"https://academic.oup.com/bioinformaticsadvances/article/2/1/vbac016/6544613?login=false\" target=\"_blank\">https://academic.oup.com/bioinformaticsadvances/article/2/1/vbac016/6544613?login=false</a></li>\n<li><a href=\"https://www.kaggle.com/code/pablormier/op2-biologically-aware-dimensionality-reduction\" target=\"_blank\">https://www.kaggle.com/code/pablormier/op2-biologically-aware-dimensionality-reduction</a></li>\n</ul>",
  "messages": [
    {
      "id": "2560018",
      "postDate": "12/13/2023 07:55:13",
      "content": "<p>We finally had some time to writeup our strategy for the OP2 challenge. It was a super engaging competition, and we're really thankful to both the organizers and fellow competitors for making it such a blast! The whole experience taught us a ton, and we're happy to share what we did/discover along the way. Can't wait for the next challenge!</p>\n<h2>Context</h2>\n<p>In this competition, the main objective was to predict the effect of drug perturbations on peripheral blood mononuclear cells (PBMCs) from several patient samples. For convenience, we have created a Python package with the model here <a href=\"https://github.com/scapeML/scape\" target=\"_blank\">https://github.com/scapeML/scape</a>. </p>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data</a></li>\n</ul>\n<h2>Overview of the approach</h2>\n<p>Similar to most problems in biological research via omics data, we encountered a high-dimensional feature space (~18k genes) and a low-dimensional observation space (~614 cell/drug combinations) with a low signal-to-noise ratio, where most of the genes show random fluctuations after perturbation. The main data modality to be predicted consisted of signed and log-transformed P-values from differential expression (DE) analysis. In the DE analysis, pseudo-bulk expression profiles from drug-treated cells were compared against the profiles of cells treated with Dimethyl Sulfoxide (DMSO). In addition, challenge organizers also provided the raw data from the single-cell RNA-Seq experiment and from an accompanying ATAC-Seq experiment, conducted only in basal state.</p>\n<p>At the beginning of the challenge, we tested different models using the signed log-pvalues (“de_train” data) alone, such as simple linear models, ensembles of gradient boosting with drug and cell features, conditional variational autoencoders, etc. We soon realized that a simple Neural Network using only a small subset of genes to compute drug and cell features (median of the genes grouped by drug and cell) was enough to have a competitive model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2F72933c3ebc89980311c9824a42bffde2%2Fnn-architecture.png?generation=1702418417955250&amp;alt=media\" alt=\"\"></p>\n<p>The figure above shows the final architecture used for our submission. We used a Neural Network that takes as inputs drug and cell features and produces signed log-pvalues. Features were computed as the median of the signed log-pvalues grouped by drugs and cells, calculated from the de_train.parquet file. Additionally, we also estimated log fold-changes (LFCs) from pseudobulk expression, to produce a matrix of the same shape as the de_train data but containing LFCs instead. We also computed the median per cell/drug as features.</p>\n<p>Similar to a Conditional Variational Autoencoder (CVAE), we used cell features both in the encoding part and the decoding part of the NN. Initially, the model consisted of a CVAE that was trained using the cell features as the conditional features to learn an encoding/decoding function conditioned on the particular cell type. However, after testing different ways to train the CVAE (similar to a beta-VAE with different annealing strategies for the KL term), we finally considered a non probabilistic NN since we did not find any practical advantage in this case with respect to a simpler non-probabilistic NN. </p>\n<h3>Neural Net</h3>\n<p>We created a method to parametrize the architecture of the NN and the feature extraction from different data sources. This is the code needed to create the NN through the scAPE library specifically created for this challenge: <a href=\"https://github.com/scapeML/scape/blob/222b19f47a32afb8d157aecd5e46b23d90b73e9d/scape/_model.py#L678\" target=\"_blank\">https://github.com/scapeML/scape/blob/222b19f47a32afb8d157aecd5e46b23d90b73e9d/scape/_model.py#L678</a></p>\n<p>We used <code>n_genes=64</code> (top 64 genes sorted by variance across conditions).This generates a NN with 9637475 parameters (36.76 MB). The inputs are computed from de_train and from log-fold changes calculated from pseudobulk. Cell features are duplicated both in the encoder and decoder. We did some permutations to estimate the distribution of CV errors, permuting drug and cell features, also in the encoder and decoder part. Drug features have more impact on the final error (something to expect since there are 146 datapoints per cell type + B/Myeloid drugs). For the cell features, in general we observed that when used through the encoder and decoder, the NN places more importance in the cell features on the decoder rather than the encoder. This might suggest that cell features are more important for performing a conditional decoding of the drug features which is cell-type specific.</p>\n<h3>Model selection</h3>\n<p>Using the previous NN, we did a leave-one-drug-out for NK cells, which resulted in 146 models. We used the median of the predictions from the 146 models on the cell/drugs for the submission to generate what we call <em>base predictions</em>.</p>\n<p>The idea of using this strategy is motivated by the fact that NK cells are the most similar ones to B/Myeloid cells. The other advantage of adopting a leave-one-drug-out approach for NK-cells is that it allows us to estimate how well the model generalizes to unseen drugs, on a per-drug basis per cell type. We also observed that in general, the median was much better than the mean for aggregating the results of the 146 models, since it is more robust to outliers (some models did not generalize well on some drugs, and early stopping selected bad models in those situations).</p>\n<p>We also trained a second neural network with the same hyperparameters, but this time using only the top 256 most variable genes and focusing on the 60 most variable drugs. In this second set of predictions, instead of predicting the 18211 genes, the NN predicts only the top 256 genes used as inputs. We did this because we realized the NN was learning to decide if there was an effect on a given cell type from a small set of genes (essentially, determining where to place values close to 0 in the matrix). We reasoned that training again on only a subset of the data, where most of the changes were concentrated, would help increase performance for that subset of genes. We generated 60 models and computed the median of the predictions, which we referred to as <strong>enhanced predictions</strong>.</p>\n<p>We finally replaced the base predictions with the enhanced predictions (on the subset of genes/drugs). For the final submission, to be more conservative, we mixed the predictions in 0.80/0.20 proportions (0.80 given to the enhanced predictions). <strong>We tested this strategy with several different base predictions, and it always resulted in a boost of performance, which was also the case on the private leaderboard.</strong> A reproducible notebook for the submission is available at: <a href=\"https://github.com/scapeML/scape/blob/main/docs/notebooks/solution.ipynb\" target=\"_blank\">https://github.com/scapeML/scape/blob/main/docs/notebooks/solution.ipynb</a>.</p>\n<p>One limitation of this strategy is that most of the trained models are very similar, and the blending with the median is very conservative. We also tested different CV strategies, and we found that using blendings of models trained on a 4-CV setting with handpicked drugs on both B/Myeloid cells provided better results in the private leaderboard. However, we didn't trust this strategy that much since it was not very stable and it was hard to understand how well those models were performing in particular cell/drugs combinations.</p>\n<h3>Baselines</h3>\n<p>We think that having simple baselines is important to understand 1) if the model works, and 2) how good it is. </p>\n<p>We decide to use two simple baselines: predicting zeros (as the baseline used in the competition, which achieves a 0.666 error in the public LB), and the median of the genes grouped by drugs (computed on the training data). The second baseline is more informative.</p>\n<p>We combined those baselines with our leave-on-drug-out strategy to produce plots per drug, so we could have an upper bound estimation of the generalization for each drug. Here is an example for NK and Prednisolone:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2Fba8aa85e97594aed0ab09f324528e0fa%2Fexample-nk-prednisolone.png?generation=1702418460653383&amp;alt=media\" alt=\"\"></p>\n<h2>Specific questions</h2>\n<h3>1. Use of prior knowledge</h3>\n<p>We decided against using LINCS data in our model because it primarily focuses on cancer cell data, which tends to have a molecular state quite distinct from the PBMCs we were investigating. Despite our exploration of published work and datasets related to predicting drug-induced changes in single-cell states, none of them encompassed the vast array of drug perturbations examined in the challenge. Additionally, we chose not to integrate external data into our approach due to concerns about handling batch effects caused by differences in laboratory settings, protocols, and other related factors.</p>\n<p>We've also tried to use ATAC-seq with no success. We believe that this data would be useful in the case of not having any measurement for B/Myeloid. However, more informative than ATAC-seq are the actual perturbational profiles on the small subset of drugs on those cells.</p>\n<p>Here is a summary of different features we tested:</p>\n<ol>\n<li>Dummy binary variables for cell types and drugs.</li>\n<li>Basal omics features, including average expression in DMSO and average accessibility per the ATAC-Seq data.</li>\n<li>Summary statistics of the drug response after grouping by cell type and drug, including standard deviation, mean and median.</li>\n<li>A “raw” fold-change computed over the raw counts of the single-cell RNA-Seq data (this is, without the corrections applied by limma).</li>\n<li>Centroids of the principal component space of the drug response data, using cell-type and drug as grouping variables.</li>\n</ol>\n<p>And we obtained the best results using the median of the response after grouping by cell type and drug in combination with the raw fold changes, using only a subset with the most variable genes in the dataset.</p>\n<h3>2. Exploration of the problem</h3>\n<p>We found that the error distribution for the drugs was more or less even except for the first four drugs, which accounted for 15% of the total error. As expected, we also found out that the response of drugs that were harder to predict was very different from training cell-types in comparison to the test cell-type. For instance, the drug that accounted for the maximum proportion of the error (IN1451), produced a strong response in NK cells, but seemed to have little effect in T cells CD4+, T cells CD8+ and T regulatory cells.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2Fec079268e6ac589644cf28ae1921517f%2Ffig6.png?generation=1702418626577518&amp;alt=media\" alt=\"\"></p>\n<p>Our approach was refined to better understand cell-type errors, aiming to identify the most challenging cell type for accurate prediction. We evaluated 15 drugs across all cell types, selecting 4 at random for testing. This test set was used for cell-type cross-validation, where the model was trained on data from all 15 drugs, excluding the 4 test drugs within a specific cell type. Our method facilitated evaluation of predictive performance for each cell type. Our findings, illustrated in the figure below, indicated that myeloid cells were more difficult to predict than others, corroborating RNA-Seq PCA analysis results.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2F86646a0d4e453e58e2856a74b029b11d%2Ffig7.png?generation=1702418748195660&amp;alt=media\" alt=\"\"></p>\n<p>Regarding genes, we investigated if specific biological functions were harder to predict. An enrichment analysis of the top 5% genes with the highest average error in our local CV setup was conducted using MSigDB hallmarks and <a href=\"https://www.kaggle.com/code/pablormier/op2-biologically-aware-dimensionality-reduction\" target=\"_blank\">decoupleR</a>. This revealed that certain hallmarks, such as epithelial mesenchymal transition and TNF alpha signaling, had a significant number of genes with high error rates:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2Fcc674087faeccb96578b85de39123a7a%2Ffig8.png?generation=1702418832195389&amp;alt=media\" alt=\"\"></p>\n<h3>3. Model design</h3>\n<p>We wanted to check if simpler models could perform just as well. So, we cut down the input features in our models. Considering that our architecture was simple already, we aimed to find the fewest input features that could match the local CV performance of using the top 128 genes with the highest variance.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2Ff29dbc0d67029908ca8bcf86be308426%2Ffig9.png?generation=1702418909501736&amp;alt=media\" alt=\"\"></p>\n<p>Interestingly, we found that models with 8 to 64 input features would achieve similar performance that the model that employed 128 features.</p>\n<p>Regarding explainability, even though the model is not easily interpretable, we put some extra care in understanding better how the NN behaved through the leave-on-drug-out + baselines, and by doing permutations on the input data after training a model, to asses the impact on the validation loss.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2F309451c68977bac8fa6e310a788103df%2Ffig10.png?generation=1702418986950956&amp;alt=media\" alt=\"\"></p>\n<p>We observed that while both components had a direct impact on the performance of our model, the mean of the errors after drug features permutation was higher compared to the average error after cell features permutation. This is something to expect, as we have more data points of gene values per drug (146 data points per cell type except B/Myeloid), but we only have 6 data points grouping by cell type. We used this type of permutation tests to estimate the importance that different features had in the CV error.</p>\n<h3>4. Robustness</h3>\n<p>Our model included a Gaussian Noise layer from Keras to perturb the input data. We used this to test CV errors for different noise levels. The following figure shows that a gaussian noise of std=0.01 w. This is the value we selected for training the final models:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2F21ba9f226003e9246391f7181ba834bb%2Ffig11.png?generation=1702419081567894&amp;alt=media\" alt=\"\"></p>\n<h3>5. Documentation &amp; code style</h3>\n<p>For convenience, we refactored the code and created a package called “scape” (<a href=\"https://github.com/scapeML/scape\" target=\"_blank\">https://github.com/scapeML/scape</a>) using <a href=\"https://pdm-project.org/\" target=\"_blank\">https://pdm-project.org/</a>, which contains the code that we finally used for the submission. The code is documented using numpydoc docstrings, and we included a series of notebooks using the scape package to learn how to use it and how to manually create the setup for generating our submission. We have put effort into developing a library that allows for the comfortable configuration and parameterization of neural networks, with an automatic mode for calculating diverse features from drugs and cell lines.</p>\n<h3>6. Reproducibility</h3>\n<p>In order to improve reproducibility, we show how the tool package can be installed and used directly from Google Colab <a href=\"https://colab.research.google.com/drive/1-o_lT-ttoKS-nbozj2RQusGoi-vm0-XL?usp=sharing\" target=\"_blank\">https://colab.research.google.com/drive/1-o_lT-ttoKS-nbozj2RQusGoi-vm0-XL?usp=sharing</a>. </p>\n<p>We also included an environment.yml file to exactly recreate the environment we used for testing using conda.</p>\n<h2>Sources</h2>\n<ul>\n<li><a href=\"https://github.com/scapeML/scape\" target=\"_blank\">https://github.com/scapeML/scape</a></li>\n<li><a href=\"https://academic.oup.com/bioinformaticsadvances/article/2/1/vbac016/6544613?login=false\" target=\"_blank\">https://academic.oup.com/bioinformaticsadvances/article/2/1/vbac016/6544613?login=false</a></li>\n<li><a href=\"https://www.kaggle.com/code/pablormier/op2-biologically-aware-dimensionality-reduction\" target=\"_blank\">https://www.kaggle.com/code/pablormier/op2-biologically-aware-dimensionality-reduction</a></li>\n</ul>",
      "rawMarkdown": "We finally had some time to writeup our strategy for the OP2 challenge. It was a super engaging competition, and we're really thankful to both the organizers and fellow competitors for making it such a blast! The whole experience taught us a ton, and we're happy to share what we did/discover along the way. Can't wait for the next challenge!\n\n## Context\n\nIn this competition, the main objective was to predict the effect of drug perturbations on peripheral blood mononuclear cells (PBMCs) from several patient samples. For convenience, we have created a Python package with the model here https://github.com/scapeML/scape. \n\n- Business context: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\n- Data context: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\n\n## Overview of the approach\n\nSimilar to most problems in biological research via omics data, we encountered a high-dimensional feature space (~18k genes) and a low-dimensional observation space (~614 cell/drug combinations) with a low signal-to-noise ratio, where most of the genes show random fluctuations after perturbation. The main data modality to be predicted consisted of signed and log-transformed P-values from differential expression (DE) analysis. In the DE analysis, pseudo-bulk expression profiles from drug-treated cells were compared against the profiles of cells treated with Dimethyl Sulfoxide (DMSO). In addition, challenge organizers also provided the raw data from the single-cell RNA-Seq experiment and from an accompanying ATAC-Seq experiment, conducted only in basal state.\n\nAt the beginning of the challenge, we tested different models using the signed log-pvalues (“de_train” data) alone, such as simple linear models, ensembles of gradient boosting with drug and cell features, conditional variational autoencoders, etc. We soon realized that a simple Neural Network using only a small subset of genes to compute drug and cell features (median of the genes grouped by drug and cell) was enough to have a competitive model.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2F72933c3ebc89980311c9824a42bffde2%2Fnn-architecture.png?generation=1702418417955250&alt=media)\n\nThe figure above shows the final architecture used for our submission. We used a Neural Network that takes as inputs drug and cell features and produces signed log-pvalues. Features were computed as the median of the signed log-pvalues grouped by drugs and cells, calculated from the de_train.parquet file. Additionally, we also estimated log fold-changes (LFCs) from pseudobulk expression, to produce a matrix of the same shape as the de_train data but containing LFCs instead. We also computed the median per cell/drug as features.\n\nSimilar to a Conditional Variational Autoencoder (CVAE), we used cell features both in the encoding part and the decoding part of the NN. Initially, the model consisted of a CVAE that was trained using the cell features as the conditional features to learn an encoding/decoding function conditioned on the particular cell type. However, after testing different ways to train the CVAE (similar to a beta-VAE with different annealing strategies for the KL term), we finally considered a non probabilistic NN since we did not find any practical advantage in this case with respect to a simpler non-probabilistic NN. \n\n### Neural Net\n\nWe created a method to parametrize the architecture of the NN and the feature extraction from different data sources. This is the code needed to create the NN through the scAPE library specifically created for this challenge: https://github.com/scapeML/scape/blob/222b19f47a32afb8d157aecd5e46b23d90b73e9d/scape/_model.py#L678\n\nWe used `n_genes=64` (top 64 genes sorted by variance across conditions).This generates a NN with 9637475 parameters (36.76 MB). The inputs are computed from de_train and from log-fold changes calculated from pseudobulk. Cell features are duplicated both in the encoder and decoder. We did some permutations to estimate the distribution of CV errors, permuting drug and cell features, also in the encoder and decoder part. Drug features have more impact on the final error (something to expect since there are 146 datapoints per cell type + B/Myeloid drugs). For the cell features, in general we observed that when used through the encoder and decoder, the NN places more importance in the cell features on the decoder rather than the encoder. This might suggest that cell features are more important for performing a conditional decoding of the drug features which is cell-type specific.\n\n### Model selection\n\nUsing the previous NN, we did a leave-one-drug-out for NK cells, which resulted in 146 models. We used the median of the predictions from the 146 models on the cell/drugs for the submission to generate what we call _base predictions_.\n\nThe idea of using this strategy is motivated by the fact that NK cells are the most similar ones to B/Myeloid cells. The other advantage of adopting a leave-one-drug-out approach for NK-cells is that it allows us to estimate how well the model generalizes to unseen drugs, on a per-drug basis per cell type. We also observed that in general, the median was much better than the mean for aggregating the results of the 146 models, since it is more robust to outliers (some models did not generalize well on some drugs, and early stopping selected bad models in those situations).\n\nWe also trained a second neural network with the same hyperparameters, but this time using only the top 256 most variable genes and focusing on the 60 most variable drugs. In this second set of predictions, instead of predicting the 18211 genes, the NN predicts only the top 256 genes used as inputs. We did this because we realized the NN was learning to decide if there was an effect on a given cell type from a small set of genes (essentially, determining where to place values close to 0 in the matrix). We reasoned that training again on only a subset of the data, where most of the changes were concentrated, would help increase performance for that subset of genes. We generated 60 models and computed the median of the predictions, which we referred to as __enhanced predictions__.\n\nWe finally replaced the base predictions with the enhanced predictions (on the subset of genes/drugs). For the final submission, to be more conservative, we mixed the predictions in 0.80/0.20 proportions (0.80 given to the enhanced predictions). __We tested this strategy with several different base predictions, and it always resulted in a boost of performance, which was also the case on the private leaderboard.__ A reproducible notebook for the submission is available at: https://github.com/scapeML/scape/blob/main/docs/notebooks/solution.ipynb.\n\nOne limitation of this strategy is that most of the trained models are very similar, and the blending with the median is very conservative. We also tested different CV strategies, and we found that using blendings of models trained on a 4-CV setting with handpicked drugs on both B/Myeloid cells provided better results in the private leaderboard. However, we didn't trust this strategy that much since it was not very stable and it was hard to understand how well those models were performing in particular cell/drugs combinations.\n\n### Baselines\n\nWe think that having simple baselines is important to understand 1) if the model works, and 2) how good it is. \n\nWe decide to use two simple baselines: predicting zeros (as the baseline used in the competition, which achieves a 0.666 error in the public LB), and the median of the genes grouped by drugs (computed on the training data). The second baseline is more informative.\n\nWe combined those baselines with our leave-on-drug-out strategy to produce plots per drug, so we could have an upper bound estimation of the generalization for each drug. Here is an example for NK and Prednisolone:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2Fba8aa85e97594aed0ab09f324528e0fa%2Fexample-nk-prednisolone.png?generation=1702418460653383&alt=media)\n\n## Specific questions\n\n### 1. Use of prior knowledge\n\nWe decided against using LINCS data in our model because it primarily focuses on cancer cell data, which tends to have a molecular state quite distinct from the PBMCs we were investigating. Despite our exploration of published work and datasets related to predicting drug-induced changes in single-cell states, none of them encompassed the vast array of drug perturbations examined in the challenge. Additionally, we chose not to integrate external data into our approach due to concerns about handling batch effects caused by differences in laboratory settings, protocols, and other related factors.\n\nWe've also tried to use ATAC-seq with no success. We believe that this data would be useful in the case of not having any measurement for B/Myeloid. However, more informative than ATAC-seq are the actual perturbational profiles on the small subset of drugs on those cells.\n\nHere is a summary of different features we tested:\n\n1. Dummy binary variables for cell types and drugs.\n2. Basal omics features, including average expression in DMSO and average accessibility per the ATAC-Seq data.\n3. Summary statistics of the drug response after grouping by cell type and drug, including standard deviation, mean and median.\n4. A “raw” fold-change computed over the raw counts of the single-cell RNA-Seq data (this is, without the corrections applied by limma).\n5. Centroids of the principal component space of the drug response data, using cell-type and drug as grouping variables.\n\nAnd we obtained the best results using the median of the response after grouping by cell type and drug in combination with the raw fold changes, using only a subset with the most variable genes in the dataset.\n\n### 2. Exploration of the problem\n\nWe found that the error distribution for the drugs was more or less even except for the first four drugs, which accounted for 15% of the total error. As expected, we also found out that the response of drugs that were harder to predict was very different from training cell-types in comparison to the test cell-type. For instance, the drug that accounted for the maximum proportion of the error (IN1451), produced a strong response in NK cells, but seemed to have little effect in T cells CD4+, T cells CD8+ and T regulatory cells.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2Fec079268e6ac589644cf28ae1921517f%2Ffig6.png?generation=1702418626577518&alt=media)\n\nOur approach was refined to better understand cell-type errors, aiming to identify the most challenging cell type for accurate prediction. We evaluated 15 drugs across all cell types, selecting 4 at random for testing. This test set was used for cell-type cross-validation, where the model was trained on data from all 15 drugs, excluding the 4 test drugs within a specific cell type. Our method facilitated evaluation of predictive performance for each cell type. Our findings, illustrated in the figure below, indicated that myeloid cells were more difficult to predict than others, corroborating RNA-Seq PCA analysis results.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2F86646a0d4e453e58e2856a74b029b11d%2Ffig7.png?generation=1702418748195660&alt=media)\n\nRegarding genes, we investigated if specific biological functions were harder to predict. An enrichment analysis of the top 5% genes with the highest average error in our local CV setup was conducted using MSigDB hallmarks and [decoupleR](https://www.kaggle.com/code/pablormier/op2-biologically-aware-dimensionality-reduction). This revealed that certain hallmarks, such as epithelial mesenchymal transition and TNF alpha signaling, had a significant number of genes with high error rates:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2Fcc674087faeccb96578b85de39123a7a%2Ffig8.png?generation=1702418832195389&alt=media)\n\n### 3. Model design\n\nWe wanted to check if simpler models could perform just as well. So, we cut down the input features in our models. Considering that our architecture was simple already, we aimed to find the fewest input features that could match the local CV performance of using the top 128 genes with the highest variance.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2Ff29dbc0d67029908ca8bcf86be308426%2Ffig9.png?generation=1702418909501736&alt=media)\n\nInterestingly, we found that models with 8 to 64 input features would achieve similar performance that the model that employed 128 features.\n\nRegarding explainability, even though the model is not easily interpretable, we put some extra care in understanding better how the NN behaved through the leave-on-drug-out + baselines, and by doing permutations on the input data after training a model, to asses the impact on the validation loss.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2F309451c68977bac8fa6e310a788103df%2Ffig10.png?generation=1702418986950956&alt=media)\n\nWe observed that while both components had a direct impact on the performance of our model, the mean of the errors after drug features permutation was higher compared to the average error after cell features permutation. This is something to expect, as we have more data points of gene values per drug (146 data points per cell type except B/Myeloid), but we only have 6 data points grouping by cell type. We used this type of permutation tests to estimate the importance that different features had in the CV error.\n\n### 4. Robustness\n\nOur model included a Gaussian Noise layer from Keras to perturb the input data. We used this to test CV errors for different noise levels. The following figure shows that a gaussian noise of std=0.01 w. This is the value we selected for training the final models:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2F21ba9f226003e9246391f7181ba834bb%2Ffig11.png?generation=1702419081567894&alt=media)\n\n### 5. Documentation & code style\n\nFor convenience, we refactored the code and created a package called “scape” ([https://github.com/scapeML/scape](https://github.com/scapeML/scape)) using [https://pdm-project.org/](https://pdm-project.org/), which contains the code that we finally used for the submission. The code is documented using numpydoc docstrings, and we included a series of notebooks using the scape package to learn how to use it and how to manually create the setup for generating our submission. We have put effort into developing a library that allows for the comfortable configuration and parameterization of neural networks, with an automatic mode for calculating diverse features from drugs and cell lines.\n\n### 6. Reproducibility\n\nIn order to improve reproducibility, we show how the tool package can be installed and used directly from Google Colab https://colab.research.google.com/drive/1-o_lT-ttoKS-nbozj2RQusGoi-vm0-XL?usp=sharing. \n\nWe also included an environment.yml file to exactly recreate the environment we used for testing using conda.\n\n## Sources\n\n- https://github.com/scapeML/scape\n- https://academic.oup.com/bioinformaticsadvances/article/2/1/vbac016/6544613?login=false\n- https://www.kaggle.com/code/pablormier/op2-biologically-aware-dimensionality-reduction",
      "votes": null
    },
    {
      "id": "2562507",
      "postDate": "12/15/2023 13:28:14",
      "content": "<p>wow so 206 model outputs form your submission</p>",
      "rawMarkdown": "wow so 206 model outputs form your submission",
      "votes": null
    },
    {
      "id": "2562921",
      "postDate": "12/15/2023 19:16:31",
      "content": "<p>Thank you for sharing your solution write-up and generously sharing all those resources! Your post is very well written – I enjoyed reading it. Congratulations on your 16th place!</p>\n<p>PS: Looks like you found a pretty unique and fairly robust cross-validation scheme. </p>",
      "rawMarkdown": "Thank you for sharing your solution write-up and generously sharing all those resources! Your post is very well written – I enjoyed reading it. Congratulations on your 16th place!\n\nPS: Looks like you found a pretty unique and fairly robust cross-validation scheme.",
      "votes": null
    },
    {
      "id": "2571065",
      "postDate": "12/22/2023 20:10:56",
      "content": "<p>I read with great interest your detailed and in-depth analysis presented in your raitap. Your conclusions and research methodology, look interesting!</p>\n<p>In examining the data further, I have noticed certain groups of genes that appear to exhibit stronger responses to stressful conditions. These include genes associated with epithelial-mesenchymal transition (HALLMARK_EPITHELIAL_MESENCHYMAL_TRANSITION) and response to hypoxia (HALLMARK_HYPOXIA). In the graphs shown, these groups of genes show increased percentages of expression at various threshold levels compared to other groups, including non-specific \"not housekeeping\" genes.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15071851%2F33b0b034d44a63d1e97bb81b2f69785b%2FScreenshot%202023-12-22%20at%2022.35.27.png?generation=1703275851274143&amp;alt=media\" alt=\"\"><br>\nI would be extremely interested in your opinion on this observation. Do you think that the enhanced response of the genes in question may correlate with an error in the data, or could this reflect their actual increased sensitivity to stress? Perhaps there is some biological mechanism that accounts for this behavior of these genes, or it may be due to the peculiarities of the analysis methods?</p>\n<p>I would appreciate exchanges and additional comments on this issue.</p>\n<p><a href=\"https://www.kaggle.com/code/nikolenkosergei/op2-eda-housekeeping-genes\" target=\"_blank\"><strong>My Notebook</strong></a><br>\nSpecial thanks to Alexander Chervov: <a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-housekeeping-genes\" target=\"_blank\">Notebook</a>. His notebook served as the basis for the analysis.</p>",
      "rawMarkdown": "I read with great interest your detailed and in-depth analysis presented in your raitap. Your conclusions and research methodology, look interesting!\n\nIn examining the data further, I have noticed certain groups of genes that appear to exhibit stronger responses to stressful conditions. These include genes associated with epithelial-mesenchymal transition (HALLMARK_EPITHELIAL_MESENCHYMAL_TRANSITION) and response to hypoxia (HALLMARK_HYPOXIA). In the graphs shown, these groups of genes show increased percentages of expression at various threshold levels compared to other groups, including non-specific \"not housekeeping\" genes.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15071851%2F33b0b034d44a63d1e97bb81b2f69785b%2FScreenshot%202023-12-22%20at%2022.35.27.png?generation=1703275851274143&alt=media)\nI would be extremely interested in your opinion on this observation. Do you think that the enhanced response of the genes in question may correlate with an error in the data, or could this reflect their actual increased sensitivity to stress? Perhaps there is some biological mechanism that accounts for this behavior of these genes, or it may be due to the peculiarities of the analysis methods?\n\nI would appreciate exchanges and additional comments on this issue.\n\n[**My Notebook**](https://www.kaggle.com/code/nikolenkosergei/op2-eda-housekeeping-genes)\nSpecial thanks to Alexander Chervov: [Notebook](https://www.kaggle.com/code/alexandervc/op2-eda-housekeeping-genes). His notebook served as the basis for the analysis.",
      "votes": null
    },
    {
      "id": "2614317",
      "postDate": "01/22/2024 14:44:57",
      "content": "<p>Hi Nikolenko,</p>\n<p>Thanks for your comment! We mostly focused on exploring biological hallmarks contributing to our prediction error. We did not go further in the functional analysis of the differential profiles per se, but what you found out is definitely interesting. Answering your first question, we also found these two terms in the top of our hallmark analysis, indicating that genes belonging to these processes were harder to predict on average, probably because they have a larger variance. </p>\n<p>With regards to your second question, the truth probably lies somewhere in between of the two hypotheses. There are definitely molecular programs that allow the cell to respond to environmental stress (chaperones are a good example of this), and hence it would not be surprising to find groups of genes that prepare the cell or adapt the cell to stress conditions. On the other hand, the technology and preprocessing performed here can also influence what type of signal we observe (e.g. the pseudo-bulking step can remove some single-cell information but it is best practice at the moment). Further data would be needed to disentangle whether this signal is biological or technical (maybe using different single-cell technology or preprocessing pipeline). I hope this helps you!!! </p>\n<p>Los Rodriguez</p>",
      "rawMarkdown": "Hi Nikolenko,\n\nThanks for your comment! We mostly focused on exploring biological hallmarks contributing to our prediction error. We did not go further in the functional analysis of the differential profiles per se, but what you found out is definitely interesting. Answering your first question, we also found these two terms in the top of our hallmark analysis, indicating that genes belonging to these processes were harder to predict on average, probably because they have a larger variance. \n\nWith regards to your second question, the truth probably lies somewhere in between of the two hypotheses. There are definitely molecular programs that allow the cell to respond to environmental stress (chaperones are a good example of this), and hence it would not be surprising to find groups of genes that prepare the cell or adapt the cell to stress conditions. On the other hand, the technology and preprocessing performed here can also influence what type of signal we observe (e.g. the pseudo-bulking step can remove some single-cell information but it is best practice at the moment). Further data would be needed to disentangle whether this signal is biological or technical (maybe using different single-cell technology or preprocessing pipeline). I hope this helps you!!! \n\nLos Rodriguez",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2562507,
      "author_name": "maxleverage",
      "author_url": "",
      "post_date": "12/15/2023 13:28:14",
      "content": "<p>wow so 206 model outputs form your submission</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2562921,
      "author_name": "frenio",
      "author_url": "",
      "post_date": "12/15/2023 19:16:31",
      "content": "<p>Thank you for sharing your solution write-up and generously sharing all those resources! Your post is very well written – I enjoyed reading it. Congratulations on your 16th place!</p>\n<p>PS: Looks like you found a pretty unique and fairly robust cross-validation scheme. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2571065,
      "author_name": "nikolenkosergei",
      "author_url": "",
      "post_date": "12/22/2023 20:10:56",
      "content": "<p>I read with great interest your detailed and in-depth analysis presented in your raitap. Your conclusions and research methodology, look interesting!</p>\n<p>In examining the data further, I have noticed certain groups of genes that appear to exhibit stronger responses to stressful conditions. These include genes associated with epithelial-mesenchymal transition (HALLMARK_EPITHELIAL_MESENCHYMAL_TRANSITION) and response to hypoxia (HALLMARK_HYPOXIA). In the graphs shown, these groups of genes show increased percentages of expression at various threshold levels compared to other groups, including non-specific \"not housekeeping\" genes.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15071851%2F33b0b034d44a63d1e97bb81b2f69785b%2FScreenshot%202023-12-22%20at%2022.35.27.png?generation=1703275851274143&amp;alt=media\" alt=\"\"><br>\nI would be extremely interested in your opinion on this observation. Do you think that the enhanced response of the genes in question may correlate with an error in the data, or could this reflect their actual increased sensitivity to stress? Perhaps there is some biological mechanism that accounts for this behavior of these genes, or it may be due to the peculiarities of the analysis methods?</p>\n<p>I would appreciate exchanges and additional comments on this issue.</p>\n<p><a href=\"https://www.kaggle.com/code/nikolenkosergei/op2-eda-housekeeping-genes\" target=\"_blank\"><strong>My Notebook</strong></a><br>\nSpecial thanks to Alexander Chervov: <a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-housekeeping-genes\" target=\"_blank\">Notebook</a>. His notebook served as the basis for the analysis.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2614317,
          "author_name": "martingarridorc",
          "author_url": "",
          "post_date": "01/22/2024 14:44:57",
          "content": "<p>Hi Nikolenko,</p>\n<p>Thanks for your comment! We mostly focused on exploring biological hallmarks contributing to our prediction error. We did not go further in the functional analysis of the differential profiles per se, but what you found out is definitely interesting. Answering your first question, we also found these two terms in the top of our hallmark analysis, indicating that genes belonging to these processes were harder to predict on average, probably because they have a larger variance. </p>\n<p>With regards to your second question, the truth probably lies somewhere in between of the two hypotheses. There are definitely molecular programs that allow the cell to respond to environmental stress (chaperones are a good example of this), and hence it would not be surprising to find groups of genes that prepare the cell or adapt the cell to stress conditions. On the other hand, the technology and preprocessing performed here can also influence what type of signal we observe (e.g. the pseudo-bulking step can remove some single-cell information but it is best practice at the moment). Further data would be needed to disentangle whether this signal is biological or technical (maybe using different single-cell technology or preprocessing pipeline). I hope this helps you!!! </p>\n<p>Los Rodriguez</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2560018": "We finally had some time to writeup our strategy for the OP2 challenge. It was a super engaging competition, and we're really thankful to both the organizers and fellow competitors for making it such a blast! The whole experience taught us a ton, and we're happy to share what we did/discover along the way. Can't wait for the next challenge!\n\n## Context\n\nIn this competition, the main objective was to predict the effect of drug perturbations on peripheral blood mononuclear cells (PBMCs) from several patient samples. For convenience, we have created a Python package with the model here https://github.com/scapeML/scape. \n\n- Business context: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\n- Data context: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\n\n## Overview of the approach\n\nSimilar to most problems in biological research via omics data, we encountered a high-dimensional feature space (~18k genes) and a low-dimensional observation space (~614 cell/drug combinations) with a low signal-to-noise ratio, where most of the genes show random fluctuations after perturbation. The main data modality to be predicted consisted of signed and log-transformed P-values from differential expression (DE) analysis. In the DE analysis, pseudo-bulk expression profiles from drug-treated cells were compared against the profiles of cells treated with Dimethyl Sulfoxide (DMSO). In addition, challenge organizers also provided the raw data from the single-cell RNA-Seq experiment and from an accompanying ATAC-Seq experiment, conducted only in basal state.\n\nAt the beginning of the challenge, we tested different models using the signed log-pvalues (“de_train” data) alone, such as simple linear models, ensembles of gradient boosting with drug and cell features, conditional variational autoencoders, etc. We soon realized that a simple Neural Network using only a small subset of genes to compute drug and cell features (median of the genes grouped by drug and cell) was enough to have a competitive model.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2F72933c3ebc89980311c9824a42bffde2%2Fnn-architecture.png?generation=1702418417955250&alt=media)\n\nThe figure above shows the final architecture used for our submission. We used a Neural Network that takes as inputs drug and cell features and produces signed log-pvalues. Features were computed as the median of the signed log-pvalues grouped by drugs and cells, calculated from the de_train.parquet file. Additionally, we also estimated log fold-changes (LFCs) from pseudobulk expression, to produce a matrix of the same shape as the de_train data but containing LFCs instead. We also computed the median per cell/drug as features.\n\nSimilar to a Conditional Variational Autoencoder (CVAE), we used cell features both in the encoding part and the decoding part of the NN. Initially, the model consisted of a CVAE that was trained using the cell features as the conditional features to learn an encoding/decoding function conditioned on the particular cell type. However, after testing different ways to train the CVAE (similar to a beta-VAE with different annealing strategies for the KL term), we finally considered a non probabilistic NN since we did not find any practical advantage in this case with respect to a simpler non-probabilistic NN. \n\n### Neural Net\n\nWe created a method to parametrize the architecture of the NN and the feature extraction from different data sources. This is the code needed to create the NN through the scAPE library specifically created for this challenge: https://github.com/scapeML/scape/blob/222b19f47a32afb8d157aecd5e46b23d90b73e9d/scape/_model.py#L678\n\nWe used `n_genes=64` (top 64 genes sorted by variance across conditions).This generates a NN with 9637475 parameters (36.76 MB). The inputs are computed from de_train and from log-fold changes calculated from pseudobulk. Cell features are duplicated both in the encoder and decoder. We did some permutations to estimate the distribution of CV errors, permuting drug and cell features, also in the encoder and decoder part. Drug features have more impact on the final error (something to expect since there are 146 datapoints per cell type + B/Myeloid drugs). For the cell features, in general we observed that when used through the encoder and decoder, the NN places more importance in the cell features on the decoder rather than the encoder. This might suggest that cell features are more important for performing a conditional decoding of the drug features which is cell-type specific.\n\n### Model selection\n\nUsing the previous NN, we did a leave-one-drug-out for NK cells, which resulted in 146 models. We used the median of the predictions from the 146 models on the cell/drugs for the submission to generate what we call _base predictions_.\n\nThe idea of using this strategy is motivated by the fact that NK cells are the most similar ones to B/Myeloid cells. The other advantage of adopting a leave-one-drug-out approach for NK-cells is that it allows us to estimate how well the model generalizes to unseen drugs, on a per-drug basis per cell type. We also observed that in general, the median was much better than the mean for aggregating the results of the 146 models, since it is more robust to outliers (some models did not generalize well on some drugs, and early stopping selected bad models in those situations).\n\nWe also trained a second neural network with the same hyperparameters, but this time using only the top 256 most variable genes and focusing on the 60 most variable drugs. In this second set of predictions, instead of predicting the 18211 genes, the NN predicts only the top 256 genes used as inputs. We did this because we realized the NN was learning to decide if there was an effect on a given cell type from a small set of genes (essentially, determining where to place values close to 0 in the matrix). We reasoned that training again on only a subset of the data, where most of the changes were concentrated, would help increase performance for that subset of genes. We generated 60 models and computed the median of the predictions, which we referred to as __enhanced predictions__.\n\nWe finally replaced the base predictions with the enhanced predictions (on the subset of genes/drugs). For the final submission, to be more conservative, we mixed the predictions in 0.80/0.20 proportions (0.80 given to the enhanced predictions). __We tested this strategy with several different base predictions, and it always resulted in a boost of performance, which was also the case on the private leaderboard.__ A reproducible notebook for the submission is available at: https://github.com/scapeML/scape/blob/main/docs/notebooks/solution.ipynb.\n\nOne limitation of this strategy is that most of the trained models are very similar, and the blending with the median is very conservative. We also tested different CV strategies, and we found that using blendings of models trained on a 4-CV setting with handpicked drugs on both B/Myeloid cells provided better results in the private leaderboard. However, we didn't trust this strategy that much since it was not very stable and it was hard to understand how well those models were performing in particular cell/drugs combinations.\n\n### Baselines\n\nWe think that having simple baselines is important to understand 1) if the model works, and 2) how good it is. \n\nWe decide to use two simple baselines: predicting zeros (as the baseline used in the competition, which achieves a 0.666 error in the public LB), and the median of the genes grouped by drugs (computed on the training data). The second baseline is more informative.\n\nWe combined those baselines with our leave-on-drug-out strategy to produce plots per drug, so we could have an upper bound estimation of the generalization for each drug. Here is an example for NK and Prednisolone:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2Fba8aa85e97594aed0ab09f324528e0fa%2Fexample-nk-prednisolone.png?generation=1702418460653383&alt=media)\n\n## Specific questions\n\n### 1. Use of prior knowledge\n\nWe decided against using LINCS data in our model because it primarily focuses on cancer cell data, which tends to have a molecular state quite distinct from the PBMCs we were investigating. Despite our exploration of published work and datasets related to predicting drug-induced changes in single-cell states, none of them encompassed the vast array of drug perturbations examined in the challenge. Additionally, we chose not to integrate external data into our approach due to concerns about handling batch effects caused by differences in laboratory settings, protocols, and other related factors.\n\nWe've also tried to use ATAC-seq with no success. We believe that this data would be useful in the case of not having any measurement for B/Myeloid. However, more informative than ATAC-seq are the actual perturbational profiles on the small subset of drugs on those cells.\n\nHere is a summary of different features we tested:\n\n1. Dummy binary variables for cell types and drugs.\n2. Basal omics features, including average expression in DMSO and average accessibility per the ATAC-Seq data.\n3. Summary statistics of the drug response after grouping by cell type and drug, including standard deviation, mean and median.\n4. A “raw” fold-change computed over the raw counts of the single-cell RNA-Seq data (this is, without the corrections applied by limma).\n5. Centroids of the principal component space of the drug response data, using cell-type and drug as grouping variables.\n\nAnd we obtained the best results using the median of the response after grouping by cell type and drug in combination with the raw fold changes, using only a subset with the most variable genes in the dataset.\n\n### 2. Exploration of the problem\n\nWe found that the error distribution for the drugs was more or less even except for the first four drugs, which accounted for 15% of the total error. As expected, we also found out that the response of drugs that were harder to predict was very different from training cell-types in comparison to the test cell-type. For instance, the drug that accounted for the maximum proportion of the error (IN1451), produced a strong response in NK cells, but seemed to have little effect in T cells CD4+, T cells CD8+ and T regulatory cells.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2Fec079268e6ac589644cf28ae1921517f%2Ffig6.png?generation=1702418626577518&alt=media)\n\nOur approach was refined to better understand cell-type errors, aiming to identify the most challenging cell type for accurate prediction. We evaluated 15 drugs across all cell types, selecting 4 at random for testing. This test set was used for cell-type cross-validation, where the model was trained on data from all 15 drugs, excluding the 4 test drugs within a specific cell type. Our method facilitated evaluation of predictive performance for each cell type. Our findings, illustrated in the figure below, indicated that myeloid cells were more difficult to predict than others, corroborating RNA-Seq PCA analysis results.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2F86646a0d4e453e58e2856a74b029b11d%2Ffig7.png?generation=1702418748195660&alt=media)\n\nRegarding genes, we investigated if specific biological functions were harder to predict. An enrichment analysis of the top 5% genes with the highest average error in our local CV setup was conducted using MSigDB hallmarks and [decoupleR](https://www.kaggle.com/code/pablormier/op2-biologically-aware-dimensionality-reduction). This revealed that certain hallmarks, such as epithelial mesenchymal transition and TNF alpha signaling, had a significant number of genes with high error rates:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2Fcc674087faeccb96578b85de39123a7a%2Ffig8.png?generation=1702418832195389&alt=media)\n\n### 3. Model design\n\nWe wanted to check if simpler models could perform just as well. So, we cut down the input features in our models. Considering that our architecture was simple already, we aimed to find the fewest input features that could match the local CV performance of using the top 128 genes with the highest variance.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2Ff29dbc0d67029908ca8bcf86be308426%2Ffig9.png?generation=1702418909501736&alt=media)\n\nInterestingly, we found that models with 8 to 64 input features would achieve similar performance that the model that employed 128 features.\n\nRegarding explainability, even though the model is not easily interpretable, we put some extra care in understanding better how the NN behaved through the leave-on-drug-out + baselines, and by doing permutations on the input data after training a model, to asses the impact on the validation loss.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2F309451c68977bac8fa6e310a788103df%2Ffig10.png?generation=1702418986950956&alt=media)\n\nWe observed that while both components had a direct impact on the performance of our model, the mean of the errors after drug features permutation was higher compared to the average error after cell features permutation. This is something to expect, as we have more data points of gene values per drug (146 data points per cell type except B/Myeloid), but we only have 6 data points grouping by cell type. We used this type of permutation tests to estimate the importance that different features had in the CV error.\n\n### 4. Robustness\n\nOur model included a Gaussian Noise layer from Keras to perturb the input data. We used this to test CV errors for different noise levels. The following figure shows that a gaussian noise of std=0.01 w. This is the value we selected for training the final models:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F311361%2F21ba9f226003e9246391f7181ba834bb%2Ffig11.png?generation=1702419081567894&alt=media)\n\n### 5. Documentation & code style\n\nFor convenience, we refactored the code and created a package called “scape” ([https://github.com/scapeML/scape](https://github.com/scapeML/scape)) using [https://pdm-project.org/](https://pdm-project.org/), which contains the code that we finally used for the submission. The code is documented using numpydoc docstrings, and we included a series of notebooks using the scape package to learn how to use it and how to manually create the setup for generating our submission. We have put effort into developing a library that allows for the comfortable configuration and parameterization of neural networks, with an automatic mode for calculating diverse features from drugs and cell lines.\n\n### 6. Reproducibility\n\nIn order to improve reproducibility, we show how the tool package can be installed and used directly from Google Colab https://colab.research.google.com/drive/1-o_lT-ttoKS-nbozj2RQusGoi-vm0-XL?usp=sharing. \n\nWe also included an environment.yml file to exactly recreate the environment we used for testing using conda.\n\n## Sources\n\n- https://github.com/scapeML/scape\n- https://academic.oup.com/bioinformaticsadvances/article/2/1/vbac016/6544613?login=false\n- https://www.kaggle.com/code/pablormier/op2-biologically-aware-dimensionality-reduction",
    "2562507": "wow so 206 model outputs form your submission",
    "2562921": "Thank you for sharing your solution write-up and generously sharing all those resources! Your post is very well written – I enjoyed reading it. Congratulations on your 16th place!\n\nPS: Looks like you found a pretty unique and fairly robust cross-validation scheme.",
    "2571065": "I read with great interest your detailed and in-depth analysis presented in your raitap. Your conclusions and research methodology, look interesting!\n\nIn examining the data further, I have noticed certain groups of genes that appear to exhibit stronger responses to stressful conditions. These include genes associated with epithelial-mesenchymal transition (HALLMARK_EPITHELIAL_MESENCHYMAL_TRANSITION) and response to hypoxia (HALLMARK_HYPOXIA). In the graphs shown, these groups of genes show increased percentages of expression at various threshold levels compared to other groups, including non-specific \"not housekeeping\" genes.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15071851%2F33b0b034d44a63d1e97bb81b2f69785b%2FScreenshot%202023-12-22%20at%2022.35.27.png?generation=1703275851274143&alt=media)\nI would be extremely interested in your opinion on this observation. Do you think that the enhanced response of the genes in question may correlate with an error in the data, or could this reflect their actual increased sensitivity to stress? Perhaps there is some biological mechanism that accounts for this behavior of these genes, or it may be due to the peculiarities of the analysis methods?\n\nI would appreciate exchanges and additional comments on this issue.\n\n[**My Notebook**](https://www.kaggle.com/code/nikolenkosergei/op2-eda-housekeeping-genes)\nSpecial thanks to Alexander Chervov: [Notebook](https://www.kaggle.com/code/alexandervc/op2-eda-housekeeping-genes). His notebook served as the basis for the analysis.",
    "2614317": "Hi Nikolenko,\n\nThanks for your comment! We mostly focused on exploring biological hallmarks contributing to our prediction error. We did not go further in the functional analysis of the differential profiles per se, but what you found out is definitely interesting. Answering your first question, we also found these two terms in the top of our hallmark analysis, indicating that genes belonging to these processes were harder to predict on average, probably because they have a larger variance. \n\nWith regards to your second question, the truth probably lies somewhere in between of the two hypotheses. There are definitely molecular programs that allow the cell to respond to environmental stress (chaperones are a good example of this), and hence it would not be surprising to find groups of genes that prepare the cell or adapt the cell to stress conditions. On the other hand, the technology and preprocessing performed here can also influence what type of signal we observe (e.g. the pseudo-bulking step can remove some single-cell information but it is best practice at the moment). Further data would be needed to disentangle whether this signal is biological or technical (maybe using different single-cell technology or preprocessing pipeline). I hope this helps you!!! \n\nLos Rodriguez"
  },
  "source": "meta"
}