{
  "id": 460946,
  "title": "XGBoost in Compressed Space | 425th Place Solution Writeup for Open Problems – Single-Cell Perturbations",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/460946",
  "author_name": "Anastasiia Popova",
  "post_date": "2023-12-11T22:27:30.401000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>We appreciate the organizers and Kaggle for hosting this interesting competition. We are also grateful to the participants who shared notebooks and ideas. Below is a detailed write-up of our solution, which did not achieve a high score, but we believe it gives a distinctive perspective on the problem. </p>\n<h1>Context</h1>\n<ul>\n<li><p>Competition context: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview</a></p></li>\n<li><p>Data context: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data</a></p></li>\n<li><p>Our notebook with details: <a href=\"https://www.kaggle.com/popovanastya/scp-xgboost-in-compressed-space-for-single-cell-p\" target=\"_blank\">https://www.kaggle.com/popovanastya/scp-xgboost-in-compressed-space-for-single-cell-p</a></p></li>\n</ul>\n<h1>Introduction</h1>\n<p>In this project, we built an XGBoost model to tackle a multi-output regression challenge<em>. Noteworthy for its computational efficiency and robustness to dataset noise, our approach operates effectively within a compressed space.</em>  Our model allows estimation of compound impact on gene expression in a target cell type by leveraging averages from other cells in the compressed space. This approximation can serve as a valuable baseline for training advanced models without prior biological knowledge.</p>\n<h2>Goal</h2>\n<p>In this project, we aimed to build a simple and accurate ML model to predict differential gene expression for different cell types affected by various chemical substances (referred to further as \"drugs\"). </p>\n<h2>Feature Selection</h2>\n<p>We did not integrate any prior biological knowledge for feature augmentation, i.e. feature engineering was done using only cell type, chemical compound name, and 18211 gene differential expression (DE)  from <code>de_train.parquet</code>. </p>\n<h2>Performance</h2>\n<p><strong>Without ensembling</strong> with other models, our method gives <code>0.594</code> for public and <code>0.777</code> for private scores (time of computing is around 3 min). Bagging of XGB will give an additional but insignificant improvement in the scores. </p>\n<h1>Brief Exploratory Data Analysis</h1>\n<p>At first glance, the idea of predicting 18,211 genes using only 2 features and 614 observations seems unsolvable (the combinations of cell and drug do not repeat, we have 6 cells and 146 drugs, and 4 drugs for T cells CD8+ are missed), however, the distributions of most genes are Gaussian with about zero means. These distributions also show numerous outliers that resemble noise, making it challenging to determine their biological relevance. Since the metric is very sensitive to them by definition, we lost hope of figuring this out (for example, for 3 points from zero-mean  Gaussian distribution and 1 outlier the error  <em>RMSE = (0.99-0.98)^2 + (0.01-0.03)^2 + (170 - 168)^2 = 4.0005</em>, which is mostly caused by the outlier even in the case that we predicted it almost correctly). Thus this metric would push any ML model to reward noise prediction instead of signal. Even if we clean the dataset, our prediction will be evaluated using the noisy test set. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Fae23cf95314e73cd203cdbd3eee0d9fd%2Ftypical.png?generation=1702332388281071&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F161b35727ba1a49208002e737cc038cb%2Fnon-typical2.png?generation=1702332414333546&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F525c2f2711afaf85984e9beabf1d46ca%2Fnon-typical3.png?generation=1702332076289987&amp;alt=media\" alt=\"\"></p>\n<p>Based on this logic we decided to work in compressed space (we applied TruncatedSVD). The surprisingly significant (in terms of the public score) result gave the simple averaging cell types for each drug after decompressing (median), and we continued working within compressed space. </p>\n<p>The most useful visualization for us was considering the data as a \"signal\" along rows or columns. It was also a way to assess what exactly was going on with a model's prediction for the test data.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F316867dcb963e504f03567120a747411%2Frow_num.png?generation=1702332457069523&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F35a138ef2baa99eb1c93c53d454bc6fb%2Fgene_num.png?generation=1702332475260578&amp;alt=media\" alt=\"\"></p>\n<h1>Modelling</h1>\n<p>Originally, we believed that some form of averaging over clusters must work in compressed space, and most of the time of the competition was spent on this approach. Sadly, the biggest score was obtained by a simple approximation: <strong>again average (mean) of all cells for each drug, but now in the compressed space</strong>. After decompression, it gives a suspiciously good prediction (<code>0.615</code> was the high score for public notebooks with much more complex models), and the figures of DE look like this averaging captures the main features of the dataset. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F73e2fa374df7c0af73dbf8a86e0bbf56%2Fav_pred.png?generation=1702332573559453&amp;alt=media\" alt=\"\"></p>\n<h2>Data Preprocessing</h2>\n<p>The final model just exploits these observations about average in compressed space. What the average could obviously lose is characteristics specific to a cell type. So we decided to do a feature augmentation in the following way: For each observation, we take the average over the rest of the cells (excluding the cell type for the observation) for a specific drug and use it as features (if we have 36 dimensions after compressing, then we obtain 36 new features). The target was the compressed DE signal for this cell and the drug for the XGBRegressor (also 36 data points). </p>\n<p>It worked well, but still, the model could not capture the difference in cell types treated with different drugs (signal along columns). Therefore, we added more features using TargetEncoder, which gives the numerical representation of the categorical features.  Finally, we added one-hot encoding (we put 1 for cell types across which we did averaging and 0 for the target cell), providing information on what types of cells were averaged. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Fcaa2ab99b6c21dd9e3de0fb98a6a8ebe%2Fdemo_1.png?generation=1702332625709048&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Fb6eee63ad3ef4969780a39f54b923592%2Fdemo_2.png?generation=1702332643555645&amp;alt=media\" alt=\"\"></p>\n<h2>Hyperparameters</h2>\n<p><code>tree_method</code> specifies the tree construction algorithm, we set it to 'hist', which means XGBoost will use a histogram-based algorithm for tree construction. </p>\n<p><code>eval_metric</code> defines the evaluation metric used to assess the model's performance during training, we set it to 'rmse' (Root Mean Squared Error). </p>\n<p><code>max_depth</code> controls the maximum depth of each tree in the ensemble, it's set to 2 to prevent overfitting. </p>\n<p><code>learning_rate</code> is set to 0.2.</p>\n<p><code>n_estimators</code> determines the number of boosting rounds or trees to be built. In our case, the XGBoost model will create an ensemble of 1000 trees.</p>\n<h2>Validation</h2>\n<p>We used 6-fold cross-validation to test our model, but we reduced space before train-test split. Our CV score correlates with public/privite scores. </p>\n<p>The accuracy of our model does not depend on cell types in the dataset. To show this, we computed the metric for cell types excluding one of 4 cell types from the training set (\"NK cells\", \"T cells CD4+\", \"T cells CD8+\", \"T regulatory cells\"). We believe that it happens because our prediction is based on the average of cell types in compressed space, therefore training on data with more cell types gives a better prediction. </p>\n<h2>Scores</h2>\n<p><strong>Baseline model #1:</strong> gives all zeros. </p>\n<p><strong>Model #2:</strong> TruncatedSVD (35 components) -&gt; Inverse TruncatedSVD -&gt; Aggregate Drugs -&gt; <code>.quantile(0.54)</code></p>\n<p><strong>Model #3:</strong> TruncatedSVD (36 components) -&gt; Average for each drug -&gt;  inverse TruncatedSVD </p>\n<p><strong>Model #4:</strong> TruncatedSVD (n_components=36, n_iter=7) -&gt;  TargetEncoder (smooting=8)  -&gt; XGBoost (1000 estimators, feature augmentation with average for each drug) -&gt; inverse TruncatedSVD </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F9b2fc4af41a944a01761f23c21218ced%2Fscores.png?generation=1702332495533180&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Fc7db57c6b82932593aa4a26b829f37b8%2Fxgb1.png?generation=1702332971425025&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Feed8526c3cb7504eb6b79e0354cf7ccc%2Fxgb2.png?generation=1702332986435903&amp;alt=media\" alt=\"\"></p>\n<h1>Conclusions</h1>\n<p>We employed an XGBoost model to address a high-dimensional multi-output regression problem to predict the differential expression of 18,211 genes across 6 cell types affected by 146 chemical substances. Our approach stands out for its computational efficiency and robustness to noise in the dataset, owing to its operation within a compressed space. Simultaneously, it is important to note that the accuracy of our method is constrained by information loss due to compression.</p>\n<p>As demonstrated, the impact of compounds on gene expression in a target cell type can be estimated by leveraging the average values for the other cells within the compressed space. This approximation can serve as a baseline model and be used for training advanced models in subsequent experiments, and it doesn't require any prior biological knowledge.</p>\n<h2>Reproducibility</h2>\n<p>Code will be available and documented on Github at <a href=\"https://github.com/anastasiia-popova/kaggle_OP_SCP_2023\" target=\"_blank\">https://github.com/anastasiia-popova/kaggle_OP_SCP_2023</a>.</p>\n<h1>References</h1>\n<p>We appreciate the authors of the following notebooks for sharing their work</p>\n<p>[1] <a href=\"https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense/notebook\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense/notebook</a></p>\n<p>[2] <a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s</a></p>\n<p>[3] <a href=\"https://www.kaggle.com/code/alexandervc/op2-models-cv-tuning\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-models-cv-tuning</a></p>\n<p>[4] <a href=\"https://www.kaggle.com/code/ambrosm/scp-quickstart\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/scp-quickstart</a></p>",
  "messages": [
    {
      "id": 2558114,
      "postDate": "2023-12-11T22:27:30.400Z",
      "content": "<p>We appreciate the organizers and Kaggle for hosting this interesting competition. We are also grateful to the participants who shared notebooks and ideas. Below is a detailed write-up of our solution, which did not achieve a high score, but we believe it gives a distinctive perspective on the problem. </p>\n<h1>Context</h1>\n<ul>\n<li><p>Competition context: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview</a></p></li>\n<li><p>Data context: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data</a></p></li>\n<li><p>Our notebook with details: <a href=\"https://www.kaggle.com/popovanastya/scp-xgboost-in-compressed-space-for-single-cell-p\" target=\"_blank\">https://www.kaggle.com/popovanastya/scp-xgboost-in-compressed-space-for-single-cell-p</a></p></li>\n</ul>\n<h1>Introduction</h1>\n<p>In this project, we built an XGBoost model to tackle a multi-output regression challenge<em>. Noteworthy for its computational efficiency and robustness to dataset noise, our approach operates effectively within a compressed space.</em>  Our model allows estimation of compound impact on gene expression in a target cell type by leveraging averages from other cells in the compressed space. This approximation can serve as a valuable baseline for training advanced models without prior biological knowledge.</p>\n<h2>Goal</h2>\n<p>In this project, we aimed to build a simple and accurate ML model to predict differential gene expression for different cell types affected by various chemical substances (referred to further as \"drugs\"). </p>\n<h2>Feature Selection</h2>\n<p>We did not integrate any prior biological knowledge for feature augmentation, i.e. feature engineering was done using only cell type, chemical compound name, and 18211 gene differential expression (DE)  from <code>de_train.parquet</code>. </p>\n<h2>Performance</h2>\n<p><strong>Without ensembling</strong> with other models, our method gives <code>0.594</code> for public and <code>0.777</code> for private scores (time of computing is around 3 min). Bagging of XGB will give an additional but insignificant improvement in the scores. </p>\n<h1>Brief Exploratory Data Analysis</h1>\n<p>At first glance, the idea of predicting 18,211 genes using only 2 features and 614 observations seems unsolvable (the combinations of cell and drug do not repeat, we have 6 cells and 146 drugs, and 4 drugs for T cells CD8+ are missed), however, the distributions of most genes are Gaussian with about zero means. These distributions also show numerous outliers that resemble noise, making it challenging to determine their biological relevance. Since the metric is very sensitive to them by definition, we lost hope of figuring this out (for example, for 3 points from zero-mean  Gaussian distribution and 1 outlier the error  <em>RMSE = (0.99-0.98)^2 + (0.01-0.03)^2 + (170 - 168)^2 = 4.0005</em>, which is mostly caused by the outlier even in the case that we predicted it almost correctly). Thus this metric would push any ML model to reward noise prediction instead of signal. Even if we clean the dataset, our prediction will be evaluated using the noisy test set. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Fae23cf95314e73cd203cdbd3eee0d9fd%2Ftypical.png?generation=1702332388281071&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F161b35727ba1a49208002e737cc038cb%2Fnon-typical2.png?generation=1702332414333546&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F525c2f2711afaf85984e9beabf1d46ca%2Fnon-typical3.png?generation=1702332076289987&amp;alt=media\" alt=\"\"></p>\n<p>Based on this logic we decided to work in compressed space (we applied TruncatedSVD). The surprisingly significant (in terms of the public score) result gave the simple averaging cell types for each drug after decompressing (median), and we continued working within compressed space. </p>\n<p>The most useful visualization for us was considering the data as a \"signal\" along rows or columns. It was also a way to assess what exactly was going on with a model's prediction for the test data.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F316867dcb963e504f03567120a747411%2Frow_num.png?generation=1702332457069523&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F35a138ef2baa99eb1c93c53d454bc6fb%2Fgene_num.png?generation=1702332475260578&amp;alt=media\" alt=\"\"></p>\n<h1>Modelling</h1>\n<p>Originally, we believed that some form of averaging over clusters must work in compressed space, and most of the time of the competition was spent on this approach. Sadly, the biggest score was obtained by a simple approximation: <strong>again average (mean) of all cells for each drug, but now in the compressed space</strong>. After decompression, it gives a suspiciously good prediction (<code>0.615</code> was the high score for public notebooks with much more complex models), and the figures of DE look like this averaging captures the main features of the dataset. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F73e2fa374df7c0af73dbf8a86e0bbf56%2Fav_pred.png?generation=1702332573559453&amp;alt=media\" alt=\"\"></p>\n<h2>Data Preprocessing</h2>\n<p>The final model just exploits these observations about average in compressed space. What the average could obviously lose is characteristics specific to a cell type. So we decided to do a feature augmentation in the following way: For each observation, we take the average over the rest of the cells (excluding the cell type for the observation) for a specific drug and use it as features (if we have 36 dimensions after compressing, then we obtain 36 new features). The target was the compressed DE signal for this cell and the drug for the XGBRegressor (also 36 data points). </p>\n<p>It worked well, but still, the model could not capture the difference in cell types treated with different drugs (signal along columns). Therefore, we added more features using TargetEncoder, which gives the numerical representation of the categorical features.  Finally, we added one-hot encoding (we put 1 for cell types across which we did averaging and 0 for the target cell), providing information on what types of cells were averaged. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Fcaa2ab99b6c21dd9e3de0fb98a6a8ebe%2Fdemo_1.png?generation=1702332625709048&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Fb6eee63ad3ef4969780a39f54b923592%2Fdemo_2.png?generation=1702332643555645&amp;alt=media\" alt=\"\"></p>\n<h2>Hyperparameters</h2>\n<p><code>tree_method</code> specifies the tree construction algorithm, we set it to 'hist', which means XGBoost will use a histogram-based algorithm for tree construction. </p>\n<p><code>eval_metric</code> defines the evaluation metric used to assess the model's performance during training, we set it to 'rmse' (Root Mean Squared Error). </p>\n<p><code>max_depth</code> controls the maximum depth of each tree in the ensemble, it's set to 2 to prevent overfitting. </p>\n<p><code>learning_rate</code> is set to 0.2.</p>\n<p><code>n_estimators</code> determines the number of boosting rounds or trees to be built. In our case, the XGBoost model will create an ensemble of 1000 trees.</p>\n<h2>Validation</h2>\n<p>We used 6-fold cross-validation to test our model, but we reduced space before train-test split. Our CV score correlates with public/privite scores. </p>\n<p>The accuracy of our model does not depend on cell types in the dataset. To show this, we computed the metric for cell types excluding one of 4 cell types from the training set (\"NK cells\", \"T cells CD4+\", \"T cells CD8+\", \"T regulatory cells\"). We believe that it happens because our prediction is based on the average of cell types in compressed space, therefore training on data with more cell types gives a better prediction. </p>\n<h2>Scores</h2>\n<p><strong>Baseline model #1:</strong> gives all zeros. </p>\n<p><strong>Model #2:</strong> TruncatedSVD (35 components) -&gt; Inverse TruncatedSVD -&gt; Aggregate Drugs -&gt; <code>.quantile(0.54)</code></p>\n<p><strong>Model #3:</strong> TruncatedSVD (36 components) -&gt; Average for each drug -&gt;  inverse TruncatedSVD </p>\n<p><strong>Model #4:</strong> TruncatedSVD (n_components=36, n_iter=7) -&gt;  TargetEncoder (smooting=8)  -&gt; XGBoost (1000 estimators, feature augmentation with average for each drug) -&gt; inverse TruncatedSVD </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F9b2fc4af41a944a01761f23c21218ced%2Fscores.png?generation=1702332495533180&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Fc7db57c6b82932593aa4a26b829f37b8%2Fxgb1.png?generation=1702332971425025&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Feed8526c3cb7504eb6b79e0354cf7ccc%2Fxgb2.png?generation=1702332986435903&amp;alt=media\" alt=\"\"></p>\n<h1>Conclusions</h1>\n<p>We employed an XGBoost model to address a high-dimensional multi-output regression problem to predict the differential expression of 18,211 genes across 6 cell types affected by 146 chemical substances. Our approach stands out for its computational efficiency and robustness to noise in the dataset, owing to its operation within a compressed space. Simultaneously, it is important to note that the accuracy of our method is constrained by information loss due to compression.</p>\n<p>As demonstrated, the impact of compounds on gene expression in a target cell type can be estimated by leveraging the average values for the other cells within the compressed space. This approximation can serve as a baseline model and be used for training advanced models in subsequent experiments, and it doesn't require any prior biological knowledge.</p>\n<h2>Reproducibility</h2>\n<p>Code will be available and documented on Github at <a href=\"https://github.com/anastasiia-popova/kaggle_OP_SCP_2023\" target=\"_blank\">https://github.com/anastasiia-popova/kaggle_OP_SCP_2023</a>.</p>\n<h1>References</h1>\n<p>We appreciate the authors of the following notebooks for sharing their work</p>\n<p>[1] <a href=\"https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense/notebook\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense/notebook</a></p>\n<p>[2] <a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s</a></p>\n<p>[3] <a href=\"https://www.kaggle.com/code/alexandervc/op2-models-cv-tuning\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-models-cv-tuning</a></p>\n<p>[4] <a href=\"https://www.kaggle.com/code/ambrosm/scp-quickstart\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/scp-quickstart</a></p>",
      "rawMarkdown": "We appreciate the organizers and Kaggle for hosting this interesting competition. We are also grateful to the participants who shared notebooks and ideas. Below is a detailed write-up of our solution, which did not achieve a high score, but we believe it gives a distinctive perspective on the problem. \n\n# Context\n\n- Competition context: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\n\n- Data context: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\n\n- Our notebook with details: https://www.kaggle.com/popovanastya/scp-xgboost-in-compressed-space-for-single-cell-p\n\n# Introduction\n\nIn this project, we built an XGBoost model to tackle a multi-output regression challenge*. Noteworthy for its computational efficiency and robustness to dataset noise, our approach operates effectively within a compressed space.*  Our model allows estimation of compound impact on gene expression in a target cell type by leveraging averages from other cells in the compressed space. This approximation can serve as a valuable baseline for training advanced models without prior biological knowledge.\n\n## Goal \n\nIn this project, we aimed to build a simple and accurate ML model to predict differential gene expression for different cell types affected by various chemical substances (referred to further as \"drugs\"). \n\n\n## Feature Selection\n\nWe did not integrate any prior biological knowledge for feature augmentation, i.e. feature engineering was done using only cell type, chemical compound name, and 18211 gene differential expression (DE)  from `de_train.parquet`. \n\n\n## Performance \n\n**Without ensembling** with other models, our method gives `0.594` for public and `0.777` for private scores (time of computing is around 3 min). Bagging of XGB will give an additional but insignificant improvement in the scores. \n\n\n# Brief Exploratory Data Analysis\n\nAt first glance, the idea of predicting 18,211 genes using only 2 features and 614 observations seems unsolvable (the combinations of cell and drug do not repeat, we have 6 cells and 146 drugs, and 4 drugs for T cells CD8+ are missed), however, the distributions of most genes are Gaussian with about zero means. These distributions also show numerous outliers that resemble noise, making it challenging to determine their biological relevance. Since the metric is very sensitive to them by definition, we lost hope of figuring this out (for example, for 3 points from zero-mean  Gaussian distribution and 1 outlier the error  *RMSE = (0.99-0.98)^2 + (0.01-0.03)^2 + (170 - 168)^2 = 4.0005*, which is mostly caused by the outlier even in the case that we predicted it almost correctly). Thus this metric would push any ML model to reward noise prediction instead of signal. Even if we clean the dataset, our prediction will be evaluated using the noisy test set. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Fae23cf95314e73cd203cdbd3eee0d9fd%2Ftypical.png?generation=1702332388281071&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F161b35727ba1a49208002e737cc038cb%2Fnon-typical2.png?generation=1702332414333546&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F525c2f2711afaf85984e9beabf1d46ca%2Fnon-typical3.png?generation=1702332076289987&alt=media)\n\nBased on this logic we decided to work in compressed space (we applied TruncatedSVD). The surprisingly significant (in terms of the public score) result gave the simple averaging cell types for each drug after decompressing (median), and we continued working within compressed space. \n\nThe most useful visualization for us was considering the data as a \"signal\" along rows or columns. It was also a way to assess what exactly was going on with a model's prediction for the test data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F316867dcb963e504f03567120a747411%2Frow_num.png?generation=1702332457069523&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F35a138ef2baa99eb1c93c53d454bc6fb%2Fgene_num.png?generation=1702332475260578&alt=media)\n\n# Modelling \n\nOriginally, we believed that some form of averaging over clusters must work in compressed space, and most of the time of the competition was spent on this approach. Sadly, the biggest score was obtained by a simple approximation: **again average (mean) of all cells for each drug, but now in the compressed space**. After decompression, it gives a suspiciously good prediction (`0.615` was the high score for public notebooks with much more complex models), and the figures of DE look like this averaging captures the main features of the dataset. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F73e2fa374df7c0af73dbf8a86e0bbf56%2Fav_pred.png?generation=1702332573559453&alt=media)\n\n\n## Data Preprocessing\n\nThe final model just exploits these observations about average in compressed space. What the average could obviously lose is characteristics specific to a cell type. So we decided to do a feature augmentation in the following way: For each observation, we take the average over the rest of the cells (excluding the cell type for the observation) for a specific drug and use it as features (if we have 36 dimensions after compressing, then we obtain 36 new features). The target was the compressed DE signal for this cell and the drug for the XGBRegressor (also 36 data points). \n\nIt worked well, but still, the model could not capture the difference in cell types treated with different drugs (signal along columns). Therefore, we added more features using TargetEncoder, which gives the numerical representation of the categorical features.  Finally, we added one-hot encoding (we put 1 for cell types across which we did averaging and 0 for the target cell), providing information on what types of cells were averaged. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Fcaa2ab99b6c21dd9e3de0fb98a6a8ebe%2Fdemo_1.png?generation=1702332625709048&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Fb6eee63ad3ef4969780a39f54b923592%2Fdemo_2.png?generation=1702332643555645&alt=media)\n\n\n## Hyperparameters\n\n`tree_method` specifies the tree construction algorithm, we set it to 'hist', which means XGBoost will use a histogram-based algorithm for tree construction. \n\n`eval_metric` defines the evaluation metric used to assess the model's performance during training, we set it to 'rmse' (Root Mean Squared Error). \n\n`max_depth` controls the maximum depth of each tree in the ensemble, it's set to 2 to prevent overfitting. \n\n`learning_rate` is set to 0.2.\n\n`n_estimators` determines the number of boosting rounds or trees to be built. In our case, the XGBoost model will create an ensemble of 1000 trees.\n\n\n## Validation \n\nWe used 6-fold cross-validation to test our model, but we reduced space before train-test split. Our CV score correlates with public/privite scores. \n\nThe accuracy of our model does not depend on cell types in the dataset. To show this, we computed the metric for cell types excluding one of 4 cell types from the training set (\"NK cells\", \"T cells CD4+\", \"T cells CD8+\", \"T regulatory cells\"). We believe that it happens because our prediction is based on the average of cell types in compressed space, therefore training on data with more cell types gives a better prediction. \n\n## Scores\n\n**Baseline model #1:** gives all zeros. \n\n**Model #2:** TruncatedSVD (35 components) -> Inverse TruncatedSVD -> Aggregate Drugs -> `.quantile(0.54)`\n\n**Model #3:** TruncatedSVD (36 components) -> Average for each drug ->  inverse TruncatedSVD \n\n**Model #4:** TruncatedSVD (n_components=36, n_iter=7) ->  TargetEncoder (smooting=8)  -> XGBoost (1000 estimators, feature augmentation with average for each drug) -> inverse TruncatedSVD \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F9b2fc4af41a944a01761f23c21218ced%2Fscores.png?generation=1702332495533180&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Fc7db57c6b82932593aa4a26b829f37b8%2Fxgb1.png?generation=1702332971425025&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Feed8526c3cb7504eb6b79e0354cf7ccc%2Fxgb2.png?generation=1702332986435903&alt=media)\n# Conclusions\n\nWe employed an XGBoost model to address a high-dimensional multi-output regression problem to predict the differential expression of 18,211 genes across 6 cell types affected by 146 chemical substances. Our approach stands out for its computational efficiency and robustness to noise in the dataset, owing to its operation within a compressed space. Simultaneously, it is important to note that the accuracy of our method is constrained by information loss due to compression.\n\nAs demonstrated, the impact of compounds on gene expression in a target cell type can be estimated by leveraging the average values for the other cells within the compressed space. This approximation can serve as a baseline model and be used for training advanced models in subsequent experiments, and it doesn't require any prior biological knowledge.\n\n## Reproducibility \nCode will be available and documented on Github at https://github.com/anastasiia-popova/kaggle_OP_SCP_2023.\n\n# References \n\nWe appreciate the authors of the following notebooks for sharing their work\n\n[1] https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense/notebook\n\n[2] https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s\n\n[3] https://www.kaggle.com/code/alexandervc/op2-models-cv-tuning\n\n[4] https://www.kaggle.com/code/ambrosm/scp-quickstart\n",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2558114": "We appreciate the organizers and Kaggle for hosting this interesting competition. We are also grateful to the participants who shared notebooks and ideas. Below is a detailed write-up of our solution, which did not achieve a high score, but we believe it gives a distinctive perspective on the problem. \n\n# Context\n\n- Competition context: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\n\n- Data context: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\n\n- Our notebook with details: https://www.kaggle.com/popovanastya/scp-xgboost-in-compressed-space-for-single-cell-p\n\n# Introduction\n\nIn this project, we built an XGBoost model to tackle a multi-output regression challenge*. Noteworthy for its computational efficiency and robustness to dataset noise, our approach operates effectively within a compressed space.*  Our model allows estimation of compound impact on gene expression in a target cell type by leveraging averages from other cells in the compressed space. This approximation can serve as a valuable baseline for training advanced models without prior biological knowledge.\n\n## Goal \n\nIn this project, we aimed to build a simple and accurate ML model to predict differential gene expression for different cell types affected by various chemical substances (referred to further as \"drugs\"). \n\n\n## Feature Selection\n\nWe did not integrate any prior biological knowledge for feature augmentation, i.e. feature engineering was done using only cell type, chemical compound name, and 18211 gene differential expression (DE)  from `de_train.parquet`. \n\n\n## Performance \n\n**Without ensembling** with other models, our method gives `0.594` for public and `0.777` for private scores (time of computing is around 3 min). Bagging of XGB will give an additional but insignificant improvement in the scores. \n\n\n# Brief Exploratory Data Analysis\n\nAt first glance, the idea of predicting 18,211 genes using only 2 features and 614 observations seems unsolvable (the combinations of cell and drug do not repeat, we have 6 cells and 146 drugs, and 4 drugs for T cells CD8+ are missed), however, the distributions of most genes are Gaussian with about zero means. These distributions also show numerous outliers that resemble noise, making it challenging to determine their biological relevance. Since the metric is very sensitive to them by definition, we lost hope of figuring this out (for example, for 3 points from zero-mean  Gaussian distribution and 1 outlier the error  *RMSE = (0.99-0.98)^2 + (0.01-0.03)^2 + (170 - 168)^2 = 4.0005*, which is mostly caused by the outlier even in the case that we predicted it almost correctly). Thus this metric would push any ML model to reward noise prediction instead of signal. Even if we clean the dataset, our prediction will be evaluated using the noisy test set. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Fae23cf95314e73cd203cdbd3eee0d9fd%2Ftypical.png?generation=1702332388281071&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F161b35727ba1a49208002e737cc038cb%2Fnon-typical2.png?generation=1702332414333546&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F525c2f2711afaf85984e9beabf1d46ca%2Fnon-typical3.png?generation=1702332076289987&alt=media)\n\nBased on this logic we decided to work in compressed space (we applied TruncatedSVD). The surprisingly significant (in terms of the public score) result gave the simple averaging cell types for each drug after decompressing (median), and we continued working within compressed space. \n\nThe most useful visualization for us was considering the data as a \"signal\" along rows or columns. It was also a way to assess what exactly was going on with a model's prediction for the test data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F316867dcb963e504f03567120a747411%2Frow_num.png?generation=1702332457069523&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F35a138ef2baa99eb1c93c53d454bc6fb%2Fgene_num.png?generation=1702332475260578&alt=media)\n\n# Modelling \n\nOriginally, we believed that some form of averaging over clusters must work in compressed space, and most of the time of the competition was spent on this approach. Sadly, the biggest score was obtained by a simple approximation: **again average (mean) of all cells for each drug, but now in the compressed space**. After decompression, it gives a suspiciously good prediction (`0.615` was the high score for public notebooks with much more complex models), and the figures of DE look like this averaging captures the main features of the dataset. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F73e2fa374df7c0af73dbf8a86e0bbf56%2Fav_pred.png?generation=1702332573559453&alt=media)\n\n\n## Data Preprocessing\n\nThe final model just exploits these observations about average in compressed space. What the average could obviously lose is characteristics specific to a cell type. So we decided to do a feature augmentation in the following way: For each observation, we take the average over the rest of the cells (excluding the cell type for the observation) for a specific drug and use it as features (if we have 36 dimensions after compressing, then we obtain 36 new features). The target was the compressed DE signal for this cell and the drug for the XGBRegressor (also 36 data points). \n\nIt worked well, but still, the model could not capture the difference in cell types treated with different drugs (signal along columns). Therefore, we added more features using TargetEncoder, which gives the numerical representation of the categorical features.  Finally, we added one-hot encoding (we put 1 for cell types across which we did averaging and 0 for the target cell), providing information on what types of cells were averaged. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Fcaa2ab99b6c21dd9e3de0fb98a6a8ebe%2Fdemo_1.png?generation=1702332625709048&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Fb6eee63ad3ef4969780a39f54b923592%2Fdemo_2.png?generation=1702332643555645&alt=media)\n\n\n## Hyperparameters\n\n`tree_method` specifies the tree construction algorithm, we set it to 'hist', which means XGBoost will use a histogram-based algorithm for tree construction. \n\n`eval_metric` defines the evaluation metric used to assess the model's performance during training, we set it to 'rmse' (Root Mean Squared Error). \n\n`max_depth` controls the maximum depth of each tree in the ensemble, it's set to 2 to prevent overfitting. \n\n`learning_rate` is set to 0.2.\n\n`n_estimators` determines the number of boosting rounds or trees to be built. In our case, the XGBoost model will create an ensemble of 1000 trees.\n\n\n## Validation \n\nWe used 6-fold cross-validation to test our model, but we reduced space before train-test split. Our CV score correlates with public/privite scores. \n\nThe accuracy of our model does not depend on cell types in the dataset. To show this, we computed the metric for cell types excluding one of 4 cell types from the training set (\"NK cells\", \"T cells CD4+\", \"T cells CD8+\", \"T regulatory cells\"). We believe that it happens because our prediction is based on the average of cell types in compressed space, therefore training on data with more cell types gives a better prediction. \n\n## Scores\n\n**Baseline model #1:** gives all zeros. \n\n**Model #2:** TruncatedSVD (35 components) -> Inverse TruncatedSVD -> Aggregate Drugs -> `.quantile(0.54)`\n\n**Model #3:** TruncatedSVD (36 components) -> Average for each drug ->  inverse TruncatedSVD \n\n**Model #4:** TruncatedSVD (n_components=36, n_iter=7) ->  TargetEncoder (smooting=8)  -> XGBoost (1000 estimators, feature augmentation with average for each drug) -> inverse TruncatedSVD \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2F9b2fc4af41a944a01761f23c21218ced%2Fscores.png?generation=1702332495533180&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Fc7db57c6b82932593aa4a26b829f37b8%2Fxgb1.png?generation=1702332971425025&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13819485%2Feed8526c3cb7504eb6b79e0354cf7ccc%2Fxgb2.png?generation=1702332986435903&alt=media)\n# Conclusions\n\nWe employed an XGBoost model to address a high-dimensional multi-output regression problem to predict the differential expression of 18,211 genes across 6 cell types affected by 146 chemical substances. Our approach stands out for its computational efficiency and robustness to noise in the dataset, owing to its operation within a compressed space. Simultaneously, it is important to note that the accuracy of our method is constrained by information loss due to compression.\n\nAs demonstrated, the impact of compounds on gene expression in a target cell type can be estimated by leveraging the average values for the other cells within the compressed space. This approximation can serve as a baseline model and be used for training advanced models in subsequent experiments, and it doesn't require any prior biological knowledge.\n\n## Reproducibility \nCode will be available and documented on Github at https://github.com/anastasiia-popova/kaggle_OP_SCP_2023.\n\n# References \n\nWe appreciate the authors of the following notebooks for sharing their work\n\n[1] https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense/notebook\n\n[2] https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s\n\n[3] https://www.kaggle.com/code/alexandervc/op2-models-cv-tuning\n\n[4] https://www.kaggle.com/code/ambrosm/scp-quickstart\n"
  }
}