{
  "id": 458661,
  "title": "#18: Py-boost predicting t-scores",
  "url": "/competitions/open-problems-single-cell-perturbations/writeups/ambrosm-18-py-boost-predicting-t-scores",
  "author_name": "",
  "post_date": "2023-12-11T22:08:00.973Z",
  "votes": 54,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Did you notice that in this competition  few real EDA notebooks have been published? Besides explaining my machine learning model, I'd like to share some observations which help understand the data and the intricacies of Limma.</p>\n<h1>Integration of biological knowledge</h1>\n<h2>Don't trust the cell types!</h2>\n<p>Let's recapitulate the course of the experiment in a simplified form. We can imagine an experimenter who is in front of a large pot of human blood cells. The pot contains a mixture of six cell types in certain proportions. T cells CD4+ take the largest share (42 %), only 2 % are T regulatory cells:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F26d3a6c5971cf2a5a441cb550424728e%2Fpie-chart.png?generation=1701390049641476&amp;alt=media\" alt=\"\"><br>\nThe experimenter now takes 145 droplets out of the large pot. Every droplet contains 1550 ± 240 cells (normally distributed). If we counted the cells per cell type in the droplets, we'd see a multinomial distribution. The 145 droplets might be composed like in the following bar chart (fictitious data, sorted from smallest to largest droplet):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F838eb6fe789104225c5bf6cb0f73e060%2Fdrops-before.png?generation=1701390069789098&amp;alt=media\" alt=\"\"><br>\nIn the next step, the experimenter adds 145 substances to the 145 droplets and waits 24 hours. After 24 hours the cells are analyzed. If we count the cells again, we get the following picture, as taken from the competition's training data (cell counts for the test data are hidden):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F9a57da671b14ec3198d61c54861cc27f%2Fdrops-after.png?generation=1701390083658870&amp;alt=media\" alt=\"\"><br>\nIn this diagram we first see that some compounds are so toxic that in some droplets less than 100 cells survive. These droplets are represented by the leftmost bars in the bar chart.</p>\n<p>The second observation is much more important: The long red part in the bars for Oprozomib and IN1451 show that these droplets contain several hundred T regulatory cells — much more than at the start of the experiment. Other compounds (e.g., CGM-079) have too many T cells CD8+ (green bar). How can we interpret this observation?</p>\n<ol>\n<li>Does IN1451 incite the T regulatory cells to multiply so that we have five times more of them after 24 hours? No.</li>\n<li>Does IN1451 magically convert NK cells into T regulatory cells? No.</li>\n<li>Does IN1451 affect the cells in such a way that they are misclassified? Maybe.</li>\n</ol>\n<p>Discussing differential gene expression for specific cell types becomes pointless if the cells change their type during the experiment. For the Kaggle competition this means that we have to deal with many outliers: Beyond the at least five toxic compounds, there are at least seven compounds which change the cells' types. Differential expression for these outliers is hard to model. They make cross-validation unreliable, and the outliers in the private leaderboard can't even be predicted by probing the public leaderboard.</p>\n<h2>Cell count shouldn't affect differential gene expression</h2>\n<p>Does gene expression in a cell depend on how many cells are in the experiment? Theoretically, it doesn't. A cell behaves the same way whether there are 10 cells in the experiment or 10000. We'd expect, however, a difference in the significance of the experimental results: An experiment with 10000 cells should give more precise measurements than a 10-cell experiment: As the cell count grows, variance of the measurements should decrease, t-score should be farther away from zero, and pvalues should decrease.</p>\n<p>The competition data don't fulfill this expectation. If we plot the mean t-scores versus the cell count for the 602 cell type–compound combinations (excluding the control compounds), we see a linear relationship: For every cell type, compounds with lower cell counts have positive t-score means, and compounds with higher cell counts have negative t-score means. This correlation between cell counts and t-scores shouldn't exist. It is an artefact of Limma rather than a biological effect.</p>\n<p>You can plot the diagram with median or variance instead of mean — it will look similar. You can even compare the cell counts to the first principal component of the t-scores and see the same correlation. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F50f71507c0241429e92b4ae50b092222%2Fcell-t-before.png?generation=1701390101863675&amp;alt=media\" alt=\"\"></p>\n<p>We can now put together a list of 20 compounds which are to be considered outliers because of low cell counts. Notice that we don't declare single rows of the dataset to be outliers, but all 86 rows related to the 20 compounds:</p>\n<pre><code>Outliers\n\nAT13387                             T regulatory cells\nAlvocidib                         ≤  several cell \nBAY                        mean t-score  CD8+ cells &gt; \nBMS                          T cells CD8+,  Myeloid cells\nBelinostat                        control compound  too many cells\nCEP (Delanzomib)            ≤  several cell \nCGM                           too many T cells CD8+\nCGP                          ≤  several cell \nDabrafenib                        control compound  too many cells\nGanetespib (STA)               T regulatory cells, too many NK cells\nI-BET151                          too many T cells CD8+\nIN1451                            ≤  several cell \nLY2090314                           T cells CD8+\nMLN                           ≤  several cell \nOprozomib (ONX )              ≤  several cell \nProscillaridin A;Proscillaridin-A ≤  several cell \nResminostat                        T cells CD8+\nScriptaid                           T regulatory cells\nUNII-BXU45ZH6LI                     T cells CD8+\nVorinostat                          T regulatory cell\n</code></pre>\n<p>After removing the outliers, the diagram looks much cleaner. The variance of the cell counts remains. It is a source of noise which impedes the correct interpretation (and prediction) of differential expressions. Maybe we'd get cleaner data if we equalized the cell counts before library size normalization. This would amount to throwing away a part of the measurements, which isn't desirable either.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2Ffd906cb718a3308dd60f9dc5e56e1c3d%2Fcell-t-after.png?generation=1701390162014901&amp;alt=media\" alt=\"\"></p>\n<p>After considering the small size of the dataset, the amount of noise and the Limma artefacts (more of them will be shown in the next section), I didn't try to integrate any external biological data into my model. </p>\n<h1>Exploration of the problem</h1>\n<h2>A mixture of probability distributions</h2>\n<p>A histogram of a single row of the training data (18211 t-scores for T cells CD8+ treated with Scriptaid) shows that the distribution is multimodal.</p>\n<p>The highest mode consists of 269 genes which all have an identical t-score of -3.769. It turns out that these are the 269 genes which are never expressed in T cells CD8+, neither with the negative control nor with any other compound. Isn't this strange? A gene which is never expressed in the whole experiment should have a log-fold change of zero and should not get a t-score at all (because t-score computation involves a division by the variance, and the variance of a never-expressed gene is zero).</p>\n<p>For Myeloid cells treated with Foretinib, 3856 genes are not expressed (RNA count of zero), yet most of them have a positive t-score. Their highest t-score is 6.228 (resulting in a pvalue of 4e-10 and a log10pvalue of 9.33). If an RNA count is zero, the corresponding log-fold-change (and t-score) should never be positive.</p>\n<p>We may say that the distribution of the values is a mixture of two distributions:</p>\n<ol>\n<li>The values for the genes which are expressed (blue) have a more or less bell-shaped distribution.</li>\n<li>The values for the genes which are not expressed (orange) have a distribution with an unusual shape, and it is strange that positive differential expressions are reported when not a single piece of RNA is counted.</li>\n</ol>\n<p>What we see here is an artefact of Limma, which affects every row of the datset. It suggests that Limma output can be biased and is not ideal for investigating cell-type translation of differential expressions.</p>\n<pre><code> expressed in T cells CD8+ Scriptaid:     \n not expressed in T cells CD8+ Scriptaid:  \n: -. for  genes not expressed at  in T cells CD8+\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F7134e68a91b05970eba9705209a7ed46%2Fmixture1.png?generation=1701390193470010&amp;alt=media\" alt=\"\"></p>\n<pre><code> expressed in Myeloid cells Foretinib:     \n not expressed in Myeloid cells Foretinib:  \n: . for  genes not expressed at  in Myeloid cells\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2Ff2d00ede4d7735ddf76f1c16aec6fe79%2Fmixture2.png?generation=1701390213955949&amp;alt=media\" alt=\"\"></p>\n<h2>An ideal training set</h2>\n<p>In the competition overview, the organizers ask: <em>Do you have any evidence to suggest how you might develop an ideal training set for cell type translation beyond random sampling of compounds in cell types? What is the relationship between the number of compounds measured in the held-out cell types and model performance?</em></p>\n<p>I think we are not yet ready to answer these questions. We first need cleaner data (and more of it):</p>\n<ul>\n<li>Cell types must be classified correctly. This may imply that we limit the scope of the work to compounds which do not hamper cell type classification.</li>\n<li>Samples containing too few cells must be eliminated from the dataset. These samples just add hay to the haystack where we want to find the needle.</li>\n<li>Even if we have many cells, genes with low rna counts may need to be eliminated. Otherwise they add even more hay to the haystack.</li>\n</ul>\n<p>Second, modeling strange t-scores of genes which are never expressed is a waste of time. We need to define a machine-learning task and a metric which reward biological insight rather than forcing people into modeling the noise created by upstream processing steps:</p>\n<ul>\n<li>As t-scores are always affected by cell counts and variance estimates, a metric based on less highly-processed data (i.e., log-fold changes or rna counts rather than log10pvalues or t-scores) may lead research into a better direction.</li>\n<li>Even with log-fold changes, genes with low rna count make more noise than genes with high rna count. A suitable metric should account for this fact.</li>\n</ul>\n<h1>Model design</h1>\n<h2>T-scores are better than log10pvalues</h2>\n<p>Limma performs t-tests. t-scores are (almost) normally distributed, which is good for machine learning inputs. For this competition, the t-scores were nonlinearly transformed to log10pvalues. The transformation squeezes the nice bell shape into a distribution with a much higher kurtosis.</p>\n<p>My machine learning models perform better if I transform the log10pvalues into t-score in a preprocessing step, predict t-scores, and transform the predictions back afterwards. Perhaps working with log-fold changes or RNA counts would be even better.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2Fad0c81caaf72f36d9a859296565e5882%2Ft-score-is-better.png?generation=1701390233521477&amp;alt=media\" alt=\"\"></p>\n<h2>The models</h2>\n<p>I developed four models:</p>\n<ul>\n<li>Py-boost</li>\n<li>A recommender system based on ridge regression</li>\n<li>A recommender system based on k nearest neighbors</li>\n<li>ExtraTrees</li>\n</ul>\n<p>I first implemented the Py-boost model, derived from <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a>'s public notebook.</p>\n<p>I then implemented the ExtraTrees model, which resembles <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a>'s Py-boost model. All the decision trees are fully grown (i.e., overfitted). The model gets its generalization capability from noise which is added to the target-encoded features deliberately.</p>\n<p>I then implemented the knn <a href=\"https://en.wikipedia.org/wiki/Recommender_system\" target=\"_blank\">recommender system</a> to have some diversity in the ensemble. Cell types and compounds are identified with users and items, respectively; gene expression is identified with item ratings by users.</p>\n<p>ExtraTrees and k-nearest-neighbors share the weakness that they cannot extrapolate. Even after dimensionality reduction, our training dataset essentially consists of 614 points in a high-dimensional space, so that most of the points will lie on the convex hull. Of the 255 test points, many will lie outside the convex hull of the training points, which means that the model must extrapolate. To bring the extrapolation capability into the game, I implemented the ridge regression model. </p>\n<p>The models have cv scores between 0.878 (ExtraTrees) and 0.906 (Py-boost). Py-boost, which was the worst in cross-validation, has the best public and private lb scores (0.572 and 0.748, respectively).</p>\n<h2>Data augmentation</h2>\n<p>One of the models (k nearest neighbors) is fed with <strong>data augmentation</strong>: If we know the differential expressions for two compounds, we may assume that a mixture of the two compounds will produce a differential expression which is the average of the two single-compound differential expressions.</p>\n<p>I experimented with another kind of data augmentation a well: Because there are more than twice as many T cells CD4+ as either Myeloid or B cells and I knew that the cell count biases the results of Limma, I reduced the cell count of the T cells CD4+, pseudobulked them, ran them through Limma and added the results to the training data as another cell type. This augmentation improved the scores of ExtraTrees, but not to the level of Py-boost. Perhaps I should have combined the additional cell type with Py-boost…</p>\n<h1>Robustness</h1>\n<p>The robustness of my models is demonstrated in two ways:</p>\n<p>(1) The models are fully cross-validated. The cross-validation strategy, first documented in <a href=\"https://www.kaggle.com/code/ambrosm/scp-quickstart\" target=\"_blank\">SCP Quickstart</a>, ensures that the model is validated on predicting cell_type–sm_name combinations so that it knows only 17 other compounds for the same cell type. This cross-validation strategy is more robust than the ordinary shuffled KFold, where the model knows 4/5 of all compounds for the same cell type. (And it is much more robust than a simple train-test-split.)</p>\n<p>I have to admit, though, that I'm not happy with the cv–lb correspondence.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F424a99a8b518dd08b96b0026f598c557%2Fcv-scheme.png?generation=1701390265025626&amp;alt=media\" alt=\"\"></p>\n<p>(2) For all models the performance was tested after adding Gaussian noise to the input t-scores. All models are robust against small noise. When the noise gets stronger, the knn and ExtraTrees models suffer more than Py-boost and the ridge recommender system.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F9f9538ca1d28e570ae930339c9490145%2Fnoise.png?generation=1701409934387211&amp;alt=media\" alt=\"\"></p>\n<h1>Documentation and code style</h1>\n<p>The code is documented in the notebooks.</p>\n<h1>Reproducibility</h1>\n<p>Source code is here:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense\" target=\"_blank\">EDA which makes sense ⭐️⭐️⭐️⭐️⭐️</a></li>\n<li><a href=\"https://www.kaggle.com/code/ambrosm/scp-26-py-boost-recommender-system-and-et\" target=\"_blank\">SCP #26: Py-boost, recommender system and ET</a></li>\n<li><a href=\"https://github.com/Ambros-M/Single-Cell-Perturbations-2023\" target=\"_blank\">GitHub</a></li>\n</ul>\n<h1>Conclusion</h1>\n<p>Let me conclude by summarizing the four main messages of this post:</p>\n<ol>\n<li>Recommender systems are a promising starting point for developing models for cross-cell-type differential gene expression prediction. Because of commercial interests, recommender systems are a well-researched topic, and a lot of information is available.</li>\n<li>Data augmentation is useful, and mixtures of compounds are a natural approach to data augmentation.</li>\n<li>Although Kaggle competitions with data cleaning, outlier removal and unusual metrics are entertaining, the research objective would profit from another setting. Providing clean data and scoring with a well-understood metric would help participants focus on the real topic rather than the noise in the data.</li>\n<li>We have seen that Limma in certain situations produces biased outputs. I hope that professional Limma users are aware of these effects and account for them when interpreting results in their research.</li>\n</ol>",
  "messages": [
    {
      "id": "2544595",
      "postDate": "12/01/2023 00:34:21",
      "content": "<p>Did you notice that in this competition  few real EDA notebooks have been published? Besides explaining my machine learning model, I'd like to share some observations which help understand the data and the intricacies of Limma.</p>\n<h1>Integration of biological knowledge</h1>\n<h2>Don't trust the cell types!</h2>\n<p>Let's recapitulate the course of the experiment in a simplified form. We can imagine an experimenter who is in front of a large pot of human blood cells. The pot contains a mixture of six cell types in certain proportions. T cells CD4+ take the largest share (42 %), only 2 % are T regulatory cells:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F26d3a6c5971cf2a5a441cb550424728e%2Fpie-chart.png?generation=1701390049641476&amp;alt=media\" alt=\"\"><br>\nThe experimenter now takes 145 droplets out of the large pot. Every droplet contains 1550 ± 240 cells (normally distributed). If we counted the cells per cell type in the droplets, we'd see a multinomial distribution. The 145 droplets might be composed like in the following bar chart (fictitious data, sorted from smallest to largest droplet):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F838eb6fe789104225c5bf6cb0f73e060%2Fdrops-before.png?generation=1701390069789098&amp;alt=media\" alt=\"\"><br>\nIn the next step, the experimenter adds 145 substances to the 145 droplets and waits 24 hours. After 24 hours the cells are analyzed. If we count the cells again, we get the following picture, as taken from the competition's training data (cell counts for the test data are hidden):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F9a57da671b14ec3198d61c54861cc27f%2Fdrops-after.png?generation=1701390083658870&amp;alt=media\" alt=\"\"><br>\nIn this diagram we first see that some compounds are so toxic that in some droplets less than 100 cells survive. These droplets are represented by the leftmost bars in the bar chart.</p>\n<p>The second observation is much more important: The long red part in the bars for Oprozomib and IN1451 show that these droplets contain several hundred T regulatory cells — much more than at the start of the experiment. Other compounds (e.g., CGM-079) have too many T cells CD8+ (green bar). How can we interpret this observation?</p>\n<ol>\n<li>Does IN1451 incite the T regulatory cells to multiply so that we have five times more of them after 24 hours? No.</li>\n<li>Does IN1451 magically convert NK cells into T regulatory cells? No.</li>\n<li>Does IN1451 affect the cells in such a way that they are misclassified? Maybe.</li>\n</ol>\n<p>Discussing differential gene expression for specific cell types becomes pointless if the cells change their type during the experiment. For the Kaggle competition this means that we have to deal with many outliers: Beyond the at least five toxic compounds, there are at least seven compounds which change the cells' types. Differential expression for these outliers is hard to model. They make cross-validation unreliable, and the outliers in the private leaderboard can't even be predicted by probing the public leaderboard.</p>\n<h2>Cell count shouldn't affect differential gene expression</h2>\n<p>Does gene expression in a cell depend on how many cells are in the experiment? Theoretically, it doesn't. A cell behaves the same way whether there are 10 cells in the experiment or 10000. We'd expect, however, a difference in the significance of the experimental results: An experiment with 10000 cells should give more precise measurements than a 10-cell experiment: As the cell count grows, variance of the measurements should decrease, t-score should be farther away from zero, and pvalues should decrease.</p>\n<p>The competition data don't fulfill this expectation. If we plot the mean t-scores versus the cell count for the 602 cell type–compound combinations (excluding the control compounds), we see a linear relationship: For every cell type, compounds with lower cell counts have positive t-score means, and compounds with higher cell counts have negative t-score means. This correlation between cell counts and t-scores shouldn't exist. It is an artefact of Limma rather than a biological effect.</p>\n<p>You can plot the diagram with median or variance instead of mean — it will look similar. You can even compare the cell counts to the first principal component of the t-scores and see the same correlation. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F50f71507c0241429e92b4ae50b092222%2Fcell-t-before.png?generation=1701390101863675&amp;alt=media\" alt=\"\"></p>\n<p>We can now put together a list of 20 compounds which are to be considered outliers because of low cell counts. Notice that we don't declare single rows of the dataset to be outliers, but all 86 rows related to the 20 compounds:</p>\n<pre><code>Outliers\n\nAT13387                             T regulatory cells\nAlvocidib                         ≤  several cell \nBAY                        mean t-score  CD8+ cells &gt; \nBMS                          T cells CD8+,  Myeloid cells\nBelinostat                        control compound  too many cells\nCEP (Delanzomib)            ≤  several cell \nCGM                           too many T cells CD8+\nCGP                          ≤  several cell \nDabrafenib                        control compound  too many cells\nGanetespib (STA)               T regulatory cells, too many NK cells\nI-BET151                          too many T cells CD8+\nIN1451                            ≤  several cell \nLY2090314                           T cells CD8+\nMLN                           ≤  several cell \nOprozomib (ONX )              ≤  several cell \nProscillaridin A;Proscillaridin-A ≤  several cell \nResminostat                        T cells CD8+\nScriptaid                           T regulatory cells\nUNII-BXU45ZH6LI                     T cells CD8+\nVorinostat                          T regulatory cell\n</code></pre>\n<p>After removing the outliers, the diagram looks much cleaner. The variance of the cell counts remains. It is a source of noise which impedes the correct interpretation (and prediction) of differential expressions. Maybe we'd get cleaner data if we equalized the cell counts before library size normalization. This would amount to throwing away a part of the measurements, which isn't desirable either.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2Ffd906cb718a3308dd60f9dc5e56e1c3d%2Fcell-t-after.png?generation=1701390162014901&amp;alt=media\" alt=\"\"></p>\n<p>After considering the small size of the dataset, the amount of noise and the Limma artefacts (more of them will be shown in the next section), I didn't try to integrate any external biological data into my model. </p>\n<h1>Exploration of the problem</h1>\n<h2>A mixture of probability distributions</h2>\n<p>A histogram of a single row of the training data (18211 t-scores for T cells CD8+ treated with Scriptaid) shows that the distribution is multimodal.</p>\n<p>The highest mode consists of 269 genes which all have an identical t-score of -3.769. It turns out that these are the 269 genes which are never expressed in T cells CD8+, neither with the negative control nor with any other compound. Isn't this strange? A gene which is never expressed in the whole experiment should have a log-fold change of zero and should not get a t-score at all (because t-score computation involves a division by the variance, and the variance of a never-expressed gene is zero).</p>\n<p>For Myeloid cells treated with Foretinib, 3856 genes are not expressed (RNA count of zero), yet most of them have a positive t-score. Their highest t-score is 6.228 (resulting in a pvalue of 4e-10 and a log10pvalue of 9.33). If an RNA count is zero, the corresponding log-fold-change (and t-score) should never be positive.</p>\n<p>We may say that the distribution of the values is a mixture of two distributions:</p>\n<ol>\n<li>The values for the genes which are expressed (blue) have a more or less bell-shaped distribution.</li>\n<li>The values for the genes which are not expressed (orange) have a distribution with an unusual shape, and it is strange that positive differential expressions are reported when not a single piece of RNA is counted.</li>\n</ol>\n<p>What we see here is an artefact of Limma, which affects every row of the datset. It suggests that Limma output can be biased and is not ideal for investigating cell-type translation of differential expressions.</p>\n<pre><code> expressed in T cells CD8+ Scriptaid:     \n not expressed in T cells CD8+ Scriptaid:  \n: -. for  genes not expressed at  in T cells CD8+\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F7134e68a91b05970eba9705209a7ed46%2Fmixture1.png?generation=1701390193470010&amp;alt=media\" alt=\"\"></p>\n<pre><code> expressed in Myeloid cells Foretinib:     \n not expressed in Myeloid cells Foretinib:  \n: . for  genes not expressed at  in Myeloid cells\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2Ff2d00ede4d7735ddf76f1c16aec6fe79%2Fmixture2.png?generation=1701390213955949&amp;alt=media\" alt=\"\"></p>\n<h2>An ideal training set</h2>\n<p>In the competition overview, the organizers ask: <em>Do you have any evidence to suggest how you might develop an ideal training set for cell type translation beyond random sampling of compounds in cell types? What is the relationship between the number of compounds measured in the held-out cell types and model performance?</em></p>\n<p>I think we are not yet ready to answer these questions. We first need cleaner data (and more of it):</p>\n<ul>\n<li>Cell types must be classified correctly. This may imply that we limit the scope of the work to compounds which do not hamper cell type classification.</li>\n<li>Samples containing too few cells must be eliminated from the dataset. These samples just add hay to the haystack where we want to find the needle.</li>\n<li>Even if we have many cells, genes with low rna counts may need to be eliminated. Otherwise they add even more hay to the haystack.</li>\n</ul>\n<p>Second, modeling strange t-scores of genes which are never expressed is a waste of time. We need to define a machine-learning task and a metric which reward biological insight rather than forcing people into modeling the noise created by upstream processing steps:</p>\n<ul>\n<li>As t-scores are always affected by cell counts and variance estimates, a metric based on less highly-processed data (i.e., log-fold changes or rna counts rather than log10pvalues or t-scores) may lead research into a better direction.</li>\n<li>Even with log-fold changes, genes with low rna count make more noise than genes with high rna count. A suitable metric should account for this fact.</li>\n</ul>\n<h1>Model design</h1>\n<h2>T-scores are better than log10pvalues</h2>\n<p>Limma performs t-tests. t-scores are (almost) normally distributed, which is good for machine learning inputs. For this competition, the t-scores were nonlinearly transformed to log10pvalues. The transformation squeezes the nice bell shape into a distribution with a much higher kurtosis.</p>\n<p>My machine learning models perform better if I transform the log10pvalues into t-score in a preprocessing step, predict t-scores, and transform the predictions back afterwards. Perhaps working with log-fold changes or RNA counts would be even better.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2Fad0c81caaf72f36d9a859296565e5882%2Ft-score-is-better.png?generation=1701390233521477&amp;alt=media\" alt=\"\"></p>\n<h2>The models</h2>\n<p>I developed four models:</p>\n<ul>\n<li>Py-boost</li>\n<li>A recommender system based on ridge regression</li>\n<li>A recommender system based on k nearest neighbors</li>\n<li>ExtraTrees</li>\n</ul>\n<p>I first implemented the Py-boost model, derived from <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a>'s public notebook.</p>\n<p>I then implemented the ExtraTrees model, which resembles <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a>'s Py-boost model. All the decision trees are fully grown (i.e., overfitted). The model gets its generalization capability from noise which is added to the target-encoded features deliberately.</p>\n<p>I then implemented the knn <a href=\"https://en.wikipedia.org/wiki/Recommender_system\" target=\"_blank\">recommender system</a> to have some diversity in the ensemble. Cell types and compounds are identified with users and items, respectively; gene expression is identified with item ratings by users.</p>\n<p>ExtraTrees and k-nearest-neighbors share the weakness that they cannot extrapolate. Even after dimensionality reduction, our training dataset essentially consists of 614 points in a high-dimensional space, so that most of the points will lie on the convex hull. Of the 255 test points, many will lie outside the convex hull of the training points, which means that the model must extrapolate. To bring the extrapolation capability into the game, I implemented the ridge regression model. </p>\n<p>The models have cv scores between 0.878 (ExtraTrees) and 0.906 (Py-boost). Py-boost, which was the worst in cross-validation, has the best public and private lb scores (0.572 and 0.748, respectively).</p>\n<h2>Data augmentation</h2>\n<p>One of the models (k nearest neighbors) is fed with <strong>data augmentation</strong>: If we know the differential expressions for two compounds, we may assume that a mixture of the two compounds will produce a differential expression which is the average of the two single-compound differential expressions.</p>\n<p>I experimented with another kind of data augmentation a well: Because there are more than twice as many T cells CD4+ as either Myeloid or B cells and I knew that the cell count biases the results of Limma, I reduced the cell count of the T cells CD4+, pseudobulked them, ran them through Limma and added the results to the training data as another cell type. This augmentation improved the scores of ExtraTrees, but not to the level of Py-boost. Perhaps I should have combined the additional cell type with Py-boost…</p>\n<h1>Robustness</h1>\n<p>The robustness of my models is demonstrated in two ways:</p>\n<p>(1) The models are fully cross-validated. The cross-validation strategy, first documented in <a href=\"https://www.kaggle.com/code/ambrosm/scp-quickstart\" target=\"_blank\">SCP Quickstart</a>, ensures that the model is validated on predicting cell_type–sm_name combinations so that it knows only 17 other compounds for the same cell type. This cross-validation strategy is more robust than the ordinary shuffled KFold, where the model knows 4/5 of all compounds for the same cell type. (And it is much more robust than a simple train-test-split.)</p>\n<p>I have to admit, though, that I'm not happy with the cv–lb correspondence.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F424a99a8b518dd08b96b0026f598c557%2Fcv-scheme.png?generation=1701390265025626&amp;alt=media\" alt=\"\"></p>\n<p>(2) For all models the performance was tested after adding Gaussian noise to the input t-scores. All models are robust against small noise. When the noise gets stronger, the knn and ExtraTrees models suffer more than Py-boost and the ridge recommender system.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F9f9538ca1d28e570ae930339c9490145%2Fnoise.png?generation=1701409934387211&amp;alt=media\" alt=\"\"></p>\n<h1>Documentation and code style</h1>\n<p>The code is documented in the notebooks.</p>\n<h1>Reproducibility</h1>\n<p>Source code is here:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense\" target=\"_blank\">EDA which makes sense ⭐️⭐️⭐️⭐️⭐️</a></li>\n<li><a href=\"https://www.kaggle.com/code/ambrosm/scp-26-py-boost-recommender-system-and-et\" target=\"_blank\">SCP #26: Py-boost, recommender system and ET</a></li>\n<li><a href=\"https://github.com/Ambros-M/Single-Cell-Perturbations-2023\" target=\"_blank\">GitHub</a></li>\n</ul>\n<h1>Conclusion</h1>\n<p>Let me conclude by summarizing the four main messages of this post:</p>\n<ol>\n<li>Recommender systems are a promising starting point for developing models for cross-cell-type differential gene expression prediction. Because of commercial interests, recommender systems are a well-researched topic, and a lot of information is available.</li>\n<li>Data augmentation is useful, and mixtures of compounds are a natural approach to data augmentation.</li>\n<li>Although Kaggle competitions with data cleaning, outlier removal and unusual metrics are entertaining, the research objective would profit from another setting. Providing clean data and scoring with a well-understood metric would help participants focus on the real topic rather than the noise in the data.</li>\n<li>We have seen that Limma in certain situations produces biased outputs. I hope that professional Limma users are aware of these effects and account for them when interpreting results in their research.</li>\n</ol>",
      "rawMarkdown": "Did you notice that in this competition ~~no real EDA notebook has~~ few real EDA notebooks have been published? Besides explaining my machine learning model, I'd like to share some observations which help understand the data and the intricacies of Limma.\n\n# Integration of biological knowledge\n\n## Don't trust the cell types!\n\nLet's recapitulate the course of the experiment in a simplified form. We can imagine an experimenter who is in front of a large pot of human blood cells. The pot contains a mixture of six cell types in certain proportions. T cells CD4+ take the largest share (42 %), only 2 % are T regulatory cells:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F26d3a6c5971cf2a5a441cb550424728e%2Fpie-chart.png?generation=1701390049641476&alt=media)\nThe experimenter now takes 145 droplets out of the large pot. Every droplet contains 1550 ± 240 cells (normally distributed). If we counted the cells per cell type in the droplets, we'd see a multinomial distribution. The 145 droplets might be composed like in the following bar chart (fictitious data, sorted from smallest to largest droplet):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F838eb6fe789104225c5bf6cb0f73e060%2Fdrops-before.png?generation=1701390069789098&alt=media)\nIn the next step, the experimenter adds 145 substances to the 145 droplets and waits 24 hours. After 24 hours the cells are analyzed. If we count the cells again, we get the following picture, as taken from the competition's training data (cell counts for the test data are hidden):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F9a57da671b14ec3198d61c54861cc27f%2Fdrops-after.png?generation=1701390083658870&alt=media)\nIn this diagram we first see that some compounds are so toxic that in some droplets less than 100 cells survive. These droplets are represented by the leftmost bars in the bar chart.\n\nThe second observation is much more important: The long red part in the bars for Oprozomib and IN1451 show that these droplets contain several hundred T regulatory cells — much more than at the start of the experiment. Other compounds (e.g., CGM-079) have too many T cells CD8+ (green bar). How can we interpret this observation?\n1. Does IN1451 incite the T regulatory cells to multiply so that we have five times more of them after 24 hours? No.\n1. Does IN1451 magically convert NK cells into T regulatory cells? No.\n1. Does IN1451 affect the cells in such a way that they are misclassified? Maybe.\n\nDiscussing differential gene expression for specific cell types becomes pointless if the cells change their type during the experiment. For the Kaggle competition this means that we have to deal with many outliers: Beyond the at least five toxic compounds, there are at least seven compounds which change the cells' types. Differential expression for these outliers is hard to model. They make cross-validation unreliable, and the outliers in the private leaderboard can't even be predicted by probing the public leaderboard.\n\n## Cell count shouldn't affect differential gene expression\n\nDoes gene expression in a cell depend on how many cells are in the experiment? Theoretically, it doesn't. A cell behaves the same way whether there are 10 cells in the experiment or 10000. We'd expect, however, a difference in the significance of the experimental results: An experiment with 10000 cells should give more precise measurements than a 10-cell experiment: As the cell count grows, variance of the measurements should decrease, t-score should be farther away from zero, and pvalues should decrease.\n\nThe competition data don't fulfill this expectation. If we plot the mean t-scores versus the cell count for the 602 cell type–compound combinations (excluding the control compounds), we see a linear relationship: For every cell type, compounds with lower cell counts have positive t-score means, and compounds with higher cell counts have negative t-score means. This correlation between cell counts and t-scores shouldn't exist. It is an artefact of Limma rather than a biological effect.\n\nYou can plot the diagram with median or variance instead of mean — it will look similar. You can even compare the cell counts to the first principal component of the t-scores and see the same correlation. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F50f71507c0241429e92b4ae50b092222%2Fcell-t-before.png?generation=1701390101863675&alt=media)\n\nWe can now put together a list of 20 compounds which are to be considered outliers because of low cell counts. Notice that we don't declare single rows of the dataset to be outliers, but all 86 rows related to the 20 compounds:\n\n```\nOutliers\n--------\nAT13387                           only 7 T regulatory cells\nAlvocidib                         ≤10 for several cell types\nBAY 61-3606                       mean t-score for CD8+ cells > 2\nBMS-387032                        only 10 T cells CD8+, no Myeloid cells\nBelinostat                        control compound with too many cells\nCEP-18770 (Delanzomib)            ≤10 for several cell types\nCGM-097                           too many T cells CD8+\nCGP 60474                         ≤10 for several cell types\nDabrafenib                        control compound with too many cells\nGanetespib (STA-9090)             only 4 T regulatory cells, too many NK cells\nI-BET151                          too many T cells CD8+\nIN1451                            ≤10 for several cell types\nLY2090314                         only 6 T cells CD8+\nMLN 2238                          ≤10 for several cell types\nOprozomib (ONX 0912)              ≤10 for several cell types\nProscillaridin A;Proscillaridin-A ≤10 for several cell types\nResminostat                       no T cells CD8+\nScriptaid                         only 2 T regulatory cells\nUNII-BXU45ZH6LI                   only 6 T cells CD8+\nVorinostat                        only 1 T regulatory cell\n```\n\nAfter removing the outliers, the diagram looks much cleaner. The variance of the cell counts remains. It is a source of noise which impedes the correct interpretation (and prediction) of differential expressions. Maybe we'd get cleaner data if we equalized the cell counts before library size normalization. This would amount to throwing away a part of the measurements, which isn't desirable either.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2Ffd906cb718a3308dd60f9dc5e56e1c3d%2Fcell-t-after.png?generation=1701390162014901&alt=media)\n\nAfter considering the small size of the dataset, the amount of noise and the Limma artefacts (more of them will be shown in the next section), I didn't try to integrate any external biological data into my model. \n\n# Exploration of the problem\n\n## A mixture of probability distributions\n\nA histogram of a single row of the training data (18211 t-scores for T cells CD8+ treated with Scriptaid) shows that the distribution is multimodal.\n\nThe highest mode consists of 269 genes which all have an identical t-score of -3.769. It turns out that these are the 269 genes which are never expressed in T cells CD8+, neither with the negative control nor with any other compound. Isn't this strange? A gene which is never expressed in the whole experiment should have a log-fold change of zero and should not get a t-score at all (because t-score computation involves a division by the variance, and the variance of a never-expressed gene is zero).\n\nFor Myeloid cells treated with Foretinib, 3856 genes are not expressed (RNA count of zero), yet most of them have a positive t-score. Their highest t-score is 6.228 (resulting in a pvalue of 4e-10 and a log10pvalue of 9.33). If an RNA count is zero, the corresponding log-fold-change (and t-score) should never be positive.\n\nWe may say that the distribution of the values is a mixture of two distributions:\n1. The values for the genes which are expressed (blue) have a more or less bell-shaped distribution.\n2. The values for the genes which are not expressed (orange) have a distribution with an unusual shape, and it is strange that positive differential expressions are reported when not a single piece of RNA is counted.\n\nWhat we see here is an artefact of Limma, which affects every row of the datset. It suggests that Limma output can be biased and is not ideal for investigating cell-type translation of differential expressions.\n\n```\nGenes expressed in T cells CD8+ Scriptaid:     13560\nGenes not expressed in T cells CD8+ Scriptaid:  4651\nMode: -3.804 for 269 genes not expressed at all in T cells CD8+\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F7134e68a91b05970eba9705209a7ed46%2Fmixture1.png?generation=1701390193470010&alt=media)\n\n```\nGenes expressed in Myeloid cells Foretinib:     14355\nGenes not expressed in Myeloid cells Foretinib:  3856\nMode: 6.379 for 81 genes not expressed at all in Myeloid cells\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2Ff2d00ede4d7735ddf76f1c16aec6fe79%2Fmixture2.png?generation=1701390213955949&alt=media)\n## An ideal training set\n\nIn the competition overview, the organizers ask: *Do you have any evidence to suggest how you might develop an ideal training set for cell type translation beyond random sampling of compounds in cell types? What is the relationship between the number of compounds measured in the held-out cell types and model performance?*\n\nI think we are not yet ready to answer these questions. We first need cleaner data (and more of it):\n- Cell types must be classified correctly. This may imply that we limit the scope of the work to compounds which do not hamper cell type classification.\n- Samples containing too few cells must be eliminated from the dataset. These samples just add hay to the haystack where we want to find the needle.\n- Even if we have many cells, genes with low rna counts may need to be eliminated. Otherwise they add even more hay to the haystack.\n\n\nSecond, modeling strange t-scores of genes which are never expressed is a waste of time. We need to define a machine-learning task and a metric which reward biological insight rather than forcing people into modeling the noise created by upstream processing steps:\n- As t-scores are always affected by cell counts and variance estimates, a metric based on less highly-processed data (i.e., log-fold changes or rna counts rather than log10pvalues or t-scores) may lead research into a better direction.\n- Even with log-fold changes, genes with low rna count make more noise than genes with high rna count. A suitable metric should account for this fact.\n\n# Model design\n\n## T-scores are better than log10pvalues\n\nLimma performs t-tests. t-scores are (almost) normally distributed, which is good for machine learning inputs. For this competition, the t-scores were nonlinearly transformed to log10pvalues. The transformation squeezes the nice bell shape into a distribution with a much higher kurtosis.\n\nMy machine learning models perform better if I transform the log10pvalues into t-score in a preprocessing step, predict t-scores, and transform the predictions back afterwards. Perhaps working with log-fold changes or RNA counts would be even better.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2Fad0c81caaf72f36d9a859296565e5882%2Ft-score-is-better.png?generation=1701390233521477&alt=media)\n\n## The models\n\nI developed four models:\n- Py-boost\n- A recommender system based on ridge regression\n- A recommender system based on k nearest neighbors\n- ExtraTrees\n\nI first implemented the Py-boost model, derived from @alexandervc's public notebook.\n\nI then implemented the ExtraTrees model, which resembles @alexandervc's Py-boost model. All the decision trees are fully grown (i.e., overfitted). The model gets its generalization capability from noise which is added to the target-encoded features deliberately.\n\nI then implemented the knn [recommender system](https://en.wikipedia.org/wiki/Recommender_system) to have some diversity in the ensemble. Cell types and compounds are identified with users and items, respectively; gene expression is identified with item ratings by users.\n\nExtraTrees and k-nearest-neighbors share the weakness that they cannot extrapolate. Even after dimensionality reduction, our training dataset essentially consists of 614 points in a high-dimensional space, so that most of the points will lie on the convex hull. Of the 255 test points, many will lie outside the convex hull of the training points, which means that the model must extrapolate. To bring the extrapolation capability into the game, I implemented the ridge regression model. \n\nThe models have cv scores between 0.878 (ExtraTrees) and 0.906 (Py-boost). Py-boost, which was the worst in cross-validation, has the best public and private lb scores (0.572 and 0.748, respectively).\n\n## Data augmentation\n\nOne of the models (k nearest neighbors) is fed with **data augmentation**: If we know the differential expressions for two compounds, we may assume that a mixture of the two compounds will produce a differential expression which is the average of the two single-compound differential expressions.\n\nI experimented with another kind of data augmentation a well: Because there are more than twice as many T cells CD4+ as either Myeloid or B cells and I knew that the cell count biases the results of Limma, I reduced the cell count of the T cells CD4+, pseudobulked them, ran them through Limma and added the results to the training data as another cell type. This augmentation improved the scores of ExtraTrees, but not to the level of Py-boost. Perhaps I should have combined the additional cell type with Py-boost...\n\n# Robustness\n\nThe robustness of my models is demonstrated in two ways:\n\n(1) The models are fully cross-validated. The cross-validation strategy, first documented in [SCP Quickstart](https://www.kaggle.com/code/ambrosm/scp-quickstart), ensures that the model is validated on predicting cell_type–sm_name combinations so that it knows only 17 other compounds for the same cell type. This cross-validation strategy is more robust than the ordinary shuffled KFold, where the model knows 4/5 of all compounds for the same cell type. (And it is much more robust than a simple train-test-split.)\n\nI have to admit, though, that I'm not happy with the cv–lb correspondence.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F424a99a8b518dd08b96b0026f598c557%2Fcv-scheme.png?generation=1701390265025626&alt=media)\n\n(2) For all models the performance was tested after adding Gaussian noise to the input t-scores. All models are robust against small noise. When the noise gets stronger, the knn and ExtraTrees models suffer more than Py-boost and the ridge recommender system.\n \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F9f9538ca1d28e570ae930339c9490145%2Fnoise.png?generation=1701409934387211&alt=media)\n\n# Documentation and code style\n\nThe code is documented in the notebooks.\n\n# Reproducibility\n\nSource code is here:\n- [EDA which makes sense ⭐️⭐️⭐️⭐️⭐️](https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense)\n- [SCP #26: Py-boost, recommender system and ET](https://www.kaggle.com/code/ambrosm/scp-26-py-boost-recommender-system-and-et)\n- [GitHub](https://github.com/Ambros-M/Single-Cell-Perturbations-2023)\n\n# Conclusion\n\nLet me conclude by summarizing the four main messages of this post:\n\n1. Recommender systems are a promising starting point for developing models for cross-cell-type differential gene expression prediction. Because of commercial interests, recommender systems are a well-researched topic, and a lot of information is available.\n2. Data augmentation is useful, and mixtures of compounds are a natural approach to data augmentation.\n3. Although Kaggle competitions with data cleaning, outlier removal and unusual metrics are entertaining, the research objective would profit from another setting. Providing clean data and scoring with a well-understood metric would help participants focus on the real topic rather than the noise in the data.\n4. We have seen that Limma in certain situations produces biased outputs. I hope that professional Limma users are aware of these effects and account for them when interpreting results in their research.",
      "votes": null
    },
    {
      "id": "2544634",
      "postDate": "12/01/2023 01:59:37",
      "content": "<p>\"Did you notice that in this competition no real EDA notebook has been published? \" at the beginning of your topic.</p>\n<p>That question is remarkable, mostly for those that call every/anything EDA.</p>\n<p>Now, my PySmiles is laughing at mine Ridiculous Rdkit scPerturbations. </p>\n<p>Thank you for the topic and the Kaggle Notebook where we can read the snippets. </p>\n<p>Indescreet question: How long (years of study/formation) to arrive in your professional level?</p>\n<p>Very likely more than 20 years. In my previous careers, More than 20 years is what is required to achieve such level of knowledge, skills and mostly Experience, leaving the Ego behind to help those that are starting in any field.</p>\n<p>Thank you so much AmbrosM. </p>",
      "rawMarkdown": "\"Did you notice that in this competition no real EDA notebook has been published? \" at the beginning of your topic.\n\nThat question is remarkable, mostly for those that call every/anything EDA.\n\nNow, my PySmiles is laughing at mine Ridiculous Rdkit scPerturbations. \n\nThank you for the topic and the Kaggle Notebook where we can read the snippets. \n\nIndescreet question: How long (years of study/formation) to arrive in your professional level?\n\nVery likely more than 20 years. In my previous careers, More than 20 years is what is required to achieve such level of knowledge, skills and mostly Experience, leaving the Ego behind to help those that are starting in any field.\n\nThank you so much AmbrosM.",
      "votes": null
    },
    {
      "id": "2544745",
      "postDate": "12/01/2023 04:06:22",
      "content": "<p>Congratulations. Thanks for sharing the nice writeup with colorful graphs and charts. <br>\nI think if you had provided useful references with your article, it would be helpful. </p>",
      "rawMarkdown": "Congratulations. Thanks for sharing the nice writeup with colorful graphs and charts. \nI think if you had provided useful references with your article, it would be helpful.",
      "votes": null
    },
    {
      "id": "2544896",
      "postDate": "12/01/2023 06:49:43",
      "content": "<p>I sincerely thank you for perfect sharing, <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> ! You have provided us with great help, and your method is very impressive! And I have learned a lot of knowledge from your sharing.</p>",
      "rawMarkdown": "I sincerely thank you for perfect sharing, @ambrosm ! You have provided us with great help, and your method is very impressive! And I have learned a lot of knowledge from your sharing.",
      "votes": null
    },
    {
      "id": "2545121",
      "postDate": "12/01/2023 10:06:25",
      "content": "<p>Heads Up, and thanks for sharing this.</p>",
      "rawMarkdown": "Heads Up, and thanks for sharing this.",
      "votes": null
    },
    {
      "id": "2545444",
      "postDate": "12/01/2023 13:41:17",
      "content": "<blockquote>\n  <p>Did you notice that in this competition no real EDA notebook has been publishedt</p>\n</blockquote>\n<p>two of mine are crying in the corner ashamed 😅</p>\n<p>Thank you for the great description, I learned a lot, it's very interesting too) And congratulations with the silver!</p>",
      "rawMarkdown": ">Did you notice that in this competition no real EDA notebook has been publishedt\n\ntwo of mine are crying in the corner ashamed 😅\n\nThank you for the great description, I learned a lot, it's very interesting too) And congratulations with the silver!",
      "votes": null
    },
    {
      "id": "2545578",
      "postDate": "12/01/2023 15:23:19",
      "content": "<p>Great writeup ! <br>\nIt is always pleasure to learn from your work  !<br>\nVery happy that you find my notebook on pyboost useful. <br>\nCongratulations with the medal !</p>",
      "rawMarkdown": "Great writeup ! \nIt is always pleasure to learn from your work  !\nVery happy that you find my notebook on pyboost useful. \nCongratulations with the medal !",
      "votes": null
    },
    {
      "id": "2545669",
      "postDate": "12/01/2023 16:38:14",
      "content": "<p>Sorry, <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a>, I overlooked them (maybe because I'm not fluent in R).</p>\n<p>Congratulations to you, too! </p>",
      "rawMarkdown": "Sorry, @antoninadolgorukova, I overlooked them (maybe because I'm not fluent in R).\n\nCongratulations to you, too!",
      "votes": null
    },
    {
      "id": "2546231",
      "postDate": "12/02/2023 08:31:50",
      "content": "<p>HI <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a>, thanks for bringing py-boost to my attention, and congratulations to you, too!</p>",
      "rawMarkdown": "HI @alexandervc, thanks for bringing py-boost to my attention, and congratulations to you, too!",
      "votes": null
    },
    {
      "id": "2546273",
      "postDate": "12/02/2023 09:17:30",
      "content": "<p>Don't be sorry I just meant what I said) Ashamed I saw strange things in the data but did not dig deeper to find out the reasons. So thank you a lot for the explanations and sharing!</p>",
      "rawMarkdown": "Don't be sorry I just meant what I said) Ashamed I saw strange things in the data but did not dig deeper to find out the reasons. So thank you a lot for the explanations and sharing!",
      "votes": null
    },
    {
      "id": "2546961",
      "postDate": "12/03/2023 04:10:13",
      "content": "<p>Great writeup! And great visualizations indeed!</p>",
      "rawMarkdown": "Great writeup! And great visualizations indeed!",
      "votes": null
    },
    {
      "id": "2548681",
      "postDate": "12/04/2023 15:51:50",
      "content": "<p>I don't think the conclusion that limma is biased is evident. Limma is being used incorrectly/poorly in this competition.<br>\nAn important step of the usual limma application is to remove genes that \"consistently have zero or very low counts.\" using the <code>filterByExpr</code> function from edgeR (see limma userguide). To avoid having NaNs in the data the organizers skipped this step which probably interfered with the TMM normalization. Further the Empirical Bayes procedure in limma is susceptible to genes with very small (e.g. all zero) or very large variance. Phipson et al. 16 suggest a more robust EB method which can be activated using <code>robust=TRUE</code> in the <code>eBayes</code> function but this was not used by the competition organizers. You et al. 23 suggest a modification to voom which accounts for heteroscedasticity between groups in scRNA-seq data and suggest that not doing so leads to problems. This also was not used. Nevertheless I am also very suspect of DE results in general but especially for scRNA-seq data. </p>\n<p>DE analysis is very tricky and many intransparent assumptions are made which can often lead to problems (e.g. Squair et al. 21 show that many methods produce lots of false positives, Li et al 22 suggest high false discovery rates when there are large sample sizes)<br>\nAll in all I don't think DE results are good prediction targets and I believe that this lead to the competition being focused on accounting for DE computation quirks rather than on predicting biologically meaningful change.</p>\n<p><em>Robust hyperparameter estimation protects against hypervariable genes and improves power to detect differential expression</em> <a href=\"https://projecteuclid.org/journals/annals-of-applied-statistics/volume-10/issue-2/Robust-hyperparameter-estimation-protects-against-hypervariable-genes-and-improves-power/10.1214/16-AOAS920.full\" target=\"_blank\">Phipson et al 16</a></p>\n<p><em>Modeling group heteroscedasticity in single-cell RNA-seq pseudo-bulk data</em> <a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-023-02949-2\" target=\"_blank\">You et al. 23</a></p>\n<p><em>Confronting false discoveries in single-cell differential expression</em> <a href=\"https://www.nature.com/articles/s41467-021-25960-2\" target=\"_blank\">Squair et al. 21</a></p>\n<p><em>Exaggerated false positives by popular differential expression methods when analyzing human population samples</em> <a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-022-02648-4\" target=\"_blank\">Li et al. 22</a></p>",
      "rawMarkdown": "I don't think the conclusion that limma is biased is evident. Limma is being used incorrectly/poorly in this competition.\nAn important step of the usual limma application is to remove genes that \"consistently have zero or very low counts.\" using the `filterByExpr` function from edgeR (see limma userguide). To avoid having NaNs in the data the organizers skipped this step which probably interfered with the TMM normalization. Further the Empirical Bayes procedure in limma is susceptible to genes with very small (e.g. all zero) or very large variance. Phipson et al. 16 suggest a more robust EB method which can be activated using `robust=TRUE` in the `eBayes` function but this was not used by the competition organizers. You et al. 23 suggest a modification to voom which accounts for heteroscedasticity between groups in scRNA-seq data and suggest that not doing so leads to problems. This also was not used. Nevertheless I am also very suspect of DE results in general but especially for scRNA-seq data. \n\nDE analysis is very tricky and many intransparent assumptions are made which can often lead to problems (e.g. Squair et al. 21 show that many methods produce lots of false positives, Li et al 22 suggest high false discovery rates when there are large sample sizes)\nAll in all I don't think DE results are good prediction targets and I believe that this lead to the competition being focused on accounting for DE computation quirks rather than on predicting biologically meaningful change.\n\n*Robust hyperparameter estimation protects against hypervariable genes and improves power to detect differential expression* [Phipson et al 16](https://projecteuclid.org/journals/annals-of-applied-statistics/volume-10/issue-2/Robust-hyperparameter-estimation-protects-against-hypervariable-genes-and-improves-power/10.1214/16-AOAS920.full)\n\n*Modeling group heteroscedasticity in single-cell RNA-seq pseudo-bulk data* [You et al. 23](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-023-02949-2)\n\n*Confronting false discoveries in single-cell differential expression* [Squair et al. 21](https://www.nature.com/articles/s41467-021-25960-2)\n\n*Exaggerated false positives by popular differential expression methods when analyzing human population samples* [Li et al. 22](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-022-02648-4)",
      "votes": null
    },
    {
      "id": "2548941",
      "postDate": "12/04/2023 21:28:45",
      "content": "<p>Combined with magic-top4 trick - your solution gets 0.718 on private LB - which is better than current top1.<br>\nSee <br>\n<a href=\"https://www.kaggle.com/alexandervc/op2-explore-4th-place-magic\" target=\"_blank\">https://www.kaggle.com/alexandervc/op2-explore-4th-place-magic</a></p>",
      "rawMarkdown": "Combined with magic-top4 trick - your solution gets 0.718 on private LB - which is better than current top1.\nSee \nhttps://www.kaggle.com/alexandervc/op2-explore-4th-place-magic",
      "votes": null
    },
    {
      "id": "2549107",
      "postDate": "12/05/2023 03:03:16",
      "content": "<p>Congratulations! I like your idea about cell counts and data augmentations.</p>\n<blockquote>\n  <p>Although Kaggle competitions with data cleaning, outlier removal and unusual metrics are entertaining, the research objective would profit from another setting. Providing clean data and scoring with a well-understood metric would help participants focus on the real topic rather than the noise in the data.</p>\n</blockquote>\n<p>I'm the one who focused on the noise in the data and then give up 😅.<br>\nAfter reading your post, I deeply regret that I did not even try to find a solution for it.<br>\nBeyond the tech part, I also learned a lot from your post, thank you for sharing.</p>",
      "rawMarkdown": "Congratulations! I like your idea about cell counts and data augmentations.\n\n> Although Kaggle competitions with data cleaning, outlier removal and unusual metrics are entertaining, the research objective would profit from another setting. Providing clean data and scoring with a well-understood metric would help participants focus on the real topic rather than the noise in the data.\n\nI'm the one who focused on the noise in the data and then give up 😅.\nAfter reading your post, I deeply regret that I did not even try to find a solution for it.\nBeyond the tech part, I also learned a lot from your post, thank you for sharing.",
      "votes": null
    },
    {
      "id": "2556042",
      "postDate": "12/10/2023 12:23:46",
      "content": "<p>Thank you, it was interesting to read the detailed description of the data analysis.<br>\nYour notebooks carried a lot of ideas  in this competition</p>",
      "rawMarkdown": "Thank you, it was interesting to read the detailed description of the data analysis.\nYour notebooks carried a lot of ideas  in this competition",
      "votes": null
    },
    {
      "id": "2566894",
      "postDate": "12/19/2023 08:43:38",
      "content": "<p>Thanks again for the great write-up !<br>\nTo complement your table of suspicious samples.</p>\n<p>Here is another evidence that something is wrong with them.<br>\nLet me look on count of genes which are strongly perturbed for each sample<br>\nwe can expect it should not be too much.<br>\nBut there are samples where we see about HALF(!) genes are perturbed strongly - that is suspicious.<br>\nAnd that list quite corresponds to yours:</p>\n<p>(Also note: that these samples seems to be the WORST predicted - see <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461663\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461663</a> )</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F4e1c24530cbda8ef941825e5c464268c%2FScreenshot%202023-12-19%20092655.png?generation=1702975111733165&amp;alt=media\" alt=\"\"><br>\n<a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155628891&amp;cellId=10\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155628891&amp;cellId=10</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F4e84fbb28aade383843e4a7ec4cdcf66%2FScreenshot%202023-12-19%20100930.png?generation=1702977039679503&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks again for the great write-up !\nTo complement your table of suspicious samples.\n\nHere is another evidence that something is wrong with them.\nLet me look on count of genes which are strongly perturbed for each sample\nwe can expect it should not be too much.\nBut there are samples where we see about HALF(!) genes are perturbed strongly - that is suspicious.\nAnd that list quite corresponds to yours:\n\n(Also note: that these samples seems to be the WORST predicted - see https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461663 )\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F4e1c24530cbda8ef941825e5c464268c%2FScreenshot%202023-12-19%20092655.png?generation=1702975111733165&alt=media)\nhttps://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155628891&cellId=10\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F4e84fbb28aade383843e4a7ec4cdcf66%2FScreenshot%202023-12-19%20100930.png?generation=1702977039679503&alt=media)",
      "votes": null
    },
    {
      "id": "2576513",
      "postDate": "12/27/2023 19:52:58",
      "content": "<p>To complement your analysis on drugs which potentially lead to misclassification of cell types:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fb25870b8e55dabe6f074c05749eebfe9%2FScreenshot%202023-12-27%20204555.png?generation=1703706506138338&amp;alt=media\" alt=\"\"></p>\n<p>We can see that: 'Oprozomib (ONX 0912)','MLN 2238','CEP-18770 (Delanzomib)'<br>\nare concentrated in the same umap-seen cluster.  It is classified by T CD4/8 cells, but <br>\nwe see that cluster is so far-separate from the other cells type clusters, that we can easily imagine that it can be mis-classified <br>\n(since it is hard to imagine classification algorithms seen such kind of data - they were trained on normal cells , not drug-treated): <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F6def8f7d0ef31e73370e412481bb8144%2Fdrugs_ambrosM_1.png?generation=1703706523222615&amp;alt=media\" alt=\"\"><br>\nSee: <a href=\"https://www.kaggle.com/code/alexandervc/op2-rna-seq-data-scanpy-adata-cell-cycle?scriptVersionId=156757371&amp;cellId=45\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-rna-seq-data-scanpy-adata-cell-cycle?scriptVersionId=156757371&amp;cellId=45</a></p>\n<p>Similar other drugs you found: 'CGM-097', 'LY2090314','Ganetespib (STA-9090)','IN1451' - generate stand-alone clusters, so again it is reasonable to suspect that cell type classification may fail: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F0fc77b7f30d7b64ea0f5f08565a004db%2FScreenshot%202023-12-27%20205550.png?generation=1703706999692919&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "To complement your analysis on drugs which potentially lead to misclassification of cell types:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fb25870b8e55dabe6f074c05749eebfe9%2FScreenshot%202023-12-27%20204555.png?generation=1703706506138338&alt=media)\n\nWe can see that: 'Oprozomib (ONX 0912)','MLN 2238','CEP-18770 (Delanzomib)'\nare concentrated in the same umap-seen cluster.  It is classified by T CD4/8 cells, but \nwe see that cluster is so far-separate from the other cells type clusters, that we can easily imagine that it can be mis-classified \n(since it is hard to imagine classification algorithms seen such kind of data - they were trained on normal cells , not drug-treated): \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F6def8f7d0ef31e73370e412481bb8144%2Fdrugs_ambrosM_1.png?generation=1703706523222615&alt=media)\nSee: https://www.kaggle.com/code/alexandervc/op2-rna-seq-data-scanpy-adata-cell-cycle?scriptVersionId=156757371&cellId=45\n\nSimilar other drugs you found: 'CGM-097', 'LY2090314','Ganetespib (STA-9090)','IN1451' - generate stand-alone clusters, so again it is reasonable to suspect that cell type classification may fail: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F0fc77b7f30d7b64ea0f5f08565a004db%2FScreenshot%202023-12-27%20205550.png?generation=1703706999692919&alt=media)",
      "votes": null
    },
    {
      "id": "2576581",
      "postDate": "12/27/2023 22:30:25",
      "content": "<p>And to add more details/confirmation about <br>\n'Ganetespib (STA-9090)'<br>\nYour analysis suggest that we have too many NK-cells <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F4eaefc6acfe9777248419923b415c516%2FScreenshot%202023-12-27%20231609.png?generation=1703715458929376&amp;alt=media\" alt=\"\"></p>\n<p>That indeed quite corresponds to misclassification:<br>\nwe see cluster - it is expected to be one cell-type, but we see two (and one of them - NK-cells - the one appeared in your analysis). <br>\nMoreover - normally we can expect each cell type contains three donors - but that does not happen here. Here we see one cell type contains two donors another cell type just one donors. That is another  clear indication, that cell types were misclassified.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F4177d3d4ac38d4f08b5ac77badf551be%2FScreenshot%202023-12-27%20232630.png?generation=1703716036563754&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F48309d9847154aa4dff094779a1d8382%2FScreenshot%202023-12-27%20232951.png?generation=1703716220412047&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "And to add more details/confirmation about \n'Ganetespib (STA-9090)'\nYour analysis suggest that we have too many NK-cells \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F4eaefc6acfe9777248419923b415c516%2FScreenshot%202023-12-27%20231609.png?generation=1703715458929376&alt=media)\n\nThat indeed quite corresponds to misclassification:\nwe see cluster - it is expected to be one cell-type, but we see two (and one of them - NK-cells - the one appeared in your analysis). \nMoreover - normally we can expect each cell type contains three donors - but that does not happen here. Here we see one cell type contains two donors another cell type just one donors. That is another  clear indication, that cell types were misclassified.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F4177d3d4ac38d4f08b5ac77badf551be%2FScreenshot%202023-12-27%20232630.png?generation=1703716036563754&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F48309d9847154aa4dff094779a1d8382%2FScreenshot%202023-12-27%20232951.png?generation=1703716220412047&alt=media)",
      "votes": null
    },
    {
      "id": "2981852",
      "postDate": "09/07/2024 07:21:43",
      "content": "<p>Do you use Py-boost now?</p>",
      "rawMarkdown": "Do you use Py-boost now?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2544634,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "12/01/2023 01:59:37",
      "content": "<p>\"Did you notice that in this competition no real EDA notebook has been published? \" at the beginning of your topic.</p>\n<p>That question is remarkable, mostly for those that call every/anything EDA.</p>\n<p>Now, my PySmiles is laughing at mine Ridiculous Rdkit scPerturbations. </p>\n<p>Thank you for the topic and the Kaggle Notebook where we can read the snippets. </p>\n<p>Indescreet question: How long (years of study/formation) to arrive in your professional level?</p>\n<p>Very likely more than 20 years. In my previous careers, More than 20 years is what is required to achieve such level of knowledge, skills and mostly Experience, leaving the Ego behind to help those that are starting in any field.</p>\n<p>Thank you so much AmbrosM. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2544745,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "12/01/2023 04:06:22",
      "content": "<p>Congratulations. Thanks for sharing the nice writeup with colorful graphs and charts. <br>\nI think if you had provided useful references with your article, it would be helpful. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2544896,
      "author_name": "songqizhou",
      "author_url": "",
      "post_date": "12/01/2023 06:49:43",
      "content": "<p>I sincerely thank you for perfect sharing, <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> ! You have provided us with great help, and your method is very impressive! And I have learned a lot of knowledge from your sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2545121,
      "author_name": "vikrampython",
      "author_url": "",
      "post_date": "12/01/2023 10:06:25",
      "content": "<p>Heads Up, and thanks for sharing this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2545444,
      "author_name": "antoninadolgorukova",
      "author_url": "",
      "post_date": "12/01/2023 13:41:17",
      "content": "<blockquote>\n  <p>Did you notice that in this competition no real EDA notebook has been publishedt</p>\n</blockquote>\n<p>two of mine are crying in the corner ashamed 😅</p>\n<p>Thank you for the great description, I learned a lot, it's very interesting too) And congratulations with the silver!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2545669,
          "author_name": "ambrosm",
          "author_url": "",
          "post_date": "12/01/2023 16:38:14",
          "content": "<p>Sorry, <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a>, I overlooked them (maybe because I'm not fluent in R).</p>\n<p>Congratulations to you, too! </p>",
          "votes": null,
          "replies": [
            {
              "id": 2546273,
              "author_name": "antoninadolgorukova",
              "author_url": "",
              "post_date": "12/02/2023 09:17:30",
              "content": "<p>Don't be sorry I just meant what I said) Ashamed I saw strange things in the data but did not dig deeper to find out the reasons. So thank you a lot for the explanations and sharing!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2545578,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "12/01/2023 15:23:19",
      "content": "<p>Great writeup ! <br>\nIt is always pleasure to learn from your work  !<br>\nVery happy that you find my notebook on pyboost useful. <br>\nCongratulations with the medal !</p>",
      "votes": null,
      "replies": [
        {
          "id": 2546231,
          "author_name": "ambrosm",
          "author_url": "",
          "post_date": "12/02/2023 08:31:50",
          "content": "<p>HI <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a>, thanks for bringing py-boost to my attention, and congratulations to you, too!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2546961,
      "author_name": "abhasmalguri",
      "author_url": "",
      "post_date": "12/03/2023 04:10:13",
      "content": "<p>Great writeup! And great visualizations indeed!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2548681,
      "author_name": "samuelrutz",
      "author_url": "",
      "post_date": "12/04/2023 15:51:50",
      "content": "<p>I don't think the conclusion that limma is biased is evident. Limma is being used incorrectly/poorly in this competition.<br>\nAn important step of the usual limma application is to remove genes that \"consistently have zero or very low counts.\" using the <code>filterByExpr</code> function from edgeR (see limma userguide). To avoid having NaNs in the data the organizers skipped this step which probably interfered with the TMM normalization. Further the Empirical Bayes procedure in limma is susceptible to genes with very small (e.g. all zero) or very large variance. Phipson et al. 16 suggest a more robust EB method which can be activated using <code>robust=TRUE</code> in the <code>eBayes</code> function but this was not used by the competition organizers. You et al. 23 suggest a modification to voom which accounts for heteroscedasticity between groups in scRNA-seq data and suggest that not doing so leads to problems. This also was not used. Nevertheless I am also very suspect of DE results in general but especially for scRNA-seq data. </p>\n<p>DE analysis is very tricky and many intransparent assumptions are made which can often lead to problems (e.g. Squair et al. 21 show that many methods produce lots of false positives, Li et al 22 suggest high false discovery rates when there are large sample sizes)<br>\nAll in all I don't think DE results are good prediction targets and I believe that this lead to the competition being focused on accounting for DE computation quirks rather than on predicting biologically meaningful change.</p>\n<p><em>Robust hyperparameter estimation protects against hypervariable genes and improves power to detect differential expression</em> <a href=\"https://projecteuclid.org/journals/annals-of-applied-statistics/volume-10/issue-2/Robust-hyperparameter-estimation-protects-against-hypervariable-genes-and-improves-power/10.1214/16-AOAS920.full\" target=\"_blank\">Phipson et al 16</a></p>\n<p><em>Modeling group heteroscedasticity in single-cell RNA-seq pseudo-bulk data</em> <a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-023-02949-2\" target=\"_blank\">You et al. 23</a></p>\n<p><em>Confronting false discoveries in single-cell differential expression</em> <a href=\"https://www.nature.com/articles/s41467-021-25960-2\" target=\"_blank\">Squair et al. 21</a></p>\n<p><em>Exaggerated false positives by popular differential expression methods when analyzing human population samples</em> <a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-022-02648-4\" target=\"_blank\">Li et al. 22</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2548941,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "12/04/2023 21:28:45",
      "content": "<p>Combined with magic-top4 trick - your solution gets 0.718 on private LB - which is better than current top1.<br>\nSee <br>\n<a href=\"https://www.kaggle.com/alexandervc/op2-explore-4th-place-magic\" target=\"_blank\">https://www.kaggle.com/alexandervc/op2-explore-4th-place-magic</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2549107,
      "author_name": "awater1223",
      "author_url": "",
      "post_date": "12/05/2023 03:03:16",
      "content": "<p>Congratulations! I like your idea about cell counts and data augmentations.</p>\n<blockquote>\n  <p>Although Kaggle competitions with data cleaning, outlier removal and unusual metrics are entertaining, the research objective would profit from another setting. Providing clean data and scoring with a well-understood metric would help participants focus on the real topic rather than the noise in the data.</p>\n</blockquote>\n<p>I'm the one who focused on the noise in the data and then give up 😅.<br>\nAfter reading your post, I deeply regret that I did not even try to find a solution for it.<br>\nBeyond the tech part, I also learned a lot from your post, thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2556042,
      "author_name": "erotar",
      "author_url": "",
      "post_date": "12/10/2023 12:23:46",
      "content": "<p>Thank you, it was interesting to read the detailed description of the data analysis.<br>\nYour notebooks carried a lot of ideas  in this competition</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2566894,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "12/19/2023 08:43:38",
      "content": "<p>Thanks again for the great write-up !<br>\nTo complement your table of suspicious samples.</p>\n<p>Here is another evidence that something is wrong with them.<br>\nLet me look on count of genes which are strongly perturbed for each sample<br>\nwe can expect it should not be too much.<br>\nBut there are samples where we see about HALF(!) genes are perturbed strongly - that is suspicious.<br>\nAnd that list quite corresponds to yours:</p>\n<p>(Also note: that these samples seems to be the WORST predicted - see <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461663\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461663</a> )</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F4e1c24530cbda8ef941825e5c464268c%2FScreenshot%202023-12-19%20092655.png?generation=1702975111733165&amp;alt=media\" alt=\"\"><br>\n<a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155628891&amp;cellId=10\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155628891&amp;cellId=10</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F4e84fbb28aade383843e4a7ec4cdcf66%2FScreenshot%202023-12-19%20100930.png?generation=1702977039679503&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2576513,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "12/27/2023 19:52:58",
      "content": "<p>To complement your analysis on drugs which potentially lead to misclassification of cell types:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fb25870b8e55dabe6f074c05749eebfe9%2FScreenshot%202023-12-27%20204555.png?generation=1703706506138338&amp;alt=media\" alt=\"\"></p>\n<p>We can see that: 'Oprozomib (ONX 0912)','MLN 2238','CEP-18770 (Delanzomib)'<br>\nare concentrated in the same umap-seen cluster.  It is classified by T CD4/8 cells, but <br>\nwe see that cluster is so far-separate from the other cells type clusters, that we can easily imagine that it can be mis-classified <br>\n(since it is hard to imagine classification algorithms seen such kind of data - they were trained on normal cells , not drug-treated): <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F6def8f7d0ef31e73370e412481bb8144%2Fdrugs_ambrosM_1.png?generation=1703706523222615&amp;alt=media\" alt=\"\"><br>\nSee: <a href=\"https://www.kaggle.com/code/alexandervc/op2-rna-seq-data-scanpy-adata-cell-cycle?scriptVersionId=156757371&amp;cellId=45\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-rna-seq-data-scanpy-adata-cell-cycle?scriptVersionId=156757371&amp;cellId=45</a></p>\n<p>Similar other drugs you found: 'CGM-097', 'LY2090314','Ganetespib (STA-9090)','IN1451' - generate stand-alone clusters, so again it is reasonable to suspect that cell type classification may fail: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F0fc77b7f30d7b64ea0f5f08565a004db%2FScreenshot%202023-12-27%20205550.png?generation=1703706999692919&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2576581,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "12/27/2023 22:30:25",
          "content": "<p>And to add more details/confirmation about <br>\n'Ganetespib (STA-9090)'<br>\nYour analysis suggest that we have too many NK-cells <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F4eaefc6acfe9777248419923b415c516%2FScreenshot%202023-12-27%20231609.png?generation=1703715458929376&amp;alt=media\" alt=\"\"></p>\n<p>That indeed quite corresponds to misclassification:<br>\nwe see cluster - it is expected to be one cell-type, but we see two (and one of them - NK-cells - the one appeared in your analysis). <br>\nMoreover - normally we can expect each cell type contains three donors - but that does not happen here. Here we see one cell type contains two donors another cell type just one donors. That is another  clear indication, that cell types were misclassified.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F4177d3d4ac38d4f08b5ac77badf551be%2FScreenshot%202023-12-27%20232630.png?generation=1703716036563754&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F48309d9847154aa4dff094779a1d8382%2FScreenshot%202023-12-27%20232951.png?generation=1703716220412047&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2981852,
      "author_name": "",
      "author_url": "",
      "post_date": "09/07/2024 07:21:43",
      "content": "<p>Do you use Py-boost now?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2544595": "Did you notice that in this competition ~~no real EDA notebook has~~ few real EDA notebooks have been published? Besides explaining my machine learning model, I'd like to share some observations which help understand the data and the intricacies of Limma.\n\n# Integration of biological knowledge\n\n## Don't trust the cell types!\n\nLet's recapitulate the course of the experiment in a simplified form. We can imagine an experimenter who is in front of a large pot of human blood cells. The pot contains a mixture of six cell types in certain proportions. T cells CD4+ take the largest share (42 %), only 2 % are T regulatory cells:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F26d3a6c5971cf2a5a441cb550424728e%2Fpie-chart.png?generation=1701390049641476&alt=media)\nThe experimenter now takes 145 droplets out of the large pot. Every droplet contains 1550 ± 240 cells (normally distributed). If we counted the cells per cell type in the droplets, we'd see a multinomial distribution. The 145 droplets might be composed like in the following bar chart (fictitious data, sorted from smallest to largest droplet):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F838eb6fe789104225c5bf6cb0f73e060%2Fdrops-before.png?generation=1701390069789098&alt=media)\nIn the next step, the experimenter adds 145 substances to the 145 droplets and waits 24 hours. After 24 hours the cells are analyzed. If we count the cells again, we get the following picture, as taken from the competition's training data (cell counts for the test data are hidden):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F9a57da671b14ec3198d61c54861cc27f%2Fdrops-after.png?generation=1701390083658870&alt=media)\nIn this diagram we first see that some compounds are so toxic that in some droplets less than 100 cells survive. These droplets are represented by the leftmost bars in the bar chart.\n\nThe second observation is much more important: The long red part in the bars for Oprozomib and IN1451 show that these droplets contain several hundred T regulatory cells — much more than at the start of the experiment. Other compounds (e.g., CGM-079) have too many T cells CD8+ (green bar). How can we interpret this observation?\n1. Does IN1451 incite the T regulatory cells to multiply so that we have five times more of them after 24 hours? No.\n1. Does IN1451 magically convert NK cells into T regulatory cells? No.\n1. Does IN1451 affect the cells in such a way that they are misclassified? Maybe.\n\nDiscussing differential gene expression for specific cell types becomes pointless if the cells change their type during the experiment. For the Kaggle competition this means that we have to deal with many outliers: Beyond the at least five toxic compounds, there are at least seven compounds which change the cells' types. Differential expression for these outliers is hard to model. They make cross-validation unreliable, and the outliers in the private leaderboard can't even be predicted by probing the public leaderboard.\n\n## Cell count shouldn't affect differential gene expression\n\nDoes gene expression in a cell depend on how many cells are in the experiment? Theoretically, it doesn't. A cell behaves the same way whether there are 10 cells in the experiment or 10000. We'd expect, however, a difference in the significance of the experimental results: An experiment with 10000 cells should give more precise measurements than a 10-cell experiment: As the cell count grows, variance of the measurements should decrease, t-score should be farther away from zero, and pvalues should decrease.\n\nThe competition data don't fulfill this expectation. If we plot the mean t-scores versus the cell count for the 602 cell type–compound combinations (excluding the control compounds), we see a linear relationship: For every cell type, compounds with lower cell counts have positive t-score means, and compounds with higher cell counts have negative t-score means. This correlation between cell counts and t-scores shouldn't exist. It is an artefact of Limma rather than a biological effect.\n\nYou can plot the diagram with median or variance instead of mean — it will look similar. You can even compare the cell counts to the first principal component of the t-scores and see the same correlation. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F50f71507c0241429e92b4ae50b092222%2Fcell-t-before.png?generation=1701390101863675&alt=media)\n\nWe can now put together a list of 20 compounds which are to be considered outliers because of low cell counts. Notice that we don't declare single rows of the dataset to be outliers, but all 86 rows related to the 20 compounds:\n\n```\nOutliers\n--------\nAT13387                           only 7 T regulatory cells\nAlvocidib                         ≤10 for several cell types\nBAY 61-3606                       mean t-score for CD8+ cells > 2\nBMS-387032                        only 10 T cells CD8+, no Myeloid cells\nBelinostat                        control compound with too many cells\nCEP-18770 (Delanzomib)            ≤10 for several cell types\nCGM-097                           too many T cells CD8+\nCGP 60474                         ≤10 for several cell types\nDabrafenib                        control compound with too many cells\nGanetespib (STA-9090)             only 4 T regulatory cells, too many NK cells\nI-BET151                          too many T cells CD8+\nIN1451                            ≤10 for several cell types\nLY2090314                         only 6 T cells CD8+\nMLN 2238                          ≤10 for several cell types\nOprozomib (ONX 0912)              ≤10 for several cell types\nProscillaridin A;Proscillaridin-A ≤10 for several cell types\nResminostat                       no T cells CD8+\nScriptaid                         only 2 T regulatory cells\nUNII-BXU45ZH6LI                   only 6 T cells CD8+\nVorinostat                        only 1 T regulatory cell\n```\n\nAfter removing the outliers, the diagram looks much cleaner. The variance of the cell counts remains. It is a source of noise which impedes the correct interpretation (and prediction) of differential expressions. Maybe we'd get cleaner data if we equalized the cell counts before library size normalization. This would amount to throwing away a part of the measurements, which isn't desirable either.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2Ffd906cb718a3308dd60f9dc5e56e1c3d%2Fcell-t-after.png?generation=1701390162014901&alt=media)\n\nAfter considering the small size of the dataset, the amount of noise and the Limma artefacts (more of them will be shown in the next section), I didn't try to integrate any external biological data into my model. \n\n# Exploration of the problem\n\n## A mixture of probability distributions\n\nA histogram of a single row of the training data (18211 t-scores for T cells CD8+ treated with Scriptaid) shows that the distribution is multimodal.\n\nThe highest mode consists of 269 genes which all have an identical t-score of -3.769. It turns out that these are the 269 genes which are never expressed in T cells CD8+, neither with the negative control nor with any other compound. Isn't this strange? A gene which is never expressed in the whole experiment should have a log-fold change of zero and should not get a t-score at all (because t-score computation involves a division by the variance, and the variance of a never-expressed gene is zero).\n\nFor Myeloid cells treated with Foretinib, 3856 genes are not expressed (RNA count of zero), yet most of them have a positive t-score. Their highest t-score is 6.228 (resulting in a pvalue of 4e-10 and a log10pvalue of 9.33). If an RNA count is zero, the corresponding log-fold-change (and t-score) should never be positive.\n\nWe may say that the distribution of the values is a mixture of two distributions:\n1. The values for the genes which are expressed (blue) have a more or less bell-shaped distribution.\n2. The values for the genes which are not expressed (orange) have a distribution with an unusual shape, and it is strange that positive differential expressions are reported when not a single piece of RNA is counted.\n\nWhat we see here is an artefact of Limma, which affects every row of the datset. It suggests that Limma output can be biased and is not ideal for investigating cell-type translation of differential expressions.\n\n```\nGenes expressed in T cells CD8+ Scriptaid:     13560\nGenes not expressed in T cells CD8+ Scriptaid:  4651\nMode: -3.804 for 269 genes not expressed at all in T cells CD8+\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F7134e68a91b05970eba9705209a7ed46%2Fmixture1.png?generation=1701390193470010&alt=media)\n\n```\nGenes expressed in Myeloid cells Foretinib:     14355\nGenes not expressed in Myeloid cells Foretinib:  3856\nMode: 6.379 for 81 genes not expressed at all in Myeloid cells\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2Ff2d00ede4d7735ddf76f1c16aec6fe79%2Fmixture2.png?generation=1701390213955949&alt=media)\n## An ideal training set\n\nIn the competition overview, the organizers ask: *Do you have any evidence to suggest how you might develop an ideal training set for cell type translation beyond random sampling of compounds in cell types? What is the relationship between the number of compounds measured in the held-out cell types and model performance?*\n\nI think we are not yet ready to answer these questions. We first need cleaner data (and more of it):\n- Cell types must be classified correctly. This may imply that we limit the scope of the work to compounds which do not hamper cell type classification.\n- Samples containing too few cells must be eliminated from the dataset. These samples just add hay to the haystack where we want to find the needle.\n- Even if we have many cells, genes with low rna counts may need to be eliminated. Otherwise they add even more hay to the haystack.\n\n\nSecond, modeling strange t-scores of genes which are never expressed is a waste of time. We need to define a machine-learning task and a metric which reward biological insight rather than forcing people into modeling the noise created by upstream processing steps:\n- As t-scores are always affected by cell counts and variance estimates, a metric based on less highly-processed data (i.e., log-fold changes or rna counts rather than log10pvalues or t-scores) may lead research into a better direction.\n- Even with log-fold changes, genes with low rna count make more noise than genes with high rna count. A suitable metric should account for this fact.\n\n# Model design\n\n## T-scores are better than log10pvalues\n\nLimma performs t-tests. t-scores are (almost) normally distributed, which is good for machine learning inputs. For this competition, the t-scores were nonlinearly transformed to log10pvalues. The transformation squeezes the nice bell shape into a distribution with a much higher kurtosis.\n\nMy machine learning models perform better if I transform the log10pvalues into t-score in a preprocessing step, predict t-scores, and transform the predictions back afterwards. Perhaps working with log-fold changes or RNA counts would be even better.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2Fad0c81caaf72f36d9a859296565e5882%2Ft-score-is-better.png?generation=1701390233521477&alt=media)\n\n## The models\n\nI developed four models:\n- Py-boost\n- A recommender system based on ridge regression\n- A recommender system based on k nearest neighbors\n- ExtraTrees\n\nI first implemented the Py-boost model, derived from @alexandervc's public notebook.\n\nI then implemented the ExtraTrees model, which resembles @alexandervc's Py-boost model. All the decision trees are fully grown (i.e., overfitted). The model gets its generalization capability from noise which is added to the target-encoded features deliberately.\n\nI then implemented the knn [recommender system](https://en.wikipedia.org/wiki/Recommender_system) to have some diversity in the ensemble. Cell types and compounds are identified with users and items, respectively; gene expression is identified with item ratings by users.\n\nExtraTrees and k-nearest-neighbors share the weakness that they cannot extrapolate. Even after dimensionality reduction, our training dataset essentially consists of 614 points in a high-dimensional space, so that most of the points will lie on the convex hull. Of the 255 test points, many will lie outside the convex hull of the training points, which means that the model must extrapolate. To bring the extrapolation capability into the game, I implemented the ridge regression model. \n\nThe models have cv scores between 0.878 (ExtraTrees) and 0.906 (Py-boost). Py-boost, which was the worst in cross-validation, has the best public and private lb scores (0.572 and 0.748, respectively).\n\n## Data augmentation\n\nOne of the models (k nearest neighbors) is fed with **data augmentation**: If we know the differential expressions for two compounds, we may assume that a mixture of the two compounds will produce a differential expression which is the average of the two single-compound differential expressions.\n\nI experimented with another kind of data augmentation a well: Because there are more than twice as many T cells CD4+ as either Myeloid or B cells and I knew that the cell count biases the results of Limma, I reduced the cell count of the T cells CD4+, pseudobulked them, ran them through Limma and added the results to the training data as another cell type. This augmentation improved the scores of ExtraTrees, but not to the level of Py-boost. Perhaps I should have combined the additional cell type with Py-boost...\n\n# Robustness\n\nThe robustness of my models is demonstrated in two ways:\n\n(1) The models are fully cross-validated. The cross-validation strategy, first documented in [SCP Quickstart](https://www.kaggle.com/code/ambrosm/scp-quickstart), ensures that the model is validated on predicting cell_type–sm_name combinations so that it knows only 17 other compounds for the same cell type. This cross-validation strategy is more robust than the ordinary shuffled KFold, where the model knows 4/5 of all compounds for the same cell type. (And it is much more robust than a simple train-test-split.)\n\nI have to admit, though, that I'm not happy with the cv–lb correspondence.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F424a99a8b518dd08b96b0026f598c557%2Fcv-scheme.png?generation=1701390265025626&alt=media)\n\n(2) For all models the performance was tested after adding Gaussian noise to the input t-scores. All models are robust against small noise. When the noise gets stronger, the knn and ExtraTrees models suffer more than Py-boost and the ridge recommender system.\n \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F9f9538ca1d28e570ae930339c9490145%2Fnoise.png?generation=1701409934387211&alt=media)\n\n# Documentation and code style\n\nThe code is documented in the notebooks.\n\n# Reproducibility\n\nSource code is here:\n- [EDA which makes sense ⭐️⭐️⭐️⭐️⭐️](https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense)\n- [SCP #26: Py-boost, recommender system and ET](https://www.kaggle.com/code/ambrosm/scp-26-py-boost-recommender-system-and-et)\n- [GitHub](https://github.com/Ambros-M/Single-Cell-Perturbations-2023)\n\n# Conclusion\n\nLet me conclude by summarizing the four main messages of this post:\n\n1. Recommender systems are a promising starting point for developing models for cross-cell-type differential gene expression prediction. Because of commercial interests, recommender systems are a well-researched topic, and a lot of information is available.\n2. Data augmentation is useful, and mixtures of compounds are a natural approach to data augmentation.\n3. Although Kaggle competitions with data cleaning, outlier removal and unusual metrics are entertaining, the research objective would profit from another setting. Providing clean data and scoring with a well-understood metric would help participants focus on the real topic rather than the noise in the data.\n4. We have seen that Limma in certain situations produces biased outputs. I hope that professional Limma users are aware of these effects and account for them when interpreting results in their research.",
    "2544634": "\"Did you notice that in this competition no real EDA notebook has been published? \" at the beginning of your topic.\n\nThat question is remarkable, mostly for those that call every/anything EDA.\n\nNow, my PySmiles is laughing at mine Ridiculous Rdkit scPerturbations. \n\nThank you for the topic and the Kaggle Notebook where we can read the snippets. \n\nIndescreet question: How long (years of study/formation) to arrive in your professional level?\n\nVery likely more than 20 years. In my previous careers, More than 20 years is what is required to achieve such level of knowledge, skills and mostly Experience, leaving the Ego behind to help those that are starting in any field.\n\nThank you so much AmbrosM.",
    "2544745": "Congratulations. Thanks for sharing the nice writeup with colorful graphs and charts. \nI think if you had provided useful references with your article, it would be helpful.",
    "2544896": "I sincerely thank you for perfect sharing, @ambrosm ! You have provided us with great help, and your method is very impressive! And I have learned a lot of knowledge from your sharing.",
    "2545121": "Heads Up, and thanks for sharing this.",
    "2545444": ">Did you notice that in this competition no real EDA notebook has been publishedt\n\ntwo of mine are crying in the corner ashamed 😅\n\nThank you for the great description, I learned a lot, it's very interesting too) And congratulations with the silver!",
    "2545578": "Great writeup ! \nIt is always pleasure to learn from your work  !\nVery happy that you find my notebook on pyboost useful. \nCongratulations with the medal !",
    "2545669": "Sorry, @antoninadolgorukova, I overlooked them (maybe because I'm not fluent in R).\n\nCongratulations to you, too!",
    "2546231": "HI @alexandervc, thanks for bringing py-boost to my attention, and congratulations to you, too!",
    "2546273": "Don't be sorry I just meant what I said) Ashamed I saw strange things in the data but did not dig deeper to find out the reasons. So thank you a lot for the explanations and sharing!",
    "2546961": "Great writeup! And great visualizations indeed!",
    "2548681": "I don't think the conclusion that limma is biased is evident. Limma is being used incorrectly/poorly in this competition.\nAn important step of the usual limma application is to remove genes that \"consistently have zero or very low counts.\" using the `filterByExpr` function from edgeR (see limma userguide). To avoid having NaNs in the data the organizers skipped this step which probably interfered with the TMM normalization. Further the Empirical Bayes procedure in limma is susceptible to genes with very small (e.g. all zero) or very large variance. Phipson et al. 16 suggest a more robust EB method which can be activated using `robust=TRUE` in the `eBayes` function but this was not used by the competition organizers. You et al. 23 suggest a modification to voom which accounts for heteroscedasticity between groups in scRNA-seq data and suggest that not doing so leads to problems. This also was not used. Nevertheless I am also very suspect of DE results in general but especially for scRNA-seq data. \n\nDE analysis is very tricky and many intransparent assumptions are made which can often lead to problems (e.g. Squair et al. 21 show that many methods produce lots of false positives, Li et al 22 suggest high false discovery rates when there are large sample sizes)\nAll in all I don't think DE results are good prediction targets and I believe that this lead to the competition being focused on accounting for DE computation quirks rather than on predicting biologically meaningful change.\n\n*Robust hyperparameter estimation protects against hypervariable genes and improves power to detect differential expression* [Phipson et al 16](https://projecteuclid.org/journals/annals-of-applied-statistics/volume-10/issue-2/Robust-hyperparameter-estimation-protects-against-hypervariable-genes-and-improves-power/10.1214/16-AOAS920.full)\n\n*Modeling group heteroscedasticity in single-cell RNA-seq pseudo-bulk data* [You et al. 23](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-023-02949-2)\n\n*Confronting false discoveries in single-cell differential expression* [Squair et al. 21](https://www.nature.com/articles/s41467-021-25960-2)\n\n*Exaggerated false positives by popular differential expression methods when analyzing human population samples* [Li et al. 22](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-022-02648-4)",
    "2548941": "Combined with magic-top4 trick - your solution gets 0.718 on private LB - which is better than current top1.\nSee \nhttps://www.kaggle.com/alexandervc/op2-explore-4th-place-magic",
    "2549107": "Congratulations! I like your idea about cell counts and data augmentations.\n\n> Although Kaggle competitions with data cleaning, outlier removal and unusual metrics are entertaining, the research objective would profit from another setting. Providing clean data and scoring with a well-understood metric would help participants focus on the real topic rather than the noise in the data.\n\nI'm the one who focused on the noise in the data and then give up 😅.\nAfter reading your post, I deeply regret that I did not even try to find a solution for it.\nBeyond the tech part, I also learned a lot from your post, thank you for sharing.",
    "2556042": "Thank you, it was interesting to read the detailed description of the data analysis.\nYour notebooks carried a lot of ideas  in this competition",
    "2566894": "Thanks again for the great write-up !\nTo complement your table of suspicious samples.\n\nHere is another evidence that something is wrong with them.\nLet me look on count of genes which are strongly perturbed for each sample\nwe can expect it should not be too much.\nBut there are samples where we see about HALF(!) genes are perturbed strongly - that is suspicious.\nAnd that list quite corresponds to yours:\n\n(Also note: that these samples seems to be the WORST predicted - see https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461663 )\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F4e1c24530cbda8ef941825e5c464268c%2FScreenshot%202023-12-19%20092655.png?generation=1702975111733165&alt=media)\nhttps://www.kaggle.com/code/alexandervc/op2-eda-new?scriptVersionId=155628891&cellId=10\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F4e84fbb28aade383843e4a7ec4cdcf66%2FScreenshot%202023-12-19%20100930.png?generation=1702977039679503&alt=media)",
    "2576513": "To complement your analysis on drugs which potentially lead to misclassification of cell types:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fb25870b8e55dabe6f074c05749eebfe9%2FScreenshot%202023-12-27%20204555.png?generation=1703706506138338&alt=media)\n\nWe can see that: 'Oprozomib (ONX 0912)','MLN 2238','CEP-18770 (Delanzomib)'\nare concentrated in the same umap-seen cluster.  It is classified by T CD4/8 cells, but \nwe see that cluster is so far-separate from the other cells type clusters, that we can easily imagine that it can be mis-classified \n(since it is hard to imagine classification algorithms seen such kind of data - they were trained on normal cells , not drug-treated): \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F6def8f7d0ef31e73370e412481bb8144%2Fdrugs_ambrosM_1.png?generation=1703706523222615&alt=media)\nSee: https://www.kaggle.com/code/alexandervc/op2-rna-seq-data-scanpy-adata-cell-cycle?scriptVersionId=156757371&cellId=45\n\nSimilar other drugs you found: 'CGM-097', 'LY2090314','Ganetespib (STA-9090)','IN1451' - generate stand-alone clusters, so again it is reasonable to suspect that cell type classification may fail: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F0fc77b7f30d7b64ea0f5f08565a004db%2FScreenshot%202023-12-27%20205550.png?generation=1703706999692919&alt=media)",
    "2576581": "And to add more details/confirmation about \n'Ganetespib (STA-9090)'\nYour analysis suggest that we have too many NK-cells \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F4eaefc6acfe9777248419923b415c516%2FScreenshot%202023-12-27%20231609.png?generation=1703715458929376&alt=media)\n\nThat indeed quite corresponds to misclassification:\nwe see cluster - it is expected to be one cell-type, but we see two (and one of them - NK-cells - the one appeared in your analysis). \nMoreover - normally we can expect each cell type contains three donors - but that does not happen here. Here we see one cell type contains two donors another cell type just one donors. That is another  clear indication, that cell types were misclassified.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F4177d3d4ac38d4f08b5ac77badf551be%2FScreenshot%202023-12-27%20232630.png?generation=1703716036563754&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F48309d9847154aa4dff094779a1d8382%2FScreenshot%202023-12-27%20232951.png?generation=1703716220412047&alt=media)",
    "2981852": "Do you use Py-boost now?"
  },
  "source": "meta"
}