{
  "id": 461663,
  "title": "Insights from the comparison of different models (13th place)",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/461663",
  "author_name": "",
  "post_date": "2023-12-15T18:09:57.477228400Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello community! I'd like to share some insights from the analysis of the predictions of different models used in our solution (13th place) - made in <a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-analysis-of-different-models-predictions\" target=\"_blank\">this notebook</a>.</p>\n<p>Would be happy if you could share any comments or thoughts about the approach used and the findings.</p>\n<p>The analysis is mainly based on comparing the prediction variabilities of the models as a measure of uncertainty. </p>\n<p>Using all versions of MLP NN (10) and Pyboost (3), which were included in the final ensemble, and 4 versions each of Catboost and NLP, which are outside the final solution but have similar LB scores, the prediction variabilities were obtained as follows:</p>\n<ul>\n<li>calculated the standard deviation (SD) of the predictions of the same gene for the same id </li>\n<li>calculated the median of these standard deviations by id (sm_name*cell_type pair) or by gene and cell type.</li>\n</ul>\n<p>1) All models are less confident in their predictions for myeloid cells compared to B cells (medians of prediction variability across genes and samples are higher).</p>\n<p>2) However, the highest bias (differences between predicted and true values) and variability of gene expression change predictions are associated with individual drugs rather than cell types. </p>\n<p>If we plot intersections between the drugs that are hardest to predict for each model, we see that almost half (42.9%) of them overlap for all 4 models.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F3102542797ecd61c923227233a237616%2FScreenshot%202023-12-15%20204109.png?generation=1702662117938915&amp;alt=media\" alt=\"\"></p>\n<p>These drugs are mostly outliers identified by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> in his <a href=\"https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense\" target=\"_blank\">Excellent EDA</a> as compounds whose data were obtained from the lowest number of cells (≤10 cells), or drugs that affected the cells in such a way that they were misclassified.</p>\n<p>Interestingly, in opposite to drugs for which the models give the most certain predictions of gene expression change, most of the hard-to-predict drugs have evidenced or potential cytotoxic effects which could explain the lower number of cells (toxic effect) or misclassification (overlapping gene expression patterns across different cell types?).</p>\n<p>3) Now, if we compare the prediction variability across samples (the medians for each gene, the highest medians indicate genes for which it is most hard to predict expression change in general), we can see that:</p>\n<ul>\n<li><p>a very small proportion of the hard and easy-to-predict genes is the same for all 4 models. Still, if we do despite this GO enrichment analysis comparing these subsets of genes, it appears, as opposed to most certainly predicted genes, the hard-to-predict genes are often related to immune cell activities, cytotoxicity, and cell death, but I'm not sure how reliably is this (the intersection is very small and might be just by chance). </p></li>\n<li><p>most of the genes, for which all the models have the lowest prediction confidence are also the ones displaying the most variable responses to the compounds ( the higher the standard deviations (SD) across samples, the higher the prediction variance), but there are some exclusions - genes with high standard deviations and relatively low prediction variance, as well as genes for which the models give highly variable predictions despite their comparatively stable expression profiles.</p></li>\n</ul>\n<p>However, only a limited subset of genes of both groups is common across all four models, indicating that such genes are model-specific and rather attributed to computational differences than biological factors.</p>",
  "messages": [
    {
      "id": "2562867",
      "postDate": "12/15/2023 18:09:57",
      "content": "<p>Hello community! I'd like to share some insights from the analysis of the predictions of different models used in our solution (13th place) - made in <a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-analysis-of-different-models-predictions\" target=\"_blank\">this notebook</a>.</p>\n<p>Would be happy if you could share any comments or thoughts about the approach used and the findings.</p>\n<p>The analysis is mainly based on comparing the prediction variabilities of the models as a measure of uncertainty. </p>\n<p>Using all versions of MLP NN (10) and Pyboost (3), which were included in the final ensemble, and 4 versions each of Catboost and NLP, which are outside the final solution but have similar LB scores, the prediction variabilities were obtained as follows:</p>\n<ul>\n<li>calculated the standard deviation (SD) of the predictions of the same gene for the same id </li>\n<li>calculated the median of these standard deviations by id (sm_name*cell_type pair) or by gene and cell type.</li>\n</ul>\n<p>1) All models are less confident in their predictions for myeloid cells compared to B cells (medians of prediction variability across genes and samples are higher).</p>\n<p>2) However, the highest bias (differences between predicted and true values) and variability of gene expression change predictions are associated with individual drugs rather than cell types. </p>\n<p>If we plot intersections between the drugs that are hardest to predict for each model, we see that almost half (42.9%) of them overlap for all 4 models.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F3102542797ecd61c923227233a237616%2FScreenshot%202023-12-15%20204109.png?generation=1702662117938915&amp;alt=media\" alt=\"\"></p>\n<p>These drugs are mostly outliers identified by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> in his <a href=\"https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense\" target=\"_blank\">Excellent EDA</a> as compounds whose data were obtained from the lowest number of cells (≤10 cells), or drugs that affected the cells in such a way that they were misclassified.</p>\n<p>Interestingly, in opposite to drugs for which the models give the most certain predictions of gene expression change, most of the hard-to-predict drugs have evidenced or potential cytotoxic effects which could explain the lower number of cells (toxic effect) or misclassification (overlapping gene expression patterns across different cell types?).</p>\n<p>3) Now, if we compare the prediction variability across samples (the medians for each gene, the highest medians indicate genes for which it is most hard to predict expression change in general), we can see that:</p>\n<ul>\n<li><p>a very small proportion of the hard and easy-to-predict genes is the same for all 4 models. Still, if we do despite this GO enrichment analysis comparing these subsets of genes, it appears, as opposed to most certainly predicted genes, the hard-to-predict genes are often related to immune cell activities, cytotoxicity, and cell death, but I'm not sure how reliably is this (the intersection is very small and might be just by chance). </p></li>\n<li><p>most of the genes, for which all the models have the lowest prediction confidence are also the ones displaying the most variable responses to the compounds ( the higher the standard deviations (SD) across samples, the higher the prediction variance), but there are some exclusions - genes with high standard deviations and relatively low prediction variance, as well as genes for which the models give highly variable predictions despite their comparatively stable expression profiles.</p></li>\n</ul>\n<p>However, only a limited subset of genes of both groups is common across all four models, indicating that such genes are model-specific and rather attributed to computational differences than biological factors.</p>",
      "rawMarkdown": "Hello community! I'd like to share some insights from the analysis of the predictions of different models used in our solution (13th place) - made in [this notebook](https://www.kaggle.com/code/antoninadolgorukova/op2-analysis-of-different-models-predictions).\n\nWould be happy if you could share any comments or thoughts about the approach used and the findings.\n\nThe analysis is mainly based on comparing the prediction variabilities of the models as a measure of uncertainty. \n\nUsing all versions of MLP NN (10) and Pyboost (3), which were included in the final ensemble, and 4 versions each of Catboost and NLP, which are outside the final solution but have similar LB scores, the prediction variabilities were obtained as follows:\n\n- calculated the standard deviation (SD) of the predictions of the same gene for the same id \n- calculated the median of these standard deviations by id (sm_name*cell_type pair) or by gene and cell type.\n\n1) All models are less confident in their predictions for myeloid cells compared to B cells (medians of prediction variability across genes and samples are higher).\n\n2) However, the highest bias (differences between predicted and true values) and variability of gene expression change predictions are associated with individual drugs rather than cell types. \n\nIf we plot intersections between the drugs that are hardest to predict for each model, we see that almost half (42.9%) of them overlap for all 4 models.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F3102542797ecd61c923227233a237616%2FScreenshot%202023-12-15%20204109.png?generation=1702662117938915&alt=media)\n\nThese drugs are mostly outliers identified by @ambrosm in his [Excellent EDA](https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense) as compounds whose data were obtained from the lowest number of cells (≤10 cells), or drugs that affected the cells in such a way that they were misclassified.\n\nInterestingly, in opposite to drugs for which the models give the most certain predictions of gene expression change, most of the hard-to-predict drugs have evidenced or potential cytotoxic effects which could explain the lower number of cells (toxic effect) or misclassification (overlapping gene expression patterns across different cell types?).\n\n3) Now, if we compare the prediction variability across samples (the medians for each gene, the highest medians indicate genes for which it is most hard to predict expression change in general), we can see that:\n\n- a very small proportion of the hard and easy-to-predict genes is the same for all 4 models. Still, if we do despite this GO enrichment analysis comparing these subsets of genes, it appears, as opposed to most certainly predicted genes, the hard-to-predict genes are often related to immune cell activities, cytotoxicity, and cell death, but I'm not sure how reliably is this (the intersection is very small and might be just by chance). \n\n- most of the genes, for which all the models have the lowest prediction confidence are also the ones displaying the most variable responses to the compounds ( the higher the standard deviations (SD) across samples, the higher the prediction variance), but there are some exclusions - genes with high standard deviations and relatively low prediction variance, as well as genes for which the models give highly variable predictions despite their comparatively stable expression profiles.\n\nHowever, only a limited subset of genes of both groups is common across all four models, indicating that such genes are model-specific and rather attributed to computational differences than biological factors.",
      "votes": null
    },
    {
      "id": "2738899",
      "postDate": "04/06/2024 17:30:15",
      "content": "<p>Hi Antonina,<br>\nAlthough I'm late 4 months: Congratulations for your 13th place on Open Problems/Single Cells.<br>\nThat's is really a huge, amazing achievement. </p>",
      "rawMarkdown": "Hi Antonina,\nAlthough I'm late 4 months: Congratulations for your 13th place on Open Problems/Single Cells.\nThat's is really a huge, amazing achievement.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2738899,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "04/06/2024 17:30:15",
      "content": "<p>Hi Antonina,<br>\nAlthough I'm late 4 months: Congratulations for your 13th place on Open Problems/Single Cells.<br>\nThat's is really a huge, amazing achievement. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2562867": "Hello community! I'd like to share some insights from the analysis of the predictions of different models used in our solution (13th place) - made in [this notebook](https://www.kaggle.com/code/antoninadolgorukova/op2-analysis-of-different-models-predictions).\n\nWould be happy if you could share any comments or thoughts about the approach used and the findings.\n\nThe analysis is mainly based on comparing the prediction variabilities of the models as a measure of uncertainty. \n\nUsing all versions of MLP NN (10) and Pyboost (3), which were included in the final ensemble, and 4 versions each of Catboost and NLP, which are outside the final solution but have similar LB scores, the prediction variabilities were obtained as follows:\n\n- calculated the standard deviation (SD) of the predictions of the same gene for the same id \n- calculated the median of these standard deviations by id (sm_name*cell_type pair) or by gene and cell type.\n\n1) All models are less confident in their predictions for myeloid cells compared to B cells (medians of prediction variability across genes and samples are higher).\n\n2) However, the highest bias (differences between predicted and true values) and variability of gene expression change predictions are associated with individual drugs rather than cell types. \n\nIf we plot intersections between the drugs that are hardest to predict for each model, we see that almost half (42.9%) of them overlap for all 4 models.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F3102542797ecd61c923227233a237616%2FScreenshot%202023-12-15%20204109.png?generation=1702662117938915&alt=media)\n\nThese drugs are mostly outliers identified by @ambrosm in his [Excellent EDA](https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense) as compounds whose data were obtained from the lowest number of cells (≤10 cells), or drugs that affected the cells in such a way that they were misclassified.\n\nInterestingly, in opposite to drugs for which the models give the most certain predictions of gene expression change, most of the hard-to-predict drugs have evidenced or potential cytotoxic effects which could explain the lower number of cells (toxic effect) or misclassification (overlapping gene expression patterns across different cell types?).\n\n3) Now, if we compare the prediction variability across samples (the medians for each gene, the highest medians indicate genes for which it is most hard to predict expression change in general), we can see that:\n\n- a very small proportion of the hard and easy-to-predict genes is the same for all 4 models. Still, if we do despite this GO enrichment analysis comparing these subsets of genes, it appears, as opposed to most certainly predicted genes, the hard-to-predict genes are often related to immune cell activities, cytotoxicity, and cell death, but I'm not sure how reliably is this (the intersection is very small and might be just by chance). \n\n- most of the genes, for which all the models have the lowest prediction confidence are also the ones displaying the most variable responses to the compounds ( the higher the standard deviations (SD) across samples, the higher the prediction variance), but there are some exclusions - genes with high standard deviations and relatively low prediction variance, as well as genes for which the models give highly variable predictions despite their comparatively stable expression profiles.\n\nHowever, only a limited subset of genes of both groups is common across all four models, indicating that such genes are model-specific and rather attributed to computational differences than biological factors.",
    "2738899": "Hi Antonina,\nAlthough I'm late 4 months: Congratulations for your 13th place on Open Problems/Single Cells.\nThat's is really a huge, amazing achievement."
  },
  "source": "meta"
}