{
  "id": 455259,
  "title": "Challenge of Myeloid cells...",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/455259",
  "author_name": "",
  "post_date": "2023-11-14T02:01:43.152120600Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I am having trouble with capturing a good predictive model for Myeloid cells. They are the most distinct, at least in how the gene expressions correlate with other cell types, and reading further into their biology I see they are not as specific a category of a cell type - but rather board class of cells including Neutrophils, Monocytes, Macrophages, Dendritic Cells, Mast cells etc. </p>\n<p>Suppose an instance of a Myleoid cell data is actually from a Dendritic cell and another instance is a Macrophage - two different cell types both treated with compound C. The differential gene expression response, Yd, of Dendritic cell to compound C is going to be different than response of the Macrophage to compound C, Ym, but we have no way to differentiate our inputs Xd and Xm and outputs Yd and Ym - They have not provided that us with that information in the training -  both go as instances of the same cell category input and outputs into training of F(.), as in Y=F(X,C), where X is the cell type descriptor and C is the compound. So is the best F(.) can do here NOT learn the true response for Xd or Xm but somehow a less accurate version that is closer to the average response to C?</p>\n<p>Myeloid cells make up a significant portion of the test data.  Could this be part of why the most successful strategy so far in this competition is coming down to averaging ensembles? Any insight on how to  deal with diversity of cell types lumped into one category of Myeloid cells? Thank you!</p>",
  "messages": [
    {
      "id": "2524104",
      "postDate": "11/14/2023 02:01:43",
      "content": "<p>I am having trouble with capturing a good predictive model for Myeloid cells. They are the most distinct, at least in how the gene expressions correlate with other cell types, and reading further into their biology I see they are not as specific a category of a cell type - but rather board class of cells including Neutrophils, Monocytes, Macrophages, Dendritic Cells, Mast cells etc. </p>\n<p>Suppose an instance of a Myleoid cell data is actually from a Dendritic cell and another instance is a Macrophage - two different cell types both treated with compound C. The differential gene expression response, Yd, of Dendritic cell to compound C is going to be different than response of the Macrophage to compound C, Ym, but we have no way to differentiate our inputs Xd and Xm and outputs Yd and Ym - They have not provided that us with that information in the training -  both go as instances of the same cell category input and outputs into training of F(.), as in Y=F(X,C), where X is the cell type descriptor and C is the compound. So is the best F(.) can do here NOT learn the true response for Xd or Xm but somehow a less accurate version that is closer to the average response to C?</p>\n<p>Myeloid cells make up a significant portion of the test data.  Could this be part of why the most successful strategy so far in this competition is coming down to averaging ensembles? Any insight on how to  deal with diversity of cell types lumped into one category of Myeloid cells? Thank you!</p>",
      "rawMarkdown": "I am having trouble with capturing a good predictive model for Myeloid cells. They are the most distinct, at least in how the gene expressions correlate with other cell types, and reading further into their biology I see they are not as specific a category of a cell type - but rather board class of cells including Neutrophils, Monocytes, Macrophages, Dendritic Cells, Mast cells etc. \n\nSuppose an instance of a Myleoid cell data is actually from a Dendritic cell and another instance is a Macrophage - two different cell types both treated with compound C. The differential gene expression response, Yd, of Dendritic cell to compound C is going to be different than response of the Macrophage to compound C, Ym, but we have no way to differentiate our inputs Xd and Xm and outputs Yd and Ym - They have not provided that us with that information in the training -  both go as instances of the same cell category input and outputs into training of F(.), as in Y=F(X,C), where X is the cell type descriptor and C is the compound. So is the best F(.) can do here NOT learn the true response for Xd or Xm but somehow a less accurate version that is closer to the average response to C?\n \nMyeloid cells make up a significant portion of the test data.  Could this be part of why the most successful strategy so far in this competition is coming down to averaging ensembles? Any insight on how to  deal with diversity of cell types lumped into one category of Myeloid cells? Thank you!",
      "votes": null
    },
    {
      "id": "2527054",
      "postDate": "11/16/2023 08:44:33",
      "content": "<p>Thanks for sharing!<br>\nIt indeed might be a key point.</p>\n<p>If you can share your analysis to show that myeloid are different - would be very kind of you.</p>\n<p>Do not have good answer to your question.<br>\nJust couple of remarks:</p>\n<p>In recent public blends(ensembles) they put different weights for B-Cell znd Myeloid - that lines with [:128] ….</p>\n<p>We may try to check what models perform better on myelods locally with our CV and give them higher weight in blend (on Myeloid cells)</p>",
      "rawMarkdown": "Thanks for sharing!\nIt indeed might be a key point.\n\nIf you can share your analysis to show that myeloid are different - would be very kind of you.\n\nDo not have good answer to your question.\nJust couple of remarks:\n\nIn recent public blends(ensembles) they put different weights for B-Cell znd Myeloid - that lines with [:128] ....\n\nWe may try to check what models perform better on myelods locally with our CV and give them higher weight in blend (on Myeloid cells)",
      "votes": null
    },
    {
      "id": "2527685",
      "postDate": "11/16/2023 18:09:16",
      "content": "<p>Thank you for your comment! I think I was able to answer part of my own question as I did some analysis on the makeup of the cells based on gene expressions (96 negative control samples on all cell types) via UMAP.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1556503%2Fc33d3bd8fe94476047e1f892d0b7dec4%2FkaggleUmap.JPG?generation=1700157936971863&amp;alt=media\" alt=\"\"></p>\n<p>You can see in the UMAP clusters, the Myeloid cell clusters are tightly coupled. Perhaps, the organizers picked a very specific type of cell (only one of Monocytes,  Dendritic cells etc. and chose to not disclose what specific cell type).  Based on UMAP analysis, if indeed they focused on ONE TYPE of Myeloid Cell, I would be less concerned about Myleoid cells throwing the learning process off. The remaining question would be the indirect p-value based output generation process, the  -log(Prob( Gene_Expression), Nuisance parameter) transformation model that I speculate must be adding the noise into the output we are trying to predict in this model: Y=F(X,C)+ noise. So, the question would remain as to how big is the noise (nuisance factor) compared to the actual mode F(X,C). Any thoughts?</p>",
      "rawMarkdown": "Thank you for your comment! I think I was able to answer part of my own question as I did some analysis on the makeup of the cells based on gene expressions (96 negative control samples on all cell types) via UMAP.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1556503%2Fc33d3bd8fe94476047e1f892d0b7dec4%2FkaggleUmap.JPG?generation=1700157936971863&alt=media)\n\nYou can see in the UMAP clusters, the Myeloid cell clusters are tightly coupled. Perhaps, the organizers picked a very specific type of cell (only one of Monocytes,  Dendritic cells etc. and chose to not disclose what specific cell type).  Based on UMAP analysis, if indeed they focused on ONE TYPE of Myeloid Cell, I would be less concerned about Myleoid cells throwing the learning process off. The remaining question would be the indirect p-value based output generation process, the  -log(Prob( Gene_Expression), Nuisance parameter) transformation model that I speculate must be adding the noise into the output we are trying to predict in this model: Y=F(X,C)+ noise. So, the question would remain as to how big is the noise (nuisance factor) compared to the actual mode F(X,C). Any thoughts?",
      "votes": null
    },
    {
      "id": "2527833",
      "postDate": "11/16/2023 21:03:28",
      "content": "<p>Thanks for sharing ! <br>\nWhat data you use for that umap  ? Raw data - from \"adata\" ? I guess it is NOT differential expression \"de_train\" - our main input file - is that correct ?</p>",
      "rawMarkdown": "Thanks for sharing ! \nWhat data you use for that umap  ? Raw data - from \"adata\" ? I guess it is NOT differential expression \"de_train\" - our main input file - is that correct ?",
      "votes": null
    },
    {
      "id": "2527955",
      "postDate": "11/17/2023 02:00:50",
      "content": "<p>Precisely…  it's from adata and therefore NOT differential expression in the de_train!</p>",
      "rawMarkdown": "Precisely...  it's from adata and therefore NOT differential expression in the de_train!",
      "votes": null
    },
    {
      "id": "2528341",
      "postDate": "11/17/2023 09:33:51",
      "content": "<p>On a related note, B Cells and NK cells also have subtypes like CD56bright NK Cells, CD56dim NK Cells, Naive B Cells, Memory B cells, etc. It's a common challenge, and strategies addressing this diversity could lead to more effective modeling</p>",
      "rawMarkdown": "On a related note, B Cells and NK cells also have subtypes like CD56bright NK Cells, CD56dim NK Cells, Naive B Cells, Memory B cells, etc. It's a common challenge, and strategies addressing this diversity could lead to more effective modeling",
      "votes": null
    },
    {
      "id": "2529049",
      "postDate": "11/17/2023 21:30:59",
      "content": "<p>By the way here is some clustering of cell types based on Differential expression data:</p>\n<p><a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=147818286&amp;cellId=21\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=147818286&amp;cellId=21</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Feca255be612094a8bd4c4682f93c846e%2FScreenshot%202023-11-17%20222806.png?generation=1700256605674386&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "By the way here is some clustering of cell types based on Differential expression data:\n\nhttps://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=147818286&cellId=21\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Feca255be612094a8bd4c4682f93c846e%2FScreenshot%202023-11-17%20222806.png?generation=1700256605674386&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2527054,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "11/16/2023 08:44:33",
      "content": "<p>Thanks for sharing!<br>\nIt indeed might be a key point.</p>\n<p>If you can share your analysis to show that myeloid are different - would be very kind of you.</p>\n<p>Do not have good answer to your question.<br>\nJust couple of remarks:</p>\n<p>In recent public blends(ensembles) they put different weights for B-Cell znd Myeloid - that lines with [:128] ….</p>\n<p>We may try to check what models perform better on myelods locally with our CV and give them higher weight in blend (on Myeloid cells)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2527685,
          "author_name": "nanocipher",
          "author_url": "",
          "post_date": "11/16/2023 18:09:16",
          "content": "<p>Thank you for your comment! I think I was able to answer part of my own question as I did some analysis on the makeup of the cells based on gene expressions (96 negative control samples on all cell types) via UMAP.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1556503%2Fc33d3bd8fe94476047e1f892d0b7dec4%2FkaggleUmap.JPG?generation=1700157936971863&amp;alt=media\" alt=\"\"></p>\n<p>You can see in the UMAP clusters, the Myeloid cell clusters are tightly coupled. Perhaps, the organizers picked a very specific type of cell (only one of Monocytes,  Dendritic cells etc. and chose to not disclose what specific cell type).  Based on UMAP analysis, if indeed they focused on ONE TYPE of Myeloid Cell, I would be less concerned about Myleoid cells throwing the learning process off. The remaining question would be the indirect p-value based output generation process, the  -log(Prob( Gene_Expression), Nuisance parameter) transformation model that I speculate must be adding the noise into the output we are trying to predict in this model: Y=F(X,C)+ noise. So, the question would remain as to how big is the noise (nuisance factor) compared to the actual mode F(X,C). Any thoughts?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2527833,
              "author_name": "alexandervc",
              "author_url": "",
              "post_date": "11/16/2023 21:03:28",
              "content": "<p>Thanks for sharing ! <br>\nWhat data you use for that umap  ? Raw data - from \"adata\" ? I guess it is NOT differential expression \"de_train\" - our main input file - is that correct ?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2527955,
                  "author_name": "nanocipher",
                  "author_url": "",
                  "post_date": "11/17/2023 02:00:50",
                  "content": "<p>Precisely…  it's from adata and therefore NOT differential expression in the de_train!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2528341,
      "author_name": "kishanvavdara",
      "author_url": "",
      "post_date": "11/17/2023 09:33:51",
      "content": "<p>On a related note, B Cells and NK cells also have subtypes like CD56bright NK Cells, CD56dim NK Cells, Naive B Cells, Memory B cells, etc. It's a common challenge, and strategies addressing this diversity could lead to more effective modeling</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2529049,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "11/17/2023 21:30:59",
      "content": "<p>By the way here is some clustering of cell types based on Differential expression data:</p>\n<p><a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=147818286&amp;cellId=21\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=147818286&amp;cellId=21</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Feca255be612094a8bd4c4682f93c846e%2FScreenshot%202023-11-17%20222806.png?generation=1700256605674386&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2524104": "I am having trouble with capturing a good predictive model for Myeloid cells. They are the most distinct, at least in how the gene expressions correlate with other cell types, and reading further into their biology I see they are not as specific a category of a cell type - but rather board class of cells including Neutrophils, Monocytes, Macrophages, Dendritic Cells, Mast cells etc. \n\nSuppose an instance of a Myleoid cell data is actually from a Dendritic cell and another instance is a Macrophage - two different cell types both treated with compound C. The differential gene expression response, Yd, of Dendritic cell to compound C is going to be different than response of the Macrophage to compound C, Ym, but we have no way to differentiate our inputs Xd and Xm and outputs Yd and Ym - They have not provided that us with that information in the training -  both go as instances of the same cell category input and outputs into training of F(.), as in Y=F(X,C), where X is the cell type descriptor and C is the compound. So is the best F(.) can do here NOT learn the true response for Xd or Xm but somehow a less accurate version that is closer to the average response to C?\n \nMyeloid cells make up a significant portion of the test data.  Could this be part of why the most successful strategy so far in this competition is coming down to averaging ensembles? Any insight on how to  deal with diversity of cell types lumped into one category of Myeloid cells? Thank you!",
    "2527054": "Thanks for sharing!\nIt indeed might be a key point.\n\nIf you can share your analysis to show that myeloid are different - would be very kind of you.\n\nDo not have good answer to your question.\nJust couple of remarks:\n\nIn recent public blends(ensembles) they put different weights for B-Cell znd Myeloid - that lines with [:128] ....\n\nWe may try to check what models perform better on myelods locally with our CV and give them higher weight in blend (on Myeloid cells)",
    "2527685": "Thank you for your comment! I think I was able to answer part of my own question as I did some analysis on the makeup of the cells based on gene expressions (96 negative control samples on all cell types) via UMAP.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1556503%2Fc33d3bd8fe94476047e1f892d0b7dec4%2FkaggleUmap.JPG?generation=1700157936971863&alt=media)\n\nYou can see in the UMAP clusters, the Myeloid cell clusters are tightly coupled. Perhaps, the organizers picked a very specific type of cell (only one of Monocytes,  Dendritic cells etc. and chose to not disclose what specific cell type).  Based on UMAP analysis, if indeed they focused on ONE TYPE of Myeloid Cell, I would be less concerned about Myleoid cells throwing the learning process off. The remaining question would be the indirect p-value based output generation process, the  -log(Prob( Gene_Expression), Nuisance parameter) transformation model that I speculate must be adding the noise into the output we are trying to predict in this model: Y=F(X,C)+ noise. So, the question would remain as to how big is the noise (nuisance factor) compared to the actual mode F(X,C). Any thoughts?",
    "2527833": "Thanks for sharing ! \nWhat data you use for that umap  ? Raw data - from \"adata\" ? I guess it is NOT differential expression \"de_train\" - our main input file - is that correct ?",
    "2527955": "Precisely...  it's from adata and therefore NOT differential expression in the de_train!",
    "2528341": "On a related note, B Cells and NK cells also have subtypes like CD56bright NK Cells, CD56dim NK Cells, Naive B Cells, Memory B cells, etc. It's a common challenge, and strategies addressing this diversity could lead to more effective modeling",
    "2529049": "By the way here is some clustering of cell types based on Differential expression data:\n\nhttps://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=147818286&cellId=21\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Feca255be612094a8bd4c4682f93c846e%2FScreenshot%202023-11-17%20222806.png?generation=1700256605674386&alt=media)"
  },
  "source": "meta"
}