{
  "id": 346210,
  "title": "Negative values in DSB-normalized CITE-seq data",
  "url": "/competitions/open-problems-multimodal/discussion/346210",
  "author_name": "",
  "post_date": "2022-08-18T12:44:54.193744Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>A hopefully reasonable comment/question on the choice of normalization for CITE-seq data.  DSB normalization can generate negative values.  This is not intuitive, because cells should never express negative amounts of proteins.  Predicting or fitting to negative protein expression is probably not a good idea in theory or practice, especially when using activation functions such as ReLU or when fitting models with NNLS which do better at imputing signal dropout than unconstrained least squares, etc..</p>\n<p>Note that just over 25% of values in \"train_cite_targets\" are less than 0.</p>\n<p>I understand that normalization is an ongoing challenge without consensus in single-cell data analysis, but I want to at least comment that suboptimal pre-processing can result in suboptimal approaches to prediction through integration.</p>",
  "messages": [
    {
      "id": "1904728",
      "postDate": "08/18/2022 12:44:54",
      "content": "<p>A hopefully reasonable comment/question on the choice of normalization for CITE-seq data.  DSB normalization can generate negative values.  This is not intuitive, because cells should never express negative amounts of proteins.  Predicting or fitting to negative protein expression is probably not a good idea in theory or practice, especially when using activation functions such as ReLU or when fitting models with NNLS which do better at imputing signal dropout than unconstrained least squares, etc..</p>\n<p>Note that just over 25% of values in \"train_cite_targets\" are less than 0.</p>\n<p>I understand that normalization is an ongoing challenge without consensus in single-cell data analysis, but I want to at least comment that suboptimal pre-processing can result in suboptimal approaches to prediction through integration.</p>",
      "rawMarkdown": "A hopefully reasonable comment/question on the choice of normalization for CITE-seq data.  DSB normalization can generate negative values.  This is not intuitive, because cells should never express negative amounts of proteins.  Predicting or fitting to negative protein expression is probably not a good idea in theory or practice, especially when using activation functions such as ReLU or when fitting models with NNLS which do better at imputing signal dropout than unconstrained least squares, etc..\n\nNote that just over 25% of values in \"train_cite_targets\" are less than 0.\n\nI understand that normalization is an ongoing challenge without consensus in single-cell data analysis, but I want to at least comment that suboptimal pre-processing can result in suboptimal approaches to prediction through integration.",
      "votes": null
    },
    {
      "id": "1906032",
      "postDate": "08/19/2022 14:16:13",
      "content": "<p>I agree this is weird! We've found that when dealing with samples from different sites, DSB normalization offers the most consistent patterns of protein expression that match our understanding of the biological system.</p>\n<p>That said, I agree that cells don't have negative expression, and my understanding is this is just a random sampling problem where the amount of protein measured in that cell is less than what we observe in the background (which is expected).</p>\n<p>I'm happy to hear other suggestions for normalizing the data, and as I mentioned in a few other threads, we're hard at work preparing the un-normalized counts for folks who want them.</p>",
      "rawMarkdown": "I agree this is weird! We've found that when dealing with samples from different sites, DSB normalization offers the most consistent patterns of protein expression that match our understanding of the biological system.\n\nThat said, I agree that cells don't have negative expression, and my understanding is this is just a random sampling problem where the amount of protein measured in that cell is less than what we observe in the background (which is expected).\n\nI'm happy to hear other suggestions for normalizing the data, and as I mentioned in a few other threads, we're hard at work preparing the un-normalized counts for folks who want them.",
      "votes": null
    },
    {
      "id": "1906080",
      "postDate": "08/19/2022 14:54:59",
      "content": "<p>Awesome.  We've quite well discussed this in a different discussion thread, so thanks for sharing your perspective there.  I did discuss this with a few folks at my institute and we're kind of wondering whether a hard cutoff at zero after DSB normalization wouldn't be most appropriate (this is beyond the scope of the competition, more thinking through the open questions in the field).  It does seem that in poor-quality clinical samples (especially those that weren't frozen down quickly) DSB normalization can introduce batch effect-like discrepancies between samples, as an artifact of total signal and variance of that signal.  But I haven't worked much with this and have no better suggestions.</p>",
      "rawMarkdown": "Awesome.  We've quite well discussed this in a different discussion thread, so thanks for sharing your perspective there.  I did discuss this with a few folks at my institute and we're kind of wondering whether a hard cutoff at zero after DSB normalization wouldn't be most appropriate (this is beyond the scope of the competition, more thinking through the open questions in the field).  It does seem that in poor-quality clinical samples (especially those that weren't frozen down quickly) DSB normalization can introduce batch effect-like discrepancies between samples, as an artifact of total signal and variance of that signal.  But I haven't worked much with this and have no better suggestions.",
      "votes": null
    },
    {
      "id": "1916335",
      "postDate": "08/27/2022 19:14:51",
      "content": "<p><a href=\"https://www.kaggle.com/zdebruine\" target=\"_blank\">@zdebruine</a> <br>\nPlease look at:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293</a><br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348294\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348294</a><br>\nWould be so kind to join and/or give a webinar on your work ? </p>",
      "rawMarkdown": "zdebruine \nPlease look at:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/348294\nWould be so kind to join and/or give a webinar on your work ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1906032,
      "author_name": "danielburkhardt",
      "author_url": "",
      "post_date": "08/19/2022 14:16:13",
      "content": "<p>I agree this is weird! We've found that when dealing with samples from different sites, DSB normalization offers the most consistent patterns of protein expression that match our understanding of the biological system.</p>\n<p>That said, I agree that cells don't have negative expression, and my understanding is this is just a random sampling problem where the amount of protein measured in that cell is less than what we observe in the background (which is expected).</p>\n<p>I'm happy to hear other suggestions for normalizing the data, and as I mentioned in a few other threads, we're hard at work preparing the un-normalized counts for folks who want them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1906080,
          "author_name": "zdebruine",
          "author_url": "",
          "post_date": "08/19/2022 14:54:59",
          "content": "<p>Awesome.  We've quite well discussed this in a different discussion thread, so thanks for sharing your perspective there.  I did discuss this with a few folks at my institute and we're kind of wondering whether a hard cutoff at zero after DSB normalization wouldn't be most appropriate (this is beyond the scope of the competition, more thinking through the open questions in the field).  It does seem that in poor-quality clinical samples (especially those that weren't frozen down quickly) DSB normalization can introduce batch effect-like discrepancies between samples, as an artifact of total signal and variance of that signal.  But I haven't worked much with this and have no better suggestions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1916335,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "08/27/2022 19:14:51",
      "content": "<p><a href=\"https://www.kaggle.com/zdebruine\" target=\"_blank\">@zdebruine</a> <br>\nPlease look at:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293</a><br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348294\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348294</a><br>\nWould be so kind to join and/or give a webinar on your work ? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1904728": "A hopefully reasonable comment/question on the choice of normalization for CITE-seq data.  DSB normalization can generate negative values.  This is not intuitive, because cells should never express negative amounts of proteins.  Predicting or fitting to negative protein expression is probably not a good idea in theory or practice, especially when using activation functions such as ReLU or when fitting models with NNLS which do better at imputing signal dropout than unconstrained least squares, etc..\n\nNote that just over 25% of values in \"train_cite_targets\" are less than 0.\n\nI understand that normalization is an ongoing challenge without consensus in single-cell data analysis, but I want to at least comment that suboptimal pre-processing can result in suboptimal approaches to prediction through integration.",
    "1906032": "I agree this is weird! We've found that when dealing with samples from different sites, DSB normalization offers the most consistent patterns of protein expression that match our understanding of the biological system.\n\nThat said, I agree that cells don't have negative expression, and my understanding is this is just a random sampling problem where the amount of protein measured in that cell is less than what we observe in the background (which is expected).\n\nI'm happy to hear other suggestions for normalizing the data, and as I mentioned in a few other threads, we're hard at work preparing the un-normalized counts for folks who want them.",
    "1906080": "Awesome.  We've quite well discussed this in a different discussion thread, so thanks for sharing your perspective there.  I did discuss this with a few folks at my institute and we're kind of wondering whether a hard cutoff at zero after DSB normalization wouldn't be most appropriate (this is beyond the scope of the competition, more thinking through the open questions in the field).  It does seem that in poor-quality clinical samples (especially those that weren't frozen down quickly) DSB normalization can introduce batch effect-like discrepancies between samples, as an artifact of total signal and variance of that signal.  But I haven't worked much with this and have no better suggestions.",
    "1916335": "zdebruine \nPlease look at:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/348294\nWould be so kind to join and/or give a webinar on your work ?"
  },
  "source": "meta"
}