{
  "id": 349132,
  "title": "Postprocessing Tips: normalize Y to 1e6 (Multiome), etc",
  "url": "/competitions/open-problems-multimodal/discussion/349132",
  "author_name": "",
  "post_date": "2022-08-31T10:35:12.579081600Z",
  "votes": 12,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Suggestion for Multiome prediction:</p>\n<p>1)</p>\n<p>Do normalization to achieve : (np.exp(Y)-1).sum(axis = 1) = 1e6 - because true targets satisfy that. (Well, \"satisfy\" - up to machine precision, but nevertheless. You can check - see below.) Why satisfy - that is standard biological processing - it is mentioned in data info tab, and well-known for everyone in the field.</p>\n<p>How ? Can be different ways, but the most straighforward one:</p>\n<p>1) calculate predictions \"Y\" 2) calculate normalizer Z = sum(exp(Y)) 3) renorm: Y_i -&gt; Y_i + (log((1e6+22050 )/Z))</p>\n<p>2)</p>\n<p>Do not forget that true answers are non-negative - that is why it might be useful to clip negative predictions to zero - even before step 1.</p>\n<p>3)</p>\n<p>It might be exp(Y) are INTs - but that is not clear for the moment - depends what soft organizers used. Will investigate later. If it so: then use processing to achieve ints.</p>\n<p>=== </p>\n<p>See notebook for checks:</p>\n<p><a href=\"https://www.kaggle.com/code/alexandervc/postprocessing-tips-normalize-y-to-1e6-multiome\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/postprocessing-tips-normalize-y-to-1e6-multiome</a></p>\n<p>Analysis as usually on Kaggle.</p>",
  "messages": [
    {
      "id": "1920703",
      "postDate": "08/31/2022 10:35:12",
      "content": "<p>Suggestion for Multiome prediction:</p>\n<p>1)</p>\n<p>Do normalization to achieve : (np.exp(Y)-1).sum(axis = 1) = 1e6 - because true targets satisfy that. (Well, \"satisfy\" - up to machine precision, but nevertheless. You can check - see below.) Why satisfy - that is standard biological processing - it is mentioned in data info tab, and well-known for everyone in the field.</p>\n<p>How ? Can be different ways, but the most straighforward one:</p>\n<p>1) calculate predictions \"Y\" 2) calculate normalizer Z = sum(exp(Y)) 3) renorm: Y_i -&gt; Y_i + (log((1e6+22050 )/Z))</p>\n<p>2)</p>\n<p>Do not forget that true answers are non-negative - that is why it might be useful to clip negative predictions to zero - even before step 1.</p>\n<p>3)</p>\n<p>It might be exp(Y) are INTs - but that is not clear for the moment - depends what soft organizers used. Will investigate later. If it so: then use processing to achieve ints.</p>\n<p>=== </p>\n<p>See notebook for checks:</p>\n<p><a href=\"https://www.kaggle.com/code/alexandervc/postprocessing-tips-normalize-y-to-1e6-multiome\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/postprocessing-tips-normalize-y-to-1e6-multiome</a></p>\n<p>Analysis as usually on Kaggle.</p>",
      "rawMarkdown": "Suggestion for Multiome prediction:\n\n1)\n\nDo normalization to achieve : (np.exp(Y)-1).sum(axis = 1) = 1e6 - because true targets satisfy that. (Well, \"satisfy\" - up to machine precision, but nevertheless. You can check - see below.) Why satisfy - that is standard biological processing - it is mentioned in data info tab, and well-known for everyone in the field.\n\nHow ? Can be different ways, but the most straighforward one:\n\n1) calculate predictions \"Y\" 2) calculate normalizer Z = sum(exp(Y)) 3) renorm: Y_i -> Y_i + (log((1e6+22050 )/Z))\n\n2)\n\nDo not forget that true answers are non-negative - that is why it might be useful to clip negative predictions to zero - even before step 1.\n\n3)\n\nIt might be exp(Y) are INTs - but that is not clear for the moment - depends what soft organizers used. Will investigate later. If it so: then use processing to achieve ints.\n\n=== \n\nSee notebook for checks:\n\nhttps://www.kaggle.com/code/alexandervc/postprocessing-tips-normalize-y-to-1e6-multiome\n\nAnalysis as usually on Kaggle.",
      "votes": null
    },
    {
      "id": "1921752",
      "postDate": "09/01/2022 03:25:43",
      "content": "<p>Thanks for the info. Cliping to zero is not helpful so far. Can someone help explain why in standard biological processing, it is necessary to do above normalization? Cheers</p>",
      "rawMarkdown": "Thanks for the info. Cliping to zero is not helpful so far. Can someone help explain why in standard biological processing, it is necessary to do above normalization? Cheers",
      "votes": null
    },
    {
      "id": "1922059",
      "postDate": "09/01/2022 08:41:39",
      "content": "<p>The short answer is : because of PCR - Polymerase chain reaction, and other technical problems to get that data</p>\n<p>The counts what we see are NOT the real counts in the cell, but counts after PCR reaction which is random - for one cell you might get much more counts than for another cell, despite real number of counts is the same - but just PCR was going on differently and may be some other randomness in  the process of getting these numbers. </p>\n<p>I am not biologist , if any biologist can comment - would be great. </p>\n<p>So to avoid these technical problem people use normalization to some constant. Orgs choose 1e6, others choose other number e.g. 1e4.</p>\n<p>Is it the perfect way to do such normalization ? <br>\nNo. <br>\nBut that is the most simple, common and stable one for the moment. <br>\nPS <br>\nThere are research papers with proposals of different normalization in particular we also discuss it and even have some examples on Kaggle.  <a href=\"https://arxiv.org/abs/2208.05229\" target=\"_blank\">https://arxiv.org/abs/2208.05229</a></p>\n<p>Organizers mentioned they plan to release so called \"raw\" counts - that mean the counts before normalization - I hope for that - but it will not change things drastically. </p>",
      "rawMarkdown": "The short answer is : because of PCR - Polymerase chain reaction, and other technical problems to get that data\n\nThe counts what we see are NOT the real counts in the cell, but counts after PCR reaction which is random - for one cell you might get much more counts than for another cell, despite real number of counts is the same - but just PCR was going on differently and may be some other randomness in  the process of getting these numbers. \n\nI am not biologist , if any biologist can comment - would be great. \n\nSo to avoid these technical problem people use normalization to some constant. Orgs choose 1e6, others choose other number e.g. 1e4.\n\nIs it the perfect way to do such normalization ? \nNo. \nBut that is the most simple, common and stable one for the moment. \nPS \nThere are research papers with proposals of different normalization in particular we also discuss it and even have some examples on Kaggle.  https://arxiv.org/abs/2208.05229\n\nOrganizers mentioned they plan to release so called \"raw\" counts - that mean the counts before normalization - I hope for that - but it will not change things drastically.",
      "votes": null
    },
    {
      "id": "1922862",
      "postDate": "09/01/2022 19:00:36",
      "content": "<p>Actually I think it is the best way to handle the raw data, see <a href=\"https://scanpy.readthedocs.io/en/stable/generated/scanpy.pp.normalize_total.html\" target=\"_blank\">https://scanpy.readthedocs.io/en/stable/generated/scanpy.pp.normalize_total.html</a><br>\nWe need to ensure the gene expression level of different cells is comparable and non negative, and because of the data scale, we use CPM (counts per million, as 10^6).</p>",
      "rawMarkdown": "Actually I think it is the best way to handle the raw data, see https://scanpy.readthedocs.io/en/stable/generated/scanpy.pp.normalize_total.html\nWe need to ensure the gene expression level of different cells is comparable and non negative, and because of the data scale, we use CPM (counts per million, as 10^6).",
      "votes": null
    },
    {
      "id": "1923069",
      "postDate": "09/02/2022 01:34:40",
      "content": "<p>Thanks! Since, understanding what \"library-size normalized\" in the data page was on my agenda, this post and notebook is super useful!</p>\n<p></p>\n<p>Added: I understand that Y is normalized to sum(exp(Y) - 1) = 1e6, but I am not sure how to normalize unnormalized Y'. I am wondering if the normalization is a multiplicative factor on N = exp(Y) - 1, not on exp(Y),</p>\n<ol>\n<li>N' = exp(Y') - 1 and Z = sum(N')</li>\n<li>Normalized number N = (1e6/Z) N'</li>\n<li>Normalized Y = log(1 + N)</li>\n</ol>\n<p>In this case, the transformation is nonlinear and is not a simple offset on Y.</p>\n<p><strong>Integerness/discreteness</strong></p>\n<p>Indeed, N = exp(Y) - 1 (for the normalized data we are provided), N are <em>proportional to</em> integers,</p>\n<pre><code>n = np.exp(y) - 1   # sample_size x 23418 for multiome target\nnn = np.unique(np.sort(n[0]))  # one random cell\nnp.diff(nn) =&gt; 200.52135, 401.04297, 601.5654,  802.08203, ... excluding machine precision differences\n</code></pre>\n<p>The unit of discreetness depends on each row (cell)</p>",
      "rawMarkdown": "Thanks! Since, understanding what \"library-size normalized\" in the data page was on my agenda, this post and notebook is super useful!\n\n~~One comment is that the evaluation metric, Pearson correlation coefficient, does not depend on the overall offset and factor; the mean is subtracted first, and both numerator and denominator are proportional to the overall factor. So the bottom line is that we don't have to worry about the normalization? (\"for each observation\" in the metric means each cell_id?)~~\n\nAdded: I understand that Y is normalized to sum(exp(Y) - 1) = 1e6, but I am not sure how to normalize unnormalized Y'. I am wondering if the normalization is a multiplicative factor on N = exp(Y) - 1, not on exp(Y),\n\n1. N' = exp(Y') - 1 and Z = sum(N')\n2. Normalized number N = (1e6/Z) N'\n3. Normalized Y = log(1 + N)\n\nIn this case, the transformation is nonlinear and is not a simple offset on Y.\n\n\n**Integerness/discreteness**\n\nIndeed, N = exp(Y) - 1 (for the normalized data we are provided), N are *proportional to* integers,\n\n```\nn = np.exp(y) - 1   # sample_size x 23418 for multiome target\nnn = np.unique(np.sort(n[0]))  # one random cell\nnp.diff(nn) => 200.52135, 401.04297, 601.5654,  802.08203, ... excluding machine precision differences\n```\n\nThe unit of discreetness depends on each row (cell)",
      "votes": null
    },
    {
      "id": "1928882",
      "postDate": "09/06/2022 17:11:35",
      "content": "<p>Nice! Thanks! The crazy idea can be  to optimize such that np.diff(nn) is constant,<br>\nI mean adding penalty for descrepancy,<br>\nhowever it seems it is not quite possible directly,<br>\nonly may be at the region of the correct solution.</p>\n<p>It is just a fantasy - most probably due to high level of noise it would not work… </p>",
      "rawMarkdown": "Nice! Thanks! The crazy idea can be  to optimize such that np.diff(nn) is constant,\nI mean adding penalty for descrepancy,\nhowever it seems it is not quite possible directly,\nonly may be at the region of the correct solution.\n\nIt is just a fantasy - most probably due to high level of noise it would not work...",
      "votes": null
    },
    {
      "id": "1929532",
      "postDate": "09/07/2022 06:33:04",
      "content": "<p>It might that SE question is related <br>\n<a href=\"https://stackoverflow.com/questions/445113/approximate-greatest-common-divisor\" target=\"_blank\">https://stackoverflow.com/questions/445113/approximate-greatest-common-divisor</a></p>",
      "rawMarkdown": "It might that SE question is related \nhttps://stackoverflow.com/questions/445113/approximate-greatest-common-divisor",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1921752,
      "author_name": "drpatrickchan",
      "author_url": "",
      "post_date": "09/01/2022 03:25:43",
      "content": "<p>Thanks for the info. Cliping to zero is not helpful so far. Can someone help explain why in standard biological processing, it is necessary to do above normalization? Cheers</p>",
      "votes": null,
      "replies": [
        {
          "id": 1922059,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "09/01/2022 08:41:39",
          "content": "<p>The short answer is : because of PCR - Polymerase chain reaction, and other technical problems to get that data</p>\n<p>The counts what we see are NOT the real counts in the cell, but counts after PCR reaction which is random - for one cell you might get much more counts than for another cell, despite real number of counts is the same - but just PCR was going on differently and may be some other randomness in  the process of getting these numbers. </p>\n<p>I am not biologist , if any biologist can comment - would be great. </p>\n<p>So to avoid these technical problem people use normalization to some constant. Orgs choose 1e6, others choose other number e.g. 1e4.</p>\n<p>Is it the perfect way to do such normalization ? <br>\nNo. <br>\nBut that is the most simple, common and stable one for the moment. <br>\nPS <br>\nThere are research papers with proposals of different normalization in particular we also discuss it and even have some examples on Kaggle.  <a href=\"https://arxiv.org/abs/2208.05229\" target=\"_blank\">https://arxiv.org/abs/2208.05229</a></p>\n<p>Organizers mentioned they plan to release so called \"raw\" counts - that mean the counts before normalization - I hope for that - but it will not change things drastically. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1922862,
      "author_name": "llttyy",
      "author_url": "",
      "post_date": "09/01/2022 19:00:36",
      "content": "<p>Actually I think it is the best way to handle the raw data, see <a href=\"https://scanpy.readthedocs.io/en/stable/generated/scanpy.pp.normalize_total.html\" target=\"_blank\">https://scanpy.readthedocs.io/en/stable/generated/scanpy.pp.normalize_total.html</a><br>\nWe need to ensure the gene expression level of different cells is comparable and non negative, and because of the data scale, we use CPM (counts per million, as 10^6).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1923069,
      "author_name": "junkoda",
      "author_url": "",
      "post_date": "09/02/2022 01:34:40",
      "content": "<p>Thanks! Since, understanding what \"library-size normalized\" in the data page was on my agenda, this post and notebook is super useful!</p>\n<p></p>\n<p>Added: I understand that Y is normalized to sum(exp(Y) - 1) = 1e6, but I am not sure how to normalize unnormalized Y'. I am wondering if the normalization is a multiplicative factor on N = exp(Y) - 1, not on exp(Y),</p>\n<ol>\n<li>N' = exp(Y') - 1 and Z = sum(N')</li>\n<li>Normalized number N = (1e6/Z) N'</li>\n<li>Normalized Y = log(1 + N)</li>\n</ol>\n<p>In this case, the transformation is nonlinear and is not a simple offset on Y.</p>\n<p><strong>Integerness/discreteness</strong></p>\n<p>Indeed, N = exp(Y) - 1 (for the normalized data we are provided), N are <em>proportional to</em> integers,</p>\n<pre><code>n = np.exp(y) - 1   # sample_size x 23418 for multiome target\nnn = np.unique(np.sort(n[0]))  # one random cell\nnp.diff(nn) =&gt; 200.52135, 401.04297, 601.5654,  802.08203, ... excluding machine precision differences\n</code></pre>\n<p>The unit of discreetness depends on each row (cell)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1928882,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "09/06/2022 17:11:35",
          "content": "<p>Nice! Thanks! The crazy idea can be  to optimize such that np.diff(nn) is constant,<br>\nI mean adding penalty for descrepancy,<br>\nhowever it seems it is not quite possible directly,<br>\nonly may be at the region of the correct solution.</p>\n<p>It is just a fantasy - most probably due to high level of noise it would not work… </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1929532,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "09/07/2022 06:33:04",
          "content": "<p>It might that SE question is related <br>\n<a href=\"https://stackoverflow.com/questions/445113/approximate-greatest-common-divisor\" target=\"_blank\">https://stackoverflow.com/questions/445113/approximate-greatest-common-divisor</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1920703": "Suggestion for Multiome prediction:\n\n1)\n\nDo normalization to achieve : (np.exp(Y)-1).sum(axis = 1) = 1e6 - because true targets satisfy that. (Well, \"satisfy\" - up to machine precision, but nevertheless. You can check - see below.) Why satisfy - that is standard biological processing - it is mentioned in data info tab, and well-known for everyone in the field.\n\nHow ? Can be different ways, but the most straighforward one:\n\n1) calculate predictions \"Y\" 2) calculate normalizer Z = sum(exp(Y)) 3) renorm: Y_i -> Y_i + (log((1e6+22050 )/Z))\n\n2)\n\nDo not forget that true answers are non-negative - that is why it might be useful to clip negative predictions to zero - even before step 1.\n\n3)\n\nIt might be exp(Y) are INTs - but that is not clear for the moment - depends what soft organizers used. Will investigate later. If it so: then use processing to achieve ints.\n\n=== \n\nSee notebook for checks:\n\nhttps://www.kaggle.com/code/alexandervc/postprocessing-tips-normalize-y-to-1e6-multiome\n\nAnalysis as usually on Kaggle.",
    "1921752": "Thanks for the info. Cliping to zero is not helpful so far. Can someone help explain why in standard biological processing, it is necessary to do above normalization? Cheers",
    "1922059": "The short answer is : because of PCR - Polymerase chain reaction, and other technical problems to get that data\n\nThe counts what we see are NOT the real counts in the cell, but counts after PCR reaction which is random - for one cell you might get much more counts than for another cell, despite real number of counts is the same - but just PCR was going on differently and may be some other randomness in  the process of getting these numbers. \n\nI am not biologist , if any biologist can comment - would be great. \n\nSo to avoid these technical problem people use normalization to some constant. Orgs choose 1e6, others choose other number e.g. 1e4.\n\nIs it the perfect way to do such normalization ? \nNo. \nBut that is the most simple, common and stable one for the moment. \nPS \nThere are research papers with proposals of different normalization in particular we also discuss it and even have some examples on Kaggle.  https://arxiv.org/abs/2208.05229\n\nOrganizers mentioned they plan to release so called \"raw\" counts - that mean the counts before normalization - I hope for that - but it will not change things drastically.",
    "1922862": "Actually I think it is the best way to handle the raw data, see https://scanpy.readthedocs.io/en/stable/generated/scanpy.pp.normalize_total.html\nWe need to ensure the gene expression level of different cells is comparable and non negative, and because of the data scale, we use CPM (counts per million, as 10^6).",
    "1923069": "Thanks! Since, understanding what \"library-size normalized\" in the data page was on my agenda, this post and notebook is super useful!\n\n~~One comment is that the evaluation metric, Pearson correlation coefficient, does not depend on the overall offset and factor; the mean is subtracted first, and both numerator and denominator are proportional to the overall factor. So the bottom line is that we don't have to worry about the normalization? (\"for each observation\" in the metric means each cell_id?)~~\n\nAdded: I understand that Y is normalized to sum(exp(Y) - 1) = 1e6, but I am not sure how to normalize unnormalized Y'. I am wondering if the normalization is a multiplicative factor on N = exp(Y) - 1, not on exp(Y),\n\n1. N' = exp(Y') - 1 and Z = sum(N')\n2. Normalized number N = (1e6/Z) N'\n3. Normalized Y = log(1 + N)\n\nIn this case, the transformation is nonlinear and is not a simple offset on Y.\n\n\n**Integerness/discreteness**\n\nIndeed, N = exp(Y) - 1 (for the normalized data we are provided), N are *proportional to* integers,\n\n```\nn = np.exp(y) - 1   # sample_size x 23418 for multiome target\nnn = np.unique(np.sort(n[0]))  # one random cell\nnp.diff(nn) => 200.52135, 401.04297, 601.5654,  802.08203, ... excluding machine precision differences\n```\n\nThe unit of discreetness depends on each row (cell)",
    "1928882": "Nice! Thanks! The crazy idea can be  to optimize such that np.diff(nn) is constant,\nI mean adding penalty for descrepancy,\nhowever it seems it is not quite possible directly,\nonly may be at the region of the correct solution.\n\nIt is just a fantasy - most probably due to high level of noise it would not work...",
    "1929532": "It might that SE question is related \nhttps://stackoverflow.com/questions/445113/approximate-greatest-common-divisor"
  },
  "source": "meta"
}