{
  "id": 360384,
  "title": "How to reproduce competition normalized data from raw count data ?",
  "url": "/competitions/open-problems-multimodal/discussion/360384",
  "author_name": "senkin13",
  "post_date": "2022-10-16T13:02:34.162000",
  "votes": 14,
  "comment_count": 8,
  "views": 0,
  "content": "<p>As we have raw count data now, could we reproduce provided competition normalized data(library-size normalized and log1p transformed, dsb normalized ) from raw count data ?<br>\nI tried to use scanpy to transform raw count to library-size normalized and log1p transformed data,but very different from  provided data.And I can't find easy way to  transform to dsb normalized,any suggestions?</p>",
  "messages": [
    {
      "id": 1990310,
      "postDate": "2022-10-16T13:02:34.163Z",
      "content": "<p>As we have raw count data now, could we reproduce provided competition normalized data(library-size normalized and log1p transformed, dsb normalized ) from raw count data ?<br>\nI tried to use scanpy to transform raw count to library-size normalized and log1p transformed data,but very different from  provided data.And I can't find easy way to  transform to dsb normalized,any suggestions?</p>",
      "rawMarkdown": "As we have raw count data now, could we reproduce provided competition normalized data(library-size normalized and log1p transformed, dsb normalized ) from raw count data ?\nI tried to use scanpy to transform raw count to library-size normalized and log1p transformed data,but very different from  provided data.And I can't find easy way to  transform to dsb normalized,any suggestions?",
      "votes": 14
    },
    {
      "id": 1990317,
      "postDate": "2022-10-16T13:07:57.590Z",
      "content": "<p>for RNA-seq part of data <br>\nit should be log(1+ v/sum(v)*1e6), modula details like adding 35 new genes.<br>\nAt least normalized data (initial input) satisfy:<br>\nsum( exp(v)-1 ) = 1e6. See checks here: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349132\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349132</a></p>\n<p>For ATAC-seq, and CD-Proteins - I do not know. </p>\n<p>PS</p>\n<p>It seems NUMBER  of  raw counts for CITE seq is differnt from initial input - what can it mean ?  (just different number of samples by about 460)</p>",
      "rawMarkdown": "for RNA-seq part of data \nit should be log(1+ v/sum(v)*1e6), modula details like adding 35 new genes.\nAt least normalized data (initial input) satisfy:\nsum( exp(v)-1 ) = 1e6. See checks here: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349132\n\nFor ATAC-seq, and CD-Proteins - I do not know. \n\nPS\n\nIt seems NUMBER  of  raw counts for CITE seq is differnt from initial input - what can it mean ?  (just different number of samples by about 460)\n\n ",
      "votes": 2,
      "replies": [
        {
          "id": 1991434,
          "postDate": "2022-10-17T06:45:31.660Z",
          "content": "<p>Thanks your insight,it's useful.I hope organizer can tell us the method of transformation in case of some teams know and others don't know.</p>",
          "rawMarkdown": "Thanks your insight,it's useful.I hope organizer can tell us the method of transformation in case of some teams know and others don't know.",
          "votes": 1
        },
        {
          "id": 1991461,
          "postDate": "2022-10-17T07:04:26.437Z",
          "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> <br>\nCan you please share preprocessing algo and code, please ?</p>\n<p>PS<br>\nAlso there is problem with number of test samples for CITEseq. It is different from initial one. How to match? </p>",
          "rawMarkdown": "@danielburkhardt \nCan you please share preprocessing algo and code, please ?\n\nPS\nAlso there is problem with number of test samples for CITEseq. It is different from initial one. How to match? "
        },
        {
          "id": 1991964,
          "postDate": "2022-10-17T13:24:07.910Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1992247,
          "postDate": "2022-10-17T15:45:50.333Z",
          "content": "<p>Hi! A few items here:</p>\n<p><strong>Discrepancies in cell counts or total cells between raw and train / test data</strong><br>\nSome of the filtering steps we use are stochastic, and unfortunately we did not save the raw output for the competition data before preparing the raw counts data. As a result, there are differences in terms of the cells present or the total counts. They should be minor.</p>\n<p><strong>Processing</strong><br>\nWe will release our cell filtering, preprocessing, normalization, and transformation notebooks after competition close, but the main code you want to go from raw to normalized counts is:</p>\n<p>ATAC</p>\n<pre><code>from muon import atac as ac\n\nac.pp.tfidf(atac, scale_factor=1e4)\nac.tl.lsi(atac)\n</code></pre>\n<p>RNA</p>\n<pre><code>import scanpy as sc\nsc.pp.normalize_per_cell(rna, counts_per_cell_after = 1e6)\nsc.pp.log1p(rna\n</code></pre>\n<p>PROT</p>\n<pre><code>from muon import prot as pt\nDSB_LOWER=1.5 #These are set manually for each sample\nDSB_UPPER=2.8\n\nmdata['prot'].layers['counts'] = mdata['prot'].X  # keep count data\npt.pp.dsb(mdata, \n          data_raw = mdata_raw, \n          empty_counts_range=(DSB_LOWER, DSB_UPPER), \n          isotype_controls=isotypes, \n          random_state=1)\n\nprot = mdata['prot']\n</code></pre>",
          "rawMarkdown": "Hi! A few items here:\n\n**Discrepancies in cell counts or total cells between raw and train / test data**\nSome of the filtering steps we use are stochastic, and unfortunately we did not save the raw output for the competition data before preparing the raw counts data. As a result, there are differences in terms of the cells present or the total counts. They should be minor.\n\n **Processing**\nWe will release our cell filtering, preprocessing, normalization, and transformation notebooks after competition close, but the main code you want to go from raw to normalized counts is:\n\nATAC\n```python\nfrom muon import atac as ac\n\nac.pp.tfidf(atac, scale_factor=1e4)\nac.tl.lsi(atac)\n```\n\nRNA\n```python\nimport scanpy as sc\nsc.pp.normalize_per_cell(rna, counts_per_cell_after = 1e6)\nsc.pp.log1p(rna\n```\n\nPROT\n```python\nfrom muon import prot as pt\nDSB_LOWER=1.5 #These are set manually for each sample\nDSB_UPPER=2.8\n\nmdata['prot'].layers['counts'] = mdata['prot'].X  # keep count data\npt.pp.dsb(mdata, \n          data_raw = mdata_raw, \n          empty_counts_range=(DSB_LOWER, DSB_UPPER), \n          isotype_controls=isotypes, \n          random_state=1)\n\nprot = mdata['prot']\n```\n",
          "votes": 2
        },
        {
          "id": 1992321,
          "postDate": "2022-10-17T16:25:06.303Z",
          "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> thanks for your information, I am not familiar with muon, could you share more details about PROT, how to define mdata,mdata_raw,isotypes?</p>",
          "rawMarkdown": "@danielburkhardt thanks for your information, I am not familiar with muon, could you share more details about PROT, how to define mdata,mdata_raw,isotypes?",
          "votes": 1
        },
        {
          "id": 1992338,
          "postDate": "2022-10-17T16:38:36.263Z",
          "content": "<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> happy to help! Can you please help me understand more how you want to use this? </p>\n<p>The DSB normalization uses both the raw filtered (released) and raw unfiltered (not released) data. Because the unfiltered data is very large, you won't be able to recreate <code>mdata_raw</code> with the files we released. I'm worried we won't have time to prepare that data for release before the competition closes. The challenge is there are many different stages of data processing, and as much as I wish it were easy, it takes some time to prep the intermediate data for public release.</p>\n<p>However, if there's another way we can help you here, I would like to do that!</p>",
          "rawMarkdown": "@senkin13 happy to help! Can you please help me understand more how you want to use this? \n\nThe DSB normalization uses both the raw filtered (released) and raw unfiltered (not released) data. Because the unfiltered data is very large, you won't be able to recreate `mdata_raw` with the files we released. I'm worried we won't have time to prepare that data for release before the competition closes. The challenge is there are many different stages of data processing, and as much as I wish it were easy, it takes some time to prep the intermediate data for public release.\n\nHowever, if there's another way we can help you here, I would like to do that!",
          "votes": 1
        },
        {
          "id": 1992375,
          "postDate": "2022-10-17T16:56:27.060Z",
          "content": "<p>Totally understood, I think you have provided enough information for me, thanks!</p>",
          "rawMarkdown": "Totally understood, I think you have provided enough information for me, thanks!",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1990317,
      "author_name": "Alexander Chervov",
      "author_url": "",
      "post_date": "2022-10-16T13:07:57.590000",
      "content": "<p>for RNA-seq part of data <br>\nit should be log(1+ v/sum(v)*1e6), modula details like adding 35 new genes.<br>\nAt least normalized data (initial input) satisfy:<br>\nsum( exp(v)-1 ) = 1e6. See checks here: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349132\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349132</a></p>\n<p>For ATAC-seq, and CD-Proteins - I do not know. </p>\n<p>PS</p>\n<p>It seems NUMBER  of  raw counts for CITE seq is differnt from initial input - what can it mean ?  (just different number of samples by about 460)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1991434,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2022-10-17T06:45:31.660000",
          "content": "<p>Thanks your insight,it's useful.I hope organizer can tell us the method of transformation in case of some teams know and others don't know.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1991461,
          "author_name": "Alexander Chervov",
          "author_url": "",
          "post_date": "2022-10-17T07:04:26.437000",
          "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> <br>\nCan you please share preprocessing algo and code, please ?</p>\n<p>PS<br>\nAlso there is problem with number of test samples for CITEseq. It is different from initial one. How to match? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1991964,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-10-17T13:24:07.910000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1992247,
          "author_name": "Daniel Burkhardt",
          "author_url": "",
          "post_date": "2022-10-17T15:45:50.333000",
          "content": "<p>Hi! A few items here:</p>\n<p><strong>Discrepancies in cell counts or total cells between raw and train / test data</strong><br>\nSome of the filtering steps we use are stochastic, and unfortunately we did not save the raw output for the competition data before preparing the raw counts data. As a result, there are differences in terms of the cells present or the total counts. They should be minor.</p>\n<p><strong>Processing</strong><br>\nWe will release our cell filtering, preprocessing, normalization, and transformation notebooks after competition close, but the main code you want to go from raw to normalized counts is:</p>\n<p>ATAC</p>\n<pre><code>from muon import atac as ac\n\nac.pp.tfidf(atac, scale_factor=1e4)\nac.tl.lsi(atac)\n</code></pre>\n<p>RNA</p>\n<pre><code>import scanpy as sc\nsc.pp.normalize_per_cell(rna, counts_per_cell_after = 1e6)\nsc.pp.log1p(rna\n</code></pre>\n<p>PROT</p>\n<pre><code>from muon import prot as pt\nDSB_LOWER=1.5 #These are set manually for each sample\nDSB_UPPER=2.8\n\nmdata['prot'].layers['counts'] = mdata['prot'].X  # keep count data\npt.pp.dsb(mdata, \n          data_raw = mdata_raw, \n          empty_counts_range=(DSB_LOWER, DSB_UPPER), \n          isotype_controls=isotypes, \n          random_state=1)\n\nprot = mdata['prot']\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1992321,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2022-10-17T16:25:06.303000",
          "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> thanks for your information, I am not familiar with muon, could you share more details about PROT, how to define mdata,mdata_raw,isotypes?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1992338,
          "author_name": "Daniel Burkhardt",
          "author_url": "",
          "post_date": "2022-10-17T16:38:36.263000",
          "content": "<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> happy to help! Can you please help me understand more how you want to use this? </p>\n<p>The DSB normalization uses both the raw filtered (released) and raw unfiltered (not released) data. Because the unfiltered data is very large, you won't be able to recreate <code>mdata_raw</code> with the files we released. I'm worried we won't have time to prepare that data for release before the competition closes. The challenge is there are many different stages of data processing, and as much as I wish it were easy, it takes some time to prep the intermediate data for public release.</p>\n<p>However, if there's another way we can help you here, I would like to do that!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1992375,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2022-10-17T16:56:27.060000",
          "content": "<p>Totally understood, I think you have provided enough information for me, thanks!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1990310": "As we have raw count data now, could we reproduce provided competition normalized data(library-size normalized and log1p transformed, dsb normalized ) from raw count data ?\nI tried to use scanpy to transform raw count to library-size normalized and log1p transformed data,but very different from  provided data.And I can't find easy way to  transform to dsb normalized,any suggestions?",
    "1990317": "for RNA-seq part of data \nit should be log(1+ v/sum(v)*1e6), modula details like adding 35 new genes.\nAt least normalized data (initial input) satisfy:\nsum( exp(v)-1 ) = 1e6. See checks here: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349132\n\nFor ATAC-seq, and CD-Proteins - I do not know. \n\nPS\n\nIt seems NUMBER  of  raw counts for CITE seq is differnt from initial input - what can it mean ?  (just different number of samples by about 460)\n\n "
  }
}