{
  "id": 346341,
  "title": "Will the unnormalized values be released?",
  "url": "/competitions/open-problems-multimodal/discussion/346341",
  "author_name": "",
  "post_date": "2022-08-19T01:16:24.815205Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi there,</p>\n<p>I think its common practice for the counts in single cell field to be modelled as a negative binomial distribution. Will the unnormalized values be released? Furthermore, I am not sure that log(TF)-log(IDF) is the best method to normalize ATAC data. Here is a benchmark from snapatac2, which shows that log(TF)-log(IDF) often performs poorly in ATAC data. </p>\n<p><a href=\"https://kzhang.org/SnapATAC2/\" target=\"_blank\">https://kzhang.org/SnapATAC2/</a></p>",
  "messages": [
    {
      "id": "1905336",
      "postDate": "08/19/2022 01:16:24",
      "content": "<p>Hi there,</p>\n<p>I think its common practice for the counts in single cell field to be modelled as a negative binomial distribution. Will the unnormalized values be released? Furthermore, I am not sure that log(TF)-log(IDF) is the best method to normalize ATAC data. Here is a benchmark from snapatac2, which shows that log(TF)-log(IDF) often performs poorly in ATAC data. </p>\n<p><a href=\"https://kzhang.org/SnapATAC2/\" target=\"_blank\">https://kzhang.org/SnapATAC2/</a></p>",
      "rawMarkdown": "Hi there,\n\nI think its common practice for the counts in single cell field to be modelled as a negative binomial distribution. Will the unnormalized values be released? Furthermore, I am not sure that log(TF)-log(IDF) is the best method to normalize ATAC data. Here is a benchmark from snapatac2, which shows that log(TF)-log(IDF) often performs poorly in ATAC data. \n\nhttps://kzhang.org/SnapATAC2/",
      "votes": null
    },
    {
      "id": "1906512",
      "postDate": "08/20/2022 00:12:18",
      "content": "<p>Another point on the evaluation. Why not just do pearson correlation on the original counts data, instead of the normalized data? It makes much more sense to me that way, since normalization is still a kinda of open problem. </p>",
      "rawMarkdown": "Another point on the evaluation. Why not just do pearson correlation on the original counts data, instead of the normalized data? It makes much more sense to me that way, since normalization is still a kinda of open problem.",
      "votes": null
    },
    {
      "id": "1918355",
      "postDate": "08/29/2022 14:22:25",
      "content": "<p>Hi Chengwei,</p>\n<p>Thanks a lot for your question and apologies for replying bit late. I was involved especially in the ATAC part of the competition data preparation and think I have a few answers to your question.</p>\n<p>Regarding <strong>RNA data normalization</strong>: Yes, you're very right that RNA-seq data is often modeled using a negative binomial distribution such as in sctransform or when it comes to testing for differential gene expression (e.g. in DEseq2). However, benchmarks evaluating normalization methods are rather conflicting (e.g. <a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-02136-7\" target=\"_blank\">Germain et al. </a> versus <a href=\"https://www.biorxiv.org/content/10.1101/2022.05.06.490859v1.full\" target=\"_blank\">Booeshaghi et al.</a> making it still an open problem as you already mentioned in your comment. However, since we're using Pearson correlation as an evaluation metric, which is sensitive to outliers, we did want to make use of one of the normalization methods also for variance stabilization. This way, for example very high outlier gene expression values don't have such a strong effect on our metric anymore.</p>\n<p>Regarding the <strong>TF-IDF transformation</strong>: The benchmark from snapatac2 is indeed very interesting and I'm following it very closely at the moment because it might be a very interesting preprocessing strategy for scATAC data in general. The issue for our task here is that in snapatac2 (and also in the previous version, which has shown to perform very similar to the TF-IDF approaches in an independent benchmark from <a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-019-1854-5\" target=\"_blank\">Chen et al.</a> ), the count matrix on windows is binarizes and then transformed into a cell-by-cell similarity matrix as part of their spectral embedding (see <a href=\"https://kzhang.org/SnapATAC2/algorithms/dimension_reduction.html\" target=\"_blank\">snapatac2 docs algorithm</a>). This cell-by-cell matrix is not very useful for our task anymore, since the goal is to predict one modality from the other. I do agree, that it would be interesting to see, whether starting from the raw scATAC counts competitors would be able to even better predict RNA expression values, since as you said the TF-IDF transformation is already a preprocessing method, which has shown to perform well, but as indicated in the benchmark, might not capture all biological signal ideally. I will double check with the other organizers, whether we will provide this data as an optional input  and get back to you soon!</p>\n<p>Hope that already helped and best regards,<br>\nChris</p>",
      "rawMarkdown": "Hi Chengwei,\n\nThanks a lot for your question and apologies for replying bit late. I was involved especially in the ATAC part of the competition data preparation and think I have a few answers to your question.\n\nRegarding **RNA data normalization**: Yes, you're very right that RNA-seq data is often modeled using a negative binomial distribution such as in sctransform or when it comes to testing for differential gene expression (e.g. in DEseq2). However, benchmarks evaluating normalization methods are rather conflicting (e.g. [Germain et al. ](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-02136-7) versus [Booeshaghi et al.](https://www.biorxiv.org/content/10.1101/2022.05.06.490859v1.full) making it still an open problem as you already mentioned in your comment. However, since we're using Pearson correlation as an evaluation metric, which is sensitive to outliers, we did want to make use of one of the normalization methods also for variance stabilization. This way, for example very high outlier gene expression values don't have such a strong effect on our metric anymore.\n\nRegarding the **TF-IDF transformation**: The benchmark from snapatac2 is indeed very interesting and I'm following it very closely at the moment because it might be a very interesting preprocessing strategy for scATAC data in general. The issue for our task here is that in snapatac2 (and also in the previous version, which has shown to perform very similar to the TF-IDF approaches in an independent benchmark from [Chen et al.](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-019-1854-5) ), the count matrix on windows is binarizes and then transformed into a cell-by-cell similarity matrix as part of their spectral embedding (see [snapatac2 docs algorithm](https://kzhang.org/SnapATAC2/algorithms/dimension_reduction.html)). This cell-by-cell matrix is not very useful for our task anymore, since the goal is to predict one modality from the other. I do agree, that it would be interesting to see, whether starting from the raw scATAC counts competitors would be able to even better predict RNA expression values, since as you said the TF-IDF transformation is already a preprocessing method, which has shown to perform well, but as indicated in the benchmark, might not capture all biological signal ideally. I will double check with the other organizers, whether we will provide this data as an optional input  and get back to you soon!\n\nHope that already helped and best regards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1906512,
      "author_name": "kjwandzcw",
      "author_url": "",
      "post_date": "08/20/2022 00:12:18",
      "content": "<p>Another point on the evaluation. Why not just do pearson correlation on the original counts data, instead of the normalized data? It makes much more sense to me that way, since normalization is still a kinda of open problem. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1918355,
      "author_name": "lancerunstats",
      "author_url": "",
      "post_date": "08/29/2022 14:22:25",
      "content": "<p>Hi Chengwei,</p>\n<p>Thanks a lot for your question and apologies for replying bit late. I was involved especially in the ATAC part of the competition data preparation and think I have a few answers to your question.</p>\n<p>Regarding <strong>RNA data normalization</strong>: Yes, you're very right that RNA-seq data is often modeled using a negative binomial distribution such as in sctransform or when it comes to testing for differential gene expression (e.g. in DEseq2). However, benchmarks evaluating normalization methods are rather conflicting (e.g. <a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-02136-7\" target=\"_blank\">Germain et al. </a> versus <a href=\"https://www.biorxiv.org/content/10.1101/2022.05.06.490859v1.full\" target=\"_blank\">Booeshaghi et al.</a> making it still an open problem as you already mentioned in your comment. However, since we're using Pearson correlation as an evaluation metric, which is sensitive to outliers, we did want to make use of one of the normalization methods also for variance stabilization. This way, for example very high outlier gene expression values don't have such a strong effect on our metric anymore.</p>\n<p>Regarding the <strong>TF-IDF transformation</strong>: The benchmark from snapatac2 is indeed very interesting and I'm following it very closely at the moment because it might be a very interesting preprocessing strategy for scATAC data in general. The issue for our task here is that in snapatac2 (and also in the previous version, which has shown to perform very similar to the TF-IDF approaches in an independent benchmark from <a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-019-1854-5\" target=\"_blank\">Chen et al.</a> ), the count matrix on windows is binarizes and then transformed into a cell-by-cell similarity matrix as part of their spectral embedding (see <a href=\"https://kzhang.org/SnapATAC2/algorithms/dimension_reduction.html\" target=\"_blank\">snapatac2 docs algorithm</a>). This cell-by-cell matrix is not very useful for our task anymore, since the goal is to predict one modality from the other. I do agree, that it would be interesting to see, whether starting from the raw scATAC counts competitors would be able to even better predict RNA expression values, since as you said the TF-IDF transformation is already a preprocessing method, which has shown to perform well, but as indicated in the benchmark, might not capture all biological signal ideally. I will double check with the other organizers, whether we will provide this data as an optional input  and get back to you soon!</p>\n<p>Hope that already helped and best regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1905336": "Hi there,\n\nI think its common practice for the counts in single cell field to be modelled as a negative binomial distribution. Will the unnormalized values be released? Furthermore, I am not sure that log(TF)-log(IDF) is the best method to normalize ATAC data. Here is a benchmark from snapatac2, which shows that log(TF)-log(IDF) often performs poorly in ATAC data. \n\nhttps://kzhang.org/SnapATAC2/",
    "1906512": "Another point on the evaluation. Why not just do pearson correlation on the original counts data, instead of the normalized data? It makes much more sense to me that way, since normalization is still a kinda of open problem.",
    "1918355": "Hi Chengwei,\n\nThanks a lot for your question and apologies for replying bit late. I was involved especially in the ATAC part of the competition data preparation and think I have a few answers to your question.\n\nRegarding **RNA data normalization**: Yes, you're very right that RNA-seq data is often modeled using a negative binomial distribution such as in sctransform or when it comes to testing for differential gene expression (e.g. in DEseq2). However, benchmarks evaluating normalization methods are rather conflicting (e.g. [Germain et al. ](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-02136-7) versus [Booeshaghi et al.](https://www.biorxiv.org/content/10.1101/2022.05.06.490859v1.full) making it still an open problem as you already mentioned in your comment. However, since we're using Pearson correlation as an evaluation metric, which is sensitive to outliers, we did want to make use of one of the normalization methods also for variance stabilization. This way, for example very high outlier gene expression values don't have such a strong effect on our metric anymore.\n\nRegarding the **TF-IDF transformation**: The benchmark from snapatac2 is indeed very interesting and I'm following it very closely at the moment because it might be a very interesting preprocessing strategy for scATAC data in general. The issue for our task here is that in snapatac2 (and also in the previous version, which has shown to perform very similar to the TF-IDF approaches in an independent benchmark from [Chen et al.](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-019-1854-5) ), the count matrix on windows is binarizes and then transformed into a cell-by-cell similarity matrix as part of their spectral embedding (see [snapatac2 docs algorithm](https://kzhang.org/SnapATAC2/algorithms/dimension_reduction.html)). This cell-by-cell matrix is not very useful for our task anymore, since the goal is to predict one modality from the other. I do agree, that it would be interesting to see, whether starting from the raw scATAC counts competitors would be able to even better predict RNA expression values, since as you said the TF-IDF transformation is already a preprocessing method, which has shown to perform well, but as indicated in the benchmark, might not capture all biological signal ideally. I will double check with the other organizers, whether we will provide this data as an optional input  and get back to you soon!\n\nHope that already helped and best regards,\nChris"
  },
  "source": "meta"
}