{
  "id": 350856,
  "title": "BioQuestion 01: MAGIC features ",
  "url": "/competitions/open-problems-multimodal/discussion/350856",
  "author_name": "",
  "post_date": "2022-09-07T11:08:45.714953500Z",
  "votes": 21,
  "comment_count": 5,
  "views": 0,
  "content": "<p><strong>Disclaimer.</strong> That might improve score, might not, but any outcome would be of interest for research community. <br>\nSo everyone is welcome to collaborate  - hopefully produce a paper - see <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293\" target=\"_blank\">Discussion1</a>, <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293\" target=\"_blank\">Discussion2</a>, <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348661\" target=\"_blank\">Discussion3\n</a></p>\n<p><strong>Very briefly:</strong> For CITEseq part of task - try to use features stored in the dataset<br>\n<a href=\"https://www.kaggle.com/datasets/alexandervc/multimodal-singlecell-integration-related-data-01\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/multimodal-singlecell-integration-related-data-01</a> <br>\n(Currently there is only \"cite_seq_inputs_denoising_pooling_n_neighbours20.h5ad\" , but later other related will appear - it takes some time). <br>\n<strong>Will it improve the score ?</strong> If yes, how much and for what downstream methods ?</p>\n<p><strong>What is going on (briefly):</strong>  The features of CITEseq task ( and also targets of Multiome) are the so-called single cell RNA sequencing count matrices, which are widely studied last years. It is known that they suffer from certain specific type of noise - technical (not biological) zero values (called \"dropouts\" ). And there are plenty methods of \"denoising\"  were developed.  One of them called <strong>\"MAGIC\"</strong> ( <a href=\"https://github.com/KrishnaswamyLab/MAGIC\" target=\"_blank\">https://github.com/KrishnaswamyLab/MAGIC</a> ), another is denoising autoencoder - both implemented in SCANPY  <a href=\"https://scanpy.readthedocs.io/en/stable/generated/scanpy.external.pp.dca.html\" target=\"_blank\">https://scanpy.readthedocs.io/en/stable/generated/scanpy.external.pp.dca.html</a> </p>\n<p>There are other methods. For example we use another - more simple and fast method, sometimes called \"pooling\".<br>\n(See the script: <a href=\"https://www.kaggle.com/alexandervc/denoising-pooling-for-adata-citeseq-features\" target=\"_blank\">https://www.kaggle.com/alexandervc/denoising-pooling-for-adata-citeseq-features</a> ) </p>\n<p><strong>Research question:</strong> benchmark different denoising methods - will they give score improvement or not, if yes how much ? </p>\n<p><strong>Why it is of interest:</strong>   Sometimes \"denoising\" is helpful for analysis of the scRNA-seq data, but sometimes it is HARMFUL(!) - that are non-linear algorithms which may introduce too strong distortion to the data. The open research problems is to produce some guidelines: when \"denoising\" algorithm are good, when bad; how to choose/tune the params, how much improvement we may expect to achieve, how it depend on the task and data (cell types, technology, number of cells etc); and other questions. Typically analysis of the scRNA-seq is UNsupervised, so the estimation of the outcomes of the algorithms is not straightforward. The Kaggle competion is quite exceptional since it has SUPERvised task - so it allows to judge the algorithm performance based on clear supervised scores. So it is an interesting research question whether denoising helps to improve scores or not and if yes - how much.</p>\n<p><strong>So you are welcome to contribute to research and might be it will improve your score !</strong></p>\n<p>PS</p>\n<p><strong>EXAMPLES</strong></p>\n<p><strong>MAGIC gif example.</strong> Click the page <a href=\"https://github.com/KrishnaswamyLab/MAGIC\" target=\"_blank\">https://github.com/KrishnaswamyLab/MAGIC</a> and look on short gif-video - it gives a flavor what methods do. The zeroes which are not biological, but technical dissapper and biologically expexted correlations are restored.</p>\n<p><strong>Examples of the current script for the cell cycle analysis:</strong> on the frontpage image for the Kaggle dataset <a href=\"https://www.kaggle.com/datasets/alexandervc/scrnaseq-data-for-ribamolina-paper-deepcycle\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/scrnaseq-data-for-ribamolina-paper-deepcycle</a> (and scripts there in) you can see that without denoising the expected \"cyclic\" structure is NOT seen, but it appears after the denoising. That is good example. The \"bad\" example is here: <a href=\"https://www.kaggle.com/datasets/alexandervc/scrnaseq-bone-marrow-collection-of-datasets\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/scrnaseq-bone-marrow-collection-of-datasets</a> the frontpage image demonstrates too much distortion introduced by the denoising from the present script - expected \"cycle/triangle\" is lost - too much correlations are introduced.</p>\n<p><strong>Method of denoising used here (pooling).</strong> There are several methods of \"denoising\". The current script provides implementation for one of the most simple, but rather fast one. Roughly speaking it substitutes the features by the average of several neighboring features. The number of neighbors e.g. \"n_neighbours = 20\" is the key parameter of the strength of denoising. However, it is ALMOST exactly that much simple, but not really: the subtle point - in what space to calculate the distances for KNN and in what space to take averages. Here we calculate distances for KNN in the space which is 1) normalized+log 2) PCA reduced, but the averages are taken in the original space of the \"raw\"-counts , and only then we take standard normalization+log. It might it is not the optimal way and might be modifications would work better in some situations. Everyone is welcome to collaborate on improvements. Technicalities. There is also dropping out some less expressed genes - there is param - how many to keep n_top_genes_to_keep = 10000, and there is list of genes which we want to keep mandatory. There are some other params which might lead to dropping out some cells and genes, but current values - are such that it does not happen: threshold_pct_counts_mt = 15 min_count = 1_000 max_count = 10_000_000</p>\n<p>PS </p>\n<p>And sorry the MAGIC features are not yet ready, but we are thinking on that ))))) </p>\n<p>PSPS<br>\nMore results when \"denoising\" improves the solution of the  particular task -  cell cycle analysis - see in our paper: <a href=\"https://arxiv.org/abs/2208.05229\" target=\"_blank\">https://arxiv.org/abs/2208.05229</a> . All the work have been done on the Kaggle platform - look for datasets <a href=\"https://www.kaggle.com/alexandervc/datasets\" target=\"_blank\">https://www.kaggle.com/alexandervc/datasets</a>, <a href=\"https://www.kaggle.com/andreizinovyev/datasets\" target=\"_blank\">https://www.kaggle.com/andreizinovyev/datasets</a>  and scripts there in. <br>\nNote: that are NOT multimodal datasets, so NOT of direct use for the present competition. </p>",
  "messages": [
    {
      "id": "1929789",
      "postDate": "09/07/2022 11:08:45",
      "content": "<p><strong>Disclaimer.</strong> That might improve score, might not, but any outcome would be of interest for research community. <br>\nSo everyone is welcome to collaborate  - hopefully produce a paper - see <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293\" target=\"_blank\">Discussion1</a>, <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293\" target=\"_blank\">Discussion2</a>, <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348661\" target=\"_blank\">Discussion3\n</a></p>\n<p><strong>Very briefly:</strong> For CITEseq part of task - try to use features stored in the dataset<br>\n<a href=\"https://www.kaggle.com/datasets/alexandervc/multimodal-singlecell-integration-related-data-01\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/multimodal-singlecell-integration-related-data-01</a> <br>\n(Currently there is only \"cite_seq_inputs_denoising_pooling_n_neighbours20.h5ad\" , but later other related will appear - it takes some time). <br>\n<strong>Will it improve the score ?</strong> If yes, how much and for what downstream methods ?</p>\n<p><strong>What is going on (briefly):</strong>  The features of CITEseq task ( and also targets of Multiome) are the so-called single cell RNA sequencing count matrices, which are widely studied last years. It is known that they suffer from certain specific type of noise - technical (not biological) zero values (called \"dropouts\" ). And there are plenty methods of \"denoising\"  were developed.  One of them called <strong>\"MAGIC\"</strong> ( <a href=\"https://github.com/KrishnaswamyLab/MAGIC\" target=\"_blank\">https://github.com/KrishnaswamyLab/MAGIC</a> ), another is denoising autoencoder - both implemented in SCANPY  <a href=\"https://scanpy.readthedocs.io/en/stable/generated/scanpy.external.pp.dca.html\" target=\"_blank\">https://scanpy.readthedocs.io/en/stable/generated/scanpy.external.pp.dca.html</a> </p>\n<p>There are other methods. For example we use another - more simple and fast method, sometimes called \"pooling\".<br>\n(See the script: <a href=\"https://www.kaggle.com/alexandervc/denoising-pooling-for-adata-citeseq-features\" target=\"_blank\">https://www.kaggle.com/alexandervc/denoising-pooling-for-adata-citeseq-features</a> ) </p>\n<p><strong>Research question:</strong> benchmark different denoising methods - will they give score improvement or not, if yes how much ? </p>\n<p><strong>Why it is of interest:</strong>   Sometimes \"denoising\" is helpful for analysis of the scRNA-seq data, but sometimes it is HARMFUL(!) - that are non-linear algorithms which may introduce too strong distortion to the data. The open research problems is to produce some guidelines: when \"denoising\" algorithm are good, when bad; how to choose/tune the params, how much improvement we may expect to achieve, how it depend on the task and data (cell types, technology, number of cells etc); and other questions. Typically analysis of the scRNA-seq is UNsupervised, so the estimation of the outcomes of the algorithms is not straightforward. The Kaggle competion is quite exceptional since it has SUPERvised task - so it allows to judge the algorithm performance based on clear supervised scores. So it is an interesting research question whether denoising helps to improve scores or not and if yes - how much.</p>\n<p><strong>So you are welcome to contribute to research and might be it will improve your score !</strong></p>\n<p>PS</p>\n<p><strong>EXAMPLES</strong></p>\n<p><strong>MAGIC gif example.</strong> Click the page <a href=\"https://github.com/KrishnaswamyLab/MAGIC\" target=\"_blank\">https://github.com/KrishnaswamyLab/MAGIC</a> and look on short gif-video - it gives a flavor what methods do. The zeroes which are not biological, but technical dissapper and biologically expexted correlations are restored.</p>\n<p><strong>Examples of the current script for the cell cycle analysis:</strong> on the frontpage image for the Kaggle dataset <a href=\"https://www.kaggle.com/datasets/alexandervc/scrnaseq-data-for-ribamolina-paper-deepcycle\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/scrnaseq-data-for-ribamolina-paper-deepcycle</a> (and scripts there in) you can see that without denoising the expected \"cyclic\" structure is NOT seen, but it appears after the denoising. That is good example. The \"bad\" example is here: <a href=\"https://www.kaggle.com/datasets/alexandervc/scrnaseq-bone-marrow-collection-of-datasets\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/scrnaseq-bone-marrow-collection-of-datasets</a> the frontpage image demonstrates too much distortion introduced by the denoising from the present script - expected \"cycle/triangle\" is lost - too much correlations are introduced.</p>\n<p><strong>Method of denoising used here (pooling).</strong> There are several methods of \"denoising\". The current script provides implementation for one of the most simple, but rather fast one. Roughly speaking it substitutes the features by the average of several neighboring features. The number of neighbors e.g. \"n_neighbours = 20\" is the key parameter of the strength of denoising. However, it is ALMOST exactly that much simple, but not really: the subtle point - in what space to calculate the distances for KNN and in what space to take averages. Here we calculate distances for KNN in the space which is 1) normalized+log 2) PCA reduced, but the averages are taken in the original space of the \"raw\"-counts , and only then we take standard normalization+log. It might it is not the optimal way and might be modifications would work better in some situations. Everyone is welcome to collaborate on improvements. Technicalities. There is also dropping out some less expressed genes - there is param - how many to keep n_top_genes_to_keep = 10000, and there is list of genes which we want to keep mandatory. There are some other params which might lead to dropping out some cells and genes, but current values - are such that it does not happen: threshold_pct_counts_mt = 15 min_count = 1_000 max_count = 10_000_000</p>\n<p>PS </p>\n<p>And sorry the MAGIC features are not yet ready, but we are thinking on that ))))) </p>\n<p>PSPS<br>\nMore results when \"denoising\" improves the solution of the  particular task -  cell cycle analysis - see in our paper: <a href=\"https://arxiv.org/abs/2208.05229\" target=\"_blank\">https://arxiv.org/abs/2208.05229</a> . All the work have been done on the Kaggle platform - look for datasets <a href=\"https://www.kaggle.com/alexandervc/datasets\" target=\"_blank\">https://www.kaggle.com/alexandervc/datasets</a>, <a href=\"https://www.kaggle.com/andreizinovyev/datasets\" target=\"_blank\">https://www.kaggle.com/andreizinovyev/datasets</a>  and scripts there in. <br>\nNote: that are NOT multimodal datasets, so NOT of direct use for the present competition. </p>",
      "rawMarkdown": "**Disclaimer.** That might improve score, might not, but any outcome would be of interest for research community. \nSo everyone is welcome to collaborate  - hopefully produce a paper - see [Discussion1](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293), [Discussion2](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293), [Discussion3\n](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348661)\n\n**Very briefly:** For CITEseq part of task - try to use features stored in the dataset\nhttps://www.kaggle.com/datasets/alexandervc/multimodal-singlecell-integration-related-data-01 \n(Currently there is only \"cite_seq_inputs_denoising_pooling_n_neighbours20.h5ad\" , but later other related will appear - it takes some time). \n**Will it improve the score ?** If yes, how much and for what downstream methods ?\n\n**What is going on (briefly):**  The features of CITEseq task ( and also targets of Multiome) are the so-called single cell RNA sequencing count matrices, which are widely studied last years. It is known that they suffer from certain specific type of noise - technical (not biological) zero values (called \"dropouts\" ). And there are plenty methods of \"denoising\"  were developed.  One of them called **\"MAGIC\"** ( https://github.com/KrishnaswamyLab/MAGIC ), another is denoising autoencoder - both implemented in SCANPY  https://scanpy.readthedocs.io/en/stable/generated/scanpy.external.pp.dca.html \n\nThere are other methods. For example we use another - more simple and fast method, sometimes called \"pooling\".\n(See the script: https://www.kaggle.com/alexandervc/denoising-pooling-for-adata-citeseq-features ) \n\n**Research question:** benchmark different denoising methods - will they give score improvement or not, if yes how much ? \n\n**Why it is of interest:**   Sometimes \"denoising\" is helpful for analysis of the scRNA-seq data, but sometimes it is HARMFUL(!) - that are non-linear algorithms which may introduce too strong distortion to the data. The open research problems is to produce some guidelines: when \"denoising\" algorithm are good, when bad; how to choose/tune the params, how much improvement we may expect to achieve, how it depend on the task and data (cell types, technology, number of cells etc); and other questions. Typically analysis of the scRNA-seq is UNsupervised, so the estimation of the outcomes of the algorithms is not straightforward. The Kaggle competion is quite exceptional since it has SUPERvised task - so it allows to judge the algorithm performance based on clear supervised scores. So it is an interesting research question whether denoising helps to improve scores or not and if yes - how much.\n\n**So you are welcome to contribute to research and might be it will improve your score !**\n\n PS\n\n**EXAMPLES**\n\n**MAGIC gif example.** Click the page https://github.com/KrishnaswamyLab/MAGIC and look on short gif-video - it gives a flavor what methods do. The zeroes which are not biological, but technical dissapper and biologically expexted correlations are restored.\n\n**Examples of the current script for the cell cycle analysis:** on the frontpage image for the Kaggle dataset https://www.kaggle.com/datasets/alexandervc/scrnaseq-data-for-ribamolina-paper-deepcycle (and scripts there in) you can see that without denoising the expected \"cyclic\" structure is NOT seen, but it appears after the denoising. That is good example. The \"bad\" example is here: https://www.kaggle.com/datasets/alexandervc/scrnaseq-bone-marrow-collection-of-datasets the frontpage image demonstrates too much distortion introduced by the denoising from the present script - expected \"cycle/triangle\" is lost - too much correlations are introduced.\n\n**Method of denoising used here (pooling).** There are several methods of \"denoising\". The current script provides implementation for one of the most simple, but rather fast one. Roughly speaking it substitutes the features by the average of several neighboring features. The number of neighbors e.g. \"n_neighbours = 20\" is the key parameter of the strength of denoising. However, it is ALMOST exactly that much simple, but not really: the subtle point - in what space to calculate the distances for KNN and in what space to take averages. Here we calculate distances for KNN in the space which is 1) normalized+log 2) PCA reduced, but the averages are taken in the original space of the \"raw\"-counts , and only then we take standard normalization+log. It might it is not the optimal way and might be modifications would work better in some situations. Everyone is welcome to collaborate on improvements. Technicalities. There is also dropping out some less expressed genes - there is param - how many to keep n_top_genes_to_keep = 10000, and there is list of genes which we want to keep mandatory. There are some other params which might lead to dropping out some cells and genes, but current values - are such that it does not happen: threshold_pct_counts_mt = 15 min_count = 1_000 max_count = 10_000_000\n\nPS \n\nAnd sorry the MAGIC features are not yet ready, but we are thinking on that ))))) \n\nPSPS\nMore results when \"denoising\" improves the solution of the  particular task -  cell cycle analysis - see in our paper: https://arxiv.org/abs/2208.05229 . All the work have been done on the Kaggle platform - look for datasets https://www.kaggle.com/alexandervc/datasets, https://www.kaggle.com/andreizinovyev/datasets  and scripts there in. \nNote: that are NOT multimodal datasets, so NOT of direct use for the present competition.",
      "votes": null
    },
    {
      "id": "1944354",
      "postDate": "09/18/2022 08:29:47",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fc1a028019955176108d8ea46b1687ea2%2Fphoto_2022-09-18_10-27-43.jpg?generation=1663489709806623&amp;alt=media\" alt=\"\"><br>\n<a href=\"https://twitter.com/dana_peer/status/1570775298844819457?t=2LmD5fPp_ngTMndgbbZwaQ&amp;s=19\" target=\"_blank\">https://twitter.com/dana_peer/status/1570775298844819457?t=2LmD5fPp_ngTMndgbbZwaQ&amp;s=19</a></p>\n<p>One the key authors of the \"MAGIC\" - Doctor Strange - oops, sorry Doctor Dana Peer )))) </p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fc1a028019955176108d8ea46b1687ea2%2Fphoto_2022-09-18_10-27-43.jpg?generation=1663489709806623&alt=media)\nhttps://twitter.com/dana_peer/status/1570775298844819457?t=2LmD5fPp_ngTMndgbbZwaQ&s=19\n\nOne the key authors of the \"MAGIC\" - Doctor Strange - oops, sorry Doctor Dana Peer ))))",
      "votes": null
    },
    {
      "id": "1958755",
      "postDate": "09/27/2022 16:06:58",
      "content": "<p>I have denoised citeseq data, here is the dataset: <br>\n<a href=\"https://www.kaggle.com/datasets/geraseva/citeseq-denoised\" target=\"_blank\">https://www.kaggle.com/datasets/geraseva/citeseq-denoised</a></p>",
      "rawMarkdown": "I have denoised citeseq data, here is the dataset: \nhttps://www.kaggle.com/datasets/geraseva/citeseq-denoised",
      "votes": null
    },
    {
      "id": "1959026",
      "postDate": "09/27/2022 20:41:04",
      "content": "<p>See script to work with it: <br>\n<a href=\"https://www.kaggle.com/alexandervc/magic-features-load-example-and-cell-cycle\" target=\"_blank\">https://www.kaggle.com/alexandervc/magic-features-load-example-and-cell-cycle</a></p>",
      "rawMarkdown": "See script to work with it: \nhttps://www.kaggle.com/alexandervc/magic-features-load-example-and-cell-cycle",
      "votes": null
    },
    {
      "id": "2032175",
      "postDate": "11/16/2022 13:10:52",
      "content": "<p>Hi, I have tried MAGIC and DeepImpute, and they actually decreased my CV. I think this problem is caused by a high false positive rate. I can submit a kernel notebook later to show my steps.</p>",
      "rawMarkdown": "Hi, I have tried MAGIC and DeepImpute, and they actually decreased my CV. I think this problem is caused by a high false positive rate. I can submit a kernel notebook later to show my steps.",
      "votes": null
    },
    {
      "id": "2032181",
      "postDate": "11/16/2022 13:18:56",
      "content": "<p><a href=\"https://www.kaggle.com/llttyy\" target=\"_blank\">@llttyy</a> Thank you very much  for your comment !</p>\n<p>We also see similar effect.  If we use these features by themselves they decrease results. In our analysis they improved CV and public LB, when some of them were concatenated with the original features. <br>\nBut it is not clear now would it be the same for private LB or not.</p>\n<p>Well in general it is kind of puzzling - we see that at least some correlations improve between CD-proteins and corresponding CD-RNA - as expected from denoising (I mean after using Magic for RNA side). Why then models performs worse ? That is puzzling </p>",
      "rawMarkdown": "llttyy Thank you very much  for your comment !\n\nWe also see similar effect.  If we use these features by themselves they decrease results. In our analysis they improved CV and public LB, when some of them were concatenated with the original features. \nBut it is not clear now would it be the same for private LB or not.\n\nWell in general it is kind of puzzling - we see that at least some correlations improve between CD-proteins and corresponding CD-RNA - as expected from denoising (I mean after using Magic for RNA side). Why then models performs worse ? That is puzzling",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1944354,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "09/18/2022 08:29:47",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fc1a028019955176108d8ea46b1687ea2%2Fphoto_2022-09-18_10-27-43.jpg?generation=1663489709806623&amp;alt=media\" alt=\"\"><br>\n<a href=\"https://twitter.com/dana_peer/status/1570775298844819457?t=2LmD5fPp_ngTMndgbbZwaQ&amp;s=19\" target=\"_blank\">https://twitter.com/dana_peer/status/1570775298844819457?t=2LmD5fPp_ngTMndgbbZwaQ&amp;s=19</a></p>\n<p>One the key authors of the \"MAGIC\" - Doctor Strange - oops, sorry Doctor Dana Peer )))) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1958755,
      "author_name": "geraseva",
      "author_url": "",
      "post_date": "09/27/2022 16:06:58",
      "content": "<p>I have denoised citeseq data, here is the dataset: <br>\n<a href=\"https://www.kaggle.com/datasets/geraseva/citeseq-denoised\" target=\"_blank\">https://www.kaggle.com/datasets/geraseva/citeseq-denoised</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1959026,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "09/27/2022 20:41:04",
          "content": "<p>See script to work with it: <br>\n<a href=\"https://www.kaggle.com/alexandervc/magic-features-load-example-and-cell-cycle\" target=\"_blank\">https://www.kaggle.com/alexandervc/magic-features-load-example-and-cell-cycle</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2032175,
      "author_name": "llttyy",
      "author_url": "",
      "post_date": "11/16/2022 13:10:52",
      "content": "<p>Hi, I have tried MAGIC and DeepImpute, and they actually decreased my CV. I think this problem is caused by a high false positive rate. I can submit a kernel notebook later to show my steps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2032181,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "11/16/2022 13:18:56",
          "content": "<p><a href=\"https://www.kaggle.com/llttyy\" target=\"_blank\">@llttyy</a> Thank you very much  for your comment !</p>\n<p>We also see similar effect.  If we use these features by themselves they decrease results. In our analysis they improved CV and public LB, when some of them were concatenated with the original features. <br>\nBut it is not clear now would it be the same for private LB or not.</p>\n<p>Well in general it is kind of puzzling - we see that at least some correlations improve between CD-proteins and corresponding CD-RNA - as expected from denoising (I mean after using Magic for RNA side). Why then models performs worse ? That is puzzling </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1929789": "**Disclaimer.** That might improve score, might not, but any outcome would be of interest for research community. \nSo everyone is welcome to collaborate  - hopefully produce a paper - see [Discussion1](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293), [Discussion2](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293), [Discussion3\n](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348661)\n\n**Very briefly:** For CITEseq part of task - try to use features stored in the dataset\nhttps://www.kaggle.com/datasets/alexandervc/multimodal-singlecell-integration-related-data-01 \n(Currently there is only \"cite_seq_inputs_denoising_pooling_n_neighbours20.h5ad\" , but later other related will appear - it takes some time). \n**Will it improve the score ?** If yes, how much and for what downstream methods ?\n\n**What is going on (briefly):**  The features of CITEseq task ( and also targets of Multiome) are the so-called single cell RNA sequencing count matrices, which are widely studied last years. It is known that they suffer from certain specific type of noise - technical (not biological) zero values (called \"dropouts\" ). And there are plenty methods of \"denoising\"  were developed.  One of them called **\"MAGIC\"** ( https://github.com/KrishnaswamyLab/MAGIC ), another is denoising autoencoder - both implemented in SCANPY  https://scanpy.readthedocs.io/en/stable/generated/scanpy.external.pp.dca.html \n\nThere are other methods. For example we use another - more simple and fast method, sometimes called \"pooling\".\n(See the script: https://www.kaggle.com/alexandervc/denoising-pooling-for-adata-citeseq-features ) \n\n**Research question:** benchmark different denoising methods - will they give score improvement or not, if yes how much ? \n\n**Why it is of interest:**   Sometimes \"denoising\" is helpful for analysis of the scRNA-seq data, but sometimes it is HARMFUL(!) - that are non-linear algorithms which may introduce too strong distortion to the data. The open research problems is to produce some guidelines: when \"denoising\" algorithm are good, when bad; how to choose/tune the params, how much improvement we may expect to achieve, how it depend on the task and data (cell types, technology, number of cells etc); and other questions. Typically analysis of the scRNA-seq is UNsupervised, so the estimation of the outcomes of the algorithms is not straightforward. The Kaggle competion is quite exceptional since it has SUPERvised task - so it allows to judge the algorithm performance based on clear supervised scores. So it is an interesting research question whether denoising helps to improve scores or not and if yes - how much.\n\n**So you are welcome to contribute to research and might be it will improve your score !**\n\n PS\n\n**EXAMPLES**\n\n**MAGIC gif example.** Click the page https://github.com/KrishnaswamyLab/MAGIC and look on short gif-video - it gives a flavor what methods do. The zeroes which are not biological, but technical dissapper and biologically expexted correlations are restored.\n\n**Examples of the current script for the cell cycle analysis:** on the frontpage image for the Kaggle dataset https://www.kaggle.com/datasets/alexandervc/scrnaseq-data-for-ribamolina-paper-deepcycle (and scripts there in) you can see that without denoising the expected \"cyclic\" structure is NOT seen, but it appears after the denoising. That is good example. The \"bad\" example is here: https://www.kaggle.com/datasets/alexandervc/scrnaseq-bone-marrow-collection-of-datasets the frontpage image demonstrates too much distortion introduced by the denoising from the present script - expected \"cycle/triangle\" is lost - too much correlations are introduced.\n\n**Method of denoising used here (pooling).** There are several methods of \"denoising\". The current script provides implementation for one of the most simple, but rather fast one. Roughly speaking it substitutes the features by the average of several neighboring features. The number of neighbors e.g. \"n_neighbours = 20\" is the key parameter of the strength of denoising. However, it is ALMOST exactly that much simple, but not really: the subtle point - in what space to calculate the distances for KNN and in what space to take averages. Here we calculate distances for KNN in the space which is 1) normalized+log 2) PCA reduced, but the averages are taken in the original space of the \"raw\"-counts , and only then we take standard normalization+log. It might it is not the optimal way and might be modifications would work better in some situations. Everyone is welcome to collaborate on improvements. Technicalities. There is also dropping out some less expressed genes - there is param - how many to keep n_top_genes_to_keep = 10000, and there is list of genes which we want to keep mandatory. There are some other params which might lead to dropping out some cells and genes, but current values - are such that it does not happen: threshold_pct_counts_mt = 15 min_count = 1_000 max_count = 10_000_000\n\nPS \n\nAnd sorry the MAGIC features are not yet ready, but we are thinking on that ))))) \n\nPSPS\nMore results when \"denoising\" improves the solution of the  particular task -  cell cycle analysis - see in our paper: https://arxiv.org/abs/2208.05229 . All the work have been done on the Kaggle platform - look for datasets https://www.kaggle.com/alexandervc/datasets, https://www.kaggle.com/andreizinovyev/datasets  and scripts there in. \nNote: that are NOT multimodal datasets, so NOT of direct use for the present competition.",
    "1944354": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fc1a028019955176108d8ea46b1687ea2%2Fphoto_2022-09-18_10-27-43.jpg?generation=1663489709806623&alt=media)\nhttps://twitter.com/dana_peer/status/1570775298844819457?t=2LmD5fPp_ngTMndgbbZwaQ&s=19\n\nOne the key authors of the \"MAGIC\" - Doctor Strange - oops, sorry Doctor Dana Peer ))))",
    "1958755": "I have denoised citeseq data, here is the dataset: \nhttps://www.kaggle.com/datasets/geraseva/citeseq-denoised",
    "1959026": "See script to work with it: \nhttps://www.kaggle.com/alexandervc/magic-features-load-example-and-cell-cycle",
    "2032175": "Hi, I have tried MAGIC and DeepImpute, and they actually decreased my CV. I think this problem is caused by a high false positive rate. I can submit a kernel notebook later to show my steps.",
    "2032181": "llttyy Thank you very much  for your comment !\n\nWe also see similar effect.  If we use these features by themselves they decrease results. In our analysis they improved CV and public LB, when some of them were concatenated with the original features. \nBut it is not clear now would it be the same for private LB or not.\n\nWell in general it is kind of puzzling - we see that at least some correlations improve between CD-proteins and corresponding CD-RNA - as expected from denoising (I mean after using Magic for RNA side). Why then models performs worse ? That is puzzling"
  },
  "source": "meta"
}