{
  "id": 348295,
  "title": "Let us list biological/bioinformatics research open questions which we can explore using that or similar data (except competition goal) ",
  "url": "/competitions/open-problems-multimodal/discussion/348295",
  "author_name": "",
  "post_date": "2022-08-27T19:11:29.451455600Z",
  "votes": 13,
  "comment_count": 1,
  "views": 0,
  "content": "<p>We are planning to organize some research and educational activity around (widely around - means not about getting top score) that remarkable competition.  See the discussion: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293</a></p>\n<p>The data provided in the competition is cutting-edge technology and not so many similar datasets available. <br>\nSuch data potentially can lead to new insights related to molecular biology, drug design, etc and are of interest to many researchers. May be such questions might be of interest to Kaggle community and some people might try to analyze them. </p>\n<p>So that question is to those people who already have some experience in that research area. <br>\nLet us create some list of biologically/bioinformatically open questions which can be resolved with such or similar datasets.</p>\n<p>Let me start.<br>\nWe have just finished some paper <a href=\"https://arxiv.org/abs/2208.05229\" target=\"_blank\">https://arxiv.org/abs/2208.05229</a> where we discuss some open questions related to cell cycle analysis, based on single cell RNA sequencing data (just one modality - scRNA-seq). <br>\nSo let me list a couple of them, which might interesting to research withing current competition.</p>\n<p>1)  Use/not use/how use  denoising algorithms (MAGIC, etc). <br>\nscRNA-seq data has specific kind of noise - technical zeroes, sometimes denoising algorithms can help to cure that,<br>\ne.g. see image on front <a href=\"https://www.kaggle.com/datasets/alexandervc/scrnaseq-data-for-ribamolina-paper-deepcycle\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/scrnaseq-data-for-ribamolina-paper-deepcycle</a><br>\nbut sometimes they create distortation which leads to incorrect bio conclusions<br>\nsee image on from <a href=\"https://www.kaggle.com/datasets/alexandervc/scrnaseq-bone-marrow-collection-of-datasets\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/scrnaseq-bone-marrow-collection-of-datasets</a></p>\n<p>The question is to provide some guidence/benchmarking how use/or not use such algorithms. </p>\n<p>In particular in the context of the current competition:<br>\nif we apply denoising algorithms to scRNAseq features at CITEseq - will it improve prediction score ? <br>\nwhat algorithm to use , what params to choose ? </p>\n<p>We might explore that question for current competition, for cell cycle task, for cell type identifications etc.</p>\n<p>I guess something is already known for some tasks - we should review the literature - but for sure the problem is not resolved completely </p>\n<p>2) Cell types identification quality control. <br>\nCell type identitification is one the main tasks in single cell RNA seq analysis. <br>\nIt is very widely explored and quite good results obtained, but seems still the methods are not perfect.<br>\nIn particular for the current dataset, for Tabula Muris datasets, our analysis based on cell cycle<br>\nseems to produce certain doubts that identification is fully correct. <br>\nThe reason is quite simple - we look on cell cycle for one cell type - expect to see \"cycle\" , but see \"half-cycle\",<br>\nthen look on the other cell type - and see the rest of the \"half for it\". <br>\nThus it is tempting to think that cells were identified as different cell type, but actually they are of the same type,<br>\nonly different by cell cycle state. </p>\n<p>Now we have multi-modal dataset - so we can look not only on RNA-seq , but on the othe modalities,<br>\nand that seems to give new doubts that cell types were identified correctly.<br>\nBecause some CD** markers behave strangely. <br>\nI will recheck that and provide the picture later. </p>\n<p>Thus the task is to use cell cycle analysis and/or multiomics data  to provide some quality control for cell type identification and possible notify about the mistakes for some known annotations. </p>",
  "messages": [
    {
      "id": "1916331",
      "postDate": "08/27/2022 19:11:29",
      "content": "<p>We are planning to organize some research and educational activity around (widely around - means not about getting top score) that remarkable competition.  See the discussion: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293</a></p>\n<p>The data provided in the competition is cutting-edge technology and not so many similar datasets available. <br>\nSuch data potentially can lead to new insights related to molecular biology, drug design, etc and are of interest to many researchers. May be such questions might be of interest to Kaggle community and some people might try to analyze them. </p>\n<p>So that question is to those people who already have some experience in that research area. <br>\nLet us create some list of biologically/bioinformatically open questions which can be resolved with such or similar datasets.</p>\n<p>Let me start.<br>\nWe have just finished some paper <a href=\"https://arxiv.org/abs/2208.05229\" target=\"_blank\">https://arxiv.org/abs/2208.05229</a> where we discuss some open questions related to cell cycle analysis, based on single cell RNA sequencing data (just one modality - scRNA-seq). <br>\nSo let me list a couple of them, which might interesting to research withing current competition.</p>\n<p>1)  Use/not use/how use  denoising algorithms (MAGIC, etc). <br>\nscRNA-seq data has specific kind of noise - technical zeroes, sometimes denoising algorithms can help to cure that,<br>\ne.g. see image on front <a href=\"https://www.kaggle.com/datasets/alexandervc/scrnaseq-data-for-ribamolina-paper-deepcycle\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/scrnaseq-data-for-ribamolina-paper-deepcycle</a><br>\nbut sometimes they create distortation which leads to incorrect bio conclusions<br>\nsee image on from <a href=\"https://www.kaggle.com/datasets/alexandervc/scrnaseq-bone-marrow-collection-of-datasets\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/scrnaseq-bone-marrow-collection-of-datasets</a></p>\n<p>The question is to provide some guidence/benchmarking how use/or not use such algorithms. </p>\n<p>In particular in the context of the current competition:<br>\nif we apply denoising algorithms to scRNAseq features at CITEseq - will it improve prediction score ? <br>\nwhat algorithm to use , what params to choose ? </p>\n<p>We might explore that question for current competition, for cell cycle task, for cell type identifications etc.</p>\n<p>I guess something is already known for some tasks - we should review the literature - but for sure the problem is not resolved completely </p>\n<p>2) Cell types identification quality control. <br>\nCell type identitification is one the main tasks in single cell RNA seq analysis. <br>\nIt is very widely explored and quite good results obtained, but seems still the methods are not perfect.<br>\nIn particular for the current dataset, for Tabula Muris datasets, our analysis based on cell cycle<br>\nseems to produce certain doubts that identification is fully correct. <br>\nThe reason is quite simple - we look on cell cycle for one cell type - expect to see \"cycle\" , but see \"half-cycle\",<br>\nthen look on the other cell type - and see the rest of the \"half for it\". <br>\nThus it is tempting to think that cells were identified as different cell type, but actually they are of the same type,<br>\nonly different by cell cycle state. </p>\n<p>Now we have multi-modal dataset - so we can look not only on RNA-seq , but on the othe modalities,<br>\nand that seems to give new doubts that cell types were identified correctly.<br>\nBecause some CD** markers behave strangely. <br>\nI will recheck that and provide the picture later. </p>\n<p>Thus the task is to use cell cycle analysis and/or multiomics data  to provide some quality control for cell type identification and possible notify about the mistakes for some known annotations. </p>",
      "rawMarkdown": "We are planning to organize some research and educational activity around (widely around - means not about getting top score) that remarkable competition.  See the discussion: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293\n\nThe data provided in the competition is cutting-edge technology and not so many similar datasets available. \nSuch data potentially can lead to new insights related to molecular biology, drug design, etc and are of interest to many researchers. May be such questions might be of interest to Kaggle community and some people might try to analyze them. \n\nSo that question is to those people who already have some experience in that research area. \nLet us create some list of biologically/bioinformatically open questions which can be resolved with such or similar datasets.\n\nLet me start.\nWe have just finished some paper https://arxiv.org/abs/2208.05229 where we discuss some open questions related to cell cycle analysis, based on single cell RNA sequencing data (just one modality - scRNA-seq). \nSo let me list a couple of them, which might interesting to research withing current competition.\n\n1)  Use/not use/how use  denoising algorithms (MAGIC, etc). \nscRNA-seq data has specific kind of noise - technical zeroes, sometimes denoising algorithms can help to cure that,\ne.g. see image on front https://www.kaggle.com/datasets/alexandervc/scrnaseq-data-for-ribamolina-paper-deepcycle\nbut sometimes they create distortation which leads to incorrect bio conclusions\nsee image on from https://www.kaggle.com/datasets/alexandervc/scrnaseq-bone-marrow-collection-of-datasets\n\nThe question is to provide some guidence/benchmarking how use/or not use such algorithms. \n\nIn particular in the context of the current competition:\nif we apply denoising algorithms to scRNAseq features at CITEseq - will it improve prediction score ? \nwhat algorithm to use , what params to choose ? \n\nWe might explore that question for current competition, for cell cycle task, for cell type identifications etc.\n\nI guess something is already known for some tasks - we should review the literature - but for sure the problem is not resolved completely \n\n2) Cell types identification quality control. \nCell type identitification is one the main tasks in single cell RNA seq analysis. \nIt is very widely explored and quite good results obtained, but seems still the methods are not perfect.\nIn particular for the current dataset, for Tabula Muris datasets, our analysis based on cell cycle\nseems to produce certain doubts that identification is fully correct. \nThe reason is quite simple - we look on cell cycle for one cell type - expect to see \"cycle\" , but see \"half-cycle\",\nthen look on the other cell type - and see the rest of the \"half for it\". \nThus it is tempting to think that cells were identified as different cell type, but actually they are of the same type,\nonly different by cell cycle state. \n\nNow we have multi-modal dataset - so we can look not only on RNA-seq , but on the othe modalities,\nand that seems to give new doubts that cell types were identified correctly.\nBecause some CD** markers behave strangely. \nI will recheck that and provide the picture later. \n\nThus the task is to use cell cycle analysis and/or multiomics data  to provide some quality control for cell type identification and possible notify about the mistakes for some known annotations.",
      "votes": null
    },
    {
      "id": "1927860",
      "postDate": "09/06/2022 03:09:52",
      "content": "<p>i think there will be an explosion of cell analysis related competitions coming up in kaggle. It is quite difficult for one without genomics background to understand how the data is prepared and used. Such prior knowledge is actually very important to create a good model for the problem.</p>\n<p>Some course for \"understand from scratch\" will be good.</p>",
      "rawMarkdown": "i think there will be an explosion of cell analysis related competitions coming up in kaggle. It is quite difficult for one without genomics background to understand how the data is prepared and used. Such prior knowledge is actually very important to create a good model for the problem.\n\nSome course for \"understand from scratch\" will be good.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1927860,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/06/2022 03:09:52",
      "content": "<p>i think there will be an explosion of cell analysis related competitions coming up in kaggle. It is quite difficult for one without genomics background to understand how the data is prepared and used. Such prior knowledge is actually very important to create a good model for the problem.</p>\n<p>Some course for \"understand from scratch\" will be good.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1916331": "We are planning to organize some research and educational activity around (widely around - means not about getting top score) that remarkable competition.  See the discussion: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293\n\nThe data provided in the competition is cutting-edge technology and not so many similar datasets available. \nSuch data potentially can lead to new insights related to molecular biology, drug design, etc and are of interest to many researchers. May be such questions might be of interest to Kaggle community and some people might try to analyze them. \n\nSo that question is to those people who already have some experience in that research area. \nLet us create some list of biologically/bioinformatically open questions which can be resolved with such or similar datasets.\n\nLet me start.\nWe have just finished some paper https://arxiv.org/abs/2208.05229 where we discuss some open questions related to cell cycle analysis, based on single cell RNA sequencing data (just one modality - scRNA-seq). \nSo let me list a couple of them, which might interesting to research withing current competition.\n\n1)  Use/not use/how use  denoising algorithms (MAGIC, etc). \nscRNA-seq data has specific kind of noise - technical zeroes, sometimes denoising algorithms can help to cure that,\ne.g. see image on front https://www.kaggle.com/datasets/alexandervc/scrnaseq-data-for-ribamolina-paper-deepcycle\nbut sometimes they create distortation which leads to incorrect bio conclusions\nsee image on from https://www.kaggle.com/datasets/alexandervc/scrnaseq-bone-marrow-collection-of-datasets\n\nThe question is to provide some guidence/benchmarking how use/or not use such algorithms. \n\nIn particular in the context of the current competition:\nif we apply denoising algorithms to scRNAseq features at CITEseq - will it improve prediction score ? \nwhat algorithm to use , what params to choose ? \n\nWe might explore that question for current competition, for cell cycle task, for cell type identifications etc.\n\nI guess something is already known for some tasks - we should review the literature - but for sure the problem is not resolved completely \n\n2) Cell types identification quality control. \nCell type identitification is one the main tasks in single cell RNA seq analysis. \nIt is very widely explored and quite good results obtained, but seems still the methods are not perfect.\nIn particular for the current dataset, for Tabula Muris datasets, our analysis based on cell cycle\nseems to produce certain doubts that identification is fully correct. \nThe reason is quite simple - we look on cell cycle for one cell type - expect to see \"cycle\" , but see \"half-cycle\",\nthen look on the other cell type - and see the rest of the \"half for it\". \nThus it is tempting to think that cells were identified as different cell type, but actually they are of the same type,\nonly different by cell cycle state. \n\nNow we have multi-modal dataset - so we can look not only on RNA-seq , but on the othe modalities,\nand that seems to give new doubts that cell types were identified correctly.\nBecause some CD** markers behave strangely. \nI will recheck that and provide the picture later. \n\nThus the task is to use cell cycle analysis and/or multiomics data  to provide some quality control for cell type identification and possible notify about the mistakes for some known annotations.",
    "1927860": "i think there will be an explosion of cell analysis related competitions coming up in kaggle. It is quite difficult for one without genomics background to understand how the data is prepared and used. Such prior knowledge is actually very important to create a good model for the problem.\n\nSome course for \"understand from scratch\" will be good."
  },
  "source": "meta"
}