{
  "id": 344607,
  "title": "Welcome to the Open Problems Multimodal Single-Cell Integration challenge!",
  "url": "/competitions/open-problems-multimodal/discussion/344607",
  "author_name": "Daniel Burkhardt",
  "post_date": "2022-08-15T21:48:54.694000",
  "votes": 32,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Dear Kagglers,</p>\n<p>On behalf of Open Problem in Single-Cell Analysis, I would like to officially welcome you to the “Open Problems - Multimodal Single-Cell Integration” competition, part of the <a href=\"https://neurips.cc/Conferences/2022/CompetitionTrack\" target=\"_blank\">NeurIPS 2022 Competition Track</a>!</p>\n<p>This competition is about predicting the relationship between different genetic measurements (DNA, RNA, and protein) measured simultaneously in single cells. In the field of biomedicine, single-cell technologies have led to an explosion of new findings about the diversity of the 37 trillion cells in the human body and how these cell types vary between health and disease. However, the field still has several grand challenges ahead of it, and modeling dynamics of cell state is one of the biggest open problems (<a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-1926-6\" target=\"_blank\">Genome Biology 2020</a>).</p>\n<p>For this competition, Open Problems partnered with Cellarity, a cell-centric drug creation company, to generate a first-of-its-kind benchmarking dataset designed to drive advances in algorithms that can capture drivers of changes in cell state over time. We measured CD34+ hematopoietic stem and progenitor cells (HSPCs) from four healthy human donors at 5 time points using two different multimodal single-cell technologies. Your challenge will be to learn the relationship between these modalities and learn how to predict from one to another at a later unseen time point in the datasets.</p>\n<p>In the Data section, you’ll find descriptions of the features in each modality and explanations of some metadata associated with the observations, which each correspond to a single cell. Please feel free to ask us any questions!</p>\n<p>In our NeurIPS competition last year (<a href=\"https://openproblems.bio/neurips_2021/\" target=\"_blank\">link</a>), we were thrilled to see contributions from teams that specialize in genomics compete neck and neck with algorithms originally designed for completely different domains, such as image captioning. We can’t wait to see what solutions are proposed in this competition.</p>\n<p>This competition was a collaborative effort between scientists at Cellarity, the Chan Zuckerberg Biohub, Yale University, and Helmholtz Munich with sponsorship from Cellarity and the Chan Zuckerberg Initiative's Single-Cell Biology Program. </p>\n<p>If you’re interested to learn more about Open Problems in Single-Cell Analysis, please visit our homepage, <a href=\"https://openproblems.bio\" target=\"_blank\">https://openproblems.bio</a>, where you can sign up for <a href=\"https://docs.google.com/forms/d/e/1FAIpQLSe90Oky4-1b0HbdLsp5Yqo9juCd2mq-NlGHU9NHRW1ECok1xQ/viewform?usp=sf_link\" target=\"_blank\">our mailing list</a>.</p>\n<p>Best of luck! We can’t wait to see what you build.<br>\nDaniel Burkhardt and the rest of the Core Team at Open Problems</p>",
  "messages": [
    {
      "id": 1900287,
      "postDate": "2022-08-15T21:48:54.693Z",
      "content": "<p>Dear Kagglers,</p>\n<p>On behalf of Open Problem in Single-Cell Analysis, I would like to officially welcome you to the “Open Problems - Multimodal Single-Cell Integration” competition, part of the <a href=\"https://neurips.cc/Conferences/2022/CompetitionTrack\" target=\"_blank\">NeurIPS 2022 Competition Track</a>!</p>\n<p>This competition is about predicting the relationship between different genetic measurements (DNA, RNA, and protein) measured simultaneously in single cells. In the field of biomedicine, single-cell technologies have led to an explosion of new findings about the diversity of the 37 trillion cells in the human body and how these cell types vary between health and disease. However, the field still has several grand challenges ahead of it, and modeling dynamics of cell state is one of the biggest open problems (<a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-1926-6\" target=\"_blank\">Genome Biology 2020</a>).</p>\n<p>For this competition, Open Problems partnered with Cellarity, a cell-centric drug creation company, to generate a first-of-its-kind benchmarking dataset designed to drive advances in algorithms that can capture drivers of changes in cell state over time. We measured CD34+ hematopoietic stem and progenitor cells (HSPCs) from four healthy human donors at 5 time points using two different multimodal single-cell technologies. Your challenge will be to learn the relationship between these modalities and learn how to predict from one to another at a later unseen time point in the datasets.</p>\n<p>In the Data section, you’ll find descriptions of the features in each modality and explanations of some metadata associated with the observations, which each correspond to a single cell. Please feel free to ask us any questions!</p>\n<p>In our NeurIPS competition last year (<a href=\"https://openproblems.bio/neurips_2021/\" target=\"_blank\">link</a>), we were thrilled to see contributions from teams that specialize in genomics compete neck and neck with algorithms originally designed for completely different domains, such as image captioning. We can’t wait to see what solutions are proposed in this competition.</p>\n<p>This competition was a collaborative effort between scientists at Cellarity, the Chan Zuckerberg Biohub, Yale University, and Helmholtz Munich with sponsorship from Cellarity and the Chan Zuckerberg Initiative's Single-Cell Biology Program. </p>\n<p>If you’re interested to learn more about Open Problems in Single-Cell Analysis, please visit our homepage, <a href=\"https://openproblems.bio\" target=\"_blank\">https://openproblems.bio</a>, where you can sign up for <a href=\"https://docs.google.com/forms/d/e/1FAIpQLSe90Oky4-1b0HbdLsp5Yqo9juCd2mq-NlGHU9NHRW1ECok1xQ/viewform?usp=sf_link\" target=\"_blank\">our mailing list</a>.</p>\n<p>Best of luck! We can’t wait to see what you build.<br>\nDaniel Burkhardt and the rest of the Core Team at Open Problems</p>",
      "rawMarkdown": "Dear Kagglers,\n\nOn behalf of Open Problem in Single-Cell Analysis, I would like to officially welcome you to the “Open Problems - Multimodal Single-Cell Integration” competition, part of the [NeurIPS 2022 Competition Track](https://neurips.cc/Conferences/2022/CompetitionTrack)!\n\nThis competition is about predicting the relationship between different genetic measurements (DNA, RNA, and protein) measured simultaneously in single cells. In the field of biomedicine, single-cell technologies have led to an explosion of new findings about the diversity of the 37 trillion cells in the human body and how these cell types vary between health and disease. However, the field still has several grand challenges ahead of it, and modeling dynamics of cell state is one of the biggest open problems ([Genome Biology 2020](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-1926-6)).\n\nFor this competition, Open Problems partnered with Cellarity, a cell-centric drug creation company, to generate a first-of-its-kind benchmarking dataset designed to drive advances in algorithms that can capture drivers of changes in cell state over time. We measured CD34+ hematopoietic stem and progenitor cells (HSPCs) from four healthy human donors at 5 time points using two different multimodal single-cell technologies. Your challenge will be to learn the relationship between these modalities and learn how to predict from one to another at a later unseen time point in the datasets.\n\nIn the Data section, you’ll find descriptions of the features in each modality and explanations of some metadata associated with the observations, which each correspond to a single cell. Please feel free to ask us any questions!\n\nIn our NeurIPS competition last year ([link](https://openproblems.bio/neurips_2021/)), we were thrilled to see contributions from teams that specialize in genomics compete neck and neck with algorithms originally designed for completely different domains, such as image captioning. We can’t wait to see what solutions are proposed in this competition.\n\nThis competition was a collaborative effort between scientists at Cellarity, the Chan Zuckerberg Biohub, Yale University, and Helmholtz Munich with sponsorship from Cellarity and the Chan Zuckerberg Initiative's Single-Cell Biology Program. \n\nIf you’re interested to learn more about Open Problems in Single-Cell Analysis, please visit our homepage, https://openproblems.bio, where you can sign up for [our mailing list](https://docs.google.com/forms/d/e/1FAIpQLSe90Oky4-1b0HbdLsp5Yqo9juCd2mq-NlGHU9NHRW1ECok1xQ/viewform?usp=sf_link).\n\nBest of luck! We can’t wait to see what you build.\nDaniel Burkhardt and the rest of the Core Team at Open Problems\n",
      "votes": 32
    },
    {
      "id": 1918940,
      "postDate": "2022-08-30T02:04:48.853Z",
      "content": "<p>Hi Daniel <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> ,  is it allowed to use the NeurIPS 2021 Single-cell multi-modality dataset as an external dataset for this competition? Thanks.</p>",
      "rawMarkdown": "Hi Daniel @danielburkhardt ,  is it allowed to use the NeurIPS 2021 Single-cell multi-modality dataset as an external dataset for this competition? Thanks.",
      "votes": 9
    },
    {
      "id": 1900325,
      "postDate": "2022-08-15T22:57:30.327Z",
      "content": "<p>Hi Daniel, thank you.</p>\n<p>Are you (I mean, the organizers of the competition) planning to release some notebook with basic code snippets to get some of us up and running?</p>\n<p>I tried to read one h5 file using muon (more specifically, this function: <a href=\"https://muon.readthedocs.io/en/latest/api/generated/muon.read_10x_h5.html#muon.read_10x_h5\" target=\"_blank\">read_10x_h5</a>) library but I got an error: \"OSError: Can't read data (can't open directory: /usr/local/hdf5/lib/plugin)\". As far as I could see, there is no such directory. Does this have to do with running these functions in a kaggle environment? </p>\n<p>Sorry if this is a basic question. Thank you in advance!</p>",
      "rawMarkdown": "Hi Daniel, thank you.\n\nAre you (I mean, the organizers of the competition) planning to release some notebook with basic code snippets to get some of us up and running?\n\nI tried to read one h5 file using muon (more specifically, this function: [read_10x_h5](https://muon.readthedocs.io/en/latest/api/generated/muon.read_10x_h5.html#muon.read_10x_h5)) library but I got an error: \"OSError: Can't read data (can't open directory: /usr/local/hdf5/lib/plugin)\". As far as I could see, there is no such directory. Does this have to do with running these functions in a kaggle environment? \n\nSorry if this is a basic question. Thank you in advance!",
      "votes": 3,
      "replies": [
        {
          "id": 1900328,
          "postDate": "2022-08-15T23:06:54.597Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1900335,
          "postDate": "2022-08-15T23:23:11.843Z",
          "content": "<p><strong>Update: Notebook is now available as \"Getting Started - Data Loading\"</strong></p>\n<p>You beat us to it! This went live afterhours local time for us, but <a href=\"https://www.kaggle.com/peterholderrieth\" target=\"_blank\">@peterholderrieth</a> will post our code notebook later tonight.</p>\n<p>For now, I can point you towards the easiest way to load the data:</p>\n<pre><code>!pip install --quiet tables # required in the Kaggle environment\n\n# Imports\nimport os\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Data paths\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\n# Load metadata\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell\n\n# Load the data\ndf_cite_train_x = pd.read_hdf(FP_CITE_TRAIN_INPUTS)\ndf_cite_test_x = pd.read_hdf(FP_CITE_TEST_INPUTS)\ndf_cite_train_x.head()\n\ndf_cite_train_y = pd.read_hdf(FP_CITE_TRAIN_TARGETS)\ndf_cite_train_y.head()\n</code></pre>\n<p>More details to come later!</p>",
          "rawMarkdown": "**Update: Notebook is now available as \"Getting Started - Data Loading\"**\n\nYou beat us to it! This went live afterhours local time for us, but @peterholderrieth will post our code notebook later tonight.\n\nFor now, I can point you towards the easiest way to load the data:\n\n```\n!pip install --quiet tables # required in the Kaggle environment\n\n# Imports\nimport os\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Data paths\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\n# Load metadata\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell\n\n# Load the data\ndf_cite_train_x = pd.read_hdf(FP_CITE_TRAIN_INPUTS)\ndf_cite_test_x = pd.read_hdf(FP_CITE_TEST_INPUTS)\ndf_cite_train_x.head()\n\ndf_cite_train_y = pd.read_hdf(FP_CITE_TRAIN_TARGETS)\ndf_cite_train_y.head()\n```\n\nMore details to come later!\n\n",
          "votes": 5
        },
        {
          "id": 1900364,
          "postDate": "2022-08-16T00:24:03.737Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/alekeuro\" target=\"_blank\">@alekeuro</a>! I have uploaded a <a href=\"https://www.kaggle.com/code/peterholderrieth/getting-started-data-loading\" target=\"_blank\">notebook</a> that step by step loads and explains the structure of the data. I hope it is helpful to you! </p>",
          "rawMarkdown": "Hi @alekeuro! I have uploaded a [notebook](https://www.kaggle.com/code/peterholderrieth/getting-started-data-loading) that step by step loads and explains the structure of the data. I hope it is helpful to you! ",
          "votes": 10
        },
        {
          "id": 1900369,
          "postDate": "2022-08-16T01:02:02.023Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/peterholderrieth\" target=\"_blank\">@peterholderrieth</a> and <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> !</p>",
          "rawMarkdown": "Thank you @peterholderrieth and @danielburkhardt !"
        },
        {
          "id": 1903777,
          "postDate": "2022-08-17T16:39:35.763Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/peterholderrieth\" target=\"_blank\">@peterholderrieth</a> ,</p>\n<p>going through your Getting Started notebook I noticed that, according to the bar chart (cell counts per technology, day and donor), there are no cells for donor nr. 4 (donor id: 27678) on day 4, for multiome tech. This count comes from the metadata.csv file. </p>\n<p>On the other side, the figure illustrating the experimental setup shows that part of the Public Test data should include cells from donor 4, at day 4, for multiome tech.</p>\n<p>Is it possible that one of the two pieces of info is wrong? Either the metadata file or the figure? Sorry if this is a dumb question.</p>\n<p>Thanks in advance!</p>",
          "rawMarkdown": "Hi @peterholderrieth ,\n\ngoing through your Getting Started notebook I noticed that, according to the bar chart (cell counts per technology, day and donor), there are no cells for donor nr. 4 (donor id: 27678) on day 4, for multiome tech. This count comes from the metadata.csv file. \n\nOn the other side, the figure illustrating the experimental setup shows that part of the Public Test data should include cells from donor 4, at day 4, for multiome tech.\n\nIs it possible that one of the two pieces of info is wrong? Either the metadata file or the figure? Sorry if this is a dumb question.\n\nThanks in advance!",
          "votes": 2
        },
        {
          "id": 1903797,
          "postDate": "2022-08-17T16:55:47.153Z",
          "content": "<p>Ahh this is a mistake in the figure! This dataset failed, and we decided it would be okay to use for public test because it would mean 3 time points for public test for each modality. I'll fix the image!</p>",
          "rawMarkdown": "Ahh this is a mistake in the figure! This dataset failed, and we decided it would be okay to use for public test because it would mean 3 time points for public test for each modality. I'll fix the image!"
        },
        {
          "id": 1904137,
          "postDate": "2022-08-18T01:57:56.647Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    },
    {
      "id": 2030466,
      "postDate": "2022-11-15T12:48:47.890Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> , thank you for giving us a unique problem and even more unique datasets and hosting this competition. <br>\nNow that we are on the finish line of it, I wonder if you had any results in mind with respect to the correlation metric. <br>\nAnd whether the expectation has been met or not ?</p>",
      "rawMarkdown": "Hi @danielburkhardt , thank you for giving us a unique problem and even more unique datasets and hosting this competition. \nNow that we are on the finish line of it, I wonder if you had any results in mind with respect to the correlation metric. \nAnd whether the expectation has been met or not ?",
      "votes": 1
    },
    {
      "id": 2018551,
      "postDate": "2022-11-05T19:43:00.977Z",
      "content": "<p>Hi Daniel, I wonder if it is correct that we will access the private testing datasets after three days. Thanks a lot.</p>",
      "rawMarkdown": "Hi Daniel, I wonder if it is correct that we will access the private testing datasets after three days. Thanks a lot.",
      "votes": 1
    },
    {
      "id": 1992601,
      "postDate": "2022-10-17T19:07:57.130Z",
      "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> <br>\nDaniel you shared somewhere in discussion topics the code how to get positions of the genes on DNA, to match the ATAC seq data.<br>\nI cannot find it now. <br>\nWould you be so kind to share it again, please ? </p>",
      "rawMarkdown": "@danielburkhardt \nDaniel you shared somewhere in discussion topics the code how to get positions of the genes on DNA, to match the ATAC seq data.\nI cannot find it now. \nWould you be so kind to share it again, please ? ",
      "votes": 1,
      "replies": [
        {
          "id": 1992758,
          "postDate": "2022-10-17T22:11:51.097Z",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349559\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349559</a></p>\n<p>Here you go! I tried putting this into a Kaggle notebook but ran into a memory error and didn't have time to troubleshoot.</p>",
          "rawMarkdown": "https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349559\n\nHere you go! I tried putting this into a Kaggle notebook but ran into a memory error and didn't have time to troubleshoot.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1974000,
      "postDate": "2022-10-06T02:43:56.867Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> is it allowed to integrate any public database into the inference such as gene networks?</p>",
      "rawMarkdown": "Hi @danielburkhardt is it allowed to integrate any public database into the inference such as gene networks?",
      "votes": 1
    },
    {
      "id": 1960657,
      "postDate": "2022-09-28T17:40:56.523Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> , thank you for this great competition. Just to confirm: the cells were NOT stimulated in any way after plating? It would be helpful, if you could describe a little bit more about what happened to the cells after they were obtained from the donors over those different time periods.</p>",
      "rawMarkdown": "Hi @danielburkhardt , thank you for this great competition. Just to confirm: the cells were NOT stimulated in any way after plating? It would be helpful, if you could describe a little bit more about what happened to the cells after they were obtained from the donors over those different time periods.",
      "votes": 1,
      "replies": [
        {
          "id": 1960755,
          "postDate": "2022-09-28T18:35:23.943Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mxposed\" target=\"_blank\">@mxposed</a> the cells were cultured with StemSpan SFEM media supplemented with CC100 and thrombopoietin (TPO) over 10 days. Cells were incubated at 37ºC and media was changed every 2-3 days.</p>",
          "rawMarkdown": "Hi @mxposed the cells were cultured with StemSpan SFEM media supplemented with CC100 and thrombopoietin (TPO) over 10 days. Cells were incubated at 37ºC and media was changed every 2-3 days.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1946685,
      "postDate": "2022-09-20T03:25:15.637Z",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> <br>\nI want to check the detailed correlation CV/LB.<br>\nCan I see one more digit LB?</p>",
      "rawMarkdown": "Hi, @danielburkhardt \nI want to check the detailed correlation CV/LB.\nCan I see one more digit LB?",
      "votes": 1
    },
    {
      "id": 1988300,
      "postDate": "2022-10-15T08:14:00.213Z",
      "content": "<p>Thank you for this nice information <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> </p>",
      "rawMarkdown": "Thank you for this nice information @danielburkhardt ",
      "votes": 2
    },
    {
      "id": 1943314,
      "postDate": "2022-09-17T12:55:25.330Z",
      "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a>, sorry if this has been asked on other part. What exactly mean 0 on gene expression, is not meassured or something else?</p>",
      "rawMarkdown": "@danielburkhardt, sorry if this has been asked on other part. What exactly mean 0 on gene expression, is not meassured or something else?",
      "votes": 2,
      "replies": [
        {
          "id": 1945877,
          "postDate": "2022-09-19T13:16:55.367Z",
          "content": "<p>Think of the data like counts from a Poisson process. It means that for that cell, we have 0 counts for that gene. It doesn't mean necessarily that the gene doesn't exist in that cell, just that we didn't observe any counts when we measured the cell.</p>\n<p>Note, this is not the same as \"missing data\" in a matrix completion sense. We can only capture about 30-50% of the mRNA in a cell, so there will be some molecules we don't sample when we're counting.</p>",
          "rawMarkdown": "Think of the data like counts from a Poisson process. It means that for that cell, we have 0 counts for that gene. It doesn't mean necessarily that the gene doesn't exist in that cell, just that we didn't observe any counts when we measured the cell.\n\nNote, this is not the same as \"missing data\" in a matrix completion sense. We can only capture about 30-50% of the mRNA in a cell, so there will be some molecules we don't sample when we're counting.",
          "votes": 1
        },
        {
          "id": 1945990,
          "postDate": "2022-09-19T14:44:38.750Z",
          "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a>, and there are not other technical factors that can drive the count to 0? Or in other words, can we try to denoise the 0?</p>",
          "rawMarkdown": "@danielburkhardt, and there are not other technical factors that can drive the count to 0? Or in other words, can we try to denoise the 0?"
        },
        {
          "id": 1947322,
          "postDate": "2022-09-20T12:13:43.867Z",
          "content": "<p>There certainly are approaches to denoising single-cell data. Check out these review articles:</p>\n<ul>\n<li><a href=\"https://pubmed.ncbi.nlm.nih.gov/33003202/\" target=\"_blank\">A review of computational strategies for denoising and imputation of single-cell transcriptomic data</a></li>\n<li><a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-02132-x\" target=\"_blank\">A systematic evaluation of single-cell RNA-sequencing imputation methods</a></li>\n</ul>",
          "rawMarkdown": "There certainly are approaches to denoising single-cell data. Check out these review articles:\n* [A review of computational strategies for denoising and imputation of single-cell transcriptomic data](https://pubmed.ncbi.nlm.nih.gov/33003202/)\n* [A systematic evaluation of single-cell RNA-sequencing imputation methods](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-02132-x)",
          "votes": 1
        },
        {
          "id": 1947375,
          "postDate": "2022-09-20T12:36:28.963Z",
          "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a>, thanks</p>",
          "rawMarkdown": "@danielburkhardt, thanks"
        },
        {
          "id": 1983856,
          "postDate": "2022-10-12T09:40:42.180Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1984172,
          "postDate": "2022-10-12T13:54:32.767Z",
          "content": "<p><a href=\"https://www.kaggle.com/ryumei\" target=\"_blank\">@ryumei</a>, when predicting one modality from another, there will often be zero values in the predicted targets because, for example, some genes are not expressed. Of the 20,000 protein coding genes in the genome, we often see 1k-5k have non-zero values in each cell.</p>\n<p>You should think of the predicted counts as originating from a Poisson counting process. We do not count all the transcripts in each cell, so there will be some zero values.</p>",
          "rawMarkdown": "@ryumei, when predicting one modality from another, there will often be zero values in the predicted targets because, for example, some genes are not expressed. Of the 20,000 protein coding genes in the genome, we often see 1k-5k have non-zero values in each cell.\n\nYou should think of the predicted counts as originating from a Poisson counting process. We do not count all the transcripts in each cell, so there will be some zero values."
        },
        {
          "id": 1984193,
          "postDate": "2022-10-12T14:02:11.087Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1904185,
      "postDate": "2022-08-18T03:10:43.073Z",
      "content": "<p>Quick question: Is the data publicly available in an unprocessed (.bam or .fastq) form? Or only in the processed form offered? I only ask as I don't detect a public accession for it or that it's on the 10x website, and I wanted to confirm that.</p>",
      "rawMarkdown": "Quick question: Is the data publicly available in an unprocessed (.bam or .fastq) form? Or only in the processed form offered? I only ask as I don't detect a public accession for it or that it's on the 10x website, and I wanted to confirm that.",
      "votes": 2
    },
    {
      "id": 2003434,
      "postDate": "2022-10-25T13:46:35.883Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1918940,
      "author_name": "Jiwei Liu",
      "author_url": "",
      "post_date": "2022-08-30T02:04:48.853000",
      "content": "<p>Hi Daniel <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> ,  is it allowed to use the NeurIPS 2021 Single-cell multi-modality dataset as an external dataset for this competition? Thanks.</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 1900325,
      "author_name": "Alejo Keuro",
      "author_url": "",
      "post_date": "2022-08-15T22:57:30.327000",
      "content": "<p>Hi Daniel, thank you.</p>\n<p>Are you (I mean, the organizers of the competition) planning to release some notebook with basic code snippets to get some of us up and running?</p>\n<p>I tried to read one h5 file using muon (more specifically, this function: <a href=\"https://muon.readthedocs.io/en/latest/api/generated/muon.read_10x_h5.html#muon.read_10x_h5\" target=\"_blank\">read_10x_h5</a>) library but I got an error: \"OSError: Can't read data (can't open directory: /usr/local/hdf5/lib/plugin)\". As far as I could see, there is no such directory. Does this have to do with running these functions in a kaggle environment? </p>\n<p>Sorry if this is a basic question. Thank you in advance!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1900328,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-15T23:06:54.597000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1900335,
          "author_name": "Daniel Burkhardt",
          "author_url": "",
          "post_date": "2022-08-15T23:23:11.843000",
          "content": "<p><strong>Update: Notebook is now available as \"Getting Started - Data Loading\"</strong></p>\n<p>You beat us to it! This went live afterhours local time for us, but <a href=\"https://www.kaggle.com/peterholderrieth\" target=\"_blank\">@peterholderrieth</a> will post our code notebook later tonight.</p>\n<p>For now, I can point you towards the easiest way to load the data:</p>\n<pre><code>!pip install --quiet tables # required in the Kaggle environment\n\n# Imports\nimport os\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Data paths\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\n# Load metadata\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell\n\n# Load the data\ndf_cite_train_x = pd.read_hdf(FP_CITE_TRAIN_INPUTS)\ndf_cite_test_x = pd.read_hdf(FP_CITE_TEST_INPUTS)\ndf_cite_train_x.head()\n\ndf_cite_train_y = pd.read_hdf(FP_CITE_TRAIN_TARGETS)\ndf_cite_train_y.head()\n</code></pre>\n<p>More details to come later!</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1900364,
          "author_name": "Peter Holderrieth",
          "author_url": "",
          "post_date": "2022-08-16T00:24:03.737000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/alekeuro\" target=\"_blank\">@alekeuro</a>! I have uploaded a <a href=\"https://www.kaggle.com/code/peterholderrieth/getting-started-data-loading\" target=\"_blank\">notebook</a> that step by step loads and explains the structure of the data. I hope it is helpful to you! </p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1900369,
          "author_name": "Alejo Keuro",
          "author_url": "",
          "post_date": "2022-08-16T01:02:02.023000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/peterholderrieth\" target=\"_blank\">@peterholderrieth</a> and <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1903777,
          "author_name": "Alejo Keuro",
          "author_url": "",
          "post_date": "2022-08-17T16:39:35.763000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/peterholderrieth\" target=\"_blank\">@peterholderrieth</a> ,</p>\n<p>going through your Getting Started notebook I noticed that, according to the bar chart (cell counts per technology, day and donor), there are no cells for donor nr. 4 (donor id: 27678) on day 4, for multiome tech. This count comes from the metadata.csv file. </p>\n<p>On the other side, the figure illustrating the experimental setup shows that part of the Public Test data should include cells from donor 4, at day 4, for multiome tech.</p>\n<p>Is it possible that one of the two pieces of info is wrong? Either the metadata file or the figure? Sorry if this is a dumb question.</p>\n<p>Thanks in advance!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1903797,
          "author_name": "Daniel Burkhardt",
          "author_url": "",
          "post_date": "2022-08-17T16:55:47.153000",
          "content": "<p>Ahh this is a mistake in the figure! This dataset failed, and we decided it would be okay to use for public test because it would mean 3 time points for public test for each modality. I'll fix the image!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1904137,
          "author_name": "Alejo Keuro",
          "author_url": "",
          "post_date": "2022-08-18T01:57:56.647000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2030466,
      "author_name": "tarick.morty",
      "author_url": "",
      "post_date": "2022-11-15T12:48:47.890000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> , thank you for giving us a unique problem and even more unique datasets and hosting this competition. <br>\nNow that we are on the finish line of it, I wonder if you had any results in mind with respect to the correlation metric. <br>\nAnd whether the expectation has been met or not ?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2018551,
      "author_name": "TESUZI",
      "author_url": "",
      "post_date": "2022-11-05T19:43:00.977000",
      "content": "<p>Hi Daniel, I wonder if it is correct that we will access the private testing datasets after three days. Thanks a lot.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1992601,
      "author_name": "Alexander Chervov",
      "author_url": "",
      "post_date": "2022-10-17T19:07:57.130000",
      "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> <br>\nDaniel you shared somewhere in discussion topics the code how to get positions of the genes on DNA, to match the ATAC seq data.<br>\nI cannot find it now. <br>\nWould you be so kind to share it again, please ? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1992758,
          "author_name": "Daniel Burkhardt",
          "author_url": "",
          "post_date": "2022-10-17T22:11:51.097000",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349559\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349559</a></p>\n<p>Here you go! I tried putting this into a Kaggle notebook but ran into a memory error and didn't have time to troubleshoot.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1974000,
      "author_name": "yjyang",
      "author_url": "",
      "post_date": "2022-10-06T02:43:56.867000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> is it allowed to integrate any public database into the inference such as gene networks?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1960657,
      "author_name": "Nikolay",
      "author_url": "",
      "post_date": "2022-09-28T17:40:56.523000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> , thank you for this great competition. Just to confirm: the cells were NOT stimulated in any way after plating? It would be helpful, if you could describe a little bit more about what happened to the cells after they were obtained from the donors over those different time periods.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1960755,
          "author_name": "Daniel Burkhardt",
          "author_url": "",
          "post_date": "2022-09-28T18:35:23.943000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mxposed\" target=\"_blank\">@mxposed</a> the cells were cultured with StemSpan SFEM media supplemented with CC100 and thrombopoietin (TPO) over 10 days. Cells were incubated at 37ºC and media was changed every 2-3 days.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1946685,
      "author_name": "museas",
      "author_url": "",
      "post_date": "2022-09-20T03:25:15.637000",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> <br>\nI want to check the detailed correlation CV/LB.<br>\nCan I see one more digit LB?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1988300,
      "author_name": "QINFANG LU",
      "author_url": "",
      "post_date": "2022-10-15T08:14:00.213000",
      "content": "<p>Thank you for this nice information <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1943314,
      "author_name": "agenlu",
      "author_url": "",
      "post_date": "2022-09-17T12:55:25.330000",
      "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a>, sorry if this has been asked on other part. What exactly mean 0 on gene expression, is not meassured or something else?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1945877,
          "author_name": "Daniel Burkhardt",
          "author_url": "",
          "post_date": "2022-09-19T13:16:55.367000",
          "content": "<p>Think of the data like counts from a Poisson process. It means that for that cell, we have 0 counts for that gene. It doesn't mean necessarily that the gene doesn't exist in that cell, just that we didn't observe any counts when we measured the cell.</p>\n<p>Note, this is not the same as \"missing data\" in a matrix completion sense. We can only capture about 30-50% of the mRNA in a cell, so there will be some molecules we don't sample when we're counting.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1945990,
          "author_name": "agenlu",
          "author_url": "",
          "post_date": "2022-09-19T14:44:38.750000",
          "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a>, and there are not other technical factors that can drive the count to 0? Or in other words, can we try to denoise the 0?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1947322,
          "author_name": "Daniel Burkhardt",
          "author_url": "",
          "post_date": "2022-09-20T12:13:43.867000",
          "content": "<p>There certainly are approaches to denoising single-cell data. Check out these review articles:</p>\n<ul>\n<li><a href=\"https://pubmed.ncbi.nlm.nih.gov/33003202/\" target=\"_blank\">A review of computational strategies for denoising and imputation of single-cell transcriptomic data</a></li>\n<li><a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-02132-x\" target=\"_blank\">A systematic evaluation of single-cell RNA-sequencing imputation methods</a></li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1947375,
          "author_name": "agenlu",
          "author_url": "",
          "post_date": "2022-09-20T12:36:28.963000",
          "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a>, thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1983856,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-10-12T09:40:42.180000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1984172,
          "author_name": "Daniel Burkhardt",
          "author_url": "",
          "post_date": "2022-10-12T13:54:32.767000",
          "content": "<p><a href=\"https://www.kaggle.com/ryumei\" target=\"_blank\">@ryumei</a>, when predicting one modality from another, there will often be zero values in the predicted targets because, for example, some genes are not expressed. Of the 20,000 protein coding genes in the genome, we often see 1k-5k have non-zero values in each cell.</p>\n<p>You should think of the predicted counts as originating from a Poisson counting process. We do not count all the transcripts in each cell, so there will be some zero values.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1984193,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-10-12T14:02:11.087000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1904185,
      "author_name": "SARA KNAACK",
      "author_url": "",
      "post_date": "2022-08-18T03:10:43.073000",
      "content": "<p>Quick question: Is the data publicly available in an unprocessed (.bam or .fastq) form? Or only in the processed form offered? I only ask as I don't detect a public accession for it or that it's on the 10x website, and I wanted to confirm that.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2003434,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-10-25T13:46:35.883000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1900287": "Dear Kagglers,\n\nOn behalf of Open Problem in Single-Cell Analysis, I would like to officially welcome you to the “Open Problems - Multimodal Single-Cell Integration” competition, part of the [NeurIPS 2022 Competition Track](https://neurips.cc/Conferences/2022/CompetitionTrack)!\n\nThis competition is about predicting the relationship between different genetic measurements (DNA, RNA, and protein) measured simultaneously in single cells. In the field of biomedicine, single-cell technologies have led to an explosion of new findings about the diversity of the 37 trillion cells in the human body and how these cell types vary between health and disease. However, the field still has several grand challenges ahead of it, and modeling dynamics of cell state is one of the biggest open problems ([Genome Biology 2020](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-1926-6)).\n\nFor this competition, Open Problems partnered with Cellarity, a cell-centric drug creation company, to generate a first-of-its-kind benchmarking dataset designed to drive advances in algorithms that can capture drivers of changes in cell state over time. We measured CD34+ hematopoietic stem and progenitor cells (HSPCs) from four healthy human donors at 5 time points using two different multimodal single-cell technologies. Your challenge will be to learn the relationship between these modalities and learn how to predict from one to another at a later unseen time point in the datasets.\n\nIn the Data section, you’ll find descriptions of the features in each modality and explanations of some metadata associated with the observations, which each correspond to a single cell. Please feel free to ask us any questions!\n\nIn our NeurIPS competition last year ([link](https://openproblems.bio/neurips_2021/)), we were thrilled to see contributions from teams that specialize in genomics compete neck and neck with algorithms originally designed for completely different domains, such as image captioning. We can’t wait to see what solutions are proposed in this competition.\n\nThis competition was a collaborative effort between scientists at Cellarity, the Chan Zuckerberg Biohub, Yale University, and Helmholtz Munich with sponsorship from Cellarity and the Chan Zuckerberg Initiative's Single-Cell Biology Program. \n\nIf you’re interested to learn more about Open Problems in Single-Cell Analysis, please visit our homepage, https://openproblems.bio, where you can sign up for [our mailing list](https://docs.google.com/forms/d/e/1FAIpQLSe90Oky4-1b0HbdLsp5Yqo9juCd2mq-NlGHU9NHRW1ECok1xQ/viewform?usp=sf_link).\n\nBest of luck! We can’t wait to see what you build.\nDaniel Burkhardt and the rest of the Core Team at Open Problems\n",
    "1918940": "Hi Daniel @danielburkhardt ,  is it allowed to use the NeurIPS 2021 Single-cell multi-modality dataset as an external dataset for this competition? Thanks.",
    "1900325": "Hi Daniel, thank you.\n\nAre you (I mean, the organizers of the competition) planning to release some notebook with basic code snippets to get some of us up and running?\n\nI tried to read one h5 file using muon (more specifically, this function: [read_10x_h5](https://muon.readthedocs.io/en/latest/api/generated/muon.read_10x_h5.html#muon.read_10x_h5)) library but I got an error: \"OSError: Can't read data (can't open directory: /usr/local/hdf5/lib/plugin)\". As far as I could see, there is no such directory. Does this have to do with running these functions in a kaggle environment? \n\nSorry if this is a basic question. Thank you in advance!",
    "2030466": "Hi @danielburkhardt , thank you for giving us a unique problem and even more unique datasets and hosting this competition. \nNow that we are on the finish line of it, I wonder if you had any results in mind with respect to the correlation metric. \nAnd whether the expectation has been met or not ?",
    "2018551": "Hi Daniel, I wonder if it is correct that we will access the private testing datasets after three days. Thanks a lot.",
    "1992601": "@danielburkhardt \nDaniel you shared somewhere in discussion topics the code how to get positions of the genes on DNA, to match the ATAC seq data.\nI cannot find it now. \nWould you be so kind to share it again, please ? ",
    "1974000": "Hi @danielburkhardt is it allowed to integrate any public database into the inference such as gene networks?",
    "1960657": "Hi @danielburkhardt , thank you for this great competition. Just to confirm: the cells were NOT stimulated in any way after plating? It would be helpful, if you could describe a little bit more about what happened to the cells after they were obtained from the donors over those different time periods.",
    "1946685": "Hi, @danielburkhardt \nI want to check the detailed correlation CV/LB.\nCan I see one more digit LB?",
    "1988300": "Thank you for this nice information @danielburkhardt ",
    "1943314": "@danielburkhardt, sorry if this has been asked on other part. What exactly mean 0 on gene expression, is not meassured or something else?",
    "1904185": "Quick question: Is the data publicly available in an unprocessed (.bam or .fastq) form? Or only in the processed form offered? I only ask as I don't detect a public accession for it or that it's on the 10x website, and I wanted to confirm that.",
    "2003434": ""
  }
}