{
  "id": 359355,
  "title": "Raw Counts Dataset",
  "url": "/competitions/open-problems-multimodal/discussion/359355",
  "author_name": "Ryan Holbrook",
  "post_date": "2022-10-11T18:09:16.272000",
  "votes": 32,
  "comment_count": 23,
  "views": 0,
  "content": "<p>The unnormalized raw counts data are now available here: <a href=\"https://www.kaggle.com/datasets/ryanholbrook/open-problems-raw-counts\" target=\"_blank\">Open Problems Raw Counts</a>. Please be aware that there are a few differences in the cells and genes represented compared to the competition dataset. Refer to the dataset documentation for more info.</p>\n<p>Edited to add: We're providing this new data as an <strong>optional</strong> supplement, something you can chose to use or not just as you find it helpful. Nothing in this supplementary data supersedes the original data. <strong>For scoring purposes, the only thing that counts is the original test set.</strong></p>",
  "messages": [
    {
      "id": 1982976,
      "postDate": "2022-10-11T18:09:16.273Z",
      "content": "<p>The unnormalized raw counts data are now available here: <a href=\"https://www.kaggle.com/datasets/ryanholbrook/open-problems-raw-counts\" target=\"_blank\">Open Problems Raw Counts</a>. Please be aware that there are a few differences in the cells and genes represented compared to the competition dataset. Refer to the dataset documentation for more info.</p>\n<p>Edited to add: We're providing this new data as an <strong>optional</strong> supplement, something you can chose to use or not just as you find it helpful. Nothing in this supplementary data supersedes the original data. <strong>For scoring purposes, the only thing that counts is the original test set.</strong></p>",
      "rawMarkdown": "The unnormalized raw counts data are now available here: [Open Problems Raw Counts](https://www.kaggle.com/datasets/ryanholbrook/open-problems-raw-counts). Please be aware that there are a few differences in the cells and genes represented compared to the competition dataset. Refer to the dataset documentation for more info.\n\nEdited to add: We're providing this new data as an **optional** supplement, something you can chose to use or not just as you find it helpful. Nothing in this supplementary data supersedes the original data. **For scoring purposes, the only thing that counts is the original test set.**",
      "votes": 32
    },
    {
      "id": 1983788,
      "postDate": "2022-10-12T08:30:39.713Z",
      "content": "<p>I don't know if I should be happy or sad to see that with less than one month left 🤒</p>\n<blockquote>\n  <p>Please be aware that there are a few differences in the cells and genes represented compared to the competition dataset. </p>\n</blockquote>",
      "rawMarkdown": "I don't know if I should be happy or sad to see that with less than one month left 🤒\n\n>  Please be aware that there are a few differences in the cells and genes represented compared to the competition dataset. ",
      "votes": 8
    },
    {
      "id": 2002541,
      "postDate": "2022-10-24T22:18:05.187Z",
      "content": "<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> There are a few steps in normalizing single cell seq data. I recommend using scanpy and following a tutorial like this one<br>\n<a href=\"https://scanpy-tutorials.readthedocs.io/en/latest/pbmc3k.html\" target=\"_blank\">https://scanpy-tutorials.readthedocs.io/en/latest/pbmc3k.html</a></p>",
      "rawMarkdown": "@senkin13 There are a few steps in normalizing single cell seq data. I recommend using scanpy and following a tutorial like this one\nhttps://scanpy-tutorials.readthedocs.io/en/latest/pbmc3k.html",
      "votes": 3,
      "replies": [
        {
          "id": 2002834,
          "postDate": "2022-10-25T05:40:30.323Z",
          "content": "<p>thanks, I got answer at this topic.<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360384\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360384</a></p>",
          "rawMarkdown": "thanks, I got answer at this topic.\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/360384",
          "votes": 1
        }
      ]
    },
    {
      "id": 1991437,
      "postDate": "2022-10-17T06:47:34.413Z",
      "content": "<p>Could you share how to transform row counts to provided normalized data in case of some teams know and others don't know.</p>",
      "rawMarkdown": "Could you share how to transform row counts to provided normalized data in case of some teams know and others don't know.",
      "votes": 3
    },
    {
      "id": 1996632,
      "postDate": "2022-10-20T10:41:13.863Z",
      "content": "<p>Hi, I think the raw train_multi_inputs is missing these ids:</p>\n<p>['77f077577f2b', 'b72c29d81777', '90c3c130728a', 'd5d7ed0cd9e0', '3325ba9931a4', 'adddbf5d12c2', '25bea866420d', <br>\n            '9804de13955e', '53bbcdf5cdb7', '5d30a38644a9', 'a170dc996a5c', '557821c73b37', 'de42455188ab', '9e628c6836dc',<br>\n            '8db4d231c90b', '2f8fc6a3c184', '3e6e2f98402a', '9f638d00a1e5', 'ff227e491cf5', '3a88ca134400', '573615970142',<br>\n            '29f691009d41', '24d70cb2af08', '22d4e906688e', '59757e56ae23', '54bd423decb0', '42dc1179382e', '7fb6a3050cc2', <br>\n            '702dec6be509', '3119e20f44de', '4d446228582f', '05f411cb69c4', 'b6980ce11761', '6054745d9caa', '48b969f009a2', <br>\n            '8e182af299f6', '08896c8cb420', '1aea9210034a', 'e3f0104b87dc', '027169483f73', '7d81dcd6b061', '2f8d45519d8a', <br>\n            '924880917f17', '69c6065afe9f', 'cc0e9902a477', '22567819a12d', 'b19cd014f0b1', '1e923a795d1b', '3f7dba9a2be2', <br>\n            '73e898bb17d5', '7c823af4c52b', 'b9b2a347bf73', 'ca730bd9cda0', 'aab665161bfb', 'c520321c0153', '78e8b2382acf',<br>\n            'b9f24b3fc48f', '0dfcb0dc4a80', '2c85b2857ba1', '2d40c5167f95', '9b4e6d250f3b', '22126d9ae1d7', 'a84cf34d4067', <br>\n            '19946ba5a73f', '47d3b6aa6e54', '911fb1d58afc', '26d30c892830', '4659d775e873', '6102ed065946', '5edd6b883772',<br>\n            'a785258a125b', '2659c4132b6a', '51427c01b10f', '6e78533ee708']</p>\n<p>with respect to the normalized counts, can you please double-check?</p>",
      "rawMarkdown": "Hi, I think the raw train_multi_inputs is missing these ids:\n\n ['77f077577f2b', 'b72c29d81777', '90c3c130728a', 'd5d7ed0cd9e0', '3325ba9931a4', 'adddbf5d12c2', '25bea866420d', \n            '9804de13955e', '53bbcdf5cdb7', '5d30a38644a9', 'a170dc996a5c', '557821c73b37', 'de42455188ab', '9e628c6836dc',\n            '8db4d231c90b', '2f8fc6a3c184', '3e6e2f98402a', '9f638d00a1e5', 'ff227e491cf5', '3a88ca134400', '573615970142',\n            '29f691009d41', '24d70cb2af08', '22d4e906688e', '59757e56ae23', '54bd423decb0', '42dc1179382e', '7fb6a3050cc2', \n            '702dec6be509', '3119e20f44de', '4d446228582f', '05f411cb69c4', 'b6980ce11761', '6054745d9caa', '48b969f009a2', \n            '8e182af299f6', '08896c8cb420', '1aea9210034a', 'e3f0104b87dc', '027169483f73', '7d81dcd6b061', '2f8d45519d8a', \n            '924880917f17', '69c6065afe9f', 'cc0e9902a477', '22567819a12d', 'b19cd014f0b1', '1e923a795d1b', '3f7dba9a2be2', \n            '73e898bb17d5', '7c823af4c52b', 'b9b2a347bf73', 'ca730bd9cda0', 'aab665161bfb', 'c520321c0153', '78e8b2382acf',\n            'b9f24b3fc48f', '0dfcb0dc4a80', '2c85b2857ba1', '2d40c5167f95', '9b4e6d250f3b', '22126d9ae1d7', 'a84cf34d4067', \n            '19946ba5a73f', '47d3b6aa6e54', '911fb1d58afc', '26d30c892830', '4659d775e873', '6102ed065946', '5edd6b883772',\n            'a785258a125b', '2659c4132b6a', '51427c01b10f', '6e78533ee708']\n\nwith respect to the normalized counts, can you please double-check?",
      "votes": 1
    },
    {
      "id": 1990576,
      "postDate": "2022-10-16T16:36:03.357Z",
      "content": "<p>Hello, Ryan!<br>\nI found that original <strong>test_cite_inputs.h5</strong> contains <strong>48.663 lines</strong>.<br>\nBut <strong>test_cite_inputs_raw.h5</strong> contains only <strong>48.203 lines</strong>.<br>\nCan we get 460 lines more?</p>",
      "rawMarkdown": "Hello, Ryan!\nI found that original **test_cite_inputs.h5** contains **48.663 lines**.\nBut **test_cite_inputs_raw.h5** contains only **48.203 lines**.\nCan we get 460 lines more?",
      "votes": 1,
      "replies": [
        {
          "id": 1992000,
          "postDate": "2022-10-17T13:37:52.033Z",
          "content": "<p>Hi! I found that the missing data in the raw data all belong to day 2 results of donor 27678 which are not evaluated in the LB, so they can be filled with 0 or anything you want. I did something list below and the LB result turned out that there was nothing wrong. You may check it out.</p>\n<pre><code>test_ori = pd.read_hdf(\"../test_cite_inputs.h5\")\ntest_indexes = test_ori.index\ntest_ori.shape\n# (48663, 22050)\n\ntest = pd.read_hdf(\"./test_cite_inputs_raw.h5\")\ntest.shape\n# (48203, 22085)\n\ntest = test.drop_duplicates()  # This step might not be necessary.\ntest = test.reindex(test_indexes)\ntest = test.fillna(0)\ntest.shape\n# (48663, 22085)\n</code></pre>",
          "rawMarkdown": "Hi! I found that the missing data in the raw data all belong to day 2 results of donor 27678 which are not evaluated in the LB, so they can be filled with 0 or anything you want. I did something list below and the LB result turned out that there was nothing wrong. You may check it out.\n\n\n``` python\ntest_ori = pd.read_hdf(\"../test_cite_inputs.h5\")\ntest_indexes = test_ori.index\ntest_ori.shape\n# (48663, 22050)\n\ntest = pd.read_hdf(\"./test_cite_inputs_raw.h5\")\ntest.shape\n# (48203, 22085)\n\ntest = test.drop_duplicates()  # This step might not be necessary.\ntest = test.reindex(test_indexes)\ntest = test.fillna(0)\ntest.shape\n# (48663, 22085)\n```\n ",
          "votes": 7
        }
      ]
    },
    {
      "id": 1983809,
      "postDate": "2022-10-12T08:48:47.560Z",
      "content": "<p>more data, more work 🤕</p>",
      "rawMarkdown": "more data, more work 🤕",
      "votes": 1
    },
    {
      "id": 2013864,
      "postDate": "2022-11-02T06:58:02.757Z",
      "content": "<p>nice work hope this will have a better output</p>",
      "rawMarkdown": "nice work hope this will have a better output",
      "votes": -1
    },
    {
      "id": 2006355,
      "postDate": "2022-10-27T14:33:54.837Z",
      "content": "<p>You did awesome，I get it</p>",
      "rawMarkdown": "You did awesome，I get it",
      "votes": -1
    },
    {
      "id": 2033337,
      "postDate": "2022-11-17T07:17:30.847Z",
      "content": "<p>Thans a lot~</p>",
      "rawMarkdown": "Thans a lot~"
    },
    {
      "id": 2020876,
      "postDate": "2022-11-07T20:33:23.657Z",
      "content": "<p>How can we identify empty droplets in the cite raw count targets?</p>",
      "rawMarkdown": "How can we identify empty droplets in the cite raw count targets?"
    },
    {
      "id": 1992902,
      "postDate": "2022-10-18T02:27:48.743Z",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Hello, the data columns added by cite are not used to participate in the scoring, right? </p>",
      "rawMarkdown": "@ryanholbrook Hello, the data columns added by cite are not used to participate in the scoring, right? ",
      "replies": [
        {
          "id": 1993484,
          "postDate": "2022-10-18T12:18:13.107Z",
          "content": "<p>We're providing this new data as a supplement, something you can chose to use or not just as you find it helpful. For scoring purposes, the only thing that counts is the original test set.</p>",
          "rawMarkdown": "We're providing this new data as a supplement, something you can chose to use or not just as you find it helpful. For scoring purposes, the only thing that counts is the original test set.",
          "votes": 1
        },
        {
          "id": 1995682,
          "postDate": "2022-10-19T18:52:59.367Z",
          "content": "<p>Thank you Ryan! Are the raw counts from the unfiltered or filtered feature barcode matrix files? </p>",
          "rawMarkdown": "Thank you Ryan! Are the raw counts from the unfiltered or filtered feature barcode matrix files? "
        }
      ]
    },
    {
      "id": 1989634,
      "postDate": "2022-10-16T05:44:56.980Z",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Hi! Thanks a lot for your efforts! But there are still some poblems confusing me and hope you can answer them. Thanks! </p>\n<ol>\n<li><p>Is the LB still based on the normalized the test (target) data?</p></li>\n<li><p>Why the protein levels (cite targets) of some cells (such as 8fb1ee60ecd8) are zero in the raw data while they are not in the normalized data? Is that because dsb normalizations are not carried out cell_wise (rows_wise) or due to any other reasons? </p></li>\n</ol>",
      "rawMarkdown": "@ryanholbrook Hi! Thanks a lot for your efforts! But there are still some poblems confusing me and hope you can answer them. Thanks! \n\n1. Is the LB still based on the normalized the test (target) data?\n\n2. Why the protein levels (cite targets) of some cells (such as 8fb1ee60ecd8) are zero in the raw data while they are not in the normalized data? Is that because dsb normalizations are not carried out cell_wise (rows_wise) or due to any other reasons? "
    },
    {
      "id": 1988031,
      "postDate": "2022-10-15T03:03:56.863Z",
      "content": "<p>that‘s great</p>",
      "rawMarkdown": "that‘s great"
    },
    {
      "id": 1984188,
      "postDate": "2022-10-12T13:58:54.997Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2010673,
      "postDate": "2022-10-31T02:27:51.960Z",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Awesome, thanks!</p>",
      "rawMarkdown": "@ryanholbrook Awesome, thanks!",
      "votes": -2
    },
    {
      "id": 2006000,
      "postDate": "2022-10-27T11:45:28.237Z",
      "content": "<p>thank you ,I get it</p>",
      "rawMarkdown": "thank you ,I get it",
      "votes": -1
    },
    {
      "id": 1993296,
      "postDate": "2022-10-18T09:02:58.190Z",
      "content": "<p>cool,Thanks a lot for your efforts! </p>",
      "rawMarkdown": "cool,Thanks a lot for your efforts! ",
      "votes": -6
    },
    {
      "id": 1989744,
      "postDate": "2022-10-16T06:43:18.190Z",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> thanks</p>",
      "rawMarkdown": "@ryanholbrook thanks"
    },
    {
      "id": 1983703,
      "postDate": "2022-10-12T07:33:28.473Z",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> , thanks</p>",
      "rawMarkdown": "@ryanholbrook , thanks"
    }
  ],
  "comments": [
    {
      "id": 1983788,
      "author_name": "Riccardo",
      "author_url": "",
      "post_date": "2022-10-12T08:30:39.713000",
      "content": "<p>I don't know if I should be happy or sad to see that with less than one month left 🤒</p>\n<blockquote>\n  <p>Please be aware that there are a few differences in the cells and genes represented compared to the competition dataset. </p>\n</blockquote>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 2002541,
      "author_name": "Sean Murphy",
      "author_url": "",
      "post_date": "2022-10-24T22:18:05.187000",
      "content": "<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> There are a few steps in normalizing single cell seq data. I recommend using scanpy and following a tutorial like this one<br>\n<a href=\"https://scanpy-tutorials.readthedocs.io/en/latest/pbmc3k.html\" target=\"_blank\">https://scanpy-tutorials.readthedocs.io/en/latest/pbmc3k.html</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2002834,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2022-10-25T05:40:30.323000",
          "content": "<p>thanks, I got answer at this topic.<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360384\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360384</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1991437,
      "author_name": "senkin13",
      "author_url": "",
      "post_date": "2022-10-17T06:47:34.413000",
      "content": "<p>Could you share how to transform row counts to provided normalized data in case of some teams know and others don't know.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1996632,
      "author_name": "MT",
      "author_url": "",
      "post_date": "2022-10-20T10:41:13.863000",
      "content": "<p>Hi, I think the raw train_multi_inputs is missing these ids:</p>\n<p>['77f077577f2b', 'b72c29d81777', '90c3c130728a', 'd5d7ed0cd9e0', '3325ba9931a4', 'adddbf5d12c2', '25bea866420d', <br>\n            '9804de13955e', '53bbcdf5cdb7', '5d30a38644a9', 'a170dc996a5c', '557821c73b37', 'de42455188ab', '9e628c6836dc',<br>\n            '8db4d231c90b', '2f8fc6a3c184', '3e6e2f98402a', '9f638d00a1e5', 'ff227e491cf5', '3a88ca134400', '573615970142',<br>\n            '29f691009d41', '24d70cb2af08', '22d4e906688e', '59757e56ae23', '54bd423decb0', '42dc1179382e', '7fb6a3050cc2', <br>\n            '702dec6be509', '3119e20f44de', '4d446228582f', '05f411cb69c4', 'b6980ce11761', '6054745d9caa', '48b969f009a2', <br>\n            '8e182af299f6', '08896c8cb420', '1aea9210034a', 'e3f0104b87dc', '027169483f73', '7d81dcd6b061', '2f8d45519d8a', <br>\n            '924880917f17', '69c6065afe9f', 'cc0e9902a477', '22567819a12d', 'b19cd014f0b1', '1e923a795d1b', '3f7dba9a2be2', <br>\n            '73e898bb17d5', '7c823af4c52b', 'b9b2a347bf73', 'ca730bd9cda0', 'aab665161bfb', 'c520321c0153', '78e8b2382acf',<br>\n            'b9f24b3fc48f', '0dfcb0dc4a80', '2c85b2857ba1', '2d40c5167f95', '9b4e6d250f3b', '22126d9ae1d7', 'a84cf34d4067', <br>\n            '19946ba5a73f', '47d3b6aa6e54', '911fb1d58afc', '26d30c892830', '4659d775e873', '6102ed065946', '5edd6b883772',<br>\n            'a785258a125b', '2659c4132b6a', '51427c01b10f', '6e78533ee708']</p>\n<p>with respect to the normalized counts, can you please double-check?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1990576,
      "author_name": "Oleg Khudyakov",
      "author_url": "",
      "post_date": "2022-10-16T16:36:03.357000",
      "content": "<p>Hello, Ryan!<br>\nI found that original <strong>test_cite_inputs.h5</strong> contains <strong>48.663 lines</strong>.<br>\nBut <strong>test_cite_inputs_raw.h5</strong> contains only <strong>48.203 lines</strong>.<br>\nCan we get 460 lines more?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1992000,
          "author_name": "Oliver Wang",
          "author_url": "",
          "post_date": "2022-10-17T13:37:52.033000",
          "content": "<p>Hi! I found that the missing data in the raw data all belong to day 2 results of donor 27678 which are not evaluated in the LB, so they can be filled with 0 or anything you want. I did something list below and the LB result turned out that there was nothing wrong. You may check it out.</p>\n<pre><code>test_ori = pd.read_hdf(\"../test_cite_inputs.h5\")\ntest_indexes = test_ori.index\ntest_ori.shape\n# (48663, 22050)\n\ntest = pd.read_hdf(\"./test_cite_inputs_raw.h5\")\ntest.shape\n# (48203, 22085)\n\ntest = test.drop_duplicates()  # This step might not be necessary.\ntest = test.reindex(test_indexes)\ntest = test.fillna(0)\ntest.shape\n# (48663, 22085)\n</code></pre>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 1983809,
      "author_name": "Duc Nguyen",
      "author_url": "",
      "post_date": "2022-10-12T08:48:47.560000",
      "content": "<p>more data, more work 🤕</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2013864,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-02T06:58:02.757000",
      "content": "<p>nice work hope this will have a better output</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2006355,
      "author_name": "aoyuuii",
      "author_url": "",
      "post_date": "2022-10-27T14:33:54.837000",
      "content": "<p>You did awesome，I get it</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2033337,
      "author_name": "zzh",
      "author_url": "",
      "post_date": "2022-11-17T07:17:30.847000",
      "content": "<p>Thans a lot~</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2020876,
      "author_name": "Nat",
      "author_url": "",
      "post_date": "2022-11-07T20:33:23.657000",
      "content": "<p>How can we identify empty droplets in the cite raw count targets?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1992902,
      "author_name": "蒹葭",
      "author_url": "",
      "post_date": "2022-10-18T02:27:48.743000",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Hello, the data columns added by cite are not used to participate in the scoring, right? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1993484,
          "author_name": "Ryan Holbrook",
          "author_url": "",
          "post_date": "2022-10-18T12:18:13.107000",
          "content": "<p>We're providing this new data as a supplement, something you can chose to use or not just as you find it helpful. For scoring purposes, the only thing that counts is the original test set.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1995682,
          "author_name": "Insiya Jafferji",
          "author_url": "",
          "post_date": "2022-10-19T18:52:59.367000",
          "content": "<p>Thank you Ryan! Are the raw counts from the unfiltered or filtered feature barcode matrix files? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1989634,
      "author_name": "Oliver Wang",
      "author_url": "",
      "post_date": "2022-10-16T05:44:56.980000",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Hi! Thanks a lot for your efforts! But there are still some poblems confusing me and hope you can answer them. Thanks! </p>\n<ol>\n<li><p>Is the LB still based on the normalized the test (target) data?</p></li>\n<li><p>Why the protein levels (cite targets) of some cells (such as 8fb1ee60ecd8) are zero in the raw data while they are not in the normalized data? Is that because dsb normalizations are not carried out cell_wise (rows_wise) or due to any other reasons? </p></li>\n</ol>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1988031,
      "author_name": "Kuz Peng",
      "author_url": "",
      "post_date": "2022-10-15T03:03:56.863000",
      "content": "<p>that‘s great</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1984188,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-10-12T13:58:54.997000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2010673,
      "author_name": "Zhilin Jin",
      "author_url": "",
      "post_date": "2022-10-31T02:27:51.960000",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Awesome, thanks!</p>",
      "votes": -2,
      "replies": []
    },
    {
      "id": 2006000,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-10-27T11:45:28.237000",
      "content": "<p>thank you ,I get it</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 1993296,
      "author_name": "wangyueping1115",
      "author_url": "",
      "post_date": "2022-10-18T09:02:58.190000",
      "content": "<p>cool,Thanks a lot for your efforts! </p>",
      "votes": -6,
      "replies": []
    },
    {
      "id": 1989744,
      "author_name": "thj2333",
      "author_url": "",
      "post_date": "2022-10-16T06:43:18.190000",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1983703,
      "author_name": "agenlu",
      "author_url": "",
      "post_date": "2022-10-12T07:33:28.473000",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> , thanks</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1982976": "The unnormalized raw counts data are now available here: [Open Problems Raw Counts](https://www.kaggle.com/datasets/ryanholbrook/open-problems-raw-counts). Please be aware that there are a few differences in the cells and genes represented compared to the competition dataset. Refer to the dataset documentation for more info.\n\nEdited to add: We're providing this new data as an **optional** supplement, something you can chose to use or not just as you find it helpful. Nothing in this supplementary data supersedes the original data. **For scoring purposes, the only thing that counts is the original test set.**",
    "1983788": "I don't know if I should be happy or sad to see that with less than one month left 🤒\n\n>  Please be aware that there are a few differences in the cells and genes represented compared to the competition dataset. ",
    "2002541": "@senkin13 There are a few steps in normalizing single cell seq data. I recommend using scanpy and following a tutorial like this one\nhttps://scanpy-tutorials.readthedocs.io/en/latest/pbmc3k.html",
    "1991437": "Could you share how to transform row counts to provided normalized data in case of some teams know and others don't know.",
    "1996632": "Hi, I think the raw train_multi_inputs is missing these ids:\n\n ['77f077577f2b', 'b72c29d81777', '90c3c130728a', 'd5d7ed0cd9e0', '3325ba9931a4', 'adddbf5d12c2', '25bea866420d', \n            '9804de13955e', '53bbcdf5cdb7', '5d30a38644a9', 'a170dc996a5c', '557821c73b37', 'de42455188ab', '9e628c6836dc',\n            '8db4d231c90b', '2f8fc6a3c184', '3e6e2f98402a', '9f638d00a1e5', 'ff227e491cf5', '3a88ca134400', '573615970142',\n            '29f691009d41', '24d70cb2af08', '22d4e906688e', '59757e56ae23', '54bd423decb0', '42dc1179382e', '7fb6a3050cc2', \n            '702dec6be509', '3119e20f44de', '4d446228582f', '05f411cb69c4', 'b6980ce11761', '6054745d9caa', '48b969f009a2', \n            '8e182af299f6', '08896c8cb420', '1aea9210034a', 'e3f0104b87dc', '027169483f73', '7d81dcd6b061', '2f8d45519d8a', \n            '924880917f17', '69c6065afe9f', 'cc0e9902a477', '22567819a12d', 'b19cd014f0b1', '1e923a795d1b', '3f7dba9a2be2', \n            '73e898bb17d5', '7c823af4c52b', 'b9b2a347bf73', 'ca730bd9cda0', 'aab665161bfb', 'c520321c0153', '78e8b2382acf',\n            'b9f24b3fc48f', '0dfcb0dc4a80', '2c85b2857ba1', '2d40c5167f95', '9b4e6d250f3b', '22126d9ae1d7', 'a84cf34d4067', \n            '19946ba5a73f', '47d3b6aa6e54', '911fb1d58afc', '26d30c892830', '4659d775e873', '6102ed065946', '5edd6b883772',\n            'a785258a125b', '2659c4132b6a', '51427c01b10f', '6e78533ee708']\n\nwith respect to the normalized counts, can you please double-check?",
    "1990576": "Hello, Ryan!\nI found that original **test_cite_inputs.h5** contains **48.663 lines**.\nBut **test_cite_inputs_raw.h5** contains only **48.203 lines**.\nCan we get 460 lines more?",
    "1983809": "more data, more work 🤕",
    "2013864": "nice work hope this will have a better output",
    "2006355": "You did awesome，I get it",
    "2033337": "Thans a lot~",
    "2020876": "How can we identify empty droplets in the cite raw count targets?",
    "1992902": "@ryanholbrook Hello, the data columns added by cite are not used to participate in the scoring, right? ",
    "1989634": "@ryanholbrook Hi! Thanks a lot for your efforts! But there are still some poblems confusing me and hope you can answer them. Thanks! \n\n1. Is the LB still based on the normalized the test (target) data?\n\n2. Why the protein levels (cite targets) of some cells (such as 8fb1ee60ecd8) are zero in the raw data while they are not in the normalized data? Is that because dsb normalizations are not carried out cell_wise (rows_wise) or due to any other reasons? ",
    "1988031": "that‘s great",
    "1984188": "",
    "2010673": "@ryanholbrook Awesome, thanks!",
    "2006000": "thank you ,I get it",
    "1993296": "cool,Thanks a lot for your efforts! ",
    "1989744": "@ryanholbrook thanks",
    "1983703": "@ryanholbrook , thanks"
  }
}