{
  "id": 372609,
  "title": "Competition Wrap Up Workshop + Last Call for Survey!",
  "url": "/competitions/open-problems-multimodal/discussion/372609",
  "author_name": "",
  "post_date": "2022-12-16T21:19:44.484373200Z",
  "votes": 11,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi all, once again, we at Open Problems are thrilled with the turn out of this year's competition. I hope it's been a productive experience for you, and if this was your first time working with single-cell data, we hope you'll continue to engage with the field.</p>\n<p>Last week, we hosted a virtual workshop as part of NeurIPS 2022, and for those of you who could make it, we're happy to post a recording of the meeting on YouTube: <a href=\"https://www.youtube.com/embed/WmkbyGMPHOE\" target=\"_blank\">https://www.youtube.com/embed/WmkbyGMPHOE</a></p>\n<p>We learned a lot from this competition, and we're still going through the results as we get ready to prepare a submission to <em>Nature Methods</em>. We're still looking for survey results, so if you haven't filled out the survey, please take 10 minutes to do so here: <a href=\"https://cellarity.qualtrics.com/jfe/form/SV_0jOnimbJmxiP0Z8\" target=\"_blank\">https://cellarity.qualtrics.com/jfe/form/SV_0jOnimbJmxiP0Z8</a></p>\n<p>So what did we find so far? </p>\n<h2>We appear to have hit a limit in predictions</h2>\n<p>One of the most interesting results of the competition that was immediately apparent was the similarity between the top predictions in the leaderboard. On the private leaderboard, 584 / 1221 teams with were within 0.01 correlation of the top performing team. We didn't expect that. Take for example the <a href=\"https://www.kaggle.com/competitions/recursion-cellular-image-classification/overview/prizes\" target=\"_blank\">Recursion Cellular Image Classification competition in 2019</a> where the 50th team's accuracy was within <code>0.05</code> of the top team or the <a href=\"https://www.kaggle.com/competitions/hpa-single-cell-image-classification\" target=\"_blank\">Human Protein Atlas Single-Cell Image Classification</a> competition in 2020 where where the 50th team's segmentation precision was within <code>0.14</code> of the top team. Compared the these tasks, the performance of the top teams was much more similar--at least by our metric.</p>\n<p>You can see this looking at the  follow plot adapted from <a href=\"https://www.kaggle.com/josecarmona\" target=\"_blank\">@josecarmona</a> graphical recap notebook.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308072%2F1cf659f47153f7000639d8f9d63d00ff%2FScreen%20Shot%202022-12-16%20at%203.30.59%20PM.png?generation=1671222686784466&amp;alt=media\" alt=\"\"></p>\n<h2>Top teams predictions were very similar overall</h2>\n<p>We also wondered: are the predictions themselves similar? To assess this, we looked at the top two teams multiome predictions, and summed the predicted expression values across all cells in the held-out test set. Looking at this we find the predictions overall are very similar per-gene with a Pearson R of 0.9986.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308072%2F219793d1906ed0de0741759890ba1a37%2Fcorrelations.png?generation=1671223463918411&amp;alt=media\" alt=\"\"></p>\n<p>We can also look at the top predictions comparing the correlation of all values for both technologies (this is taking the correlation of the raw prediction files) and find that this trend holds across the top 3 teams.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308072%2Fc6bf3c16d83afb39366b84c1eb33901a%2Fmore_corr.png?generation=1671223582290746&amp;alt=media\" alt=\"\"></p>\n<h2>Genes that are poorly predicted are enriched for mitochondrial RNA and ribosomal proteins</h2>\n<p>Finally, we looked at how well each gene was predicted in the test set for the top prediction. To assess this, we calculated the correlation between the predicted and observed values per gene across cells. Note this is the opposite orientation to how we set up the competition, and some genes may simply not be important for high performance when calculating correlation across genes per cell. </p>\n<p>What we first found was there's a surprising bimodal distribution in the correlation values across genes:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308072%2Fa1aba8583281e25a452e3ced5a684011%2FScreen%20Shot%202022-12-16%20at%203.57.38%20PM.png?generation=1671224374150990&amp;alt=media\" alt=\"\"></p>\n<p>We also look at some of the top and bottom genes by prediction performance, and find that the genes that are worst predicted are often mitochondrial RNA genes or ribosomal proteins, which we generally believe aren't highly relevant for explaining biological heterogeneity. </p>\n<h2>What's next?</h2>\n<p>We're still early in our process of investigating the results of the competition in preparation for writing a manuscript summarizing the competition in <em>Nature Methods</em>. </p>\n<p>If you'd like to help, please don't hesitate to reach out and keep posting in the forums / code section of the competition!</p>\n<p>We hope this data and the competition have been useful, and once again, thanks everyone for your participation! Thanks also to Ryan and Ashley from the Kaggle team, without whom this competition would not have run so smoothly.</p>\n<p>Best,<br>\nDaniel Burkhardt on behalf of the Open Problems Core Team</p>",
  "messages": [
    {
      "id": "2067590",
      "postDate": "12/16/2022 21:19:44",
      "content": "<p>Hi all, once again, we at Open Problems are thrilled with the turn out of this year's competition. I hope it's been a productive experience for you, and if this was your first time working with single-cell data, we hope you'll continue to engage with the field.</p>\n<p>Last week, we hosted a virtual workshop as part of NeurIPS 2022, and for those of you who could make it, we're happy to post a recording of the meeting on YouTube: <a href=\"https://www.youtube.com/embed/WmkbyGMPHOE\" target=\"_blank\">https://www.youtube.com/embed/WmkbyGMPHOE</a></p>\n<p>We learned a lot from this competition, and we're still going through the results as we get ready to prepare a submission to <em>Nature Methods</em>. We're still looking for survey results, so if you haven't filled out the survey, please take 10 minutes to do so here: <a href=\"https://cellarity.qualtrics.com/jfe/form/SV_0jOnimbJmxiP0Z8\" target=\"_blank\">https://cellarity.qualtrics.com/jfe/form/SV_0jOnimbJmxiP0Z8</a></p>\n<p>So what did we find so far? </p>\n<h2>We appear to have hit a limit in predictions</h2>\n<p>One of the most interesting results of the competition that was immediately apparent was the similarity between the top predictions in the leaderboard. On the private leaderboard, 584 / 1221 teams with were within 0.01 correlation of the top performing team. We didn't expect that. Take for example the <a href=\"https://www.kaggle.com/competitions/recursion-cellular-image-classification/overview/prizes\" target=\"_blank\">Recursion Cellular Image Classification competition in 2019</a> where the 50th team's accuracy was within <code>0.05</code> of the top team or the <a href=\"https://www.kaggle.com/competitions/hpa-single-cell-image-classification\" target=\"_blank\">Human Protein Atlas Single-Cell Image Classification</a> competition in 2020 where where the 50th team's segmentation precision was within <code>0.14</code> of the top team. Compared the these tasks, the performance of the top teams was much more similar--at least by our metric.</p>\n<p>You can see this looking at the  follow plot adapted from <a href=\"https://www.kaggle.com/josecarmona\" target=\"_blank\">@josecarmona</a> graphical recap notebook.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308072%2F1cf659f47153f7000639d8f9d63d00ff%2FScreen%20Shot%202022-12-16%20at%203.30.59%20PM.png?generation=1671222686784466&amp;alt=media\" alt=\"\"></p>\n<h2>Top teams predictions were very similar overall</h2>\n<p>We also wondered: are the predictions themselves similar? To assess this, we looked at the top two teams multiome predictions, and summed the predicted expression values across all cells in the held-out test set. Looking at this we find the predictions overall are very similar per-gene with a Pearson R of 0.9986.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308072%2F219793d1906ed0de0741759890ba1a37%2Fcorrelations.png?generation=1671223463918411&amp;alt=media\" alt=\"\"></p>\n<p>We can also look at the top predictions comparing the correlation of all values for both technologies (this is taking the correlation of the raw prediction files) and find that this trend holds across the top 3 teams.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308072%2Fc6bf3c16d83afb39366b84c1eb33901a%2Fmore_corr.png?generation=1671223582290746&amp;alt=media\" alt=\"\"></p>\n<h2>Genes that are poorly predicted are enriched for mitochondrial RNA and ribosomal proteins</h2>\n<p>Finally, we looked at how well each gene was predicted in the test set for the top prediction. To assess this, we calculated the correlation between the predicted and observed values per gene across cells. Note this is the opposite orientation to how we set up the competition, and some genes may simply not be important for high performance when calculating correlation across genes per cell. </p>\n<p>What we first found was there's a surprising bimodal distribution in the correlation values across genes:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308072%2Fa1aba8583281e25a452e3ced5a684011%2FScreen%20Shot%202022-12-16%20at%203.57.38%20PM.png?generation=1671224374150990&amp;alt=media\" alt=\"\"></p>\n<p>We also look at some of the top and bottom genes by prediction performance, and find that the genes that are worst predicted are often mitochondrial RNA genes or ribosomal proteins, which we generally believe aren't highly relevant for explaining biological heterogeneity. </p>\n<h2>What's next?</h2>\n<p>We're still early in our process of investigating the results of the competition in preparation for writing a manuscript summarizing the competition in <em>Nature Methods</em>. </p>\n<p>If you'd like to help, please don't hesitate to reach out and keep posting in the forums / code section of the competition!</p>\n<p>We hope this data and the competition have been useful, and once again, thanks everyone for your participation! Thanks also to Ryan and Ashley from the Kaggle team, without whom this competition would not have run so smoothly.</p>\n<p>Best,<br>\nDaniel Burkhardt on behalf of the Open Problems Core Team</p>",
      "rawMarkdown": "Hi all, once again, we at Open Problems are thrilled with the turn out of this year's competition. I hope it's been a productive experience for you, and if this was your first time working with single-cell data, we hope you'll continue to engage with the field.\n\nLast week, we hosted a virtual workshop as part of NeurIPS 2022, and for those of you who could make it, we're happy to post a recording of the meeting on YouTube: https://www.youtube.com/embed/WmkbyGMPHOE\n\nWe learned a lot from this competition, and we're still going through the results as we get ready to prepare a submission to *Nature Methods*. We're still looking for survey results, so if you haven't filled out the survey, please take 10 minutes to do so here: https://cellarity.qualtrics.com/jfe/form/SV_0jOnimbJmxiP0Z8\n\nSo what did we find so far? \n\n## We appear to have hit a limit in predictions\n\nOne of the most interesting results of the competition that was immediately apparent was the similarity between the top predictions in the leaderboard. On the private leaderboard, 584 / 1221 teams with were within 0.01 correlation of the top performing team. We didn't expect that. Take for example the [Recursion Cellular Image Classification competition in 2019](https://www.kaggle.com/competitions/recursion-cellular-image-classification/overview/prizes) where the 50th team's accuracy was within `0.05` of the top team or the [Human Protein Atlas Single-Cell Image Classification](https://www.kaggle.com/competitions/hpa-single-cell-image-classification) competition in 2020 where where the 50th team's segmentation precision was within `0.14` of the top team. Compared the these tasks, the performance of the top teams was much more similar--at least by our metric.\n\nYou can see this looking at the  follow plot adapted from @josecarmona graphical recap notebook.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308072%2F1cf659f47153f7000639d8f9d63d00ff%2FScreen%20Shot%202022-12-16%20at%203.30.59%20PM.png?generation=1671222686784466&alt=media)\n\n## Top teams predictions were very similar overall\n\nWe also wondered: are the predictions themselves similar? To assess this, we looked at the top two teams multiome predictions, and summed the predicted expression values across all cells in the held-out test set. Looking at this we find the predictions overall are very similar per-gene with a Pearson R of 0.9986.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308072%2F219793d1906ed0de0741759890ba1a37%2Fcorrelations.png?generation=1671223463918411&alt=media)\n\nWe can also look at the top predictions comparing the correlation of all values for both technologies (this is taking the correlation of the raw prediction files) and find that this trend holds across the top 3 teams.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308072%2Fc6bf3c16d83afb39366b84c1eb33901a%2Fmore_corr.png?generation=1671223582290746&alt=media)\n\n## Genes that are poorly predicted are enriched for mitochondrial RNA and ribosomal proteins\n\nFinally, we looked at how well each gene was predicted in the test set for the top prediction. To assess this, we calculated the correlation between the predicted and observed values per gene across cells. Note this is the opposite orientation to how we set up the competition, and some genes may simply not be important for high performance when calculating correlation across genes per cell. \n\nWhat we first found was there's a surprising bimodal distribution in the correlation values across genes:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308072%2Fa1aba8583281e25a452e3ced5a684011%2FScreen%20Shot%202022-12-16%20at%203.57.38%20PM.png?generation=1671224374150990&alt=media)\n\nWe also look at some of the top and bottom genes by prediction performance, and find that the genes that are worst predicted are often mitochondrial RNA genes or ribosomal proteins, which we generally believe aren't highly relevant for explaining biological heterogeneity. \n\n## What's next?\n\nWe're still early in our process of investigating the results of the competition in preparation for writing a manuscript summarizing the competition in *Nature Methods*. \n\nIf you'd like to help, please don't hesitate to reach out and keep posting in the forums / code section of the competition!\n\nWe hope this data and the competition have been useful, and once again, thanks everyone for your participation! Thanks also to Ryan and Ashley from the Kaggle team, without whom this competition would not have run so smoothly.\n\nBest,\nDaniel Burkhardt on behalf of the Open Problems Core Team",
      "votes": null
    },
    {
      "id": "2067626",
      "postDate": "12/16/2022 23:00:45",
      "content": "<p>Thank you very much Daniel and Open Problems Core Team for hosting this competition! We look forward to seeing the manuscript! 😊 </p>",
      "rawMarkdown": "Thank you very much Daniel and Open Problems Core Team for hosting this competition! We look forward to seeing the manuscript! 😊",
      "votes": null
    },
    {
      "id": "2071069",
      "postDate": "12/20/2022 15:25:23",
      "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> </p>\n<p>Thanks again for great competition and the analysis ! <br>\nSmall remark on:<br>\n \" Recursion Cellular Image Classification competition in 2019 where the 50th team's accuracy was within 0.05 of the top team or the Human Protein Atlas Single-Cell Image Classification competition in 2020 where where the 50th team's segmentation precision was within 0.14 of the top team. Compared the these tasks, the performance of the top teams was much more similar--at least by our metric.\"</p>\n<p>It is typical for image/video/NLP/etc related competitions that gap between places can be quite big.<br>\nBut for tabular data competitions the gap between the places is almost never big on Kaggle.<br>\nThe reasons are quite clear - tabular data are more accessible and many Kagglers perform well, while image/video/nlp/etc<br>\nsometimes requires either hardware, either specific knowledge or something else, where only some experts can perform well.</p>\n<p>So the fair comparison can be \"MOA\":<br>\n<a href=\"https://www.kaggle.com/competitions/lish-moa\" target=\"_blank\">https://www.kaggle.com/competitions/lish-moa</a></p>\n<ul>\n<li>around drugs and transcriptomics, or recent \"AMEX\": <a href=\"https://www.kaggle.com/competitions/amex-default-prediction\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction</a><br>\nor plenty other tabular data competitions. </li>\n</ul>",
      "rawMarkdown": "danielburkhardt \n\nThanks again for great competition and the analysis ! \nSmall remark on:\n \" Recursion Cellular Image Classification competition in 2019 where the 50th team's accuracy was within 0.05 of the top team or the Human Protein Atlas Single-Cell Image Classification competition in 2020 where where the 50th team's segmentation precision was within 0.14 of the top team. Compared the these tasks, the performance of the top teams was much more similar--at least by our metric.\"\n\nIt is typical for image/video/NLP/etc related competitions that gap between places can be quite big.\nBut for tabular data competitions the gap between the places is almost never big on Kaggle.\nThe reasons are quite clear - tabular data are more accessible and many Kagglers perform well, while image/video/nlp/etc\nsometimes requires either hardware, either specific knowledge or something else, where only some experts can perform well.\n\nSo the fair comparison can be \"MOA\":\nhttps://www.kaggle.com/competitions/lish-moa\n- around drugs and transcriptomics, or recent \"AMEX\": https://www.kaggle.com/competitions/amex-default-prediction\nor plenty other tabular data competitions.",
      "votes": null
    },
    {
      "id": "2087400",
      "postDate": "01/05/2023 15:15:36",
      "content": "<p>Thank you, Daniel. In Icahn School of Medicine at Mount Sinai, I use Muon to analyze CITE-Seq data, while DNA + mutational profile with MissionBio. I work 100% on human CD34+ cells in patho-physiological conditions. I would be happy to be more engaged with the soul purpose of this project. Feel free to text me - <a>md.babu.mia@mssm.edu</a>   </p>",
      "rawMarkdown": "Thank you, Daniel. In Icahn School of Medicine at Mount Sinai, I use Muon to analyze CITE-Seq data, while DNA + mutational profile with MissionBio. I work 100% on human CD34+ cells in patho-physiological conditions. I would be happy to be more engaged with the soul purpose of this project. Feel free to text me - md.babu.mia@mssm.edu",
      "votes": null
    },
    {
      "id": "2318572",
      "postDate": "06/26/2023 12:56:53",
      "content": "<p>Hello, looking to learn more about this field of machine learning.  I got introduced to it about a week ago and it is quite interesting.</p>",
      "rawMarkdown": "Hello, looking to learn more about this field of machine learning.  I got introduced to it about a week ago and it is quite interesting.",
      "votes": null
    },
    {
      "id": "2615316",
      "postDate": "01/23/2024 04:12:48",
      "content": "<p>Hi,</p>\n<p>I'm working on a similar problem to the CITE-Seq prediction as a research project. The CITE-Seq data would be very useful but I'm not so keen on dsb normalisation.</p>\n<p>Are the raw unprocessed counts (both RNA and Protein) available anywhere?</p>\n<p>Thanks</p>",
      "rawMarkdown": "Hi,\n\nI'm working on a similar problem to the CITE-Seq prediction as a research project. The CITE-Seq data would be very useful but I'm not so keen on dsb normalisation.\n\nAre the raw unprocessed counts (both RNA and Protein) available anywhere?\n\nThanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2067626,
      "author_name": "ashleychow",
      "author_url": "",
      "post_date": "12/16/2022 23:00:45",
      "content": "<p>Thank you very much Daniel and Open Problems Core Team for hosting this competition! We look forward to seeing the manuscript! 😊 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2071069,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "12/20/2022 15:25:23",
      "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> </p>\n<p>Thanks again for great competition and the analysis ! <br>\nSmall remark on:<br>\n \" Recursion Cellular Image Classification competition in 2019 where the 50th team's accuracy was within 0.05 of the top team or the Human Protein Atlas Single-Cell Image Classification competition in 2020 where where the 50th team's segmentation precision was within 0.14 of the top team. Compared the these tasks, the performance of the top teams was much more similar--at least by our metric.\"</p>\n<p>It is typical for image/video/NLP/etc related competitions that gap between places can be quite big.<br>\nBut for tabular data competitions the gap between the places is almost never big on Kaggle.<br>\nThe reasons are quite clear - tabular data are more accessible and many Kagglers perform well, while image/video/nlp/etc<br>\nsometimes requires either hardware, either specific knowledge or something else, where only some experts can perform well.</p>\n<p>So the fair comparison can be \"MOA\":<br>\n<a href=\"https://www.kaggle.com/competitions/lish-moa\" target=\"_blank\">https://www.kaggle.com/competitions/lish-moa</a></p>\n<ul>\n<li>around drugs and transcriptomics, or recent \"AMEX\": <a href=\"https://www.kaggle.com/competitions/amex-default-prediction\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction</a><br>\nor plenty other tabular data competitions. </li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2087400,
      "author_name": "babumiaphd",
      "author_url": "",
      "post_date": "01/05/2023 15:15:36",
      "content": "<p>Thank you, Daniel. In Icahn School of Medicine at Mount Sinai, I use Muon to analyze CITE-Seq data, while DNA + mutational profile with MissionBio. I work 100% on human CD34+ cells in patho-physiological conditions. I would be happy to be more engaged with the soul purpose of this project. Feel free to text me - <a>md.babu.mia@mssm.edu</a>   </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2318572,
      "author_name": "infemus",
      "author_url": "",
      "post_date": "06/26/2023 12:56:53",
      "content": "<p>Hello, looking to learn more about this field of machine learning.  I got introduced to it about a week ago and it is quite interesting.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2615316,
      "author_name": "danrawl",
      "author_url": "",
      "post_date": "01/23/2024 04:12:48",
      "content": "<p>Hi,</p>\n<p>I'm working on a similar problem to the CITE-Seq prediction as a research project. The CITE-Seq data would be very useful but I'm not so keen on dsb normalisation.</p>\n<p>Are the raw unprocessed counts (both RNA and Protein) available anywhere?</p>\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2067590": "Hi all, once again, we at Open Problems are thrilled with the turn out of this year's competition. I hope it's been a productive experience for you, and if this was your first time working with single-cell data, we hope you'll continue to engage with the field.\n\nLast week, we hosted a virtual workshop as part of NeurIPS 2022, and for those of you who could make it, we're happy to post a recording of the meeting on YouTube: https://www.youtube.com/embed/WmkbyGMPHOE\n\nWe learned a lot from this competition, and we're still going through the results as we get ready to prepare a submission to *Nature Methods*. We're still looking for survey results, so if you haven't filled out the survey, please take 10 minutes to do so here: https://cellarity.qualtrics.com/jfe/form/SV_0jOnimbJmxiP0Z8\n\nSo what did we find so far? \n\n## We appear to have hit a limit in predictions\n\nOne of the most interesting results of the competition that was immediately apparent was the similarity between the top predictions in the leaderboard. On the private leaderboard, 584 / 1221 teams with were within 0.01 correlation of the top performing team. We didn't expect that. Take for example the [Recursion Cellular Image Classification competition in 2019](https://www.kaggle.com/competitions/recursion-cellular-image-classification/overview/prizes) where the 50th team's accuracy was within `0.05` of the top team or the [Human Protein Atlas Single-Cell Image Classification](https://www.kaggle.com/competitions/hpa-single-cell-image-classification) competition in 2020 where where the 50th team's segmentation precision was within `0.14` of the top team. Compared the these tasks, the performance of the top teams was much more similar--at least by our metric.\n\nYou can see this looking at the  follow plot adapted from @josecarmona graphical recap notebook.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308072%2F1cf659f47153f7000639d8f9d63d00ff%2FScreen%20Shot%202022-12-16%20at%203.30.59%20PM.png?generation=1671222686784466&alt=media)\n\n## Top teams predictions were very similar overall\n\nWe also wondered: are the predictions themselves similar? To assess this, we looked at the top two teams multiome predictions, and summed the predicted expression values across all cells in the held-out test set. Looking at this we find the predictions overall are very similar per-gene with a Pearson R of 0.9986.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308072%2F219793d1906ed0de0741759890ba1a37%2Fcorrelations.png?generation=1671223463918411&alt=media)\n\nWe can also look at the top predictions comparing the correlation of all values for both technologies (this is taking the correlation of the raw prediction files) and find that this trend holds across the top 3 teams.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308072%2Fc6bf3c16d83afb39366b84c1eb33901a%2Fmore_corr.png?generation=1671223582290746&alt=media)\n\n## Genes that are poorly predicted are enriched for mitochondrial RNA and ribosomal proteins\n\nFinally, we looked at how well each gene was predicted in the test set for the top prediction. To assess this, we calculated the correlation between the predicted and observed values per gene across cells. Note this is the opposite orientation to how we set up the competition, and some genes may simply not be important for high performance when calculating correlation across genes per cell. \n\nWhat we first found was there's a surprising bimodal distribution in the correlation values across genes:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308072%2Fa1aba8583281e25a452e3ced5a684011%2FScreen%20Shot%202022-12-16%20at%203.57.38%20PM.png?generation=1671224374150990&alt=media)\n\nWe also look at some of the top and bottom genes by prediction performance, and find that the genes that are worst predicted are often mitochondrial RNA genes or ribosomal proteins, which we generally believe aren't highly relevant for explaining biological heterogeneity. \n\n## What's next?\n\nWe're still early in our process of investigating the results of the competition in preparation for writing a manuscript summarizing the competition in *Nature Methods*. \n\nIf you'd like to help, please don't hesitate to reach out and keep posting in the forums / code section of the competition!\n\nWe hope this data and the competition have been useful, and once again, thanks everyone for your participation! Thanks also to Ryan and Ashley from the Kaggle team, without whom this competition would not have run so smoothly.\n\nBest,\nDaniel Burkhardt on behalf of the Open Problems Core Team",
    "2067626": "Thank you very much Daniel and Open Problems Core Team for hosting this competition! We look forward to seeing the manuscript! 😊",
    "2071069": "danielburkhardt \n\nThanks again for great competition and the analysis ! \nSmall remark on:\n \" Recursion Cellular Image Classification competition in 2019 where the 50th team's accuracy was within 0.05 of the top team or the Human Protein Atlas Single-Cell Image Classification competition in 2020 where where the 50th team's segmentation precision was within 0.14 of the top team. Compared the these tasks, the performance of the top teams was much more similar--at least by our metric.\"\n\nIt is typical for image/video/NLP/etc related competitions that gap between places can be quite big.\nBut for tabular data competitions the gap between the places is almost never big on Kaggle.\nThe reasons are quite clear - tabular data are more accessible and many Kagglers perform well, while image/video/nlp/etc\nsometimes requires either hardware, either specific knowledge or something else, where only some experts can perform well.\n\nSo the fair comparison can be \"MOA\":\nhttps://www.kaggle.com/competitions/lish-moa\n- around drugs and transcriptomics, or recent \"AMEX\": https://www.kaggle.com/competitions/amex-default-prediction\nor plenty other tabular data competitions.",
    "2087400": "Thank you, Daniel. In Icahn School of Medicine at Mount Sinai, I use Muon to analyze CITE-Seq data, while DNA + mutational profile with MissionBio. I work 100% on human CD34+ cells in patho-physiological conditions. I would be happy to be more engaged with the soul purpose of this project. Feel free to text me - md.babu.mia@mssm.edu",
    "2318572": "Hello, looking to learn more about this field of machine learning.  I got introduced to it about a week ago and it is quite interesting.",
    "2615316": "Hi,\n\nI'm working on a similar problem to the CITE-Seq prediction as a research project. The CITE-Seq data would be very useful but I'm not so keen on dsb normalisation.\n\nAre the raw unprocessed counts (both RNA and Protein) available anywhere?\n\nThanks"
  },
  "source": "meta"
}