{
  "id": 466181,
  "title": "Post-competition EDA: unveiling the significance of biases in the training data",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/466181",
  "author_name": "",
  "post_date": "2024-01-07T13:51:40.601511800Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p><strong>Summary</strong><br>\nIn our previous work [1] and in the findings by Amrbosm [2], several issues in the training data provided for the competition were highlighted. These problems include misclassification of cell types, inadequate cell counts, and invalid differential expression (DE) values due to missing RNA counts. This study delves deeper into these errors and quantifies their impact. We find that nearly 35 percent of the data is potentially erroneous. Interestingly, these error-prone data display more significance compared to the remaining data. This highlights the current training dataset's limitations and underlines the need for recalculating DE values from raw expression data in future research. Moreover, we observed that only 184 genes (from an approximate total of 18,000) are consistently expressed across all samples, suggesting a need for methodological adjustement to handle gene data sparsity.</p>\n<p><strong>Reproducibility</strong><br>\nThe source code is here [ <a href=\"https://www.kaggle.com/jalilnourisa/post-eda\" target=\"_blank\">https://www.kaggle.com/jalilnourisa/post-eda</a> ].<br>\n<strong>Refs</strong>: <br>\n[1][<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461159\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461159</a>]  <br>\n[2][<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458661\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458661</a>]</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11155399%2Fe6ee635b734a72db544524cddff8ac35%2Fsig.png?generation=1704635193716353&amp;alt=media\"><br>\nOriginal data versus filtered data (65 percent of the original data). The idea of this graph is taken from Chervov [<a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-housekeeping-genes?scriptVersionId=155511517&amp;cellId=11]\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-housekeeping-genes?scriptVersionId=155511517&amp;cellId=11]</a>. </p>",
  "messages": [
    {
      "id": "2590856",
      "postDate": "01/07/2024 13:51:40",
      "content": "<p><strong>Summary</strong><br>\nIn our previous work [1] and in the findings by Amrbosm [2], several issues in the training data provided for the competition were highlighted. These problems include misclassification of cell types, inadequate cell counts, and invalid differential expression (DE) values due to missing RNA counts. This study delves deeper into these errors and quantifies their impact. We find that nearly 35 percent of the data is potentially erroneous. Interestingly, these error-prone data display more significance compared to the remaining data. This highlights the current training dataset's limitations and underlines the need for recalculating DE values from raw expression data in future research. Moreover, we observed that only 184 genes (from an approximate total of 18,000) are consistently expressed across all samples, suggesting a need for methodological adjustement to handle gene data sparsity.</p>\n<p><strong>Reproducibility</strong><br>\nThe source code is here [ <a href=\"https://www.kaggle.com/jalilnourisa/post-eda\" target=\"_blank\">https://www.kaggle.com/jalilnourisa/post-eda</a> ].<br>\n<strong>Refs</strong>: <br>\n[1][<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461159\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461159</a>]  <br>\n[2][<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458661\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458661</a>]</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11155399%2Fe6ee635b734a72db544524cddff8ac35%2Fsig.png?generation=1704635193716353&amp;alt=media\"><br>\nOriginal data versus filtered data (65 percent of the original data). The idea of this graph is taken from Chervov [<a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-housekeeping-genes?scriptVersionId=155511517&amp;cellId=11]\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-housekeeping-genes?scriptVersionId=155511517&amp;cellId=11]</a>. </p>",
      "rawMarkdown": "**Summary**\nIn our previous work [1] and in the findings by Amrbosm [2], several issues in the training data provided for the competition were highlighted. These problems include misclassification of cell types, inadequate cell counts, and invalid differential expression (DE) values due to missing RNA counts. This study delves deeper into these errors and quantifies their impact. We find that nearly 35 percent of the data is potentially erroneous. Interestingly, these error-prone data display more significance compared to the remaining data. This highlights the current training dataset's limitations and underlines the need for recalculating DE values from raw expression data in future research. Moreover, we observed that only 184 genes (from an approximate total of 18,000) are consistently expressed across all samples, suggesting a need for methodological adjustement to handle gene data sparsity.\n\n**Reproducibility**\nThe source code is here [ https://www.kaggle.com/jalilnourisa/post-eda ].\n**Refs**: \n[1][https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461159]  \n[2][https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458661]\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11155399%2Fe6ee635b734a72db544524cddff8ac35%2Fsig.png?generation=1704635193716353&alt=media)\nOriginal data versus filtered data (65 percent of the original data). The idea of this graph is taken from Chervov [https://www.kaggle.com/code/alexandervc/op2-eda-housekeeping-genes?scriptVersionId=155511517&cellId=11].",
      "votes": null
    },
    {
      "id": "2591039",
      "postDate": "01/07/2024 16:01:53",
      "content": "<p>Thank you for sharing this!</p>",
      "rawMarkdown": "Thank you for sharing this!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2591039,
      "author_name": "wguesdon",
      "author_url": "",
      "post_date": "01/07/2024 16:01:53",
      "content": "<p>Thank you for sharing this!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2590856": "**Summary**\nIn our previous work [1] and in the findings by Amrbosm [2], several issues in the training data provided for the competition were highlighted. These problems include misclassification of cell types, inadequate cell counts, and invalid differential expression (DE) values due to missing RNA counts. This study delves deeper into these errors and quantifies their impact. We find that nearly 35 percent of the data is potentially erroneous. Interestingly, these error-prone data display more significance compared to the remaining data. This highlights the current training dataset's limitations and underlines the need for recalculating DE values from raw expression data in future research. Moreover, we observed that only 184 genes (from an approximate total of 18,000) are consistently expressed across all samples, suggesting a need for methodological adjustement to handle gene data sparsity.\n\n**Reproducibility**\nThe source code is here [ https://www.kaggle.com/jalilnourisa/post-eda ].\n**Refs**: \n[1][https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461159]  \n[2][https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458661]\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11155399%2Fe6ee635b734a72db544524cddff8ac35%2Fsig.png?generation=1704635193716353&alt=media)\nOriginal data versus filtered data (65 percent of the original data). The idea of this graph is taken from Chervov [https://www.kaggle.com/code/alexandervc/op2-eda-housekeeping-genes?scriptVersionId=155511517&cellId=11].",
    "2591039": "Thank you for sharing this!"
  },
  "source": "meta"
}