{
  "id": 450545,
  "title": "Pseudobulked counts from AnnData inconsistent",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/450545",
  "author_name": "",
  "post_date": "2023-10-24T18:28:07.011495300Z",
  "votes": 9,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi there,<br>\nI started working on the challenge starting from the single-cell count matrix, but there are still some differences between the DE data created from the <code>adata_train.parquet</code> file and the <code>de_train.parquet</code> file. </p>\n<p>I tried to reproduce the DE train file on the SaturnCloud instance (to avoid version issues that would cause problems). One obvious thing is the mismatch of the number of genes, which were already noted by others (e.g. <a href=\"https://github.com/openproblems-bio/neurips-2023-scripts/issues/1)\" target=\"_blank\">https://github.com/openproblems-bio/neurips-2023-scripts/issues/1)</a>. <br>\nI can live with subsetting to the number of genes in the <code>de_train.parquet</code> file, however, that does not explain the differences between the DE data created from <code>adata_train.parquet</code> and the <code>de_train.parquet</code>. I attached one of the direct comparison files created on SaturnCloud with the notebook <code>compute_de.ipynb</code>. </p>\n<p>So examined the pseudobulk data in more detail. The most striking differences between the reference pseudobulk on the SaturnCloud (<code>train_or_control_bulk_by_cell_type_adata.h5ad</code>) and the pseudobulk created from <code>adata_train.parquet</code> are the compounds Sgc-cbp30, YK 4-279, where the metadata for the plate_id and the donor_id do not match, and the pseudobulk counts seem to be entirely different. </p>\n<p>I hope that you understand my frustration with the provided data. To my mind, it is a strong disadvantage for anyone who is attempting to work on a single-cell count data approach for the challenge. It would be really great if the organizers would provide the updated version of the <code>adata_train.parquet</code> file in due time as promised already at the beginning of the month. In case they have done so, I would be grateful for a pointer to the corrected data files. <br>\nThanks. </p>",
  "messages": [
    {
      "id": "2497620",
      "postDate": "10/24/2023 18:28:07",
      "content": "<p>Hi there,<br>\nI started working on the challenge starting from the single-cell count matrix, but there are still some differences between the DE data created from the <code>adata_train.parquet</code> file and the <code>de_train.parquet</code> file. </p>\n<p>I tried to reproduce the DE train file on the SaturnCloud instance (to avoid version issues that would cause problems). One obvious thing is the mismatch of the number of genes, which were already noted by others (e.g. <a href=\"https://github.com/openproblems-bio/neurips-2023-scripts/issues/1)\" target=\"_blank\">https://github.com/openproblems-bio/neurips-2023-scripts/issues/1)</a>. <br>\nI can live with subsetting to the number of genes in the <code>de_train.parquet</code> file, however, that does not explain the differences between the DE data created from <code>adata_train.parquet</code> and the <code>de_train.parquet</code>. I attached one of the direct comparison files created on SaturnCloud with the notebook <code>compute_de.ipynb</code>. </p>\n<p>So examined the pseudobulk data in more detail. The most striking differences between the reference pseudobulk on the SaturnCloud (<code>train_or_control_bulk_by_cell_type_adata.h5ad</code>) and the pseudobulk created from <code>adata_train.parquet</code> are the compounds Sgc-cbp30, YK 4-279, where the metadata for the plate_id and the donor_id do not match, and the pseudobulk counts seem to be entirely different. </p>\n<p>I hope that you understand my frustration with the provided data. To my mind, it is a strong disadvantage for anyone who is attempting to work on a single-cell count data approach for the challenge. It would be really great if the organizers would provide the updated version of the <code>adata_train.parquet</code> file in due time as promised already at the beginning of the month. In case they have done so, I would be grateful for a pointer to the corrected data files. <br>\nThanks. </p>",
      "rawMarkdown": "Hi there,\nI started working on the challenge starting from the single-cell count matrix, but there are still some differences between the DE data created from the `adata_train.parquet` file and the `de_train.parquet` file. \n\nI tried to reproduce the DE train file on the SaturnCloud instance (to avoid version issues that would cause problems). One obvious thing is the mismatch of the number of genes, which were already noted by others (e.g. https://github.com/openproblems-bio/neurips-2023-scripts/issues/1). \nI can live with subsetting to the number of genes in the `de_train.parquet` file, however, that does not explain the differences between the DE data created from `adata_train.parquet` and the `de_train.parquet`. I attached one of the direct comparison files created on SaturnCloud with the notebook `compute_de.ipynb`. \n\nSo examined the pseudobulk data in more detail. The most striking differences between the reference pseudobulk on the SaturnCloud (`train_or_control_bulk_by_cell_type_adata.h5ad`) and the pseudobulk created from `adata_train.parquet` are the compounds Sgc-cbp30, YK 4-279, where the metadata for the plate_id and the donor_id do not match, and the pseudobulk counts seem to be entirely different. \n\nI hope that you understand my frustration with the provided data. To my mind, it is a strong disadvantage for anyone who is attempting to work on a single-cell count data approach for the challenge. It would be really great if the organizers would provide the updated version of the `adata_train.parquet` file in due time as promised already at the beginning of the month. In case they have done so, I would be grateful for a pointer to the corrected data files. \nThanks.",
      "votes": null
    },
    {
      "id": "2497932",
      "postDate": "10/25/2023 02:42:59",
      "content": "<p>Yes, I totally understand your frustration. When I found this, I gave up on the idea of using the single-cell data.</p>\n<p>Here is anthor <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/448033\" target=\"_blank\">post</a> and two of my notebooks (<a href=\"https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v3\" target=\"_blank\">op2-02-pseudobulk-v3</a> and <a href=\"https://www.kaggle.com/code/awater1223/op2-04-merge-limma-results\" target=\"_blank\">op2-04-merge-limma-results</a>) with the result and analysis about this problem.</p>\n<p>I think the organizers should fix this ASAP to avoid more people wasting their time and effort to \"debug\" this inconsistency.</p>\n<hr>\n<p>Technically, I guess the organizers forget to remove (or add) some filters at gene level when preprocessing the data, (even worse, the filters are inconsistent among samples, so they need more time to fix it, this is my conspiracy).<br>\nUsually, this kind of inconsistency can be ignored (only negligible effect about the biological findings), but for this competition, the limma's p-value calculation amplified this difference in the value, and the MRRMSE scoring function made this difference has huge impact on the score. That made my very frustrated.</p>",
      "rawMarkdown": "Yes, I totally understand your frustration. When I found this, I gave up on the idea of using the single-cell data.\n\nHere is anthor [post](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/448033) and two of my notebooks ([op2-02-pseudobulk-v3](https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v3) and [op2-04-merge-limma-results](https://www.kaggle.com/code/awater1223/op2-04-merge-limma-results)) with the result and analysis about this problem.\n\nI think the organizers should fix this ASAP to avoid more people wasting their time and effort to \"debug\" this inconsistency.\n\n---\n\nTechnically, I guess the organizers forget to remove (or add) some filters at gene level when preprocessing the data, (even worse, the filters are inconsistent among samples, so they need more time to fix it, this is my conspiracy).\nUsually, this kind of inconsistency can be ignored (only negligible effect about the biological findings), but for this competition, the limma's p-value calculation amplified this difference in the value, and the MRRMSE scoring function made this difference has huge impact on the score. That made my very frustrated.",
      "votes": null
    },
    {
      "id": "2498914",
      "postDate": "10/25/2023 15:49:49",
      "content": "<p>Thank you for sharing your insights. It's pretty discouraging to use the single-cell data, then. </p>",
      "rawMarkdown": "Thank you for sharing your insights. It's pretty discouraging to use the single-cell data, then.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2497932,
      "author_name": "awater1223",
      "author_url": "",
      "post_date": "10/25/2023 02:42:59",
      "content": "<p>Yes, I totally understand your frustration. When I found this, I gave up on the idea of using the single-cell data.</p>\n<p>Here is anthor <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/448033\" target=\"_blank\">post</a> and two of my notebooks (<a href=\"https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v3\" target=\"_blank\">op2-02-pseudobulk-v3</a> and <a href=\"https://www.kaggle.com/code/awater1223/op2-04-merge-limma-results\" target=\"_blank\">op2-04-merge-limma-results</a>) with the result and analysis about this problem.</p>\n<p>I think the organizers should fix this ASAP to avoid more people wasting their time and effort to \"debug\" this inconsistency.</p>\n<hr>\n<p>Technically, I guess the organizers forget to remove (or add) some filters at gene level when preprocessing the data, (even worse, the filters are inconsistent among samples, so they need more time to fix it, this is my conspiracy).<br>\nUsually, this kind of inconsistency can be ignored (only negligible effect about the biological findings), but for this competition, the limma's p-value calculation amplified this difference in the value, and the MRRMSE scoring function made this difference has huge impact on the score. That made my very frustrated.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2498914,
          "author_name": "marenbuettner789",
          "author_url": "",
          "post_date": "10/25/2023 15:49:49",
          "content": "<p>Thank you for sharing your insights. It's pretty discouraging to use the single-cell data, then. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2497620": "Hi there,\nI started working on the challenge starting from the single-cell count matrix, but there are still some differences between the DE data created from the `adata_train.parquet` file and the `de_train.parquet` file. \n\nI tried to reproduce the DE train file on the SaturnCloud instance (to avoid version issues that would cause problems). One obvious thing is the mismatch of the number of genes, which were already noted by others (e.g. https://github.com/openproblems-bio/neurips-2023-scripts/issues/1). \nI can live with subsetting to the number of genes in the `de_train.parquet` file, however, that does not explain the differences between the DE data created from `adata_train.parquet` and the `de_train.parquet`. I attached one of the direct comparison files created on SaturnCloud with the notebook `compute_de.ipynb`. \n\nSo examined the pseudobulk data in more detail. The most striking differences between the reference pseudobulk on the SaturnCloud (`train_or_control_bulk_by_cell_type_adata.h5ad`) and the pseudobulk created from `adata_train.parquet` are the compounds Sgc-cbp30, YK 4-279, where the metadata for the plate_id and the donor_id do not match, and the pseudobulk counts seem to be entirely different. \n\nI hope that you understand my frustration with the provided data. To my mind, it is a strong disadvantage for anyone who is attempting to work on a single-cell count data approach for the challenge. It would be really great if the organizers would provide the updated version of the `adata_train.parquet` file in due time as promised already at the beginning of the month. In case they have done so, I would be grateful for a pointer to the corrected data files. \nThanks.",
    "2497932": "Yes, I totally understand your frustration. When I found this, I gave up on the idea of using the single-cell data.\n\nHere is anthor [post](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/448033) and two of my notebooks ([op2-02-pseudobulk-v3](https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v3) and [op2-04-merge-limma-results](https://www.kaggle.com/code/awater1223/op2-04-merge-limma-results)) with the result and analysis about this problem.\n\nI think the organizers should fix this ASAP to avoid more people wasting their time and effort to \"debug\" this inconsistency.\n\n---\n\nTechnically, I guess the organizers forget to remove (or add) some filters at gene level when preprocessing the data, (even worse, the filters are inconsistent among samples, so they need more time to fix it, this is my conspiracy).\nUsually, this kind of inconsistency can be ignored (only negligible effect about the biological findings), but for this competition, the limma's p-value calculation amplified this difference in the value, and the MRRMSE scoring function made this difference has huge impact on the score. That made my very frustrated.",
    "2498914": "Thank you for sharing your insights. It's pretty discouraging to use the single-cell data, then."
  },
  "source": "meta"
}