{
  "id": 445883,
  "title": "Some strange observations about adata_train.parquet",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/445883",
  "author_name": "Antonina Dolgorukova",
  "post_date": "2023-10-09T11:12:07.422000",
  "votes": 33,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I've just examined a bit deeper the adata_train.parquet file and corresponding meta data <a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-stats-sign-de-fingerprints-eda#Unaggregated-counts\" target=\"_blank\">here</a>.</p>\n<p>In the adata_train, here are 416,442,312 rows; 240,090 observation ids, 21255 genes (this is a superset of de_train, so the number of genes is higher). </p>\n<p>In the adata_obs_meta I found 240090 rows, each id (identifiers assigned to each cell in the raw dataset) has one row. So I plotted the number of cells, from which the data were collected, per compound and per compounds-cell type pair. The last is most interesting:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2Ffbb156c7e09c6a0c31c98f6de5d26d76%2FScreenshot%202023-10-09%20175613.png?generation=1696849122129385&amp;alt=media\" alt=\"\"></p>\n<p>Here are the number of B and Myeloid cells cells from which were collected the data for 15 compounds (train data). We are supposed to predict the impact of other compounds on these cells. But it seems that there as at least three outliers, for wchich the data were collected from just a few cells (unreliable?)<br>\nThe MLN 2238 is among them, and it has been already noted by <a href=\"https://www.kaggle.com/qihuaz\" target=\"_blank\">@qihuaz</a> as an outlier based on its effect.</p>\n<p>Here is more:<br>\nThere seems to be a huge difference in the number of CD8+ T cells and T regulatory cells</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F2e176fecab41b2c708d3f3e241d6c957%2FScreenshot%202023-10-09%20175736.png?generation=1696849237309259&amp;alt=media\" alt=\"\"></p>\n<p>compared to CD4+ T cells and NK cells</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F4c15dc18ae23a6d2e61a56b5df83208c%2FScreenshot%202023-10-09%20175655.png?generation=1696849183343461&amp;alt=media\" alt=\"\"></p>\n<p>Given this, one may speculate that the effect may vary less or more depending on the number of cells in which it is measured, affecting the p-values obtained (for all three controls btw the data were collected from a much larger number of cells than for the other compounds - more robust data). What I mean is that in a small number of cells, the variance of the drug effect can be large and the p-value is high, but if you treat more cells, the variance will reduce and the effect will appear, so the p-value will be lower. So, to get reliable predictions, it may be reasonable to account for the number of cells, but the thing is, as far as I understand we do not have this information for the test compound*cell_type pairs.</p>\n<p><strong>UPD: Crazy, but there are compound*cell type pairs with only one corresponding obs_id in adata, so the drug effect on this cell type was measured in just one cell???</strong></p>\n<p>Dear organizers, <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> ? could you please tell if there is an explanation of such a discrepancy in the number of cells of different types, from which the data were collected for the drugs?</p>\n<p>Also would be great to hear any thoughts on how we can deal with this.</p>",
  "messages": [
    {
      "id": 2474649,
      "postDate": "2023-10-09T11:12:07.423Z",
      "content": "<p>I've just examined a bit deeper the adata_train.parquet file and corresponding meta data <a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-stats-sign-de-fingerprints-eda#Unaggregated-counts\" target=\"_blank\">here</a>.</p>\n<p>In the adata_train, here are 416,442,312 rows; 240,090 observation ids, 21255 genes (this is a superset of de_train, so the number of genes is higher). </p>\n<p>In the adata_obs_meta I found 240090 rows, each id (identifiers assigned to each cell in the raw dataset) has one row. So I plotted the number of cells, from which the data were collected, per compound and per compounds-cell type pair. The last is most interesting:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2Ffbb156c7e09c6a0c31c98f6de5d26d76%2FScreenshot%202023-10-09%20175613.png?generation=1696849122129385&amp;alt=media\" alt=\"\"></p>\n<p>Here are the number of B and Myeloid cells cells from which were collected the data for 15 compounds (train data). We are supposed to predict the impact of other compounds on these cells. But it seems that there as at least three outliers, for wchich the data were collected from just a few cells (unreliable?)<br>\nThe MLN 2238 is among them, and it has been already noted by <a href=\"https://www.kaggle.com/qihuaz\" target=\"_blank\">@qihuaz</a> as an outlier based on its effect.</p>\n<p>Here is more:<br>\nThere seems to be a huge difference in the number of CD8+ T cells and T regulatory cells</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F2e176fecab41b2c708d3f3e241d6c957%2FScreenshot%202023-10-09%20175736.png?generation=1696849237309259&amp;alt=media\" alt=\"\"></p>\n<p>compared to CD4+ T cells and NK cells</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F4c15dc18ae23a6d2e61a56b5df83208c%2FScreenshot%202023-10-09%20175655.png?generation=1696849183343461&amp;alt=media\" alt=\"\"></p>\n<p>Given this, one may speculate that the effect may vary less or more depending on the number of cells in which it is measured, affecting the p-values obtained (for all three controls btw the data were collected from a much larger number of cells than for the other compounds - more robust data). What I mean is that in a small number of cells, the variance of the drug effect can be large and the p-value is high, but if you treat more cells, the variance will reduce and the effect will appear, so the p-value will be lower. So, to get reliable predictions, it may be reasonable to account for the number of cells, but the thing is, as far as I understand we do not have this information for the test compound*cell_type pairs.</p>\n<p><strong>UPD: Crazy, but there are compound*cell type pairs with only one corresponding obs_id in adata, so the drug effect on this cell type was measured in just one cell???</strong></p>\n<p>Dear organizers, <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> ? could you please tell if there is an explanation of such a discrepancy in the number of cells of different types, from which the data were collected for the drugs?</p>\n<p>Also would be great to hear any thoughts on how we can deal with this.</p>",
      "rawMarkdown": "I've just examined a bit deeper the adata_train.parquet file and corresponding meta data [here](https://www.kaggle.com/code/antoninadolgorukova/op2-stats-sign-de-fingerprints-eda#Unaggregated-counts).\n\nIn the adata_train, here are 416,442,312 rows; 240,090 observation ids, 21255 genes (this is a superset of de_train, so the number of genes is higher). \n\nIn the adata_obs_meta I found 240090 rows, each id (identifiers assigned to each cell in the raw dataset) has one row. So I plotted the number of cells, from which the data were collected, per compound and per compounds-cell type pair. The last is most interesting:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2Ffbb156c7e09c6a0c31c98f6de5d26d76%2FScreenshot%202023-10-09%20175613.png?generation=1696849122129385&alt=media)\n\nHere are the number of B and Myeloid cells cells from which were collected the data for 15 compounds (train data). We are supposed to predict the impact of other compounds on these cells. But it seems that there as at least three outliers, for wchich the data were collected from just a few cells (unreliable?)\nThe MLN 2238 is among them, and it has been already noted by @qihuaz as an outlier based on its effect.\n\nHere is more:\nThere seems to be a huge difference in the number of CD8+ T cells and T regulatory cells\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F2e176fecab41b2c708d3f3e241d6c957%2FScreenshot%202023-10-09%20175736.png?generation=1696849237309259&alt=media)\n\ncompared to CD4+ T cells and NK cells\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F4c15dc18ae23a6d2e61a56b5df83208c%2FScreenshot%202023-10-09%20175655.png?generation=1696849183343461&alt=media)\n\n\nGiven this, one may speculate that the effect may vary less or more depending on the number of cells in which it is measured, affecting the p-values obtained (for all three controls btw the data were collected from a much larger number of cells than for the other compounds - more robust data). What I mean is that in a small number of cells, the variance of the drug effect can be large and the p-value is high, but if you treat more cells, the variance will reduce and the effect will appear, so the p-value will be lower. So, to get reliable predictions, it may be reasonable to account for the number of cells, but the thing is, as far as I understand we do not have this information for the test compound*cell_type pairs.\n\n**UPD: Crazy, but there are compound*cell type pairs with only one corresponding obs_id in adata, so the drug effect on this cell type was measured in just one cell???**\n\n Dear organizers, @danielburkhardt ? could you please tell if there is an explanation of such a discrepancy in the number of cells of different types, from which the data were collected for the drugs?\n\nAlso would be great to hear any thoughts on how we can deal with this.",
      "votes": 33
    },
    {
      "id": 2475679,
      "postDate": "2023-10-10T05:11:06.710Z",
      "content": "<p>Very insightful analysis!</p>\n<p>Predicting P values seems strange to me anyway… I've been wondering for a long time why this is formulated as a regression problem, and not formulated as a (pseudo) classification problem (something like applying a Tanh function to DE, and use the result in the range of [-1, 1] as the target…). </p>\n<p>As to how to deal with it, it would have been an important feature if the observatino counts were available for the test set. But it is not,  as it is only availble for the train set, maybe we can somehow use the observation counts as a control variable during training? Something like rewighting the loss based on the counts?.. But weighting is going to be subjective and big risk of overfitting to the LB…</p>",
      "rawMarkdown": "Very insightful analysis!\n\nPredicting P values seems strange to me anyway... I've been wondering for a long time why this is formulated as a regression problem, and not formulated as a (pseudo) classification problem (something like applying a Tanh function to DE, and use the result in the range of [-1, 1] as the target...). \n\n\nAs to how to deal with it, it would have been an important feature if the observatino counts were available for the test set. But it is not,  as it is only availble for the train set, maybe we can somehow use the observation counts as a control variable during training? Something like rewighting the loss based on the counts?.. But weighting is going to be subjective and big risk of overfitting to the LB...\n\n\n",
      "votes": 5
    },
    {
      "id": 2475551,
      "postDate": "2023-10-10T02:09:03.157Z",
      "content": "<p>Big thanks for your investigation.<br>\nWhen I try another approach by predicting the raw count of the test data and running RE, I realize that the metadata of the test set is not provided, even the number of cells as you pointed out. It means that building a model to directly predict the p-value is the only available approach for this competition, not as the host said in this <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/440706\" target=\"_blank\">discussion</a>: \"It's true that one valid approach would be to train a generative model to simulate inputs to the limma model and then run DE\"</p>",
      "rawMarkdown": "Big thanks for your investigation.\nWhen I try another approach by predicting the raw count of the test data and running RE, I realize that the metadata of the test set is not provided, even the number of cells as you pointed out. It means that building a model to directly predict the p-value is the only available approach for this competition, not as the host said in this [discussion](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/440706): \"It's true that one valid approach would be to train a generative model to simulate inputs to the limma model and then run DE\"",
      "votes": 4,
      "replies": [
        {
          "id": 2475617,
          "postDate": "2023-10-10T03:44:42.927Z",
          "content": "<p>Thanks for this addition. Yes, I think it's actually the most important - the lack of metadata of the test set, so we cannot predict the counts and perform the subsequent aggregation/DE analysis and LIMMA. I haven't seen the LIMMA yet, maybe we still can predict aggregated counts and transform them into the p-values though…  anyway, the second big issue exists - such a huge difference in the amount of cells per compound*cell type pair… and sometimes it was just one cell… </p>",
          "rawMarkdown": "Thanks for this addition. Yes, I think it's actually the most important - the lack of metadata of the test set, so we cannot predict the counts and perform the subsequent aggregation/DE analysis and LIMMA. I haven't seen the LIMMA yet, maybe we still can predict aggregated counts and transform them into the p-values though...  anyway, the second big issue exists - such a huge difference in the amount of cells per compound*cell type pair... and sometimes it was just one cell... ",
          "votes": 3,
          "replies": [
            {
              "id": 2478450,
              "postDate": "2023-10-12T02:57:24.900Z",
              "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> can you clarify this? Without the test metadata, we cannot approach the way you said in the mentioned discussion</p>",
              "rawMarkdown": "@danielburkhardt can you clarify this? Without the test metadata, we cannot approach the way you said in the mentioned discussion",
              "votes": 1
            },
            {
              "id": 2487564,
              "postDate": "2023-10-18T16:40:28.077Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/minhtu123\" target=\"_blank\">@minhtu123</a> and <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a> , sorry for the delay, we've been preparing a new version of the adata_train that should be posted shortly that might be more amenable to this kind of generative approach. Note, for a compound to be included in the test set, we applied a cutoff that we needed to have 20 cells in at least 2 donors to include the compound. All samples were used for DE, but you may want to ignore those samples. This is good fodder for the judges prize write-up.</p>\n<p>I've been talking to <a href=\"https://www.kaggle.com/andrewbenz\" target=\"_blank\">@andrewbenz</a> and we don't think you need the number of cells or the counts to do the generative approach and then run LIMMA. You should be able to set a number of cells to generate in the treatment condition and run the DE that way. It would be very interesting for the judges award to hear more about the challenges with doing this, understanding the sensitivity of the regression to the number of cells you generate, etc.</p>",
              "rawMarkdown": "Hi @minhtu123 and @antoninadolgorukova , sorry for the delay, we've been preparing a new version of the adata_train that should be posted shortly that might be more amenable to this kind of generative approach. Note, for a compound to be included in the test set, we applied a cutoff that we needed to have 20 cells in at least 2 donors to include the compound. All samples were used for DE, but you may want to ignore those samples. This is good fodder for the judges prize write-up.\n\nI've been talking to @andrewbenz and we don't think you need the number of cells or the counts to do the generative approach and then run LIMMA. You should be able to set a number of cells to generate in the treatment condition and run the DE that way. It would be very interesting for the judges award to hear more about the challenges with doing this, understanding the sensitivity of the regression to the number of cells you generate, etc.",
              "votes": 3
            },
            {
              "id": 2488105,
              "postDate": "2023-10-19T03:29:26.010Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a>, you mentioned a new version of <code>adata_train</code> will be released, could you please explain what are the changes in the new version?</p>",
              "rawMarkdown": "Hi @danielburkhardt, you mentioned a new version of `adata_train` will be released, could you please explain what are the changes in the new version?",
              "votes": 2
            },
            {
              "id": 2488851,
              "postDate": "2023-10-19T14:50:36.357Z",
              "content": "<p><a href=\"https://www.kaggle.com/andrewbenz\" target=\"_blank\">@andrewbenz</a> will provide release notes, but it's going to be mostly just filtering out the extra genes that aren't needed to run the DE analysis.</p>",
              "rawMarkdown": "@andrewbenz will provide release notes, but it's going to be mostly just filtering out the extra genes that aren't needed to run the DE analysis."
            },
            {
              "id": 2497933,
              "postDate": "2023-10-25T02:43:46.450Z",
              "content": "<p>when did you plan to release the updated version of adata_train?</p>",
              "rawMarkdown": "when did you plan to release the updated version of adata_train?",
              "votes": 2
            },
            {
              "id": 2506159,
              "postDate": "2023-10-31T06:09:02.837Z",
              "content": "<p>Hi! Sorry if I have missed it, but why would we need filtered data? Did you update smth more important for analysis and did you already release the adata_train - I could not find a topic with the release notes? </p>",
              "rawMarkdown": "Hi! Sorry if I have missed it, but why would we need filtered data? Did you update smth more important for analysis and did you already release the adata_train - I could not find a topic with the release notes? ",
              "votes": 1
            }
          ]
        },
        {
          "id": 2488101,
          "postDate": "2023-10-19T03:24:34.960Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/minhtu123\" target=\"_blank\">@minhtu123</a>, I don't think the metadata of the test set is critical, because the DE is calculated using the aggregated counts (pseudobulk), the sample size to calculate p-value are fixed. For a (single cell level) generative model, the critical part is to predict the gene expression levels of different donors.</p>\n<p>I think the bigger issue for generative model is that the final predition target is the <strong>p-value</strong> with a <strong>MSE</strong> like metric. The value of p-value is <strong>\"meanless\"</strong> in most cases and the MSE metric is <strong>strict</strong> to that value, make it very <strong>hard</strong> and <strong>unstable</strong> to predict, especially for generative model.</p>\n<p>I think maybe <strong>correlation</strong>-like metric is more suitable for this competition?</p>",
          "rawMarkdown": "Hi @minhtu123, I don't think the metadata of the test set is critical, because the DE is calculated using the aggregated counts (pseudobulk), the sample size to calculate p-value are fixed. For a (single cell level) generative model, the critical part is to predict the gene expression levels of different donors.\n\nI think the bigger issue for generative model is that the final predition target is the **p-value** with a **MSE** like metric. The value of p-value is **\"meanless\"** in most cases and the MSE metric is **strict** to that value, make it very **hard** and **unstable** to predict, especially for generative model.\n\nI think maybe **correlation**-like metric is more suitable for this competition?\n",
          "votes": 3,
          "replies": [
            {
              "id": 2489962,
              "postDate": "2023-10-20T10:34:42.540Z",
              "content": "<p>Totally agree. Maybe <strong>MSE</strong> for <strong>expression value</strong> instead of <strong>p-value</strong> makes more sense for the generative models?</p>",
              "rawMarkdown": "Totally agree. Maybe **MSE** for **expression value** instead of **p-value** makes more sense for the generative models?",
              "votes": 2
            },
            {
              "id": 2490097,
              "postDate": "2023-10-20T12:28:55.157Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2503652,
      "postDate": "2023-10-29T10:28:12.567Z",
      "content": "<p>I am interested in the DEG content too.<br>\nLog Fold Change (LFC) is typically calculated by comparing the target compound with a control. It appears that Belinostat, Dimethyl Sulfoxide, and Dabrafenib are used as controls. However, it's unclear which of these three is used as the control when calculating the LFC for each compound's Differentially Expressed Genes (DEGs). Has anyone investigated this?</p>",
      "rawMarkdown": "I am interested in the DEG content too.\nLog Fold Change (LFC) is typically calculated by comparing the target compound with a control. It appears that Belinostat, Dimethyl Sulfoxide, and Dabrafenib are used as controls. However, it's unclear which of these three is used as the control when calculating the LFC for each compound's Differentially Expressed Genes (DEGs). Has anyone investigated this?",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2475679,
      "author_name": "qihuaz",
      "author_url": "",
      "post_date": "2023-10-10T05:11:06.710000",
      "content": "<p>Very insightful analysis!</p>\n<p>Predicting P values seems strange to me anyway… I've been wondering for a long time why this is formulated as a regression problem, and not formulated as a (pseudo) classification problem (something like applying a Tanh function to DE, and use the result in the range of [-1, 1] as the target…). </p>\n<p>As to how to deal with it, it would have been an important feature if the observatino counts were available for the test set. But it is not,  as it is only availble for the train set, maybe we can somehow use the observation counts as a control variable during training? Something like rewighting the loss based on the counts?.. But weighting is going to be subjective and big risk of overfitting to the LB…</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2475551,
      "author_name": "minhtu.mt.mt",
      "author_url": "",
      "post_date": "2023-10-10T02:09:03.157000",
      "content": "<p>Big thanks for your investigation.<br>\nWhen I try another approach by predicting the raw count of the test data and running RE, I realize that the metadata of the test set is not provided, even the number of cells as you pointed out. It means that building a model to directly predict the p-value is the only available approach for this competition, not as the host said in this <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/440706\" target=\"_blank\">discussion</a>: \"It's true that one valid approach would be to train a generative model to simulate inputs to the limma model and then run DE\"</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2475617,
          "author_name": "Antonina Dolgorukova",
          "author_url": "",
          "post_date": "2023-10-10T03:44:42.927000",
          "content": "<p>Thanks for this addition. Yes, I think it's actually the most important - the lack of metadata of the test set, so we cannot predict the counts and perform the subsequent aggregation/DE analysis and LIMMA. I haven't seen the LIMMA yet, maybe we still can predict aggregated counts and transform them into the p-values though…  anyway, the second big issue exists - such a huge difference in the amount of cells per compound*cell type pair… and sometimes it was just one cell… </p>",
          "votes": 3,
          "replies": [
            {
              "id": 2478450,
              "author_name": "minhtu.mt.mt",
              "author_url": "",
              "post_date": "2023-10-12T02:57:24.900000",
              "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> can you clarify this? Without the test metadata, we cannot approach the way you said in the mentioned discussion</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2487564,
              "author_name": "Daniel Burkhardt",
              "author_url": "",
              "post_date": "2023-10-18T16:40:28.077000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/minhtu123\" target=\"_blank\">@minhtu123</a> and <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a> , sorry for the delay, we've been preparing a new version of the adata_train that should be posted shortly that might be more amenable to this kind of generative approach. Note, for a compound to be included in the test set, we applied a cutoff that we needed to have 20 cells in at least 2 donors to include the compound. All samples were used for DE, but you may want to ignore those samples. This is good fodder for the judges prize write-up.</p>\n<p>I've been talking to <a href=\"https://www.kaggle.com/andrewbenz\" target=\"_blank\">@andrewbenz</a> and we don't think you need the number of cells or the counts to do the generative approach and then run LIMMA. You should be able to set a number of cells to generate in the treatment condition and run the DE that way. It would be very interesting for the judges award to hear more about the challenges with doing this, understanding the sensitivity of the regression to the number of cells you generate, etc.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2488105,
              "author_name": "hoho",
              "author_url": "",
              "post_date": "2023-10-19T03:29:26.010000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a>, you mentioned a new version of <code>adata_train</code> will be released, could you please explain what are the changes in the new version?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2488851,
              "author_name": "Daniel Burkhardt",
              "author_url": "",
              "post_date": "2023-10-19T14:50:36.357000",
              "content": "<p><a href=\"https://www.kaggle.com/andrewbenz\" target=\"_blank\">@andrewbenz</a> will provide release notes, but it's going to be mostly just filtering out the extra genes that aren't needed to run the DE analysis.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2497933,
              "author_name": "FGPC",
              "author_url": "",
              "post_date": "2023-10-25T02:43:46.450000",
              "content": "<p>when did you plan to release the updated version of adata_train?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2506159,
              "author_name": "Antonina Dolgorukova",
              "author_url": "",
              "post_date": "2023-10-31T06:09:02.837000",
              "content": "<p>Hi! Sorry if I have missed it, but why would we need filtered data? Did you update smth more important for analysis and did you already release the adata_train - I could not find a topic with the release notes? </p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2488101,
          "author_name": "hoho",
          "author_url": "",
          "post_date": "2023-10-19T03:24:34.960000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/minhtu123\" target=\"_blank\">@minhtu123</a>, I don't think the metadata of the test set is critical, because the DE is calculated using the aggregated counts (pseudobulk), the sample size to calculate p-value are fixed. For a (single cell level) generative model, the critical part is to predict the gene expression levels of different donors.</p>\n<p>I think the bigger issue for generative model is that the final predition target is the <strong>p-value</strong> with a <strong>MSE</strong> like metric. The value of p-value is <strong>\"meanless\"</strong> in most cases and the MSE metric is <strong>strict</strong> to that value, make it very <strong>hard</strong> and <strong>unstable</strong> to predict, especially for generative model.</p>\n<p>I think maybe <strong>correlation</strong>-like metric is more suitable for this competition?</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2489962,
              "author_name": "Daniel Shao",
              "author_url": "",
              "post_date": "2023-10-20T10:34:42.540000",
              "content": "<p>Totally agree. Maybe <strong>MSE</strong> for <strong>expression value</strong> instead of <strong>p-value</strong> makes more sense for the generative models?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2490097,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-10-20T12:28:55.157000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2503652,
      "author_name": "yoshitown",
      "author_url": "",
      "post_date": "2023-10-29T10:28:12.567000",
      "content": "<p>I am interested in the DEG content too.<br>\nLog Fold Change (LFC) is typically calculated by comparing the target compound with a control. It appears that Belinostat, Dimethyl Sulfoxide, and Dabrafenib are used as controls. However, it's unclear which of these three is used as the control when calculating the LFC for each compound's Differentially Expressed Genes (DEGs). Has anyone investigated this?</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2474649": "I've just examined a bit deeper the adata_train.parquet file and corresponding meta data [here](https://www.kaggle.com/code/antoninadolgorukova/op2-stats-sign-de-fingerprints-eda#Unaggregated-counts).\n\nIn the adata_train, here are 416,442,312 rows; 240,090 observation ids, 21255 genes (this is a superset of de_train, so the number of genes is higher). \n\nIn the adata_obs_meta I found 240090 rows, each id (identifiers assigned to each cell in the raw dataset) has one row. So I plotted the number of cells, from which the data were collected, per compound and per compounds-cell type pair. The last is most interesting:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2Ffbb156c7e09c6a0c31c98f6de5d26d76%2FScreenshot%202023-10-09%20175613.png?generation=1696849122129385&alt=media)\n\nHere are the number of B and Myeloid cells cells from which were collected the data for 15 compounds (train data). We are supposed to predict the impact of other compounds on these cells. But it seems that there as at least three outliers, for wchich the data were collected from just a few cells (unreliable?)\nThe MLN 2238 is among them, and it has been already noted by @qihuaz as an outlier based on its effect.\n\nHere is more:\nThere seems to be a huge difference in the number of CD8+ T cells and T regulatory cells\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F2e176fecab41b2c708d3f3e241d6c957%2FScreenshot%202023-10-09%20175736.png?generation=1696849237309259&alt=media)\n\ncompared to CD4+ T cells and NK cells\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F4c15dc18ae23a6d2e61a56b5df83208c%2FScreenshot%202023-10-09%20175655.png?generation=1696849183343461&alt=media)\n\n\nGiven this, one may speculate that the effect may vary less or more depending on the number of cells in which it is measured, affecting the p-values obtained (for all three controls btw the data were collected from a much larger number of cells than for the other compounds - more robust data). What I mean is that in a small number of cells, the variance of the drug effect can be large and the p-value is high, but if you treat more cells, the variance will reduce and the effect will appear, so the p-value will be lower. So, to get reliable predictions, it may be reasonable to account for the number of cells, but the thing is, as far as I understand we do not have this information for the test compound*cell_type pairs.\n\n**UPD: Crazy, but there are compound*cell type pairs with only one corresponding obs_id in adata, so the drug effect on this cell type was measured in just one cell???**\n\n Dear organizers, @danielburkhardt ? could you please tell if there is an explanation of such a discrepancy in the number of cells of different types, from which the data were collected for the drugs?\n\nAlso would be great to hear any thoughts on how we can deal with this.",
    "2475679": "Very insightful analysis!\n\nPredicting P values seems strange to me anyway... I've been wondering for a long time why this is formulated as a regression problem, and not formulated as a (pseudo) classification problem (something like applying a Tanh function to DE, and use the result in the range of [-1, 1] as the target...). \n\n\nAs to how to deal with it, it would have been an important feature if the observatino counts were available for the test set. But it is not,  as it is only availble for the train set, maybe we can somehow use the observation counts as a control variable during training? Something like rewighting the loss based on the counts?.. But weighting is going to be subjective and big risk of overfitting to the LB...\n\n\n",
    "2475551": "Big thanks for your investigation.\nWhen I try another approach by predicting the raw count of the test data and running RE, I realize that the metadata of the test set is not provided, even the number of cells as you pointed out. It means that building a model to directly predict the p-value is the only available approach for this competition, not as the host said in this [discussion](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/440706): \"It's true that one valid approach would be to train a generative model to simulate inputs to the limma model and then run DE\"",
    "2503652": "I am interested in the DEG content too.\nLog Fold Change (LFC) is typically calculated by comparing the target compound with a control. It appears that Belinostat, Dimethyl Sulfoxide, and Dabrafenib are used as controls. However, it's unclear which of these three is used as the control when calculating the LFC for each compound's Differentially Expressed Genes (DEGs). Has anyone investigated this?"
  }
}