{
  "id": 458893,
  "title": "Why scGEN, chemCPA, graphVCI, etc do not work？",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/458893",
  "author_name": "",
  "post_date": "2023-12-02T06:55:00.813980Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>I'm participating in a Kaggle competition for the first time and have been exploring various methods mentioned in papers like scGEN, CPA, graphVCI，etc. However, I'm struggling to achieve good results with these approaches.</p>\n<p>My current strategy involves using the raw data from adata_train. I've been standardizing and log-transforming this data before feeding it into the model. Afterward, I restore the output results according to the library size of DMSO, followed by pseudo bulking and using limma for analysis. Despite these efforts, the model's performance has not been satisfactory.</p>\n<p>I'm reaching out to see if anyone here has experience with similar methods or strategies in their projects. </p>\n<p>Thank you in advance for your help!</p>",
  "messages": [
    {
      "id": "2546160",
      "postDate": "12/02/2023 06:55:00",
      "content": "<p>Hello everyone,</p>\n<p>I'm participating in a Kaggle competition for the first time and have been exploring various methods mentioned in papers like scGEN, CPA, graphVCI，etc. However, I'm struggling to achieve good results with these approaches.</p>\n<p>My current strategy involves using the raw data from adata_train. I've been standardizing and log-transforming this data before feeding it into the model. Afterward, I restore the output results according to the library size of DMSO, followed by pseudo bulking and using limma for analysis. Despite these efforts, the model's performance has not been satisfactory.</p>\n<p>I'm reaching out to see if anyone here has experience with similar methods or strategies in their projects. </p>\n<p>Thank you in advance for your help!</p>",
      "rawMarkdown": "Hello everyone,\n\nI'm participating in a Kaggle competition for the first time and have been exploring various methods mentioned in papers like scGEN, CPA, graphVCI，etc. However, I'm struggling to achieve good results with these approaches.\n\nMy current strategy involves using the raw data from adata_train. I've been standardizing and log-transforming this data before feeding it into the model. Afterward, I restore the output results according to the library size of DMSO, followed by pseudo bulking and using limma for analysis. Despite these efforts, the model's performance has not been satisfactory.\n\nI'm reaching out to see if anyone here has experience with similar methods or strategies in their projects. \n\nThank you in advance for your help!",
      "votes": null
    },
    {
      "id": "2547256",
      "postDate": "12/03/2023 10:44:54",
      "content": "<p>I have not analyzed that much yet, so what I write below might not be fully correct.</p>\n<p>First, I would be surprised if any of that would work )))</p>\n<p>The reason is that biological data on which these models were trained seems to be different from what we have here.<br>\nThe difference  is technical - not conceptual, but in cases I met so far such difference might be crucial.</p>\n<p>For example you know about batch effect problem and despite many approaches to resolve it - not always you should trust these solutions, because  \"technical\" - vs - \"biological\"  - not a well defined thing. We may think it is technical, but actually it is biological and vice versa.</p>\n<p>So it seems to me: <br>\n it is very hard to predict drug effects, the predictions \"FRAGILE\" - if you change the input data - may expect crucial degradadtion of the performance.<br>\nAnd that what probably happens - concerning for example ChemCPA - it was (mainly) trained on L1000 data - which different measurment technology - not RNA-seq and measuring only 978 genes, the cell lines are different - cancer cell lines, not the normal cells. </p>\n<p>So  - even small difference - like absolutely the same technology and cell lines , but just do it in different labs - would cause a batch effect which is already may be seen and in some cases deserves special attention - typically not a great problem, but still.<br>\nAnd here we have - everything(!) is quite different - technology, cell lines,  computational  processing (Limma).<br>\nSo seems natural to expect everything will fail - at least to my restricted understanding of the situation.</p>",
      "rawMarkdown": "I have not analyzed that much yet, so what I write below might not be fully correct.\n\nFirst, I would be surprised if any of that would work )))\n\nThe reason is that biological data on which these models were trained seems to be different from what we have here.\nThe difference  is technical - not conceptual, but in cases I met so far such difference might be crucial.\n\nFor example you know about batch effect problem and despite many approaches to resolve it - not always you should trust these solutions, because  \"technical\" - vs - \"biological\"  - not a well defined thing. We may think it is technical, but actually it is biological and vice versa.\n\nSo it seems to me: \n it is very hard to predict drug effects, the predictions \"FRAGILE\" - if you change the input data - may expect crucial degradadtion of the performance.\nAnd that what probably happens - concerning for example ChemCPA - it was (mainly) trained on L1000 data - which different measurment technology - not RNA-seq and measuring only 978 genes, the cell lines are different - cancer cell lines, not the normal cells. \n\nSo  - even small difference - like absolutely the same technology and cell lines , but just do it in different labs - would cause a batch effect which is already may be seen and in some cases deserves special attention - typically not a great problem, but still.\nAnd here we have - everything(!) is quite different - technology, cell lines,  computational  processing (Limma).\nSo seems natural to expect everything will fail - at least to my restricted understanding of the situation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2547256,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "12/03/2023 10:44:54",
      "content": "<p>I have not analyzed that much yet, so what I write below might not be fully correct.</p>\n<p>First, I would be surprised if any of that would work )))</p>\n<p>The reason is that biological data on which these models were trained seems to be different from what we have here.<br>\nThe difference  is technical - not conceptual, but in cases I met so far such difference might be crucial.</p>\n<p>For example you know about batch effect problem and despite many approaches to resolve it - not always you should trust these solutions, because  \"technical\" - vs - \"biological\"  - not a well defined thing. We may think it is technical, but actually it is biological and vice versa.</p>\n<p>So it seems to me: <br>\n it is very hard to predict drug effects, the predictions \"FRAGILE\" - if you change the input data - may expect crucial degradadtion of the performance.<br>\nAnd that what probably happens - concerning for example ChemCPA - it was (mainly) trained on L1000 data - which different measurment technology - not RNA-seq and measuring only 978 genes, the cell lines are different - cancer cell lines, not the normal cells. </p>\n<p>So  - even small difference - like absolutely the same technology and cell lines , but just do it in different labs - would cause a batch effect which is already may be seen and in some cases deserves special attention - typically not a great problem, but still.<br>\nAnd here we have - everything(!) is quite different - technology, cell lines,  computational  processing (Limma).<br>\nSo seems natural to expect everything will fail - at least to my restricted understanding of the situation.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2546160": "Hello everyone,\n\nI'm participating in a Kaggle competition for the first time and have been exploring various methods mentioned in papers like scGEN, CPA, graphVCI，etc. However, I'm struggling to achieve good results with these approaches.\n\nMy current strategy involves using the raw data from adata_train. I've been standardizing and log-transforming this data before feeding it into the model. Afterward, I restore the output results according to the library size of DMSO, followed by pseudo bulking and using limma for analysis. Despite these efforts, the model's performance has not been satisfactory.\n\nI'm reaching out to see if anyone here has experience with similar methods or strategies in their projects. \n\nThank you in advance for your help!",
    "2547256": "I have not analyzed that much yet, so what I write below might not be fully correct.\n\nFirst, I would be surprised if any of that would work )))\n\nThe reason is that biological data on which these models were trained seems to be different from what we have here.\nThe difference  is technical - not conceptual, but in cases I met so far such difference might be crucial.\n\nFor example you know about batch effect problem and despite many approaches to resolve it - not always you should trust these solutions, because  \"technical\" - vs - \"biological\"  - not a well defined thing. We may think it is technical, but actually it is biological and vice versa.\n\nSo it seems to me: \n it is very hard to predict drug effects, the predictions \"FRAGILE\" - if you change the input data - may expect crucial degradadtion of the performance.\nAnd that what probably happens - concerning for example ChemCPA - it was (mainly) trained on L1000 data - which different measurment technology - not RNA-seq and measuring only 978 genes, the cell lines are different - cancer cell lines, not the normal cells. \n\nSo  - even small difference - like absolutely the same technology and cell lines , but just do it in different labs - would cause a batch effect which is already may be seen and in some cases deserves special attention - typically not a great problem, but still.\nAnd here we have - everything(!) is quite different - technology, cell lines,  computational  processing (Limma).\nSo seems natural to expect everything will fail - at least to my restricted understanding of the situation."
  },
  "source": "meta"
}