{
  "id": 444668,
  "title": "About reproducing the targets from training data",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/444668",
  "author_name": "minhtu.mt.mt",
  "post_date": "2023-10-03T04:43:03.844000",
  "votes": 10,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I am trying to understand and reproduce the process of creating the targets from the raw counts as described in the Dataset Description. <br>\nI was able to create the pseudobulked counts and fit them into a linear model using Limma model. But my result (p_value) is quite different from the given labels and I have no idea how to correctly mimic the process made by the host. To me, the description is only clear about pseudobulking. After that, the Differential gene expression analysis is puzzling, I read all the references and saw a lot of options and design matrices when doing DE analysis. Add on with Multiplexing and multiple testing correction concepts, it is impossible for me to reproduce the DE analysis to create the targets for the training data.<br>\nThe host said in <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/439017\" target=\"_blank\">this discussion</a> that they will provide the code but it has been more than 2 weeks from that time.<br>\nSo, in case running the process on Kaggle is a bit challenging, please give us the pseudo-code or detailed steps so that we can do the engineering by ourselves</p>",
  "messages": [
    {
      "id": 2465546,
      "postDate": "2023-10-03T04:43:03.843Z",
      "content": "<p>I am trying to understand and reproduce the process of creating the targets from the raw counts as described in the Dataset Description. <br>\nI was able to create the pseudobulked counts and fit them into a linear model using Limma model. But my result (p_value) is quite different from the given labels and I have no idea how to correctly mimic the process made by the host. To me, the description is only clear about pseudobulking. After that, the Differential gene expression analysis is puzzling, I read all the references and saw a lot of options and design matrices when doing DE analysis. Add on with Multiplexing and multiple testing correction concepts, it is impossible for me to reproduce the DE analysis to create the targets for the training data.<br>\nThe host said in <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/439017\" target=\"_blank\">this discussion</a> that they will provide the code but it has been more than 2 weeks from that time.<br>\nSo, in case running the process on Kaggle is a bit challenging, please give us the pseudo-code or detailed steps so that we can do the engineering by ourselves</p>",
      "rawMarkdown": "I am trying to understand and reproduce the process of creating the targets from the raw counts as described in the Dataset Description. \nI was able to create the pseudobulked counts and fit them into a linear model using Limma model. But my result (p_value) is quite different from the given labels and I have no idea how to correctly mimic the process made by the host. To me, the description is only clear about pseudobulking. After that, the Differential gene expression analysis is puzzling, I read all the references and saw a lot of options and design matrices when doing DE analysis. Add on with Multiplexing and multiple testing correction concepts, it is impossible for me to reproduce the DE analysis to create the targets for the training data.\nThe host said in [this discussion](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/439017) that they will provide the code but it has been more than 2 weeks from that time.\nSo, in case running the process on Kaggle is a bit challenging, please give us the pseudo-code or detailed steps so that we can do the engineering by ourselves",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2465546": "I am trying to understand and reproduce the process of creating the targets from the raw counts as described in the Dataset Description. \nI was able to create the pseudobulked counts and fit them into a linear model using Limma model. But my result (p_value) is quite different from the given labels and I have no idea how to correctly mimic the process made by the host. To me, the description is only clear about pseudobulking. After that, the Differential gene expression analysis is puzzling, I read all the references and saw a lot of options and design matrices when doing DE analysis. Add on with Multiplexing and multiple testing correction concepts, it is impossible for me to reproduce the DE analysis to create the targets for the training data.\nThe host said in [this discussion](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/439017) that they will provide the code but it has been more than 2 weeks from that time.\nSo, in case running the process on Kaggle is a bit challenging, please give us the pseudo-code or detailed steps so that we can do the engineering by ourselves"
  }
}