{
  "id": 448556,
  "title": "Mapping between differential gene expressions (DE) and log(p-value) s",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/448556",
  "author_name": "",
  "post_date": "2023-10-20T06:58:52.146572500Z",
  "votes": 8,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I am trying to better understand the actual generative process for the signal we are trying to construct the model for : </p>\n<p>\"Diff. Gene Expression\" y =  F(CellType,Compound) </p>\n<p>The output, y,  is defined as  (\"sign\" of the actual Log Fold Change for Differential Expression ) * ( - log(p-values)). Description is based on de-train.parquet training data is described as:</p>\n<p>**de_train.parquet - Aggregated differential expression data in dense array format.<br>\ngenes A1BG, A1BG-AS1, …, ZZEF1 (numbering 18,211 in all) - Differential expression value (-log10(p-value) * sign(LFC)) for each gene. Here, LFC is the estimated log-fold change in expression between the treatment and control condition after shrinkage as calculated by Limma. Positive LFC means the gene goes up in the treatment condition relative to the control. **</p>\n<p>Specifically, how is p-value related to the actual differential expression (DE) here? I am not able to decipher this connection from the \"Limma\" formulation explanation offered in the overview. For example with compound X,  if Gene A had DE of +2 Fold  and Gene B had DE of +10 Fold, what is the functional connection between DE(A) and p-value(A). Is it possible for Gene A and Gene B to have identical p-values while having drastically different differential expressions like +2 vs +10? </p>",
  "messages": [
    {
      "id": "2489718",
      "postDate": "10/20/2023 06:58:52",
      "content": "<p>I am trying to better understand the actual generative process for the signal we are trying to construct the model for : </p>\n<p>\"Diff. Gene Expression\" y =  F(CellType,Compound) </p>\n<p>The output, y,  is defined as  (\"sign\" of the actual Log Fold Change for Differential Expression ) * ( - log(p-values)). Description is based on de-train.parquet training data is described as:</p>\n<p>**de_train.parquet - Aggregated differential expression data in dense array format.<br>\ngenes A1BG, A1BG-AS1, …, ZZEF1 (numbering 18,211 in all) - Differential expression value (-log10(p-value) * sign(LFC)) for each gene. Here, LFC is the estimated log-fold change in expression between the treatment and control condition after shrinkage as calculated by Limma. Positive LFC means the gene goes up in the treatment condition relative to the control. **</p>\n<p>Specifically, how is p-value related to the actual differential expression (DE) here? I am not able to decipher this connection from the \"Limma\" formulation explanation offered in the overview. For example with compound X,  if Gene A had DE of +2 Fold  and Gene B had DE of +10 Fold, what is the functional connection between DE(A) and p-value(A). Is it possible for Gene A and Gene B to have identical p-values while having drastically different differential expressions like +2 vs +10? </p>",
      "rawMarkdown": "I am trying to better understand the actual generative process for the signal we are trying to construct the model for : \n\n\"Diff. Gene Expression\" y =  F(CellType,Compound) \n\nThe output, y,  is defined as  (\"sign\" of the actual Log Fold Change for Differential Expression ) * ( - log(p-values)). Description is based on de-train.parquet training data is described as:\n\n**de_train.parquet - Aggregated differential expression data in dense array format.\ngenes A1BG, A1BG-AS1, …, ZZEF1 (numbering 18,211 in all) - Differential expression value (-log10(p-value) * sign(LFC)) for each gene. Here, LFC is the estimated log-fold change in expression between the treatment and control condition after shrinkage as calculated by Limma. Positive LFC means the gene goes up in the treatment condition relative to the control. **\n\nSpecifically, how is p-value related to the actual differential expression (DE) here? I am not able to decipher this connection from the \"Limma\" formulation explanation offered in the overview. For example with compound X,  if Gene A had DE of +2 Fold  and Gene B had DE of +10 Fold, what is the functional connection between DE(A) and p-value(A). Is it possible for Gene A and Gene B to have identical p-values while having drastically different differential expressions like +2 vs +10?",
      "votes": null
    },
    {
      "id": "2489939",
      "postDate": "10/20/2023 10:13:07",
      "content": "<blockquote>\n  <ul>\n  <li>Specifically, how is p-value related to the actual differential expression (DE) here? </li>\n  </ul>\n</blockquote>\n<p>A bigger <code>abs(LFC)</code> will result in a smaller p-value, a bigger <code>-log10(p-value)</code>.</p>\n<hr>\n<blockquote>\n  <ul>\n  <li><p>What is the functional connection between DE(A) and p-value(A)?</p></li>\n  <li><p>Is it possible for Gene A and Gene B to have identical p-values while having drastically different differential expressions like +2 vs +10?</p></li>\n  </ul>\n</blockquote>\n<p>Interesting question, I think there is not a simple function that <code>f(LFC) = -logP</code>, because other factors such as the sample size and the distribution (variance) will affect the p-value. vice versa.</p>\n<p>Here are some examples:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1023301%2F9b7e0e9214266c54c890102d984891ff%2FDingtalk_20231020180013.jpg?generation=1697796028133628&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1023301%2Fbd5f0c47c4fd6e65cb300a0e5b1bb808%2FDingtalk_20231020180851.jpg?generation=1697796559157283&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p>The <strong>Volcano plot</strong> demonstrated the relationship between <code>LFC</code> and <code>-log10(p-value)</code>, and it also showed the <code>-log10(p-value)</code> are <strong>\"unstable\"</strong> when the <code>abs(LFC)</code> is big.</p>\n<p>You can find the plot and the examples in my <a href=\"https://www.kaggle.com/code/awater1223/op2-04-merge-limma-results\" target=\"_blank\">notebook</a>.</p>",
      "rawMarkdown": "> - Specifically, how is p-value related to the actual differential expression (DE) here? \n\nA bigger `abs(LFC)` will result in a smaller p-value, a bigger `-log10(p-value)`.\n\n---\n\n> - What is the functional connection between DE(A) and p-value(A)?\n>\n> - Is it possible for Gene A and Gene B to have identical p-values while having drastically different differential expressions like +2 vs +10?\n> \n\nInteresting question, I think there is not a simple function that `f(LFC) = -logP`, because other factors such as the sample size and the distribution (variance) will affect the p-value. vice versa.\n\nHere are some examples:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1023301%2F9b7e0e9214266c54c890102d984891ff%2FDingtalk_20231020180013.jpg?generation=1697796028133628&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1023301%2Fbd5f0c47c4fd6e65cb300a0e5b1bb808%2FDingtalk_20231020180851.jpg?generation=1697796559157283&alt=media)\n\n---\n\nThe **Volcano plot** demonstrated the relationship between `LFC` and `-log10(p-value)`, and it also showed the `-log10(p-value)` are **\"unstable\"** when the `abs(LFC)` is big.\n\nYou can find the plot and the examples in my [notebook](https://www.kaggle.com/code/awater1223/op2-04-merge-limma-results).",
      "votes": null
    },
    {
      "id": "2490472",
      "postDate": "10/20/2023 17:54:29",
      "content": "<p>Thank you for your helpful reply! What I gathered from your note is that the p-value can be seen as the probability of there being a differential \"change.\" Naturally, the higher the log fold change, the higher this probability (including the sign/direction of change).</p>\n<p>It appears that what we are trying to model and predict is <strong>NOT</strong> the continuous value of differential gene expression (a regression problem) but rather an indirect function of it, the probability of whether there was a change in the expression of Gene A (including the direction +/-), as a binary classification output [DE=Yes or DE=No change] . Your response, along with the useful notebook results showing instability in p-value outputs for the same or very similar DE, suggests the presence of nuisance factors (e.g., sample size) in the mapping from actual gene expression to p-value-based model output. This makes the output model appear more like y=G( F(T, C); n), where G is the additional mapping from the direct quantitative differential gene expression to our prediction variable.</p>\n<p>How do we capture and model G(F(.);n) if it's introducing unknown variability via the hidden nuisance parameter 'n'? Since the evaluation cost function assumes a regression model (rather than just predicting the categorical \"sign\" of DE: -1, 1, or 0), wouldn't this introduce a \"random\" factor in scoring?</p>\n<p>Imagine you had a magic wand that gave you the \"ideal\" direct differential expression predictor function DE=F(T,C). However, once you put it through G(DE;n), the answer gets diluted. I'm concerned that depending on the level of this dilution and considering how closely the scores stack up, attempting to produce the closest score matching the \"ground truth\" p-values of the test set might be a matter of chance rather than an accurate representation of the underlying generative biological process. Wouldn't it capture the process more truthfully to try to predict directly the differential expression rather than a p-values derivative - or define the cost function on a categorical prediction task (DE = [1, -1, 0])?   Thank you for your insights.</p>",
      "rawMarkdown": "Thank you for your helpful reply! What I gathered from your note is that the p-value can be seen as the probability of there being a differential \"change.\" Naturally, the higher the log fold change, the higher this probability (including the sign/direction of change).\n\nIt appears that what we are trying to model and predict is **NOT** the continuous value of differential gene expression (a regression problem) but rather an indirect function of it, the probability of whether there was a change in the expression of Gene A (including the direction +/-), as a binary classification output [DE=Yes or DE=No change] . Your response, along with the useful notebook results showing instability in p-value outputs for the same or very similar DE, suggests the presence of nuisance factors (e.g., sample size) in the mapping from actual gene expression to p-value-based model output. This makes the output model appear more like y=G( F(T, C); n), where G is the additional mapping from the direct quantitative differential gene expression to our prediction variable.\n\nHow do we capture and model G(F(.);n) if it's introducing unknown variability via the hidden nuisance parameter 'n'? Since the evaluation cost function assumes a regression model (rather than just predicting the categorical \"sign\" of DE: -1, 1, or 0), wouldn't this introduce a \"random\" factor in scoring?\n\nImagine you had a magic wand that gave you the \"ideal\" direct differential expression predictor function DE=F(T,C). However, once you put it through G(DE;n), the answer gets diluted. I'm concerned that depending on the level of this dilution and considering how closely the scores stack up, attempting to produce the closest score matching the \"ground truth\" p-values of the test set might be a matter of chance rather than an accurate representation of the underlying generative biological process. Wouldn't it capture the process more truthfully to try to predict directly the differential expression rather than a p-values derivative - or define the cost function on a categorical prediction task (DE = [1, -1, 0])?   Thank you for your insights.",
      "votes": null
    },
    {
      "id": "2490535",
      "postDate": "10/20/2023 19:05:35",
      "content": "<p>I actually agree with you… And have this concern for quite a while. Since I don't think the target/evaluation is going to change for this competition, we have to work with what we have.</p>\n<p>So in my opinion, a big part of the solution for this competition will be fitting the metric…</p>",
      "rawMarkdown": "I actually agree with you... And have this concern for quite a while. Since I don't think the target/evaluation is going to change for this competition, we have to work with what we have.\n\nSo in my opinion, a big part of the solution for this competition will be fitting the metric...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2489939,
      "author_name": "awater1223",
      "author_url": "",
      "post_date": "10/20/2023 10:13:07",
      "content": "<blockquote>\n  <ul>\n  <li>Specifically, how is p-value related to the actual differential expression (DE) here? </li>\n  </ul>\n</blockquote>\n<p>A bigger <code>abs(LFC)</code> will result in a smaller p-value, a bigger <code>-log10(p-value)</code>.</p>\n<hr>\n<blockquote>\n  <ul>\n  <li><p>What is the functional connection between DE(A) and p-value(A)?</p></li>\n  <li><p>Is it possible for Gene A and Gene B to have identical p-values while having drastically different differential expressions like +2 vs +10?</p></li>\n  </ul>\n</blockquote>\n<p>Interesting question, I think there is not a simple function that <code>f(LFC) = -logP</code>, because other factors such as the sample size and the distribution (variance) will affect the p-value. vice versa.</p>\n<p>Here are some examples:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1023301%2F9b7e0e9214266c54c890102d984891ff%2FDingtalk_20231020180013.jpg?generation=1697796028133628&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1023301%2Fbd5f0c47c4fd6e65cb300a0e5b1bb808%2FDingtalk_20231020180851.jpg?generation=1697796559157283&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p>The <strong>Volcano plot</strong> demonstrated the relationship between <code>LFC</code> and <code>-log10(p-value)</code>, and it also showed the <code>-log10(p-value)</code> are <strong>\"unstable\"</strong> when the <code>abs(LFC)</code> is big.</p>\n<p>You can find the plot and the examples in my <a href=\"https://www.kaggle.com/code/awater1223/op2-04-merge-limma-results\" target=\"_blank\">notebook</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2490472,
      "author_name": "nanocipher",
      "author_url": "",
      "post_date": "10/20/2023 17:54:29",
      "content": "<p>Thank you for your helpful reply! What I gathered from your note is that the p-value can be seen as the probability of there being a differential \"change.\" Naturally, the higher the log fold change, the higher this probability (including the sign/direction of change).</p>\n<p>It appears that what we are trying to model and predict is <strong>NOT</strong> the continuous value of differential gene expression (a regression problem) but rather an indirect function of it, the probability of whether there was a change in the expression of Gene A (including the direction +/-), as a binary classification output [DE=Yes or DE=No change] . Your response, along with the useful notebook results showing instability in p-value outputs for the same or very similar DE, suggests the presence of nuisance factors (e.g., sample size) in the mapping from actual gene expression to p-value-based model output. This makes the output model appear more like y=G( F(T, C); n), where G is the additional mapping from the direct quantitative differential gene expression to our prediction variable.</p>\n<p>How do we capture and model G(F(.);n) if it's introducing unknown variability via the hidden nuisance parameter 'n'? Since the evaluation cost function assumes a regression model (rather than just predicting the categorical \"sign\" of DE: -1, 1, or 0), wouldn't this introduce a \"random\" factor in scoring?</p>\n<p>Imagine you had a magic wand that gave you the \"ideal\" direct differential expression predictor function DE=F(T,C). However, once you put it through G(DE;n), the answer gets diluted. I'm concerned that depending on the level of this dilution and considering how closely the scores stack up, attempting to produce the closest score matching the \"ground truth\" p-values of the test set might be a matter of chance rather than an accurate representation of the underlying generative biological process. Wouldn't it capture the process more truthfully to try to predict directly the differential expression rather than a p-values derivative - or define the cost function on a categorical prediction task (DE = [1, -1, 0])?   Thank you for your insights.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2490535,
          "author_name": "qihuaz",
          "author_url": "",
          "post_date": "10/20/2023 19:05:35",
          "content": "<p>I actually agree with you… And have this concern for quite a while. Since I don't think the target/evaluation is going to change for this competition, we have to work with what we have.</p>\n<p>So in my opinion, a big part of the solution for this competition will be fitting the metric…</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2489718": "I am trying to better understand the actual generative process for the signal we are trying to construct the model for : \n\n\"Diff. Gene Expression\" y =  F(CellType,Compound) \n\nThe output, y,  is defined as  (\"sign\" of the actual Log Fold Change for Differential Expression ) * ( - log(p-values)). Description is based on de-train.parquet training data is described as:\n\n**de_train.parquet - Aggregated differential expression data in dense array format.\ngenes A1BG, A1BG-AS1, …, ZZEF1 (numbering 18,211 in all) - Differential expression value (-log10(p-value) * sign(LFC)) for each gene. Here, LFC is the estimated log-fold change in expression between the treatment and control condition after shrinkage as calculated by Limma. Positive LFC means the gene goes up in the treatment condition relative to the control. **\n\nSpecifically, how is p-value related to the actual differential expression (DE) here? I am not able to decipher this connection from the \"Limma\" formulation explanation offered in the overview. For example with compound X,  if Gene A had DE of +2 Fold  and Gene B had DE of +10 Fold, what is the functional connection between DE(A) and p-value(A). Is it possible for Gene A and Gene B to have identical p-values while having drastically different differential expressions like +2 vs +10?",
    "2489939": "> - Specifically, how is p-value related to the actual differential expression (DE) here? \n\nA bigger `abs(LFC)` will result in a smaller p-value, a bigger `-log10(p-value)`.\n\n---\n\n> - What is the functional connection between DE(A) and p-value(A)?\n>\n> - Is it possible for Gene A and Gene B to have identical p-values while having drastically different differential expressions like +2 vs +10?\n> \n\nInteresting question, I think there is not a simple function that `f(LFC) = -logP`, because other factors such as the sample size and the distribution (variance) will affect the p-value. vice versa.\n\nHere are some examples:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1023301%2F9b7e0e9214266c54c890102d984891ff%2FDingtalk_20231020180013.jpg?generation=1697796028133628&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1023301%2Fbd5f0c47c4fd6e65cb300a0e5b1bb808%2FDingtalk_20231020180851.jpg?generation=1697796559157283&alt=media)\n\n---\n\nThe **Volcano plot** demonstrated the relationship between `LFC` and `-log10(p-value)`, and it also showed the `-log10(p-value)` are **\"unstable\"** when the `abs(LFC)` is big.\n\nYou can find the plot and the examples in my [notebook](https://www.kaggle.com/code/awater1223/op2-04-merge-limma-results).",
    "2490472": "Thank you for your helpful reply! What I gathered from your note is that the p-value can be seen as the probability of there being a differential \"change.\" Naturally, the higher the log fold change, the higher this probability (including the sign/direction of change).\n\nIt appears that what we are trying to model and predict is **NOT** the continuous value of differential gene expression (a regression problem) but rather an indirect function of it, the probability of whether there was a change in the expression of Gene A (including the direction +/-), as a binary classification output [DE=Yes or DE=No change] . Your response, along with the useful notebook results showing instability in p-value outputs for the same or very similar DE, suggests the presence of nuisance factors (e.g., sample size) in the mapping from actual gene expression to p-value-based model output. This makes the output model appear more like y=G( F(T, C); n), where G is the additional mapping from the direct quantitative differential gene expression to our prediction variable.\n\nHow do we capture and model G(F(.);n) if it's introducing unknown variability via the hidden nuisance parameter 'n'? Since the evaluation cost function assumes a regression model (rather than just predicting the categorical \"sign\" of DE: -1, 1, or 0), wouldn't this introduce a \"random\" factor in scoring?\n\nImagine you had a magic wand that gave you the \"ideal\" direct differential expression predictor function DE=F(T,C). However, once you put it through G(DE;n), the answer gets diluted. I'm concerned that depending on the level of this dilution and considering how closely the scores stack up, attempting to produce the closest score matching the \"ground truth\" p-values of the test set might be a matter of chance rather than an accurate representation of the underlying generative biological process. Wouldn't it capture the process more truthfully to try to predict directly the differential expression rather than a p-values derivative - or define the cost function on a categorical prediction task (DE = [1, -1, 0])?   Thank you for your insights.",
    "2490535": "I actually agree with you... And have this concern for quite a while. Since I don't think the target/evaluation is going to change for this competition, we have to work with what we have.\n\nSo in my opinion, a big part of the solution for this competition will be fitting the metric..."
  },
  "source": "meta"
}