{
  "id": 440956,
  "title": "Batch effects between donors",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/440956",
  "author_name": "",
  "post_date": "2023-09-17T00:27:59.971718500Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thanks for organizing this competition.</p>\n<p>I am trying to visualize the cells by using the adata_train.parquet. When I plotted the UMAP using the normalized count, the cell types were not clustered together,</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2Fbfd6395de77237255c4bfb29b3946d00%2Fnormalized.png?generation=1694910012053915&amp;alt=media\" alt=\"\"></p>\n<p>It seems that there are some batch effects between donors. So I used BBKNN to correct the batch effect, however, the results didn't look much better:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2F37b5e9431417cd7a995e24e156af879d%2Fbbknn.png?generation=1694911139213681&amp;alt=media\" alt=\"\"></p>\n<p>May I ask how you corrected the batch effect when clustering and annotating the cells?</p>\n<p>Also, would it be possible to share the codes for pre-processing?</p>",
  "messages": [
    {
      "id": "2442362",
      "postDate": "09/17/2023 00:27:59",
      "content": "<p>Thanks for organizing this competition.</p>\n<p>I am trying to visualize the cells by using the adata_train.parquet. When I plotted the UMAP using the normalized count, the cell types were not clustered together,</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2Fbfd6395de77237255c4bfb29b3946d00%2Fnormalized.png?generation=1694910012053915&amp;alt=media\" alt=\"\"></p>\n<p>It seems that there are some batch effects between donors. So I used BBKNN to correct the batch effect, however, the results didn't look much better:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2F37b5e9431417cd7a995e24e156af879d%2Fbbknn.png?generation=1694911139213681&amp;alt=media\" alt=\"\"></p>\n<p>May I ask how you corrected the batch effect when clustering and annotating the cells?</p>\n<p>Also, would it be possible to share the codes for pre-processing?</p>",
      "rawMarkdown": "Thanks for organizing this competition.\n\nI am trying to visualize the cells by using the adata_train.parquet. When I plotted the UMAP using the normalized count, the cell types were not clustered together,\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2Fbfd6395de77237255c4bfb29b3946d00%2Fnormalized.png?generation=1694910012053915&alt=media)\n\nIt seems that there are some batch effects between donors. So I used BBKNN to correct the batch effect, however, the results didn't look much better:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2F37b5e9431417cd7a995e24e156af879d%2Fbbknn.png?generation=1694911139213681&alt=media)\n\nMay I ask how you corrected the batch effect when clustering and annotating the cells?\n\nAlso, would it be possible to share the codes for pre-processing?",
      "votes": null
    },
    {
      "id": "2443674",
      "postDate": "09/17/2023 21:04:43",
      "content": "<p>it is a bit more complex than just cell type. Keep in mind that the same cell type has been treated with different compounds. So it is hard to predict the outcome, but there could (should?) be subclusters…<br>\nSo if you apply any batch correction you might want to take the drug treatment as confounder into account.</p>",
      "rawMarkdown": "it is a bit more complex than just cell type. Keep in mind that the same cell type has been treated with different compounds. So it is hard to predict the outcome, but there could (should?) be subclusters...\nSo if you apply any batch correction you might want to take the drug treatment as confounder into account.",
      "votes": null
    },
    {
      "id": "2444955",
      "postDate": "09/18/2023 14:54:58",
      "content": "<p>We annotated cell types on each plate individually, so no donor variation affected the analysis. </p>\n<p>To get nice UMAPs, all I've had to do is run <code>sc.pp.highly_variable_genes(adata, batch_key=\"donor\")</code> and then make sure to go through <code>adata.var</code> to make sure that  I keep genes with <code>n_batches_variable</code> is greater than <code>2</code>.</p>",
      "rawMarkdown": "We annotated cell types on each plate individually, so no donor variation affected the analysis. \n\nTo get nice UMAPs, all I've had to do is run `sc.pp.highly_variable_genes(adata, batch_key=\"donor\")` and then make sure to go through `adata.var` to make sure that  I keep genes with `n_batches_variable` is greater than `2`.",
      "votes": null
    },
    {
      "id": "2445721",
      "postDate": "09/19/2023 03:42:35",
      "content": "<p>Thanks, this makes sense.</p>\n<p>One more question. Did you correct the perturbation effect by data integration when annotating the cells for each plate?</p>",
      "rawMarkdown": "Thanks, this makes sense.\n\nOne more question. Did you correct the perturbation effect by data integration when annotating the cells for each plate?",
      "votes": null
    },
    {
      "id": "2449067",
      "postDate": "09/21/2023 01:57:00",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/zhijianli\" target=\"_blank\">@zhijianli</a> and <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a>, sorry to bother</p>\n<p>I have tried getting the UMAP on Kaggle <a href=\"https://www.kaggle.com/code/awater1223/op2-01-singlecell-eda\" target=\"_blank\">notebook</a>.<br>\nAm I pre-processing the data correctly? Why the UMAP seem to be \"not clearly\"?</p>\n<p>BTW, I used the raw counts instead of the normalized counts</p>",
      "rawMarkdown": "Hi, @zhijianli and @danielburkhardt, sorry to bother\n\nI have tried getting the UMAP on Kaggle [notebook](https://www.kaggle.com/code/awater1223/op2-01-singlecell-eda).\nAm I pre-processing the data correctly? Why the UMAP seem to be \"not clearly\"?\n\nBTW, I used the raw counts instead of the normalized counts",
      "votes": null
    },
    {
      "id": "2449975",
      "postDate": "09/21/2023 14:53:16",
      "content": "<p>Hi! A couple points:</p>\n<ol>\n<li>Use normalized counts</li>\n<li>After running <code>highly_variable_genes</code> look in <code>adata.var</code> for a col names <code>n_batches</code>. This is the number of batches (provided in <code>batch_key</code>) that the gene is variable in. By default, scanpy keeps all genes variable in at least 1 batch. You want to increase that to at least 2 or 3 batches. You need to manually overwrite the <code>is_highly_variable</code> column to apply this filter.</li>\n</ol>\n<p>Hope this helps!</p>\n<p>P.S. the column names here might not be exactly right, I'm writing this from memory</p>",
      "rawMarkdown": "Hi! A couple points:\n1. Use normalized counts\n2. After running `highly_variable_genes` look in `adata.var` for a col names `n_batches`. This is the number of batches (provided in `batch_key`) that the gene is variable in. By default, scanpy keeps all genes variable in at least 1 batch. You want to increase that to at least 2 or 3 batches. You need to manually overwrite the `is_highly_variable` column to apply this filter.\n\nHope this helps!\n\nP.S. the column names here might not be exactly right, I'm writing this from memory",
      "votes": null
    },
    {
      "id": "2450562",
      "postDate": "09/22/2023 02:25:24",
      "content": "<p>Thanks for your reply!</p>\n<ol>\n<li>Sorry, I didn't make it clear. I used the raw counts and ran <code>sc.pp.normalize_total</code> then<code>sc.pp.log1p</code> to make \"normalized counts\".</li>\n<li>That really helpful! The column name is <code>highly_variable_nbatches</code>. I will try it later.</li>\n</ol>",
      "rawMarkdown": "Thanks for your reply!\n\n1. Sorry, I didn't make it clear. I used the raw counts and ran `sc.pp.normalize_total` then`sc.pp.log1p` to make \"normalized counts\".\n2. That really helpful! The column name is `highly_variable_nbatches`. I will try it later.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2443674,
      "author_name": "rahi37",
      "author_url": "",
      "post_date": "09/17/2023 21:04:43",
      "content": "<p>it is a bit more complex than just cell type. Keep in mind that the same cell type has been treated with different compounds. So it is hard to predict the outcome, but there could (should?) be subclusters…<br>\nSo if you apply any batch correction you might want to take the drug treatment as confounder into account.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2444955,
      "author_name": "danielburkhardt",
      "author_url": "",
      "post_date": "09/18/2023 14:54:58",
      "content": "<p>We annotated cell types on each plate individually, so no donor variation affected the analysis. </p>\n<p>To get nice UMAPs, all I've had to do is run <code>sc.pp.highly_variable_genes(adata, batch_key=\"donor\")</code> and then make sure to go through <code>adata.var</code> to make sure that  I keep genes with <code>n_batches_variable</code> is greater than <code>2</code>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2445721,
          "author_name": "zhijianli",
          "author_url": "",
          "post_date": "09/19/2023 03:42:35",
          "content": "<p>Thanks, this makes sense.</p>\n<p>One more question. Did you correct the perturbation effect by data integration when annotating the cells for each plate?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2449067,
      "author_name": "awater1223",
      "author_url": "",
      "post_date": "09/21/2023 01:57:00",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/zhijianli\" target=\"_blank\">@zhijianli</a> and <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a>, sorry to bother</p>\n<p>I have tried getting the UMAP on Kaggle <a href=\"https://www.kaggle.com/code/awater1223/op2-01-singlecell-eda\" target=\"_blank\">notebook</a>.<br>\nAm I pre-processing the data correctly? Why the UMAP seem to be \"not clearly\"?</p>\n<p>BTW, I used the raw counts instead of the normalized counts</p>",
      "votes": null,
      "replies": [
        {
          "id": 2449975,
          "author_name": "danielburkhardt",
          "author_url": "",
          "post_date": "09/21/2023 14:53:16",
          "content": "<p>Hi! A couple points:</p>\n<ol>\n<li>Use normalized counts</li>\n<li>After running <code>highly_variable_genes</code> look in <code>adata.var</code> for a col names <code>n_batches</code>. This is the number of batches (provided in <code>batch_key</code>) that the gene is variable in. By default, scanpy keeps all genes variable in at least 1 batch. You want to increase that to at least 2 or 3 batches. You need to manually overwrite the <code>is_highly_variable</code> column to apply this filter.</li>\n</ol>\n<p>Hope this helps!</p>\n<p>P.S. the column names here might not be exactly right, I'm writing this from memory</p>",
          "votes": null,
          "replies": [
            {
              "id": 2450562,
              "author_name": "awater1223",
              "author_url": "",
              "post_date": "09/22/2023 02:25:24",
              "content": "<p>Thanks for your reply!</p>\n<ol>\n<li>Sorry, I didn't make it clear. I used the raw counts and ran <code>sc.pp.normalize_total</code> then<code>sc.pp.log1p</code> to make \"normalized counts\".</li>\n<li>That really helpful! The column name is <code>highly_variable_nbatches</code>. I will try it later.</li>\n</ol>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2442362": "Thanks for organizing this competition.\n\nI am trying to visualize the cells by using the adata_train.parquet. When I plotted the UMAP using the normalized count, the cell types were not clustered together,\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2Fbfd6395de77237255c4bfb29b3946d00%2Fnormalized.png?generation=1694910012053915&alt=media)\n\nIt seems that there are some batch effects between donors. So I used BBKNN to correct the batch effect, however, the results didn't look much better:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2F37b5e9431417cd7a995e24e156af879d%2Fbbknn.png?generation=1694911139213681&alt=media)\n\nMay I ask how you corrected the batch effect when clustering and annotating the cells?\n\nAlso, would it be possible to share the codes for pre-processing?",
    "2443674": "it is a bit more complex than just cell type. Keep in mind that the same cell type has been treated with different compounds. So it is hard to predict the outcome, but there could (should?) be subclusters...\nSo if you apply any batch correction you might want to take the drug treatment as confounder into account.",
    "2444955": "We annotated cell types on each plate individually, so no donor variation affected the analysis. \n\nTo get nice UMAPs, all I've had to do is run `sc.pp.highly_variable_genes(adata, batch_key=\"donor\")` and then make sure to go through `adata.var` to make sure that  I keep genes with `n_batches_variable` is greater than `2`.",
    "2445721": "Thanks, this makes sense.\n\nOne more question. Did you correct the perturbation effect by data integration when annotating the cells for each plate?",
    "2449067": "Hi, @zhijianli and @danielburkhardt, sorry to bother\n\nI have tried getting the UMAP on Kaggle [notebook](https://www.kaggle.com/code/awater1223/op2-01-singlecell-eda).\nAm I pre-processing the data correctly? Why the UMAP seem to be \"not clearly\"?\n\nBTW, I used the raw counts instead of the normalized counts",
    "2449975": "Hi! A couple points:\n1. Use normalized counts\n2. After running `highly_variable_genes` look in `adata.var` for a col names `n_batches`. This is the number of batches (provided in `batch_key`) that the gene is variable in. By default, scanpy keeps all genes variable in at least 1 batch. You want to increase that to at least 2 or 3 batches. You need to manually overwrite the `is_highly_variable` column to apply this filter.\n\nHope this helps!\n\nP.S. the column names here might not be exactly right, I'm writing this from memory",
    "2450562": "Thanks for your reply!\n\n1. Sorry, I didn't make it clear. I used the raw counts and ran `sc.pp.normalize_total` then`sc.pp.log1p` to make \"normalized counts\".\n2. That really helpful! The column name is `highly_variable_nbatches`. I will try it later."
  },
  "source": "meta"
}