{
  "id": 350953,
  "title": "Explanation for homogeneity of UMAPs across donors? ",
  "url": "/competitions/open-problems-multimodal/discussion/350953",
  "author_name": "",
  "post_date": "2022-09-07T20:48:53.162176Z",
  "votes": 10,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I generated the following four UMAP plots by selecting only those cells corresponding to each donor. E.g. the first plot fits DNA accessibility data for cells labeled as belonging to donor 13176, and it contains all such cells. </p>\n<p>Donor 13176:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2F6b7033eb1a61bf5e5a5918712d39ac74%2F13176.png?generation=1662583268166134&amp;alt=media\" alt=\"\"></p>\n<p>Donor 27678:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2Fef9b1a64c63f6789756f9850012ccdad%2F27678.png?generation=1662583311439518&amp;alt=media\" alt=\"\"></p>\n<p>Donor 31800:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2Fb9b7ae68865d89643d65e7dd9cc077a1%2F31800.png?generation=1662583320508487&amp;alt=media\" alt=\"\"></p>\n<p>Donor 32606:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2F92428e26dcc95c0ba82f070d7669170d%2F32606.png?generation=1662583337403186&amp;alt=media\" alt=\"\"></p>\n<p>I'm surprised by the similarity in these plots. The only explanation I've come up with is that the datasets pertaining to each individual donor have been batch corrected against each other, but am hoping for a better explanation (or possibly a mistake on my part). Does anyone have a better idea for why this might be? </p>\n<p>Here is a UMAP plot of the entire dataset colored by donors, to further illustrate my question:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2F725119550af2a0d5c03f1202997568d6%2Fsuperimposed.png?generation=1662583467128556&amp;alt=media\" alt=\"\"></p>\n<p>Thanks! Looking forward to any and all responses. </p>\n<p>Yajit</p>",
  "messages": [
    {
      "id": "1930411",
      "postDate": "09/07/2022 20:48:53",
      "content": "<p>Hi all,</p>\n<p>I generated the following four UMAP plots by selecting only those cells corresponding to each donor. E.g. the first plot fits DNA accessibility data for cells labeled as belonging to donor 13176, and it contains all such cells. </p>\n<p>Donor 13176:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2F6b7033eb1a61bf5e5a5918712d39ac74%2F13176.png?generation=1662583268166134&amp;alt=media\" alt=\"\"></p>\n<p>Donor 27678:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2Fef9b1a64c63f6789756f9850012ccdad%2F27678.png?generation=1662583311439518&amp;alt=media\" alt=\"\"></p>\n<p>Donor 31800:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2Fb9b7ae68865d89643d65e7dd9cc077a1%2F31800.png?generation=1662583320508487&amp;alt=media\" alt=\"\"></p>\n<p>Donor 32606:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2F92428e26dcc95c0ba82f070d7669170d%2F32606.png?generation=1662583337403186&amp;alt=media\" alt=\"\"></p>\n<p>I'm surprised by the similarity in these plots. The only explanation I've come up with is that the datasets pertaining to each individual donor have been batch corrected against each other, but am hoping for a better explanation (or possibly a mistake on my part). Does anyone have a better idea for why this might be? </p>\n<p>Here is a UMAP plot of the entire dataset colored by donors, to further illustrate my question:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2F725119550af2a0d5c03f1202997568d6%2Fsuperimposed.png?generation=1662583467128556&amp;alt=media\" alt=\"\"></p>\n<p>Thanks! Looking forward to any and all responses. </p>\n<p>Yajit</p>",
      "rawMarkdown": "Hi all,\n\nI generated the following four UMAP plots by selecting only those cells corresponding to each donor. E.g. the first plot fits DNA accessibility data for cells labeled as belonging to donor 13176, and it contains all such cells. \n\nDonor 13176:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2F6b7033eb1a61bf5e5a5918712d39ac74%2F13176.png?generation=1662583268166134&alt=media)\n\nDonor 27678:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2Fef9b1a64c63f6789756f9850012ccdad%2F27678.png?generation=1662583311439518&alt=media)\n\nDonor 31800:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2Fb9b7ae68865d89643d65e7dd9cc077a1%2F31800.png?generation=1662583320508487&alt=media)\n\nDonor 32606:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2F92428e26dcc95c0ba82f070d7669170d%2F32606.png?generation=1662583337403186&alt=media)\n\nI'm surprised by the similarity in these plots. The only explanation I've come up with is that the datasets pertaining to each individual donor have been batch corrected against each other, but am hoping for a better explanation (or possibly a mistake on my part). Does anyone have a better idea for why this might be? \n\nHere is a UMAP plot of the entire dataset colored by donors, to further illustrate my question:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2F725119550af2a0d5c03f1202997568d6%2Fsuperimposed.png?generation=1662583467128556&alt=media)\n\nThanks! Looking forward to any and all responses. \n\nYajit",
      "votes": null
    },
    {
      "id": "1930436",
      "postDate": "09/07/2022 21:11:09",
      "content": "<p>Do you fit UMAP on DNA, RNA or protein data?</p>",
      "rawMarkdown": "Do you fit UMAP on DNA, RNA or protein data?",
      "votes": null
    },
    {
      "id": "1930468",
      "postDate": "09/07/2022 22:23:30",
      "content": "<p>This is on the DNA, in particular the files test_multi_inputs and train_multi_inputs together. </p>",
      "rawMarkdown": "This is on the DNA, in particular the files test_multi_inputs and train_multi_inputs together.",
      "votes": null
    },
    {
      "id": "1931340",
      "postDate": "09/08/2022 16:00:51",
      "content": "<p>Can you share the code please ?</p>",
      "rawMarkdown": "Can you share the code please ?",
      "votes": null
    },
    {
      "id": "1932717",
      "postDate": "09/09/2022 21:21:03",
      "content": "<p>Great question! We didn't do any batch correction on the data and were very impressed with the amount of similarity across donor samples.</p>\n<p>The reason there's such good overlap in the UMAPs is that there was very little variation between donors relative to global variation within each dataset. Generating this data was a carefully orchestrated effort specifically done for this competition. All the samples were prepared at the same time by the same scientists and processed according to SOPs developed at Cellarity where we work with these kinds of cell regularly as part of our research in therapeutic research and development in hematology.</p>",
      "rawMarkdown": "Great question! We didn't do any batch correction on the data and were very impressed with the amount of similarity across donor samples.\n\nThe reason there's such good overlap in the UMAPs is that there was very little variation between donors relative to global variation within each dataset. Generating this data was a carefully orchestrated effort specifically done for this competition. All the samples were prepared at the same time by the same scientists and processed according to SOPs developed at Cellarity where we work with these kinds of cell regularly as part of our research in therapeutic research and development in hematology.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1930436,
      "author_name": "simakov",
      "author_url": "",
      "post_date": "09/07/2022 21:11:09",
      "content": "<p>Do you fit UMAP on DNA, RNA or protein data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1930468,
          "author_name": "yajitjain",
          "author_url": "",
          "post_date": "09/07/2022 22:23:30",
          "content": "<p>This is on the DNA, in particular the files test_multi_inputs and train_multi_inputs together. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1931340,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "09/08/2022 16:00:51",
      "content": "<p>Can you share the code please ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1932717,
      "author_name": "danielburkhardt",
      "author_url": "",
      "post_date": "09/09/2022 21:21:03",
      "content": "<p>Great question! We didn't do any batch correction on the data and were very impressed with the amount of similarity across donor samples.</p>\n<p>The reason there's such good overlap in the UMAPs is that there was very little variation between donors relative to global variation within each dataset. Generating this data was a carefully orchestrated effort specifically done for this competition. All the samples were prepared at the same time by the same scientists and processed according to SOPs developed at Cellarity where we work with these kinds of cell regularly as part of our research in therapeutic research and development in hematology.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1930411": "Hi all,\n\nI generated the following four UMAP plots by selecting only those cells corresponding to each donor. E.g. the first plot fits DNA accessibility data for cells labeled as belonging to donor 13176, and it contains all such cells. \n\nDonor 13176:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2F6b7033eb1a61bf5e5a5918712d39ac74%2F13176.png?generation=1662583268166134&alt=media)\n\nDonor 27678:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2Fef9b1a64c63f6789756f9850012ccdad%2F27678.png?generation=1662583311439518&alt=media)\n\nDonor 31800:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2Fb9b7ae68865d89643d65e7dd9cc077a1%2F31800.png?generation=1662583320508487&alt=media)\n\nDonor 32606:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2F92428e26dcc95c0ba82f070d7669170d%2F32606.png?generation=1662583337403186&alt=media)\n\nI'm surprised by the similarity in these plots. The only explanation I've come up with is that the datasets pertaining to each individual donor have been batch corrected against each other, but am hoping for a better explanation (or possibly a mistake on my part). Does anyone have a better idea for why this might be? \n\nHere is a UMAP plot of the entire dataset colored by donors, to further illustrate my question:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F85726%2F725119550af2a0d5c03f1202997568d6%2Fsuperimposed.png?generation=1662583467128556&alt=media)\n\nThanks! Looking forward to any and all responses. \n\nYajit",
    "1930436": "Do you fit UMAP on DNA, RNA or protein data?",
    "1930468": "This is on the DNA, in particular the files test_multi_inputs and train_multi_inputs together.",
    "1931340": "Can you share the code please ?",
    "1932717": "Great question! We didn't do any batch correction on the data and were very impressed with the amount of similarity across donor samples.\n\nThe reason there's such good overlap in the UMAPs is that there was very little variation between donors relative to global variation within each dataset. Generating this data was a carefully orchestrated effort specifically done for this competition. All the samples were prepared at the same time by the same scientists and processed according to SOPs developed at Cellarity where we work with these kinds of cell regularly as part of our research in therapeutic research and development in hematology."
  },
  "source": "meta"
}