{
  "id": 354607,
  "title": "Does the information in the metadata directly helps your model?",
  "url": "/competitions/open-problems-multimodal/discussion/354607",
  "author_name": "",
  "post_date": "2022-09-23T03:59:50.441414400Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi all, </p>\n<p>I am wondering if the metadata has helps to the models, excepting using them to make kfolds. I have tried to add cell type information in the CITE prediction, but it doesn't help. I guess the reason could be the information in the CITE part already includes the information of cell types or cell types is not a strong assistant information to this problem. </p>\n<p>Would you mind sharing something you discovered, works and doesn't works all help!</p>",
  "messages": [
    {
      "id": "1951423",
      "postDate": "09/23/2022 03:59:50",
      "content": "<p>Hi all, </p>\n<p>I am wondering if the metadata has helps to the models, excepting using them to make kfolds. I have tried to add cell type information in the CITE prediction, but it doesn't help. I guess the reason could be the information in the CITE part already includes the information of cell types or cell types is not a strong assistant information to this problem. </p>\n<p>Would you mind sharing something you discovered, works and doesn't works all help!</p>",
      "rawMarkdown": "Hi all, \n\nI am wondering if the metadata has helps to the models, excepting using them to make kfolds. I have tried to add cell type information in the CITE prediction, but it doesn't help. I guess the reason could be the information in the CITE part already includes the information of cell types or cell types is not a strong assistant information to this problem. \n\nWould you mind sharing something you discovered, works and doesn't works all help!",
      "votes": null
    },
    {
      "id": "1954922",
      "postDate": "09/25/2022 14:55:08",
      "content": "<p>I introduced  cell_type embedding into multiome MLP model, which doesn't help and makes me confused. <br>\nLike ur said, maybe the cell_type  is not a strong assistant information.</p>",
      "rawMarkdown": "I introduced  cell_type embedding into multiome MLP model, which doesn't help and makes me confused. \nLike ur said, maybe the cell_type  is not a strong assistant information.",
      "votes": null
    },
    {
      "id": "1956937",
      "postDate": "09/26/2022 17:20:09",
      "content": "<p>I think ur right. Based on the transcriptomic information, we can conclude which celltype does a single cell belong to (with some confidence). So if expressiveness of the model is enough, I think do not need to include the cell type info? It's mainly based on my biological knowledge, havent try to build models. </p>",
      "rawMarkdown": "I think ur right. Based on the transcriptomic information, we can conclude which celltype does a single cell belong to (with some confidence). So if expressiveness of the model is enough, I think do not need to include the cell type info? It's mainly based on my biological knowledge, havent try to build models.",
      "votes": null
    },
    {
      "id": "1971829",
      "postDate": "10/04/2022 19:11:11",
      "content": "<p>Yup ALso seems to me for test data all cell type information is hidden sigh 😑 <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> is this intentional ?</p>",
      "rawMarkdown": "Yup ALso seems to me for test data all cell type information is hidden sigh 😑 @danielburkhardt is this intentional ?",
      "votes": null
    },
    {
      "id": "1971846",
      "postDate": "10/04/2022 19:21:48",
      "content": "<p>Yes, we intentionally decided not to include metadata for the RNA in the test data for two reasons:  </p>\n<ol>\n<li>In the Multiome, including cell type labels would constitute a data leak because the labels are based on the RNA</li>\n<li>We're not confident that the labels will be useful for prediction, as you've observed here. The labels are A) discrete and B) based on a small number of genes (approximately 5-10 features) + some neighborhood smoothing. I'm not surprised that some models don't benefit from the cell type labels. </li>\n</ol>",
      "rawMarkdown": "Yes, we intentionally decided not to include metadata for the RNA in the test data for two reasons:  \n1. In the Multiome, including cell type labels would constitute a data leak because the labels are based on the RNA\n2. We're not confident that the labels will be useful for prediction, as you've observed here. The labels are A) discrete and B) based on a small number of genes (approximately 5-10 features) + some neighborhood smoothing. I'm not surprised that some models don't benefit from the cell type labels.",
      "votes": null
    },
    {
      "id": "1971887",
      "postDate": "10/04/2022 20:01:26",
      "content": "<p>Great thanks for the clarification.. </p>",
      "rawMarkdown": "Great thanks for the clarification..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1954922,
      "author_name": "alvinai9603",
      "author_url": "",
      "post_date": "09/25/2022 14:55:08",
      "content": "<p>I introduced  cell_type embedding into multiome MLP model, which doesn't help and makes me confused. <br>\nLike ur said, maybe the cell_type  is not a strong assistant information.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1971829,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "10/04/2022 19:11:11",
          "content": "<p>Yup ALso seems to me for test data all cell type information is hidden sigh 😑 <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> is this intentional ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1971846,
          "author_name": "danielburkhardt",
          "author_url": "",
          "post_date": "10/04/2022 19:21:48",
          "content": "<p>Yes, we intentionally decided not to include metadata for the RNA in the test data for two reasons:  </p>\n<ol>\n<li>In the Multiome, including cell type labels would constitute a data leak because the labels are based on the RNA</li>\n<li>We're not confident that the labels will be useful for prediction, as you've observed here. The labels are A) discrete and B) based on a small number of genes (approximately 5-10 features) + some neighborhood smoothing. I'm not surprised that some models don't benefit from the cell type labels. </li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1971887,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "10/04/2022 20:01:26",
          "content": "<p>Great thanks for the clarification.. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1956937,
      "author_name": "jinyang18",
      "author_url": "",
      "post_date": "09/26/2022 17:20:09",
      "content": "<p>I think ur right. Based on the transcriptomic information, we can conclude which celltype does a single cell belong to (with some confidence). So if expressiveness of the model is enough, I think do not need to include the cell type info? It's mainly based on my biological knowledge, havent try to build models. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1951423": "Hi all, \n\nI am wondering if the metadata has helps to the models, excepting using them to make kfolds. I have tried to add cell type information in the CITE prediction, but it doesn't help. I guess the reason could be the information in the CITE part already includes the information of cell types or cell types is not a strong assistant information to this problem. \n\nWould you mind sharing something you discovered, works and doesn't works all help!",
    "1954922": "I introduced  cell_type embedding into multiome MLP model, which doesn't help and makes me confused. \nLike ur said, maybe the cell_type  is not a strong assistant information.",
    "1956937": "I think ur right. Based on the transcriptomic information, we can conclude which celltype does a single cell belong to (with some confidence). So if expressiveness of the model is enough, I think do not need to include the cell type info? It's mainly based on my biological knowledge, havent try to build models.",
    "1971829": "Yup ALso seems to me for test data all cell type information is hidden sigh 😑 @danielburkhardt is this intentional ?",
    "1971846": "Yes, we intentionally decided not to include metadata for the RNA in the test data for two reasons:  \n1. In the Multiome, including cell type labels would constitute a data leak because the labels are based on the RNA\n2. We're not confident that the labels will be useful for prediction, as you've observed here. The labels are A) discrete and B) based on a small number of genes (approximately 5-10 features) + some neighborhood smoothing. I'm not surprised that some models don't benefit from the cell type labels.",
    "1971887": "Great thanks for the clarification.."
  },
  "source": "meta"
}