{
  "id": 352980,
  "title": "How to make the most of the information in the \"metadata\" file?",
  "url": "/competitions/open-problems-multimodal/discussion/352980",
  "author_name": "",
  "post_date": "2022-09-16T12:13:59.382136300Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>In this challenge, in addition to train_data and test_data, there is interesting and important information in the \"metadata\" file. This information includes four items: \"<strong>day</strong>\", \"<strong>donor</strong>\", \"<strong>cell_type</strong>\" and \"<strong>technology</strong>\". </p>\n<p>The two technologies used in this challenge (<strong>CITEseq</strong> and <strong>Multiome</strong>) are different in terms of features and targets, so they should be modeled separately.</p>\n<p>Also, it seems that the importance of \"<strong>cell_type</strong>\" is much higher than \"day\" and \"donor\". Of course, this issue can be checked by clustering (or using the NearestNeighbors library) on the train_cite_targets and train_multi_targets files and then comparing the results.</p>\n<p>However, we can use \"<strong>cell_type</strong>\" in two ways. The first way is to add \"<strong>cell_type</strong>\" information (or even \"<strong>donor</strong>\" information) as a column (feature) to train_data and test_data. The second way is that we can split train_data and test_data based on each different \"<strong>cell_type</strong>\" so that completely separate calculations are possible (similar to what we usually do in cross-validation), of course, here the models will be independent. Also, in the second way, we will have smaller files that are easier to work with.</p>\n<p>In <a href=\"https://www.kaggle.com/code/mehrankazeminia/2-5-msci22-splitting-up1\" target=\"_blank\">this notebook</a>, I splitted all files related to CITEseq technology by \"<strong>cell_type</strong>\". Also, the results of this notebook are available in <a href=\"https://www.kaggle.com/datasets/mehrankazeminia/msci22-citeseqsplit\" target=\"_blank\">\"MSCI22-CITEseq-Split\" </a>dataset.</p>\n<p>Please note that \"<strong>cell-type</strong>\" is sorted according to the following order:</p>\n<blockquote>\n  <p>1 - MasP = Mast Cell Progenitor<br>\n  2 - MkP = Megakaryocyte Progenitor<br>\n  3 - NeuP = Neutrophil Progenitor<br>\n  4 - MoP = Monocyte Progenitor<br>\n  5 - EryP = Erythrocyte Progenitor<br>\n  6 - HSC = Hematoploetic Stem Cell<br>\n  7 - BP = B-Cell Progenitor</p>\n</blockquote>\n<p>The files associated with Multiome technology are larger and difficult to work with on these types of notebooks. Methods that use \"data flow\" are more suitable. For example, using \"<strong>mrjob</strong>\" which should be done outside the notebook. Of course, if there is a chance; I just publish the \"<strong>mrjob</strong>\" codes separately in a notebook.</p>",
  "messages": [
    {
      "id": "1942051",
      "postDate": "09/16/2022 12:13:59",
      "content": "<p>In this challenge, in addition to train_data and test_data, there is interesting and important information in the \"metadata\" file. This information includes four items: \"<strong>day</strong>\", \"<strong>donor</strong>\", \"<strong>cell_type</strong>\" and \"<strong>technology</strong>\". </p>\n<p>The two technologies used in this challenge (<strong>CITEseq</strong> and <strong>Multiome</strong>) are different in terms of features and targets, so they should be modeled separately.</p>\n<p>Also, it seems that the importance of \"<strong>cell_type</strong>\" is much higher than \"day\" and \"donor\". Of course, this issue can be checked by clustering (or using the NearestNeighbors library) on the train_cite_targets and train_multi_targets files and then comparing the results.</p>\n<p>However, we can use \"<strong>cell_type</strong>\" in two ways. The first way is to add \"<strong>cell_type</strong>\" information (or even \"<strong>donor</strong>\" information) as a column (feature) to train_data and test_data. The second way is that we can split train_data and test_data based on each different \"<strong>cell_type</strong>\" so that completely separate calculations are possible (similar to what we usually do in cross-validation), of course, here the models will be independent. Also, in the second way, we will have smaller files that are easier to work with.</p>\n<p>In <a href=\"https://www.kaggle.com/code/mehrankazeminia/2-5-msci22-splitting-up1\" target=\"_blank\">this notebook</a>, I splitted all files related to CITEseq technology by \"<strong>cell_type</strong>\". Also, the results of this notebook are available in <a href=\"https://www.kaggle.com/datasets/mehrankazeminia/msci22-citeseqsplit\" target=\"_blank\">\"MSCI22-CITEseq-Split\" </a>dataset.</p>\n<p>Please note that \"<strong>cell-type</strong>\" is sorted according to the following order:</p>\n<blockquote>\n  <p>1 - MasP = Mast Cell Progenitor<br>\n  2 - MkP = Megakaryocyte Progenitor<br>\n  3 - NeuP = Neutrophil Progenitor<br>\n  4 - MoP = Monocyte Progenitor<br>\n  5 - EryP = Erythrocyte Progenitor<br>\n  6 - HSC = Hematoploetic Stem Cell<br>\n  7 - BP = B-Cell Progenitor</p>\n</blockquote>\n<p>The files associated with Multiome technology are larger and difficult to work with on these types of notebooks. Methods that use \"data flow\" are more suitable. For example, using \"<strong>mrjob</strong>\" which should be done outside the notebook. Of course, if there is a chance; I just publish the \"<strong>mrjob</strong>\" codes separately in a notebook.</p>",
      "rawMarkdown": "In this challenge, in addition to train_data and test_data, there is interesting and important information in the \"metadata\" file. This information includes four items: \"**day**\", \"**donor**\", \"**cell_type**\" and \"**technology**\". \n\nThe two technologies used in this challenge (**CITEseq** and **Multiome**) are different in terms of features and targets, so they should be modeled separately.\n\nAlso, it seems that the importance of \"**cell_type**\" is much higher than \"day\" and \"donor\". Of course, this issue can be checked by clustering (or using the NearestNeighbors library) on the train_cite_targets and train_multi_targets files and then comparing the results.\n\nHowever, we can use \"**cell_type**\" in two ways. The first way is to add \"**cell_type**\" information (or even \"**donor**\" information) as a column (feature) to train_data and test_data. The second way is that we can split train_data and test_data based on each different \"**cell_type**\" so that completely separate calculations are possible (similar to what we usually do in cross-validation), of course, here the models will be independent. Also, in the second way, we will have smaller files that are easier to work with.\n\nIn [this notebook](https://www.kaggle.com/code/mehrankazeminia/2-5-msci22-splitting-up1), I splitted all files related to CITEseq technology by \"**cell_type**\". Also, the results of this notebook are available in [\"MSCI22-CITEseq-Split\" ](https://www.kaggle.com/datasets/mehrankazeminia/msci22-citeseqsplit)dataset.\n\nPlease note that \"**cell-type**\" is sorted according to the following order:\n> 1 - MasP = Mast Cell Progenitor\n2 - MkP = Megakaryocyte Progenitor\n3 - NeuP = Neutrophil Progenitor\n4 - MoP = Monocyte Progenitor\n5 - EryP = Erythrocyte Progenitor\n6 - HSC = Hematoploetic Stem Cell\n7 - BP = B-Cell Progenitor\n\nThe files associated with Multiome technology are larger and difficult to work with on these types of notebooks. Methods that use \"data flow\" are more suitable. For example, using \"**mrjob**\" which should be done outside the notebook. Of course, if there is a chance; I just publish the \"**mrjob**\" codes separately in a notebook.",
      "votes": null
    },
    {
      "id": "1942484",
      "postDate": "09/16/2022 17:17:59",
      "content": "<p>cell-type gene-type alteast for multinome seems to be the most interesting one to be leveraged </p>",
      "rawMarkdown": "cell-type gene-type alteast for multinome seems to be the most interesting one to be leveraged",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1942484,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "09/16/2022 17:17:59",
      "content": "<p>cell-type gene-type alteast for multinome seems to be the most interesting one to be leveraged </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1942051": "In this challenge, in addition to train_data and test_data, there is interesting and important information in the \"metadata\" file. This information includes four items: \"**day**\", \"**donor**\", \"**cell_type**\" and \"**technology**\". \n\nThe two technologies used in this challenge (**CITEseq** and **Multiome**) are different in terms of features and targets, so they should be modeled separately.\n\nAlso, it seems that the importance of \"**cell_type**\" is much higher than \"day\" and \"donor\". Of course, this issue can be checked by clustering (or using the NearestNeighbors library) on the train_cite_targets and train_multi_targets files and then comparing the results.\n\nHowever, we can use \"**cell_type**\" in two ways. The first way is to add \"**cell_type**\" information (or even \"**donor**\" information) as a column (feature) to train_data and test_data. The second way is that we can split train_data and test_data based on each different \"**cell_type**\" so that completely separate calculations are possible (similar to what we usually do in cross-validation), of course, here the models will be independent. Also, in the second way, we will have smaller files that are easier to work with.\n\nIn [this notebook](https://www.kaggle.com/code/mehrankazeminia/2-5-msci22-splitting-up1), I splitted all files related to CITEseq technology by \"**cell_type**\". Also, the results of this notebook are available in [\"MSCI22-CITEseq-Split\" ](https://www.kaggle.com/datasets/mehrankazeminia/msci22-citeseqsplit)dataset.\n\nPlease note that \"**cell-type**\" is sorted according to the following order:\n> 1 - MasP = Mast Cell Progenitor\n2 - MkP = Megakaryocyte Progenitor\n3 - NeuP = Neutrophil Progenitor\n4 - MoP = Monocyte Progenitor\n5 - EryP = Erythrocyte Progenitor\n6 - HSC = Hematoploetic Stem Cell\n7 - BP = B-Cell Progenitor\n\nThe files associated with Multiome technology are larger and difficult to work with on these types of notebooks. Methods that use \"data flow\" are more suitable. For example, using \"**mrjob**\" which should be done outside the notebook. Of course, if there is a chance; I just publish the \"**mrjob**\" codes separately in a notebook.",
    "1942484": "cell-type gene-type alteast for multinome seems to be the most interesting one to be leveraged"
  },
  "source": "meta"
}