{
  "id": 367578,
  "title": "Idea: use of precalculated genes embeddings - variation on Makotu (3-th place) theme",
  "url": "/competitions/open-problems-multimodal/writeups/sb2-idea-use-of-precalculated-genes-embeddings-var",
  "author_name": "",
  "post_date": "2022-11-21T10:06:00.513Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Makotu (@mhyodo)  (3-th place) described interesting ideas:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366428\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366428</a> <br>\nHere is some variation on his theme:</p>\n<p>The columns names  in  CITE-seq features  are  -  genes - we have enourmous-enourmous amounts of information in biology about them,  but how to incorporate that information into the model ? </p>\n<p>1) take precalculated embeddings for genes - (there are many many ways to do produce such embeddings - see some examples below) </p>\n<p>2) For each sample (i.e. cell) take only say 100 top expressed genes  </p>\n<p>3) Create a new feature - which is just an average (or sum) of embeddings for these 100 top genes.</p>\n<p>That is it. </p>\n<p>==========</p>\n<p>Original idea of Makotu was NOT to use precalculated embeddings - but to calculate embeddings by a clever use of word2vec approach. That might be more powerful approach for the task. However the variation might have some benefits since it allows to incorporate existing biological knowledge. </p>\n<p>==========</p>\n<p>Examples of genes embeddings: </p>\n<p>1) Simplest way. Take any dataset which  of form  genes * samples  - take any dimensional reduction to get: genes * \"metasamples\" - so it gives us some embeddings for genes. Taking some large and famous datasets like e.g. TCGA one may hope that these embeddings capture some important biological information.</p>\n<p>2) Graph based. Take any graph involving genes - PPI (Protein-protein interactions) or Knowledge Graphs (where genes are involved even Wikidata) there are plenty algorithms which produce vector embeddings from graphs - so any of them will produce embeddings for genes.</p>\n<p>3) NLP based. Take textual description of the genes. Use something like Sentence Bert  for each textual description - we got an embedding for the gene. <br>\nThat approach we used in the competition for feature selection - see  notebooks by Anton Kostin:<br>\n<a href=\"https://www.kaggle.com/code/visualcomments/genes-embeddings-clustermap\" target=\"_blank\">https://www.kaggle.com/code/visualcomments/genes-embeddings-clustermap</a></p>\n<p>PS <br>\nWe analyzed PPI as a possible way to produce features candidates:<br>\n<a href=\"https://www.kaggle.com/code/visualcomments/cd-genes-ppi-neighbors\" target=\"_blank\">https://www.kaggle.com/code/visualcomments/cd-genes-ppi-neighbors</a><br>\nSome of them worked. </p>\n<p>==========</p>\n<p>Some another more simple variation of the Makotu idea - which can be used for any tabular dataset (not genes):<br>\njust make from the original feature matrix - a new matrix of the same size with only 1 and 0. <br>\nWhere 1 will be placed at positions of the top100 features. Take PCA/SVD from the that matrix - you got new features. </p>",
  "messages": [
    {
      "id": "2038296",
      "postDate": "11/21/2022 10:00:22",
      "content": "<p>Makotu (@mhyodo)  (3-th place) described interesting ideas:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366428\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366428</a> <br>\nHere is some variation on his theme:</p>\n<p>The columns names  in  CITE-seq features  are  -  genes - we have enourmous-enourmous amounts of information in biology about them,  but how to incorporate that information into the model ? </p>\n<p>1) take precalculated embeddings for genes - (there are many many ways to do produce such embeddings - see some examples below) </p>\n<p>2) For each sample (i.e. cell) take only say 100 top expressed genes  </p>\n<p>3) Create a new feature - which is just an average (or sum) of embeddings for these 100 top genes.</p>\n<p>That is it. </p>\n<p>==========</p>\n<p>Original idea of Makotu was NOT to use precalculated embeddings - but to calculate embeddings by a clever use of word2vec approach. That might be more powerful approach for the task. However the variation might have some benefits since it allows to incorporate existing biological knowledge. </p>\n<p>==========</p>\n<p>Examples of genes embeddings: </p>\n<p>1) Simplest way. Take any dataset which  of form  genes * samples  - take any dimensional reduction to get: genes * \"metasamples\" - so it gives us some embeddings for genes. Taking some large and famous datasets like e.g. TCGA one may hope that these embeddings capture some important biological information.</p>\n<p>2) Graph based. Take any graph involving genes - PPI (Protein-protein interactions) or Knowledge Graphs (where genes are involved even Wikidata) there are plenty algorithms which produce vector embeddings from graphs - so any of them will produce embeddings for genes.</p>\n<p>3) NLP based. Take textual description of the genes. Use something like Sentence Bert  for each textual description - we got an embedding for the gene. <br>\nThat approach we used in the competition for feature selection - see  notebooks by Anton Kostin:<br>\n<a href=\"https://www.kaggle.com/code/visualcomments/genes-embeddings-clustermap\" target=\"_blank\">https://www.kaggle.com/code/visualcomments/genes-embeddings-clustermap</a></p>\n<p>PS <br>\nWe analyzed PPI as a possible way to produce features candidates:<br>\n<a href=\"https://www.kaggle.com/code/visualcomments/cd-genes-ppi-neighbors\" target=\"_blank\">https://www.kaggle.com/code/visualcomments/cd-genes-ppi-neighbors</a><br>\nSome of them worked. </p>\n<p>==========</p>\n<p>Some another more simple variation of the Makotu idea - which can be used for any tabular dataset (not genes):<br>\njust make from the original feature matrix - a new matrix of the same size with only 1 and 0. <br>\nWhere 1 will be placed at positions of the top100 features. Take PCA/SVD from the that matrix - you got new features. </p>",
      "rawMarkdown": "Makotu (@mhyodo)  (3-th place) described interesting ideas:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/366428 \nHere is some variation on his theme:\n\nThe columns names  in  CITE-seq features  are  -  genes - we have enourmous-enourmous amounts of information in biology about them,  but how to incorporate that information into the model ? \n\n1) take precalculated embeddings for genes - (there are many many ways to do produce such embeddings - see some examples below) \n\n2) For each sample (i.e. cell) take only say 100 top expressed genes  \n\n3) Create a new feature - which is just an average (or sum) of embeddings for these 100 top genes.\n\nThat is it. \n\n==========\n\nOriginal idea of Makotu was NOT to use precalculated embeddings - but to calculate embeddings by a clever use of word2vec approach. That might be more powerful approach for the task. However the variation might have some benefits since it allows to incorporate existing biological knowledge. \n\n==========\n\nExamples of genes embeddings: \n\n1) Simplest way. Take any dataset which  of form  genes * samples  - take any dimensional reduction to get: genes * \"metasamples\" - so it gives us some embeddings for genes. Taking some large and famous datasets like e.g. TCGA one may hope that these embeddings capture some important biological information.\n\n2) Graph based. Take any graph involving genes - PPI (Protein-protein interactions) or Knowledge Graphs (where genes are involved even Wikidata) there are plenty algorithms which produce vector embeddings from graphs - so any of them will produce embeddings for genes.\n\n\n3) NLP based. Take textual description of the genes. Use something like Sentence Bert  for each textual description - we got an embedding for the gene. \nThat approach we used in the competition for feature selection - see  notebooks by Anton Kostin:\nhttps://www.kaggle.com/code/visualcomments/genes-embeddings-clustermap\n\nPS \nWe analyzed PPI as a possible way to produce features candidates:\nhttps://www.kaggle.com/code/visualcomments/cd-genes-ppi-neighbors\nSome of them worked. \n\n==========\n\nSome another more simple variation of the Makotu idea - which can be used for any tabular dataset (not genes):\njust make from the original feature matrix - a new matrix of the same size with only 1 and 0. \nWhere 1 will be placed at positions of the top100 features. Take PCA/SVD from the that matrix - you got new features.",
      "votes": null
    },
    {
      "id": "2041670",
      "postDate": "11/24/2022 05:44:11",
      "content": "<p>Thanks for your sharing Alexander Chervov, your every sharing is very meaningful and wonderful! I'm so lucky to learn so much knowledge from your notebooks and discussion. It's my honor to join the same competition with you!</p>",
      "rawMarkdown": "Thanks for your sharing Alexander Chervov, your every sharing is very meaningful and wonderful! I'm so lucky to learn so much knowledge from your notebooks and discussion. It's my honor to join the same competition with you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2041670,
      "author_name": "songqizhou",
      "author_url": "",
      "post_date": "11/24/2022 05:44:11",
      "content": "<p>Thanks for your sharing Alexander Chervov, your every sharing is very meaningful and wonderful! I'm so lucky to learn so much knowledge from your notebooks and discussion. It's my honor to join the same competition with you!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2038296": "Makotu (@mhyodo)  (3-th place) described interesting ideas:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/366428 \nHere is some variation on his theme:\n\nThe columns names  in  CITE-seq features  are  -  genes - we have enourmous-enourmous amounts of information in biology about them,  but how to incorporate that information into the model ? \n\n1) take precalculated embeddings for genes - (there are many many ways to do produce such embeddings - see some examples below) \n\n2) For each sample (i.e. cell) take only say 100 top expressed genes  \n\n3) Create a new feature - which is just an average (or sum) of embeddings for these 100 top genes.\n\nThat is it. \n\n==========\n\nOriginal idea of Makotu was NOT to use precalculated embeddings - but to calculate embeddings by a clever use of word2vec approach. That might be more powerful approach for the task. However the variation might have some benefits since it allows to incorporate existing biological knowledge. \n\n==========\n\nExamples of genes embeddings: \n\n1) Simplest way. Take any dataset which  of form  genes * samples  - take any dimensional reduction to get: genes * \"metasamples\" - so it gives us some embeddings for genes. Taking some large and famous datasets like e.g. TCGA one may hope that these embeddings capture some important biological information.\n\n2) Graph based. Take any graph involving genes - PPI (Protein-protein interactions) or Knowledge Graphs (where genes are involved even Wikidata) there are plenty algorithms which produce vector embeddings from graphs - so any of them will produce embeddings for genes.\n\n\n3) NLP based. Take textual description of the genes. Use something like Sentence Bert  for each textual description - we got an embedding for the gene. \nThat approach we used in the competition for feature selection - see  notebooks by Anton Kostin:\nhttps://www.kaggle.com/code/visualcomments/genes-embeddings-clustermap\n\nPS \nWe analyzed PPI as a possible way to produce features candidates:\nhttps://www.kaggle.com/code/visualcomments/cd-genes-ppi-neighbors\nSome of them worked. \n\n==========\n\nSome another more simple variation of the Makotu idea - which can be used for any tabular dataset (not genes):\njust make from the original feature matrix - a new matrix of the same size with only 1 and 0. \nWhere 1 will be placed at positions of the top100 features. Take PCA/SVD from the that matrix - you got new features.",
    "2041670": "Thanks for your sharing Alexander Chervov, your every sharing is very meaningful and wonderful! I'm so lucky to learn so much knowledge from your notebooks and discussion. It's my honor to join the same competition with you!"
  },
  "source": "meta"
}