{
  "id": 366452,
  "title": "Feature importance is important for \"saving the world\" - do not forget to share it, please ",
  "url": "/competitions/open-problems-multimodal/discussion/366452",
  "author_name": "",
  "post_date": "2022-11-16T08:07:38.704708400Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks for the organizers and congratulations to winners and everybody who got medals/knowledge/fun/experience ! </p>\n<p>The competition is based on real cutting-edge biological data which is quite important for research community.</p>\n<p>One of the questions which is of classical interest in such problems is : feature importance - that would reveal<br>\nhow genes (and regulatory regions for \"multiome\" part) interact which other.</p>\n<p>So please do not hesitate to share any your results/ideas/hypothesis/observations on that</p>\n<hr>\n<p>To give an example: <br>\nthe very famous package is \"GENIE3\" from 2010<br>\n<a href=\"https://bioconductor.org/packages/release/bioc/html/GENIE3.html\" target=\"_blank\">https://bioconductor.org/packages/release/bioc/html/GENIE3.html</a><br>\nIt uses feature importance of random forest for the analysis of the gene-regulatory networks.<br>\nThe setup is slightly different from the competition but quite related. </p>",
  "messages": [
    {
      "id": "2031703",
      "postDate": "11/16/2022 08:07:38",
      "content": "<p>Thanks for the organizers and congratulations to winners and everybody who got medals/knowledge/fun/experience ! </p>\n<p>The competition is based on real cutting-edge biological data which is quite important for research community.</p>\n<p>One of the questions which is of classical interest in such problems is : feature importance - that would reveal<br>\nhow genes (and regulatory regions for \"multiome\" part) interact which other.</p>\n<p>So please do not hesitate to share any your results/ideas/hypothesis/observations on that</p>\n<hr>\n<p>To give an example: <br>\nthe very famous package is \"GENIE3\" from 2010<br>\n<a href=\"https://bioconductor.org/packages/release/bioc/html/GENIE3.html\" target=\"_blank\">https://bioconductor.org/packages/release/bioc/html/GENIE3.html</a><br>\nIt uses feature importance of random forest for the analysis of the gene-regulatory networks.<br>\nThe setup is slightly different from the competition but quite related. </p>",
      "rawMarkdown": "Thanks for the organizers and congratulations to winners and everybody who got medals/knowledge/fun/experience ! \n\nThe competition is based on real cutting-edge biological data which is quite important for research community.\n\nOne of the questions which is of classical interest in such problems is : feature importance - that would reveal\nhow genes (and regulatory regions for \"multiome\" part) interact which other.\n\nSo please do not hesitate to share any your results/ideas/hypothesis/observations on that\n\n-------\n\nTo give an example: \nthe very famous package is \"GENIE3\" from 2010\nhttps://bioconductor.org/packages/release/bioc/html/GENIE3.html\nIt uses feature importance of random forest for the analysis of the gene-regulatory networks.\nThe setup is slightly different from the competition but quite related.",
      "votes": null
    },
    {
      "id": "2032005",
      "postDate": "11/16/2022 11:02:21",
      "content": "<p>Thank you all for your excellent work.<br>\nI would also like to thank Alexander Chervov for his great contribution to this interesting competition (sharing notebooks, presenting Discussion, sharing knowledge, etc.).</p>\n<p>Feature selection (selecting important genes from many genes) is very important in today's science.</p>\n<p>I have been experimenting with various feature selection methods from bulk RNA-Seq data that I have prepared myself. (If you would like to take a look at the dataset or notebook, I would be happy to help.)</p>\n<p>Filter method, Wrapper method, Embedding method, \"WGCNA\" library, unsupervised feature extraction with PCA, etc.</p>\n<p>We hope to receive your valuable feedback through this Discussion.<br>\nThank you very much for introducing \"GENIE3\" to us.　<a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a></p>",
      "rawMarkdown": "Thank you all for your excellent work.\nI would also like to thank Alexander Chervov for his great contribution to this interesting competition (sharing notebooks, presenting Discussion, sharing knowledge, etc.).\n\nFeature selection (selecting important genes from many genes) is very important in today's science.\n\nI have been experimenting with various feature selection methods from bulk RNA-Seq data that I have prepared myself. (If you would like to take a look at the dataset or notebook, I would be happy to help.)\n\nFilter method, Wrapper method, Embedding method, \"WGCNA\" library, unsupervised feature extraction with PCA, etc.\n\nWe hope to receive your valuable feedback through this Discussion.\nThank you very much for introducing \"GENIE3\" to us.　@alexandervc",
      "votes": null
    },
    {
      "id": "2059827",
      "postDate": "12/09/2022 09:02:22",
      "content": "<p>Thanks a lot for Grandmaster Silogram for sharing hist lists of important features !</p>\n<p>Hopefully someone else  can follow . </p>\n<p><a href=\"https://www.kaggle.com/code/alexandervc/silogram-cd36-feature-importances\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/silogram-cd36-feature-importances</a><br>\nHere is some brief look on importances for CD36 from Silogram - they are quite biologically meaningful.</p>\n<p>PS</p>\n<p>We can see quite some biology from them - e.g. - CD36 protein is higly activated in erythroid like cells - similar to red blood cells - so you can see genes like HBD, HBB, HBA1 as important features - that various forms of the hemoglobin - so quite as expected from biology.</p>\n<p>Also look on the so-called genes enrichment analysis with KEGG pathways - again biologically reasonable results, and compared with importances by other methods - they are quite consistent.</p>\n<p>That is a first look - we need some time to get more insights from the data.</p>",
      "rawMarkdown": "Thanks a lot for Grandmaster Silogram for sharing hist lists of important features !\n\nHopefully someone else  can follow . \n\nhttps://www.kaggle.com/code/alexandervc/silogram-cd36-feature-importances\nHere is some brief look on importances for CD36 from Silogram - they are quite biologically meaningful.\n\nPS\n\nWe can see quite some biology from them - e.g. - CD36 protein is higly activated in erythroid like cells - similar to red blood cells - so you can see genes like HBD, HBB, HBA1 as important features - that various forms of the hemoglobin - so quite as expected from biology.\n\nAlso look on the so-called genes enrichment analysis with KEGG pathways - again biologically reasonable results, and compared with importances by other methods - they are quite consistent.\n\nThat is a first look - we need some time to get more insights from the data.",
      "votes": null
    },
    {
      "id": "2061696",
      "postDate": "12/11/2022 12:16:28",
      "content": "<p><a href=\"https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances\" target=\"_blank\">https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances</a></p>\n<p>Feature importances shared by Oleg Khudyakov.<br>\nNote: these importances for all targets simulatenously - NO split target / by / target.<br>\nThey were calculated by permuation importance algororithm for Ridge regression with metric - correlations for all targets - that is why here is no split target-by-target. </p>",
      "rawMarkdown": "https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances\n\nFeature importances shared by Oleg Khudyakov.\nNote: these importances for all targets simulatenously - NO split target / by / target.\nThey were calculated by permuation importance algororithm for Ridge regression with metric - correlations for all targets - that is why here is no split target-by-target.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2032005,
      "author_name": "yoshifumimiya",
      "author_url": "",
      "post_date": "11/16/2022 11:02:21",
      "content": "<p>Thank you all for your excellent work.<br>\nI would also like to thank Alexander Chervov for his great contribution to this interesting competition (sharing notebooks, presenting Discussion, sharing knowledge, etc.).</p>\n<p>Feature selection (selecting important genes from many genes) is very important in today's science.</p>\n<p>I have been experimenting with various feature selection methods from bulk RNA-Seq data that I have prepared myself. (If you would like to take a look at the dataset or notebook, I would be happy to help.)</p>\n<p>Filter method, Wrapper method, Embedding method, \"WGCNA\" library, unsupervised feature extraction with PCA, etc.</p>\n<p>We hope to receive your valuable feedback through this Discussion.<br>\nThank you very much for introducing \"GENIE3\" to us.　<a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2059827,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "12/09/2022 09:02:22",
      "content": "<p>Thanks a lot for Grandmaster Silogram for sharing hist lists of important features !</p>\n<p>Hopefully someone else  can follow . </p>\n<p><a href=\"https://www.kaggle.com/code/alexandervc/silogram-cd36-feature-importances\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/silogram-cd36-feature-importances</a><br>\nHere is some brief look on importances for CD36 from Silogram - they are quite biologically meaningful.</p>\n<p>PS</p>\n<p>We can see quite some biology from them - e.g. - CD36 protein is higly activated in erythroid like cells - similar to red blood cells - so you can see genes like HBD, HBB, HBA1 as important features - that various forms of the hemoglobin - so quite as expected from biology.</p>\n<p>Also look on the so-called genes enrichment analysis with KEGG pathways - again biologically reasonable results, and compared with importances by other methods - they are quite consistent.</p>\n<p>That is a first look - we need some time to get more insights from the data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2061696,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "12/11/2022 12:16:28",
      "content": "<p><a href=\"https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances\" target=\"_blank\">https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances</a></p>\n<p>Feature importances shared by Oleg Khudyakov.<br>\nNote: these importances for all targets simulatenously - NO split target / by / target.<br>\nThey were calculated by permuation importance algororithm for Ridge regression with metric - correlations for all targets - that is why here is no split target-by-target. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2031703": "Thanks for the organizers and congratulations to winners and everybody who got medals/knowledge/fun/experience ! \n\nThe competition is based on real cutting-edge biological data which is quite important for research community.\n\nOne of the questions which is of classical interest in such problems is : feature importance - that would reveal\nhow genes (and regulatory regions for \"multiome\" part) interact which other.\n\nSo please do not hesitate to share any your results/ideas/hypothesis/observations on that\n\n-------\n\nTo give an example: \nthe very famous package is \"GENIE3\" from 2010\nhttps://bioconductor.org/packages/release/bioc/html/GENIE3.html\nIt uses feature importance of random forest for the analysis of the gene-regulatory networks.\nThe setup is slightly different from the competition but quite related.",
    "2032005": "Thank you all for your excellent work.\nI would also like to thank Alexander Chervov for his great contribution to this interesting competition (sharing notebooks, presenting Discussion, sharing knowledge, etc.).\n\nFeature selection (selecting important genes from many genes) is very important in today's science.\n\nI have been experimenting with various feature selection methods from bulk RNA-Seq data that I have prepared myself. (If you would like to take a look at the dataset or notebook, I would be happy to help.)\n\nFilter method, Wrapper method, Embedding method, \"WGCNA\" library, unsupervised feature extraction with PCA, etc.\n\nWe hope to receive your valuable feedback through this Discussion.\nThank you very much for introducing \"GENIE3\" to us.　@alexandervc",
    "2059827": "Thanks a lot for Grandmaster Silogram for sharing hist lists of important features !\n\nHopefully someone else  can follow . \n\nhttps://www.kaggle.com/code/alexandervc/silogram-cd36-feature-importances\nHere is some brief look on importances for CD36 from Silogram - they are quite biologically meaningful.\n\nPS\n\nWe can see quite some biology from them - e.g. - CD36 protein is higly activated in erythroid like cells - similar to red blood cells - so you can see genes like HBD, HBB, HBA1 as important features - that various forms of the hemoglobin - so quite as expected from biology.\n\nAlso look on the so-called genes enrichment analysis with KEGG pathways - again biologically reasonable results, and compared with importances by other methods - they are quite consistent.\n\nThat is a first look - we need some time to get more insights from the data.",
    "2061696": "https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances\n\nFeature importances shared by Oleg Khudyakov.\nNote: these importances for all targets simulatenously - NO split target / by / target.\nThey were calculated by permuation importance algororithm for Ridge regression with metric - correlations for all targets - that is why here is no split target-by-target."
  },
  "source": "meta"
}