{
  "id": 361025,
  "title": "Some ideas, bio motivated and not only",
  "url": "/competitions/open-problems-multimodal/discussion/361025",
  "author_name": "",
  "post_date": "2022-10-19T12:27:39.019408800Z",
  "votes": 21,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Colleagues, that competition is kind of exceptional on Kaggle (imho), because the data considered here is new, but in forthcoming years dozens similar datasets will be produced and later hundreds. So insights which you may find - can be of DIRECT  use for many researchers and potentially in future may be  useful for cancer immunotherapy and, so saving people's lives. </p>\n<p>Although improvement of score on tiny epsilon is not useful for practice, but the ideas which can be found during that, might be useful. So let me share some ideas - no guarantee any will work at all for score improvement - but even negative results might be of some use for researchers. So it would great to hear your feedback and may be even collaborate in joint publication(s).</p>\n<p>See also some ideas posted before:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900</a><br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350863\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350863</a><br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856</a></p>\n<p>It is mainly about CITE-seq: </p>\n<p>Bio1) <br>\nTake other available datasets with similar data.<br>\na) Train simple models on them. Apply these models for current data - use that predictions as features - e.g. feature engineering. <br>\nb) Analyze  most correlated targets&lt;-&gt;features   - add these features to your models. </p>\n<p>Since you are using completely different data - there should be no overfit. Technical problem - other datasets might have different preprocessing  - that might be painful.</p>\n<p>Bio2)<br>\nAnother feature engineering idea. It is motivated by denoising  denoising algorithms - like MAGIC. <br>\nThe gist of these algorithms - to take averages of data for cells with similar features. <br>\nSo the simple form of the idea: <br>\nfor each cell find say 10 most close cells. And just add features for these cells as new features for your cell.<br>\nWhat metric to use for \"close\" - typically make PCA to 30-100 dims and use Euclidian after the PCA.</p>\n<p>The ultimate development of the idea might be something like graph neural network on cells * genes - but no clear idea, only feelings…</p>\n<p>Bio3)<br>\nTry to find important features from biological databases. <br>\nI.e. if target CD** gene is close to some genes A,B - try to use them as features for predictions. <br>\na) Protein-Protein Interactions networks e.g. <a href=\"https://www.kaggle.com/datasets/alexandervc/protein-protein-interactions\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/protein-protein-interactions</a><br>\nb) Knowledge graphs like WikiData, or recent from M.Zitnik team <br>\nc) Textual descriptions is similar after bert-like sentence embedding - <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360802\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360802</a></p>\n<p>Also Pathway and GO databases - with idea if CD <br>\nd)  - Reactome, WikiPathways, Gene Ontology<br>\nwith the idea if CD** target is in the same pathway with some genes - these genes might be important features. <br>\nE.g. For Reactome might be helpful:  <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360455\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360455</a></p>\n<p>Bio4) <br>\nTry to exclude the so-called housekeeping genes. <br>\n<a href=\"https://housekeeping.unicamp.br/?download\" target=\"_blank\">https://housekeeping.unicamp.br/?download</a><br>\n<a href=\"https://en.wikipedia.org/wiki/Housekeeping_gene\" target=\"_blank\">https://en.wikipedia.org/wiki/Housekeeping_gene</a><br>\nThese genes are related to basic functioning of all cells, <br>\nit quite might be that they are NOT good predictors for CD** proteins. <br>\nBut if that is not true - that would be interesting.</p>\n<p>Bio5) <br>\nThe genes can be splited to protein coding genes, and not.<br>\nProtein-coding are somewhat main players in everything,<br>\nso might be keeping only them would be enough ? <br>\nBut on the other hand long non coding RNA play essential role in regulation,<br>\nso may be not. </p>\n<p>Bio6) <br>\nGenes groups and feature engineering .<br>\nThere are many genes groups which mean that genes in the group are somewhat similar,<br>\nor appeared in some analysis serving one purpose (in such case term \"signature\" is used).<br>\nOne can try to take some group of genes,<br>\nand just take averages for all genes in groups - such a simple feature engineering.<br>\nAveraging decreases noise - thus for related genes we might get some <br>\ndenoised feature which is corresponding to some  biological \"function\".<br>\nThus it can be better predictor.</p>\n<p>Unfortunately there are too many genes groups (signatures)<br>\nthe database MSigDB contains 30 000+ signature - so that is clearly much much more that<br>\nmight be reasonable for that particular task.</p>\n<p>May be one should take some specific from it ? <br>\nMay be start from the \"Hallmark signatures\" - about 100++ ?</p>\n<p>Bio7) (DS/Bio) <br>\nFeature engineering - try to predict some features from the other features. And use these predictions as new features.<br>\nIt is a kind of \"denoising\". <br>\nWhy - because input features are full of technical and biological noise, in particular \"dropouts\" (zeros which should be zeros). When you make predictions you  smooth in some sense the data - and that smoothed data might have <br>\nmore biological relevance comparing to original ones. <br>\nThere are many \"denoising\" algorithms like \"MAGIC\" for scRNA-seq data - but it can be another approach. <br>\nEspecially for CITE-seq  - pay attention on predictions on 132 genes, which proteins we need to predict.</p>\n<p>Do not be surprised if NON of these bio ideas would work for score improvement.<br>\nBut even so it would be interesting benchmark for research.</p>\n<p>TO BE CONTINUED (hopefully)</p>\n<p>===============</p>\n<p>More DS , not bio ideas.</p>\n<p>DS1) It seems CD** targets are full of outliers - might be one can first clip targets by say 0.5, 99.5 percentiles<br>\nand train model on that. Ideally one can choose approptiate percentiles clip thresholds separately for each target.<br>\nExperiments with \"no-model\" , but just averaging of targets - show there is very small improvment.<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-analysis-of-evaluation-and-submission\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-analysis-of-evaluation-and-submission</a><br>\nBut not sure it will work  for your model.<br>\nNevertheless it might be intersting question in general how to proper treat target outliers for multitarget tasks.</p>\n<p>DS2)<br>\nThe ideas on cross-validation scheme has been already described:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860</a><br>\nExperiments show that CV is LB correspondence is not bad.<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes</a><br>\nOn the other hand it seems all \"tests\" behave similar.<br>\nSo probably CV scheme is not that much important as I thought before.<br>\nOrganizers - based on the last year experience chosen targetwise correlation as a metric for stability reasons (as they report). It seems it is indeed very stable and does not feel small details. That might be not bad for competitors. </p>\n<p>DS3) <br>\nUnderstanding and tricks with the metric:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360253\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360253</a></p>\n<p>DS4)<br>\nPostprocessing for Multiome: <br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349132\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349132</a></p>\n<p>DS5)<br>\nStacking. But pay attention on folds organization.<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/361560\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/361560</a></p>\n<p>DS6) <br>\nHow to make \"1000\" models from 1 model. And blend them all.<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/362751\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/362751</a></p>\n<p>DS7)<br>\nBuilding special CNN on  \"image\". <br>\nI.e. take sample and put under it K-th its nearest neigbours - thus you get an \"image\".<br>\nTry to make CNN on such an \"image\".<br>\nIt would be great first to \"learn the order\" of features - before arranging them into the matrix -  that was  winning idea (top2 place) by user \"tmp\" for MOA,<br>\npay attention - he is on top again. <br>\nI.e. special layer which learns how to reorder the features .<br>\n(That idea is similar to Bio2).</p>\n<p>TO BE CONTINUED (hopefully)</p>\n<p>===============</p>\n<p>Would be nice to discuss and hopefully collaborate. <br>\nWelcome to our chat :  <a href=\"https://t.me/sberlogacompete\" target=\"_blank\">https://t.me/sberlogacompete</a> ( NO private sharing please there ! )<br>\nAnd see some research scripts on Kaggle about cell cycle analysis based on single cell RNA seq data,<br>\n(one modality here) - e.g. like: <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-cell-cycle-03-daybydaychange-megakaryocyte\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-cell-cycle-03-daybydaychange-megakaryocyte</a> <br>\nNot clear how to use for score improvement, but it is related to many research questions.</p>",
  "messages": [
    {
      "id": "1995022",
      "postDate": "10/19/2022 12:27:39",
      "content": "<p>Colleagues, that competition is kind of exceptional on Kaggle (imho), because the data considered here is new, but in forthcoming years dozens similar datasets will be produced and later hundreds. So insights which you may find - can be of DIRECT  use for many researchers and potentially in future may be  useful for cancer immunotherapy and, so saving people's lives. </p>\n<p>Although improvement of score on tiny epsilon is not useful for practice, but the ideas which can be found during that, might be useful. So let me share some ideas - no guarantee any will work at all for score improvement - but even negative results might be of some use for researchers. So it would great to hear your feedback and may be even collaborate in joint publication(s).</p>\n<p>See also some ideas posted before:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900</a><br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350863\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350863</a><br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856</a></p>\n<p>It is mainly about CITE-seq: </p>\n<p>Bio1) <br>\nTake other available datasets with similar data.<br>\na) Train simple models on them. Apply these models for current data - use that predictions as features - e.g. feature engineering. <br>\nb) Analyze  most correlated targets&lt;-&gt;features   - add these features to your models. </p>\n<p>Since you are using completely different data - there should be no overfit. Technical problem - other datasets might have different preprocessing  - that might be painful.</p>\n<p>Bio2)<br>\nAnother feature engineering idea. It is motivated by denoising  denoising algorithms - like MAGIC. <br>\nThe gist of these algorithms - to take averages of data for cells with similar features. <br>\nSo the simple form of the idea: <br>\nfor each cell find say 10 most close cells. And just add features for these cells as new features for your cell.<br>\nWhat metric to use for \"close\" - typically make PCA to 30-100 dims and use Euclidian after the PCA.</p>\n<p>The ultimate development of the idea might be something like graph neural network on cells * genes - but no clear idea, only feelings…</p>\n<p>Bio3)<br>\nTry to find important features from biological databases. <br>\nI.e. if target CD** gene is close to some genes A,B - try to use them as features for predictions. <br>\na) Protein-Protein Interactions networks e.g. <a href=\"https://www.kaggle.com/datasets/alexandervc/protein-protein-interactions\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/protein-protein-interactions</a><br>\nb) Knowledge graphs like WikiData, or recent from M.Zitnik team <br>\nc) Textual descriptions is similar after bert-like sentence embedding - <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360802\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360802</a></p>\n<p>Also Pathway and GO databases - with idea if CD <br>\nd)  - Reactome, WikiPathways, Gene Ontology<br>\nwith the idea if CD** target is in the same pathway with some genes - these genes might be important features. <br>\nE.g. For Reactome might be helpful:  <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360455\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360455</a></p>\n<p>Bio4) <br>\nTry to exclude the so-called housekeeping genes. <br>\n<a href=\"https://housekeeping.unicamp.br/?download\" target=\"_blank\">https://housekeeping.unicamp.br/?download</a><br>\n<a href=\"https://en.wikipedia.org/wiki/Housekeeping_gene\" target=\"_blank\">https://en.wikipedia.org/wiki/Housekeeping_gene</a><br>\nThese genes are related to basic functioning of all cells, <br>\nit quite might be that they are NOT good predictors for CD** proteins. <br>\nBut if that is not true - that would be interesting.</p>\n<p>Bio5) <br>\nThe genes can be splited to protein coding genes, and not.<br>\nProtein-coding are somewhat main players in everything,<br>\nso might be keeping only them would be enough ? <br>\nBut on the other hand long non coding RNA play essential role in regulation,<br>\nso may be not. </p>\n<p>Bio6) <br>\nGenes groups and feature engineering .<br>\nThere are many genes groups which mean that genes in the group are somewhat similar,<br>\nor appeared in some analysis serving one purpose (in such case term \"signature\" is used).<br>\nOne can try to take some group of genes,<br>\nand just take averages for all genes in groups - such a simple feature engineering.<br>\nAveraging decreases noise - thus for related genes we might get some <br>\ndenoised feature which is corresponding to some  biological \"function\".<br>\nThus it can be better predictor.</p>\n<p>Unfortunately there are too many genes groups (signatures)<br>\nthe database MSigDB contains 30 000+ signature - so that is clearly much much more that<br>\nmight be reasonable for that particular task.</p>\n<p>May be one should take some specific from it ? <br>\nMay be start from the \"Hallmark signatures\" - about 100++ ?</p>\n<p>Bio7) (DS/Bio) <br>\nFeature engineering - try to predict some features from the other features. And use these predictions as new features.<br>\nIt is a kind of \"denoising\". <br>\nWhy - because input features are full of technical and biological noise, in particular \"dropouts\" (zeros which should be zeros). When you make predictions you  smooth in some sense the data - and that smoothed data might have <br>\nmore biological relevance comparing to original ones. <br>\nThere are many \"denoising\" algorithms like \"MAGIC\" for scRNA-seq data - but it can be another approach. <br>\nEspecially for CITE-seq  - pay attention on predictions on 132 genes, which proteins we need to predict.</p>\n<p>Do not be surprised if NON of these bio ideas would work for score improvement.<br>\nBut even so it would be interesting benchmark for research.</p>\n<p>TO BE CONTINUED (hopefully)</p>\n<p>===============</p>\n<p>More DS , not bio ideas.</p>\n<p>DS1) It seems CD** targets are full of outliers - might be one can first clip targets by say 0.5, 99.5 percentiles<br>\nand train model on that. Ideally one can choose approptiate percentiles clip thresholds separately for each target.<br>\nExperiments with \"no-model\" , but just averaging of targets - show there is very small improvment.<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-analysis-of-evaluation-and-submission\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-analysis-of-evaluation-and-submission</a><br>\nBut not sure it will work  for your model.<br>\nNevertheless it might be intersting question in general how to proper treat target outliers for multitarget tasks.</p>\n<p>DS2)<br>\nThe ideas on cross-validation scheme has been already described:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860</a><br>\nExperiments show that CV is LB correspondence is not bad.<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes</a><br>\nOn the other hand it seems all \"tests\" behave similar.<br>\nSo probably CV scheme is not that much important as I thought before.<br>\nOrganizers - based on the last year experience chosen targetwise correlation as a metric for stability reasons (as they report). It seems it is indeed very stable and does not feel small details. That might be not bad for competitors. </p>\n<p>DS3) <br>\nUnderstanding and tricks with the metric:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360253\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360253</a></p>\n<p>DS4)<br>\nPostprocessing for Multiome: <br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349132\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349132</a></p>\n<p>DS5)<br>\nStacking. But pay attention on folds organization.<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/361560\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/361560</a></p>\n<p>DS6) <br>\nHow to make \"1000\" models from 1 model. And blend them all.<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/362751\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/362751</a></p>\n<p>DS7)<br>\nBuilding special CNN on  \"image\". <br>\nI.e. take sample and put under it K-th its nearest neigbours - thus you get an \"image\".<br>\nTry to make CNN on such an \"image\".<br>\nIt would be great first to \"learn the order\" of features - before arranging them into the matrix -  that was  winning idea (top2 place) by user \"tmp\" for MOA,<br>\npay attention - he is on top again. <br>\nI.e. special layer which learns how to reorder the features .<br>\n(That idea is similar to Bio2).</p>\n<p>TO BE CONTINUED (hopefully)</p>\n<p>===============</p>\n<p>Would be nice to discuss and hopefully collaborate. <br>\nWelcome to our chat :  <a href=\"https://t.me/sberlogacompete\" target=\"_blank\">https://t.me/sberlogacompete</a> ( NO private sharing please there ! )<br>\nAnd see some research scripts on Kaggle about cell cycle analysis based on single cell RNA seq data,<br>\n(one modality here) - e.g. like: <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-cell-cycle-03-daybydaychange-megakaryocyte\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-cell-cycle-03-daybydaychange-megakaryocyte</a> <br>\nNot clear how to use for score improvement, but it is related to many research questions.</p>",
      "rawMarkdown": "Colleagues, that competition is kind of exceptional on Kaggle (imho), because the data considered here is new, but in forthcoming years dozens similar datasets will be produced and later hundreds. So insights which you may find - can be of DIRECT  use for many researchers and potentially in future may be  useful for cancer immunotherapy and, so saving people's lives. \n\nAlthough improvement of score on tiny epsilon is not useful for practice, but the ideas which can be found during that, might be useful. So let me share some ideas - no guarantee any will work at all for score improvement - but even negative results might be of some use for researchers. So it would great to hear your feedback and may be even collaborate in joint publication(s).\n\nSee also some ideas posted before:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/350863\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856\n\nIt is mainly about CITE-seq: \n\nBio1) \nTake other available datasets with similar data.\na) Train simple models on them. Apply these models for current data - use that predictions as features - e.g. feature engineering. \nb) Analyze  most correlated targets<->features   - add these features to your models. \n\nSince you are using completely different data - there should be no overfit. Technical problem - other datasets might have different preprocessing  - that might be painful.\n\nBio2)\nAnother feature engineering idea. It is motivated by denoising  denoising algorithms - like MAGIC. \nThe gist of these algorithms - to take averages of data for cells with similar features. \nSo the simple form of the idea: \nfor each cell find say 10 most close cells. And just add features for these cells as new features for your cell.\nWhat metric to use for \"close\" - typically make PCA to 30-100 dims and use Euclidian after the PCA.\n\nThe ultimate development of the idea might be something like graph neural network on cells * genes - but no clear idea, only feelings...\n\nBio3)\nTry to find important features from biological databases. \nI.e. if target CD** gene is close to some genes A,B - try to use them as features for predictions. \na) Protein-Protein Interactions networks e.g. https://www.kaggle.com/datasets/alexandervc/protein-protein-interactions\nb) Knowledge graphs like WikiData, or recent from M.Zitnik team \nc) Textual descriptions is similar after bert-like sentence embedding - https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360802\n\nAlso Pathway and GO databases - with idea if CD \nd)  - Reactome, WikiPathways, Gene Ontology\nwith the idea if CD** target is in the same pathway with some genes - these genes might be important features. \nE.g. For Reactome might be helpful:  https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360455\n\nBio4) \nTry to exclude the so-called housekeeping genes. \nhttps://housekeeping.unicamp.br/?download\nhttps://en.wikipedia.org/wiki/Housekeeping_gene\nThese genes are related to basic functioning of all cells, \nit quite might be that they are NOT good predictors for CD** proteins. \nBut if that is not true - that would be interesting.\n\nBio5) \nThe genes can be splited to protein coding genes, and not.\nProtein-coding are somewhat main players in everything,\nso might be keeping only them would be enough ? \nBut on the other hand long non coding RNA play essential role in regulation,\nso may be not. \n\nBio6) \nGenes groups and feature engineering .\nThere are many genes groups which mean that genes in the group are somewhat similar,\nor appeared in some analysis serving one purpose (in such case term \"signature\" is used).\nOne can try to take some group of genes,\nand just take averages for all genes in groups - such a simple feature engineering.\nAveraging decreases noise - thus for related genes we might get some \ndenoised feature which is corresponding to some  biological \"function\".\nThus it can be better predictor.\n\nUnfortunately there are too many genes groups (signatures)\nthe database MSigDB contains 30 000+ signature - so that is clearly much much more that\nmight be reasonable for that particular task.\n\nMay be one should take some specific from it ? \nMay be start from the \"Hallmark signatures\" - about 100++ ?\n\nBio7) (DS/Bio) \nFeature engineering - try to predict some features from the other features. And use these predictions as new features.\nIt is a kind of \"denoising\". \nWhy - because input features are full of technical and biological noise, in particular \"dropouts\" (zeros which should be zeros). When you make predictions you  smooth in some sense the data - and that smoothed data might have \nmore biological relevance comparing to original ones. \nThere are many \"denoising\" algorithms like \"MAGIC\" for scRNA-seq data - but it can be another approach. \nEspecially for CITE-seq  - pay attention on predictions on 132 genes, which proteins we need to predict.\n\n\nDo not be surprised if NON of these bio ideas would work for score improvement.\nBut even so it would be interesting benchmark for research.\n\nTO BE CONTINUED (hopefully)\n\n===============\n\nMore DS , not bio ideas.\n\nDS1) It seems CD** targets are full of outliers - might be one can first clip targets by say 0.5, 99.5 percentiles\nand train model on that. Ideally one can choose approptiate percentiles clip thresholds separately for each target.\nExperiments with \"no-model\" , but just averaging of targets - show there is very small improvment.\nhttps://www.kaggle.com/code/alexandervc/mmscel-analysis-of-evaluation-and-submission\nBut not sure it will work  for your model.\nNevertheless it might be intersting question in general how to proper treat target outliers for multitarget tasks.\n\nDS2)\nThe ideas on cross-validation scheme has been already described:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\nExperiments show that CV is LB correspondence is not bad.\nhttps://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\nOn the other hand it seems all \"tests\" behave similar.\nSo probably CV scheme is not that much important as I thought before.\nOrganizers - based on the last year experience chosen targetwise correlation as a metric for stability reasons (as they report). It seems it is indeed very stable and does not feel small details. That might be not bad for competitors. \n\nDS3) \nUnderstanding and tricks with the metric:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/360253\n\nDS4)\nPostprocessing for Multiome: \nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/349132\n\nDS5)\nStacking. But pay attention on folds organization.\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/361560\n\nDS6) \nHow to make \"1000\" models from 1 model. And blend them all.\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/362751\n\nDS7)\nBuilding special CNN on  \"image\". \nI.e. take sample and put under it K-th its nearest neigbours - thus you get an \"image\".\nTry to make CNN on such an \"image\".\nIt would be great first to \"learn the order\" of features - before arranging them into the matrix -  that was  winning idea (top2 place) by user \"tmp\" for MOA,\npay attention - he is on top again. \nI.e. special layer which learns how to reorder the features .\n(That idea is similar to Bio2).\n\nTO BE CONTINUED (hopefully)\n\n===============\n\nWould be nice to discuss and hopefully collaborate. \nWelcome to our chat :  https://t.me/sberlogacompete ( NO private sharing please there ! )\nAnd see some research scripts on Kaggle about cell cycle analysis based on single cell RNA seq data,\n(one modality here) - e.g. like: \nhttps://www.kaggle.com/code/alexandervc/mmscel-cell-cycle-03-daybydaychange-megakaryocyte \nNot clear how to use for score improvement, but it is related to many research questions.",
      "votes": null
    },
    {
      "id": "1998467",
      "postDate": "10/21/2022 16:19:59",
      "content": "<p>Nice thanks for your ideas .. need to find some magic (bio motivated) from them . For now just did basic ML stuff . </p>",
      "rawMarkdown": "Nice thanks for your ideas .. need to find some magic (bio motivated) from them . For now just did basic ML stuff .",
      "votes": null
    },
    {
      "id": "2008476",
      "postDate": "10/29/2022 05:34:13",
      "content": "<p>29 10 2022 Added:<br>\nBio7, DS5-7</p>",
      "rawMarkdown": "29 10 2022 Added:\nBio7, DS5-7",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1998467,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "10/21/2022 16:19:59",
      "content": "<p>Nice thanks for your ideas .. need to find some magic (bio motivated) from them . For now just did basic ML stuff . </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2008476,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "10/29/2022 05:34:13",
      "content": "<p>29 10 2022 Added:<br>\nBio7, DS5-7</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1995022": "Colleagues, that competition is kind of exceptional on Kaggle (imho), because the data considered here is new, but in forthcoming years dozens similar datasets will be produced and later hundreds. So insights which you may find - can be of DIRECT  use for many researchers and potentially in future may be  useful for cancer immunotherapy and, so saving people's lives. \n\nAlthough improvement of score on tiny epsilon is not useful for practice, but the ideas which can be found during that, might be useful. So let me share some ideas - no guarantee any will work at all for score improvement - but even negative results might be of some use for researchers. So it would great to hear your feedback and may be even collaborate in joint publication(s).\n\nSee also some ideas posted before:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/350863\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856\n\nIt is mainly about CITE-seq: \n\nBio1) \nTake other available datasets with similar data.\na) Train simple models on them. Apply these models for current data - use that predictions as features - e.g. feature engineering. \nb) Analyze  most correlated targets<->features   - add these features to your models. \n\nSince you are using completely different data - there should be no overfit. Technical problem - other datasets might have different preprocessing  - that might be painful.\n\nBio2)\nAnother feature engineering idea. It is motivated by denoising  denoising algorithms - like MAGIC. \nThe gist of these algorithms - to take averages of data for cells with similar features. \nSo the simple form of the idea: \nfor each cell find say 10 most close cells. And just add features for these cells as new features for your cell.\nWhat metric to use for \"close\" - typically make PCA to 30-100 dims and use Euclidian after the PCA.\n\nThe ultimate development of the idea might be something like graph neural network on cells * genes - but no clear idea, only feelings...\n\nBio3)\nTry to find important features from biological databases. \nI.e. if target CD** gene is close to some genes A,B - try to use them as features for predictions. \na) Protein-Protein Interactions networks e.g. https://www.kaggle.com/datasets/alexandervc/protein-protein-interactions\nb) Knowledge graphs like WikiData, or recent from M.Zitnik team \nc) Textual descriptions is similar after bert-like sentence embedding - https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360802\n\nAlso Pathway and GO databases - with idea if CD \nd)  - Reactome, WikiPathways, Gene Ontology\nwith the idea if CD** target is in the same pathway with some genes - these genes might be important features. \nE.g. For Reactome might be helpful:  https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360455\n\nBio4) \nTry to exclude the so-called housekeeping genes. \nhttps://housekeeping.unicamp.br/?download\nhttps://en.wikipedia.org/wiki/Housekeeping_gene\nThese genes are related to basic functioning of all cells, \nit quite might be that they are NOT good predictors for CD** proteins. \nBut if that is not true - that would be interesting.\n\nBio5) \nThe genes can be splited to protein coding genes, and not.\nProtein-coding are somewhat main players in everything,\nso might be keeping only them would be enough ? \nBut on the other hand long non coding RNA play essential role in regulation,\nso may be not. \n\nBio6) \nGenes groups and feature engineering .\nThere are many genes groups which mean that genes in the group are somewhat similar,\nor appeared in some analysis serving one purpose (in such case term \"signature\" is used).\nOne can try to take some group of genes,\nand just take averages for all genes in groups - such a simple feature engineering.\nAveraging decreases noise - thus for related genes we might get some \ndenoised feature which is corresponding to some  biological \"function\".\nThus it can be better predictor.\n\nUnfortunately there are too many genes groups (signatures)\nthe database MSigDB contains 30 000+ signature - so that is clearly much much more that\nmight be reasonable for that particular task.\n\nMay be one should take some specific from it ? \nMay be start from the \"Hallmark signatures\" - about 100++ ?\n\nBio7) (DS/Bio) \nFeature engineering - try to predict some features from the other features. And use these predictions as new features.\nIt is a kind of \"denoising\". \nWhy - because input features are full of technical and biological noise, in particular \"dropouts\" (zeros which should be zeros). When you make predictions you  smooth in some sense the data - and that smoothed data might have \nmore biological relevance comparing to original ones. \nThere are many \"denoising\" algorithms like \"MAGIC\" for scRNA-seq data - but it can be another approach. \nEspecially for CITE-seq  - pay attention on predictions on 132 genes, which proteins we need to predict.\n\n\nDo not be surprised if NON of these bio ideas would work for score improvement.\nBut even so it would be interesting benchmark for research.\n\nTO BE CONTINUED (hopefully)\n\n===============\n\nMore DS , not bio ideas.\n\nDS1) It seems CD** targets are full of outliers - might be one can first clip targets by say 0.5, 99.5 percentiles\nand train model on that. Ideally one can choose approptiate percentiles clip thresholds separately for each target.\nExperiments with \"no-model\" , but just averaging of targets - show there is very small improvment.\nhttps://www.kaggle.com/code/alexandervc/mmscel-analysis-of-evaluation-and-submission\nBut not sure it will work  for your model.\nNevertheless it might be intersting question in general how to proper treat target outliers for multitarget tasks.\n\nDS2)\nThe ideas on cross-validation scheme has been already described:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\nExperiments show that CV is LB correspondence is not bad.\nhttps://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\nOn the other hand it seems all \"tests\" behave similar.\nSo probably CV scheme is not that much important as I thought before.\nOrganizers - based on the last year experience chosen targetwise correlation as a metric for stability reasons (as they report). It seems it is indeed very stable and does not feel small details. That might be not bad for competitors. \n\nDS3) \nUnderstanding and tricks with the metric:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/360253\n\nDS4)\nPostprocessing for Multiome: \nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/349132\n\nDS5)\nStacking. But pay attention on folds organization.\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/361560\n\nDS6) \nHow to make \"1000\" models from 1 model. And blend them all.\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/362751\n\nDS7)\nBuilding special CNN on  \"image\". \nI.e. take sample and put under it K-th its nearest neigbours - thus you get an \"image\".\nTry to make CNN on such an \"image\".\nIt would be great first to \"learn the order\" of features - before arranging them into the matrix -  that was  winning idea (top2 place) by user \"tmp\" for MOA,\npay attention - he is on top again. \nI.e. special layer which learns how to reorder the features .\n(That idea is similar to Bio2).\n\nTO BE CONTINUED (hopefully)\n\n===============\n\nWould be nice to discuss and hopefully collaborate. \nWelcome to our chat :  https://t.me/sberlogacompete ( NO private sharing please there ! )\nAnd see some research scripts on Kaggle about cell cycle analysis based on single cell RNA seq data,\n(one modality here) - e.g. like: \nhttps://www.kaggle.com/code/alexandervc/mmscel-cell-cycle-03-daybydaychange-megakaryocyte \nNot clear how to use for score improvement, but it is related to many research questions.",
    "1998467": "Nice thanks for your ideas .. need to find some magic (bio motivated) from them . For now just did basic ML stuff .",
    "2008476": "29 10 2022 Added:\nBio7, DS5-7"
  },
  "source": "meta"
}