{
  "id": 366428,
  "title": "3rd place solution",
  "url": "/competitions/open-problems-multimodal/writeups/makotu-3rd-place-solution",
  "author_name": "",
  "post_date": "2022-11-24T10:16:27.720Z",
  "votes": 70,
  "comment_count": 18,
  "views": 0,
  "content": "<p>First of all, to the organizers and kaggle management, and to everyone who participated with me, thank you for organizing a great competition!</p>\n<p>Before starting this competition, I had no knowledge of the biological field.To be honest, I still don't know much about it.<br>\n(Since this was a completely unprofessional field, I may have been able to try various things without bias…？)</p>\n<p>I would like to share my solution below. Please forgive me if it is difficult to read in many respects, as my English skills are not very good.</p>\n<hr>\n<h1>Summary:</h1>\n<h2>multiome</h2>\n<h4>preprocess</h4>\n<ul>\n<li>use okapi bm25 instead of tfidf</li>\n<li>dimensionality reduction：use lsi(implemented in muon) instead of svd<br>\n<a href=\"https://muon.readthedocs.io/en/latest/api/generated/muon.atac.tl.lsi.html\" target=\"_blank\">https://muon.readthedocs.io/en/latest/api/generated/muon.atac.tl.lsi.html</a><br>\n<code>muon.atac.tl.lsi(rawdata(with okapi preprocessing), n_comps=64)</code></li>\n</ul>\n<h4>feature</h4>\n<p>Basically, pre-processing contributed greatly to the accuracy, but the following features also contributed somewhat to the accuracy.</p>\n<ul>\n<li>binary feature<br>\ntransformed 0/1 binary and reduced 16 dimensions(svd) as features</li>\n<li>w2v vector feature<br>\nFor each cell, the top100 with the highest expression levels were lined up<br>\nand vectorized by gensim to get feature vector(16dims) for each gene.<br>\n<a href=\"https://radimrehurek.com/gensim/models/word2vec.html\" target=\"_blank\">https://radimrehurek.com/gensim/models/word2vec.html</a><br>\n    Ex. CellA: geneB → geneE → geneF → …<br>\n        CellB: geneA → geneC → geneM → …<br>\ntop100 genes vector average in each cell used as features.</li>\n<li>leiden cluster mean feature<br>\nI made Clusters using muon's leiden clustering(23 cluster).<br>\n<a href=\"https://muon.readthedocs.io/en/latest/api/generated/muon.tl.leiden.html#muon.tl.leiden\" target=\"_blank\">https://muon.readthedocs.io/en/latest/api/generated/muon.tl.leiden.html#muon.tl.leiden</a><br>\nAfter taking the average of the features for each cluster(23 cluster × 228942 feat),<br>\nthey were reduced to 16 dimensions by svd and used as features(23 cluster × 16feat).<br>\nAfter that, join on clusters.</li>\n<li>connectivy matrix feature<br>\nSince muon's leiden clustering generates an adjacency matrix between cells<br>\nas a byproduct, I also use it 16-dimensional with svd as a feature.</li>\n</ul>\n<h4>model</h4>\n<ul>\n<li>mlp<ul>\n<li>Simple 4-layer mlp; no major differences from mlp in public notebooks</li>\n<li>target has been reduced to 128 dimensions with svd.</li>\n<li>use rmse loss</li></ul></li>\n<li>catboost<ul>\n<li>target has been reduced to 128 dimensions with svd.</li></ul></li>\n</ul>\n<h4>ensemble</h4>\n<ul>\n<li>I made a model of nearly 20 mlp and 3 catboosts with various feature combinations. and cv-based weighted averaging.</li>\n</ul>\n<h2>cite</h2>\n<h4>preprocess</h4>\n<ul>\n<li>The same process as in the organizer was applied.<br>\n<code>use sc.pp.normalize_per_cell and sc.pp.log1p</code> (excluding the gene that are significantly related to the target protein)</li>\n</ul>\n<h4>feature</h4>\n<p>I've made a lot of features, and here are some of them that have worked to some degree.</p>\n<ul>\n<li>leiden cluster feature<br>\nI made Clusters using muon's leiden clustering.<br>\nAverage of features per cluster and reduce dimensions with svd.<br>\n(excluding Important genes. It were not used svd, and use raw count's average for each cluster was used as features as is.)</li>\n<li>w2v vector feature<br>\nSame as multiome.</li>\n</ul>\n<h4>model</h4>\n<ul>\n<li>mlp<ul>\n<li>Simple 4-layer mlp; no major differences from mlp in public notebooks</li>\n<li>Using correlation_loss. No different from public notebook.</li></ul></li>\n<li>catboost</li>\n</ul>\n<h4>ensemble</h4>\n<ul>\n<li>I made a model of nearly 20 mlp and 2 catboosts with various feature combinations. and cv-based weighted averaging.</li>\n</ul>\n<h2>Validation</h2>\n<p>It goes without saying that one of the key elements of this competition is validation.<br>\nI tried binary classification to classify test data used in PB and others. (At this point, the classification accuracy was so high. So I think it is dangerous to trust LB.)<br>\nThe 10% of the training data that is close to the PB, is used as validation data. <br>\nThis method seemed to work well, the submit with my highest cv was highest pb score.</p>\n<h2>update: Code &amp; Model &amp; Data</h2>\n<p>share code and models in github and kaggle datasets<br>\n<a href=\"https://github.com/makotu1208/open-problems-multimodal-3rd-solution\" target=\"_blank\">https://github.com/makotu1208/open-problems-multimodal-3rd-solution</a></p>",
  "messages": [
    {
      "id": "2031553",
      "postDate": "11/16/2022 06:47:53",
      "content": "<p>First of all, to the organizers and kaggle management, and to everyone who participated with me, thank you for organizing a great competition!</p>\n<p>Before starting this competition, I had no knowledge of the biological field.To be honest, I still don't know much about it.<br>\n(Since this was a completely unprofessional field, I may have been able to try various things without bias…？)</p>\n<p>I would like to share my solution below. Please forgive me if it is difficult to read in many respects, as my English skills are not very good.</p>\n<hr>\n<h1>Summary:</h1>\n<h2>multiome</h2>\n<h4>preprocess</h4>\n<ul>\n<li>use okapi bm25 instead of tfidf</li>\n<li>dimensionality reduction：use lsi(implemented in muon) instead of svd<br>\n<a href=\"https://muon.readthedocs.io/en/latest/api/generated/muon.atac.tl.lsi.html\" target=\"_blank\">https://muon.readthedocs.io/en/latest/api/generated/muon.atac.tl.lsi.html</a><br>\n<code>muon.atac.tl.lsi(rawdata(with okapi preprocessing), n_comps=64)</code></li>\n</ul>\n<h4>feature</h4>\n<p>Basically, pre-processing contributed greatly to the accuracy, but the following features also contributed somewhat to the accuracy.</p>\n<ul>\n<li>binary feature<br>\ntransformed 0/1 binary and reduced 16 dimensions(svd) as features</li>\n<li>w2v vector feature<br>\nFor each cell, the top100 with the highest expression levels were lined up<br>\nand vectorized by gensim to get feature vector(16dims) for each gene.<br>\n<a href=\"https://radimrehurek.com/gensim/models/word2vec.html\" target=\"_blank\">https://radimrehurek.com/gensim/models/word2vec.html</a><br>\n    Ex. CellA: geneB → geneE → geneF → …<br>\n        CellB: geneA → geneC → geneM → …<br>\ntop100 genes vector average in each cell used as features.</li>\n<li>leiden cluster mean feature<br>\nI made Clusters using muon's leiden clustering(23 cluster).<br>\n<a href=\"https://muon.readthedocs.io/en/latest/api/generated/muon.tl.leiden.html#muon.tl.leiden\" target=\"_blank\">https://muon.readthedocs.io/en/latest/api/generated/muon.tl.leiden.html#muon.tl.leiden</a><br>\nAfter taking the average of the features for each cluster(23 cluster × 228942 feat),<br>\nthey were reduced to 16 dimensions by svd and used as features(23 cluster × 16feat).<br>\nAfter that, join on clusters.</li>\n<li>connectivy matrix feature<br>\nSince muon's leiden clustering generates an adjacency matrix between cells<br>\nas a byproduct, I also use it 16-dimensional with svd as a feature.</li>\n</ul>\n<h4>model</h4>\n<ul>\n<li>mlp<ul>\n<li>Simple 4-layer mlp; no major differences from mlp in public notebooks</li>\n<li>target has been reduced to 128 dimensions with svd.</li>\n<li>use rmse loss</li></ul></li>\n<li>catboost<ul>\n<li>target has been reduced to 128 dimensions with svd.</li></ul></li>\n</ul>\n<h4>ensemble</h4>\n<ul>\n<li>I made a model of nearly 20 mlp and 3 catboosts with various feature combinations. and cv-based weighted averaging.</li>\n</ul>\n<h2>cite</h2>\n<h4>preprocess</h4>\n<ul>\n<li>The same process as in the organizer was applied.<br>\n<code>use sc.pp.normalize_per_cell and sc.pp.log1p</code> (excluding the gene that are significantly related to the target protein)</li>\n</ul>\n<h4>feature</h4>\n<p>I've made a lot of features, and here are some of them that have worked to some degree.</p>\n<ul>\n<li>leiden cluster feature<br>\nI made Clusters using muon's leiden clustering.<br>\nAverage of features per cluster and reduce dimensions with svd.<br>\n(excluding Important genes. It were not used svd, and use raw count's average for each cluster was used as features as is.)</li>\n<li>w2v vector feature<br>\nSame as multiome.</li>\n</ul>\n<h4>model</h4>\n<ul>\n<li>mlp<ul>\n<li>Simple 4-layer mlp; no major differences from mlp in public notebooks</li>\n<li>Using correlation_loss. No different from public notebook.</li></ul></li>\n<li>catboost</li>\n</ul>\n<h4>ensemble</h4>\n<ul>\n<li>I made a model of nearly 20 mlp and 2 catboosts with various feature combinations. and cv-based weighted averaging.</li>\n</ul>\n<h2>Validation</h2>\n<p>It goes without saying that one of the key elements of this competition is validation.<br>\nI tried binary classification to classify test data used in PB and others. (At this point, the classification accuracy was so high. So I think it is dangerous to trust LB.)<br>\nThe 10% of the training data that is close to the PB, is used as validation data. <br>\nThis method seemed to work well, the submit with my highest cv was highest pb score.</p>\n<h2>update: Code &amp; Model &amp; Data</h2>\n<p>share code and models in github and kaggle datasets<br>\n<a href=\"https://github.com/makotu1208/open-problems-multimodal-3rd-solution\" target=\"_blank\">https://github.com/makotu1208/open-problems-multimodal-3rd-solution</a></p>",
      "rawMarkdown": "First of all, to the organizers and kaggle management, and to everyone who participated with me, thank you for organizing a great competition!\n\nBefore starting this competition, I had no knowledge of the biological field.To be honest, I still don't know much about it.\n(Since this was a completely unprofessional field, I may have been able to try various things without bias…？)\n\nI would like to share my solution below. Please forgive me if it is difficult to read in many respects, as my English skills are not very good.\n\n----------------------------------------------------------------------------------------\n# Summary:\n\n## multiome\n\n#### preprocess\n- use okapi bm25 instead of tfidf\n- dimensionality reduction：use lsi(implemented in muon) instead of svd\nhttps://muon.readthedocs.io/en/latest/api/generated/muon.atac.tl.lsi.html\n`muon.atac.tl.lsi(rawdata(with okapi preprocessing), n_comps=64)`\n\n#### feature\nBasically, pre-processing contributed greatly to the accuracy, but the following features also contributed somewhat to the accuracy.\n\n- binary feature\n   transformed 0/1 binary and reduced 16 dimensions(svd) as features\n- w2v vector feature\n   For each cell, the top100 with the highest expression levels were lined up\n   and vectorized by gensim to get feature vector(16dims) for each gene.\n   https://radimrehurek.com/gensim/models/word2vec.html\n        Ex. CellA: geneB → geneE → geneF → …\n            CellB: geneA → geneC → geneM → …\n   top100 genes vector average in each cell used as features.\n- leiden cluster mean feature\n   I made Clusters using muon's leiden clustering(23 cluster).\n   https://muon.readthedocs.io/en/latest/api/generated/muon.tl.leiden.html#muon.tl.leiden\n   After taking the average of the features for each cluster(23 cluster × 228942 feat),\n   they were reduced to 16 dimensions by svd and used as features(23 cluster × 16feat).\n   After that, join on clusters.\n- connectivy matrix feature\n   Since muon's leiden clustering generates an adjacency matrix between cells\n   as a byproduct, I also use it 16-dimensional with svd as a feature.\n\n#### model\n   - mlp\n      - Simple 4-layer mlp; no major differences from mlp in public notebooks\n      - target has been reduced to 128 dimensions with svd.\n      - use rmse loss\n   - catboost\n      - target has been reduced to 128 dimensions with svd.\n\n#### ensemble\n- I made a model of nearly 20 mlp and 3 catboosts with various feature combinations. and cv-based weighted averaging.\n\n\n## cite\n#### preprocess\n- The same process as in the organizer was applied.\n`use sc.pp.normalize_per_cell and sc.pp.log1p` (excluding the gene that are significantly related to the target protein)\n\n#### feature\n I've made a lot of features, and here are some of them that have worked to some degree.\n - leiden cluster feature\nI made Clusters using muon's leiden clustering.\nAverage of features per cluster and reduce dimensions with svd.\n(excluding Important genes. It were not used svd, and use raw count's average for each cluster was used as features as is.)\n - w2v vector feature\nSame as multiome.\n\n#### model\n  - mlp\n    - Simple 4-layer mlp; no major differences from mlp in public notebooks\n    - Using correlation_loss. No different from public notebook.\n  - catboost\n\n#### ensemble\n- I made a model of nearly 20 mlp and 2 catboosts with various feature combinations. and cv-based weighted averaging.\n\n## Validation\n\nIt goes without saying that one of the key elements of this competition is validation.\nI tried binary classification to classify test data used in PB and others. (At this point, the classification accuracy was so high. So I think it is dangerous to trust LB.)\nThe 10% of the training data that is close to the PB, is used as validation data. \nThis method seemed to work well, the submit with my highest cv was highest pb score.\n\n## update: Code & Model & Data\nshare code and models in github and kaggle datasets\nhttps://github.com/makotu1208/open-problems-multimodal-3rd-solution",
      "votes": null
    },
    {
      "id": "2031699",
      "postDate": "11/16/2022 08:06:13",
      "content": "<p>Interestingly you used word2vec instead of autoencoder, have you compared them?</p>",
      "rawMarkdown": "Interestingly you used word2vec instead of autoencoder, have you compared them?",
      "votes": null
    },
    {
      "id": "2031702",
      "postDate": "11/16/2022 08:07:36",
      "content": "<p><a href=\"https://www.kaggle.com/mhyodo\" target=\"_blank\">@mhyodo</a> adversarial validation was surely a nice way to go! Congrats!</p>",
      "rawMarkdown": "mhyodo adversarial validation was surely a nice way to go! Congrats!",
      "votes": null
    },
    {
      "id": "2031829",
      "postDate": "11/16/2022 09:02:10",
      "content": "<p>a little aside, your avatar is so charming on LB😆</p>",
      "rawMarkdown": "a little aside, your avatar is so charming on LB😆",
      "votes": null
    },
    {
      "id": "2031859",
      "postDate": "11/16/2022 09:21:01",
      "content": "<p>Thanks for your comment!  I also tried making feature using autoencoder, but it did not work that well within my experiments.</p>",
      "rawMarkdown": "Thanks for your comment!  I also tried making feature using autoencoder, but it did not work that well within my experiments.",
      "votes": null
    },
    {
      "id": "2032062",
      "postDate": "11/16/2022 11:55:35",
      "content": "<p>I think it is also interesting to fit the vectors for each position directly to an RNN layer rather than taking an average. </p>",
      "rawMarkdown": "I think it is also interesting to fit the vectors for each position directly to an RNN layer rather than taking an average.",
      "votes": null
    },
    {
      "id": "2032397",
      "postDate": "11/16/2022 15:42:41",
      "content": "<p>This is really a nice idea</p>",
      "rawMarkdown": "This is really a nice idea",
      "votes": null
    },
    {
      "id": "2032502",
      "postDate": "11/16/2022 16:35:42",
      "content": "<blockquote>\n  <p>cite<br>\n  model<br>\n  catboost<br>\n  target has been reduced to 128 dimensions with svd.</p>\n</blockquote>\n<p>Did you reduce target for CITE to 128 dimentions also?</p>",
      "rawMarkdown": ">  cite\n> model\n> catboost\n> target has been reduced to 128 dimensions with svd.\n\nDid you reduce target for CITE to 128 dimentions also?",
      "votes": null
    },
    {
      "id": "2032975",
      "postDate": "11/17/2022 00:08:25",
      "content": "<p>Congratulations for your solo gold, and thank you for your informative solutions. <br>\nI want to ask a question about w2v vector features. Did you input numerical data into w2v and converted them to 16 dimensional vector? And I couldn't surely understand the part \"top100 genes vector average in each cell used as features”. Could you explain more details? <br>\nThank you in advance!</p>",
      "rawMarkdown": "Congratulations for your solo gold, and thank you for your informative solutions. \nI want to ask a question about w2v vector features. Did you input numerical data into w2v and converted them to 16 dimensional vector? And I couldn't surely understand the part \"top100 genes vector average in each cell used as features”. Could you explain more details? \nThank you in advance!",
      "votes": null
    },
    {
      "id": "2033067",
      "postDate": "11/17/2022 02:35:33",
      "content": "<p>Thank you for comment!<br>\nI think it would be faster to have you look at code, so I created a notebook.<br>\n<a href=\"https://www.kaggle.com/code/mhyodo/w2v-feature-sample\" target=\"_blank\">https://www.kaggle.com/code/mhyodo/w2v-feature-sample</a></p>",
      "rawMarkdown": "Thank you for comment!\nI think it would be faster to have you look at code, so I created a notebook.\nhttps://www.kaggle.com/code/mhyodo/w2v-feature-sample",
      "votes": null
    },
    {
      "id": "2033070",
      "postDate": "11/17/2022 02:37:18",
      "content": "<p>Oops, sorry. cite used the 140 dimensional target as is without dimensional reduction. I will correct this. Thanks for pointing that out!</p>",
      "rawMarkdown": "Oops, sorry. cite used the 140 dimensional target as is without dimensional reduction. I will correct this. Thanks for pointing that out!",
      "votes": null
    },
    {
      "id": "2033124",
      "postDate": "11/17/2022 03:30:04",
      "content": "<p>I understood!<br>\nThank you, and again congratulations!</p>",
      "rawMarkdown": "I understood!\nThank you, and again congratulations!",
      "votes": null
    },
    {
      "id": "2033924",
      "postDate": "11/17/2022 16:57:01",
      "content": "<p>Thanks for sharing :) </p>",
      "rawMarkdown": "Thanks for sharing :)",
      "votes": null
    },
    {
      "id": "2034201",
      "postDate": "11/17/2022 22:53:13",
      "content": "<p><a href=\"https://www.kaggle.com/mhyodo\" target=\"_blank\">@mhyodo</a>  Big Congrat on your solo gold. And thanks a lot for sharing your solution.<br>\nCould you please help me to understand the validation part?</p>\n<p>Why did you start to doubt the LB with a high binary classification result? Does it mean that with a high score, the test data for PB and others have a huge variance, and this variance will lead to the final score shift?</p>",
      "rawMarkdown": "mhyodo  Big Congrat on your solo gold. And thanks a lot for sharing your solution.\nCould you please help me to understand the validation part?\n \nWhy did you start to doubt the LB with a high binary classification result? Does it mean that with a high score, the test data for PB and others have a huge variance, and this variance will lead to the final score shift?",
      "votes": null
    },
    {
      "id": "2034343",
      "postDate": "11/18/2022 04:26:06",
      "content": "<p>Thanks for sharing! As a beginner, this is really helpful for illustrating the importance of validation sets.</p>",
      "rawMarkdown": "Thanks for sharing! As a beginner, this is really helpful for illustrating the importance of validation sets.",
      "votes": null
    },
    {
      "id": "2035514",
      "postDate": "11/19/2022 00:01:59",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/mhyodo\" target=\"_blank\">@mhyodo</a> Fantastic job winning 3rd place Cash Gold Solo!</p>",
      "rawMarkdown": "Congratulations @mhyodo Fantastic job winning 3rd place Cash Gold Solo!",
      "votes": null
    },
    {
      "id": "2036995",
      "postDate": "11/20/2022 11:09:47",
      "content": "<p>Thanks for your question! <br>\nAlso, sorry for the delay in responding.<br>\nSorry, I need to add something. Precisely, I clearly felt that it might be dangerous to trust LB because when I performed the classification with the LB test data as 0 and the PB test data as 1 I felt this was dangerous because of the high accuracy of the classification.<br>\nIf the test data used in LB and the test data used in PB are similar I would believe LB, but since this was not the case, I thought it was dangerous to believe LB.</p>",
      "rawMarkdown": "Thanks for your question! \nAlso, sorry for the delay in responding.\nSorry, I need to add something. Precisely, I clearly felt that it might be dangerous to trust LB because when I performed the classification with the LB test data as 0 and the PB test data as 1 I felt this was dangerous because of the high accuracy of the classification.\nIf the test data used in LB and the test data used in PB are similar I would believe LB, but since this was not the case, I thought it was dangerous to believe LB.",
      "votes": null
    },
    {
      "id": "2037350",
      "postDate": "11/20/2022 16:03:55",
      "content": "<p>Ah， got it! thanks again for your explanation </p>",
      "rawMarkdown": "Ah， got it! thanks again for your explanation",
      "votes": null
    },
    {
      "id": "2041640",
      "postDate": "11/24/2022 05:18:33",
      "content": "<p>Congratulations and thanks for posting your great solution Makotu!</p>",
      "rawMarkdown": "Congratulations and thanks for posting your great solution Makotu!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2031699,
      "author_name": "kingychiu",
      "author_url": "",
      "post_date": "11/16/2022 08:06:13",
      "content": "<p>Interestingly you used word2vec instead of autoencoder, have you compared them?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2031859,
          "author_name": "mhyodo",
          "author_url": "",
          "post_date": "11/16/2022 09:21:01",
          "content": "<p>Thanks for your comment!  I also tried making feature using autoencoder, but it did not work that well within my experiments.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2032062,
          "author_name": "kingychiu",
          "author_url": "",
          "post_date": "11/16/2022 11:55:35",
          "content": "<p>I think it is also interesting to fit the vectors for each position directly to an RNN layer rather than taking an average. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2031702,
      "author_name": "callmeb",
      "author_url": "",
      "post_date": "11/16/2022 08:07:36",
      "content": "<p><a href=\"https://www.kaggle.com/mhyodo\" target=\"_blank\">@mhyodo</a> adversarial validation was surely a nice way to go! Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2031829,
      "author_name": "sing4itluo",
      "author_url": "",
      "post_date": "11/16/2022 09:02:10",
      "content": "<p>a little aside, your avatar is so charming on LB😆</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2032397,
      "author_name": "mosesaborisade",
      "author_url": "",
      "post_date": "11/16/2022 15:42:41",
      "content": "<p>This is really a nice idea</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2032502,
      "author_name": "artemfedorov",
      "author_url": "",
      "post_date": "11/16/2022 16:35:42",
      "content": "<blockquote>\n  <p>cite<br>\n  model<br>\n  catboost<br>\n  target has been reduced to 128 dimensions with svd.</p>\n</blockquote>\n<p>Did you reduce target for CITE to 128 dimentions also?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2033070,
          "author_name": "mhyodo",
          "author_url": "",
          "post_date": "11/17/2022 02:37:18",
          "content": "<p>Oops, sorry. cite used the 140 dimensional target as is without dimensional reduction. I will correct this. Thanks for pointing that out!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2032975,
      "author_name": "ludditep",
      "author_url": "",
      "post_date": "11/17/2022 00:08:25",
      "content": "<p>Congratulations for your solo gold, and thank you for your informative solutions. <br>\nI want to ask a question about w2v vector features. Did you input numerical data into w2v and converted them to 16 dimensional vector? And I couldn't surely understand the part \"top100 genes vector average in each cell used as features”. Could you explain more details? <br>\nThank you in advance!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2033067,
          "author_name": "mhyodo",
          "author_url": "",
          "post_date": "11/17/2022 02:35:33",
          "content": "<p>Thank you for comment!<br>\nI think it would be faster to have you look at code, so I created a notebook.<br>\n<a href=\"https://www.kaggle.com/code/mhyodo/w2v-feature-sample\" target=\"_blank\">https://www.kaggle.com/code/mhyodo/w2v-feature-sample</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2033124,
          "author_name": "ludditep",
          "author_url": "",
          "post_date": "11/17/2022 03:30:04",
          "content": "<p>I understood!<br>\nThank you, and again congratulations!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2033924,
      "author_name": "pvtrmalli",
      "author_url": "",
      "post_date": "11/17/2022 16:57:01",
      "content": "<p>Thanks for sharing :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2034201,
      "author_name": "bwhale",
      "author_url": "",
      "post_date": "11/17/2022 22:53:13",
      "content": "<p><a href=\"https://www.kaggle.com/mhyodo\" target=\"_blank\">@mhyodo</a>  Big Congrat on your solo gold. And thanks a lot for sharing your solution.<br>\nCould you please help me to understand the validation part?</p>\n<p>Why did you start to doubt the LB with a high binary classification result? Does it mean that with a high score, the test data for PB and others have a huge variance, and this variance will lead to the final score shift?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2036995,
          "author_name": "mhyodo",
          "author_url": "",
          "post_date": "11/20/2022 11:09:47",
          "content": "<p>Thanks for your question! <br>\nAlso, sorry for the delay in responding.<br>\nSorry, I need to add something. Precisely, I clearly felt that it might be dangerous to trust LB because when I performed the classification with the LB test data as 0 and the PB test data as 1 I felt this was dangerous because of the high accuracy of the classification.<br>\nIf the test data used in LB and the test data used in PB are similar I would believe LB, but since this was not the case, I thought it was dangerous to believe LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2037350,
          "author_name": "bwhale",
          "author_url": "",
          "post_date": "11/20/2022 16:03:55",
          "content": "<p>Ah， got it! thanks again for your explanation </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2034343,
      "author_name": "perrypineapple",
      "author_url": "",
      "post_date": "11/18/2022 04:26:06",
      "content": "<p>Thanks for sharing! As a beginner, this is really helpful for illustrating the importance of validation sets.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2035514,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/19/2022 00:01:59",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/mhyodo\" target=\"_blank\">@mhyodo</a> Fantastic job winning 3rd place Cash Gold Solo!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2041640,
      "author_name": "songqizhou",
      "author_url": "",
      "post_date": "11/24/2022 05:18:33",
      "content": "<p>Congratulations and thanks for posting your great solution Makotu!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2031553": "First of all, to the organizers and kaggle management, and to everyone who participated with me, thank you for organizing a great competition!\n\nBefore starting this competition, I had no knowledge of the biological field.To be honest, I still don't know much about it.\n(Since this was a completely unprofessional field, I may have been able to try various things without bias…？)\n\nI would like to share my solution below. Please forgive me if it is difficult to read in many respects, as my English skills are not very good.\n\n----------------------------------------------------------------------------------------\n# Summary:\n\n## multiome\n\n#### preprocess\n- use okapi bm25 instead of tfidf\n- dimensionality reduction：use lsi(implemented in muon) instead of svd\nhttps://muon.readthedocs.io/en/latest/api/generated/muon.atac.tl.lsi.html\n`muon.atac.tl.lsi(rawdata(with okapi preprocessing), n_comps=64)`\n\n#### feature\nBasically, pre-processing contributed greatly to the accuracy, but the following features also contributed somewhat to the accuracy.\n\n- binary feature\n   transformed 0/1 binary and reduced 16 dimensions(svd) as features\n- w2v vector feature\n   For each cell, the top100 with the highest expression levels were lined up\n   and vectorized by gensim to get feature vector(16dims) for each gene.\n   https://radimrehurek.com/gensim/models/word2vec.html\n        Ex. CellA: geneB → geneE → geneF → …\n            CellB: geneA → geneC → geneM → …\n   top100 genes vector average in each cell used as features.\n- leiden cluster mean feature\n   I made Clusters using muon's leiden clustering(23 cluster).\n   https://muon.readthedocs.io/en/latest/api/generated/muon.tl.leiden.html#muon.tl.leiden\n   After taking the average of the features for each cluster(23 cluster × 228942 feat),\n   they were reduced to 16 dimensions by svd and used as features(23 cluster × 16feat).\n   After that, join on clusters.\n- connectivy matrix feature\n   Since muon's leiden clustering generates an adjacency matrix between cells\n   as a byproduct, I also use it 16-dimensional with svd as a feature.\n\n#### model\n   - mlp\n      - Simple 4-layer mlp; no major differences from mlp in public notebooks\n      - target has been reduced to 128 dimensions with svd.\n      - use rmse loss\n   - catboost\n      - target has been reduced to 128 dimensions with svd.\n\n#### ensemble\n- I made a model of nearly 20 mlp and 3 catboosts with various feature combinations. and cv-based weighted averaging.\n\n\n## cite\n#### preprocess\n- The same process as in the organizer was applied.\n`use sc.pp.normalize_per_cell and sc.pp.log1p` (excluding the gene that are significantly related to the target protein)\n\n#### feature\n I've made a lot of features, and here are some of them that have worked to some degree.\n - leiden cluster feature\nI made Clusters using muon's leiden clustering.\nAverage of features per cluster and reduce dimensions with svd.\n(excluding Important genes. It were not used svd, and use raw count's average for each cluster was used as features as is.)\n - w2v vector feature\nSame as multiome.\n\n#### model\n  - mlp\n    - Simple 4-layer mlp; no major differences from mlp in public notebooks\n    - Using correlation_loss. No different from public notebook.\n  - catboost\n\n#### ensemble\n- I made a model of nearly 20 mlp and 2 catboosts with various feature combinations. and cv-based weighted averaging.\n\n## Validation\n\nIt goes without saying that one of the key elements of this competition is validation.\nI tried binary classification to classify test data used in PB and others. (At this point, the classification accuracy was so high. So I think it is dangerous to trust LB.)\nThe 10% of the training data that is close to the PB, is used as validation data. \nThis method seemed to work well, the submit with my highest cv was highest pb score.\n\n## update: Code & Model & Data\nshare code and models in github and kaggle datasets\nhttps://github.com/makotu1208/open-problems-multimodal-3rd-solution",
    "2031699": "Interestingly you used word2vec instead of autoencoder, have you compared them?",
    "2031702": "mhyodo adversarial validation was surely a nice way to go! Congrats!",
    "2031829": "a little aside, your avatar is so charming on LB😆",
    "2031859": "Thanks for your comment!  I also tried making feature using autoencoder, but it did not work that well within my experiments.",
    "2032062": "I think it is also interesting to fit the vectors for each position directly to an RNN layer rather than taking an average.",
    "2032397": "This is really a nice idea",
    "2032502": ">  cite\n> model\n> catboost\n> target has been reduced to 128 dimensions with svd.\n\nDid you reduce target for CITE to 128 dimentions also?",
    "2032975": "Congratulations for your solo gold, and thank you for your informative solutions. \nI want to ask a question about w2v vector features. Did you input numerical data into w2v and converted them to 16 dimensional vector? And I couldn't surely understand the part \"top100 genes vector average in each cell used as features”. Could you explain more details? \nThank you in advance!",
    "2033067": "Thank you for comment!\nI think it would be faster to have you look at code, so I created a notebook.\nhttps://www.kaggle.com/code/mhyodo/w2v-feature-sample",
    "2033070": "Oops, sorry. cite used the 140 dimensional target as is without dimensional reduction. I will correct this. Thanks for pointing that out!",
    "2033124": "I understood!\nThank you, and again congratulations!",
    "2033924": "Thanks for sharing :)",
    "2034201": "mhyodo  Big Congrat on your solo gold. And thanks a lot for sharing your solution.\nCould you please help me to understand the validation part?\n \nWhy did you start to doubt the LB with a high binary classification result? Does it mean that with a high score, the test data for PB and others have a huge variance, and this variance will lead to the final score shift?",
    "2034343": "Thanks for sharing! As a beginner, this is really helpful for illustrating the importance of validation sets.",
    "2035514": "Congratulations @mhyodo Fantastic job winning 3rd place Cash Gold Solo!",
    "2036995": "Thanks for your question! \nAlso, sorry for the delay in responding.\nSorry, I need to add something. Precisely, I clearly felt that it might be dangerous to trust LB because when I performed the classification with the LB test data as 0 and the PB test data as 1 I felt this was dangerous because of the high accuracy of the classification.\nIf the test data used in LB and the test data used in PB are similar I would believe LB, but since this was not the case, I thought it was dangerous to believe LB.",
    "2037350": "Ah， got it! thanks again for your explanation",
    "2041640": "Congratulations and thanks for posting your great solution Makotu!"
  },
  "source": "meta"
}