{
  "id": 362870,
  "title": "Breaking the 0.810 barrier",
  "url": "/competitions/open-problems-multimodal/discussion/362870",
  "author_name": "KirkDCO",
  "post_date": "2022-10-29T15:19:56.254000",
  "votes": 21,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Last weekend I was able to break into the the top 1/3 of the leaderboard with a score of 0.810.  Looking at the leaderboard today, I see that scores of 0.812 have come to dominate the board, with 322 submissions at 0.812!!  Now, with only a couple of weeks left, I'm contemplating what to try next.</p>\n<p>Some of the things I've tried:</p>\n<ul>\n<li><p>My current best submission uses truncated SVD (seems this is the most common approach) and a Keras Tuner tuned MLP for CITE and Multiome using Negative Log Correlation Loss.  I do 3-fold cross-validation by donor, and combine the predictions from each fold for the test set.</p></li>\n<li><p>I've tried combining more than 1 model per fold with as many as 15 or 50 models per fold - this doesn't improve my score.</p></li>\n<li><p>XGBoost and LightGBM did not improve performance for me - 0.809 score.</p></li>\n<li><p>CV-tuned elastic net only got a score of 0.803.</p></li>\n<li><p>KNN came in with 0.802. </p></li>\n<li><p>I've also tried developing independent models for each cell type, but performance went down compared to general models.</p></li>\n<li><p>I have not tried considering the day as a factor.</p></li>\n</ul>\n<p>All of these are single model approaches, and my guess is that to see improvement on the leaderboard, I need to try stacking or blending.  I haven't tried those approaches yet.  I also have not used any public submissions as additions to my model.</p>\n<p>I'm curious to know what types of techniques have shown promise for others.  Is it really a stacking/blending world out there?  </p>\n<p><strong>UPDATE</strong><br>\nToday (Nov 4th) I see there are 120 submissions at 0.811 and 432 submissions at 0.812.  Seems everyone is moving up the boards into the 0.812 range!  </p>\n<p>Given I have the time, I'm hoping to follow advice from <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> to try adding most correlated features to the truncated SVD.  I've seen this elsewhere, too, and it seems the best option to try next.  I'll update here with the results once I have them…</p>\n<p><strong>UPDATE</strong><br>\nI <em>FINALLY</em> made it into the .811 range.  🙌</p>\n<p>Over the weekend I changed my CITE dataset to use the top 100 most correlated features for each output with no other dimensionality reduction.  This left me at 0.810, but yesterday I combined this dataset with the truncated SVD of the remaining features not in my correlated list.  Same model - MLP tuned with Keras-tuner.  </p>\n<p>I'm still nowhere near where I would like to be, but I did move up the LB by about 100.</p>\n<p>Less than a week left…more work to be done and more approaches to try….</p>\n<p><strong>UPDATE</strong><br>\nI just now made it to the .812 group!  Woot!  I'm testing a couple more things and will update here again later.  Very exciting!</p>\n<p>Since last night, my standing on the LB has improved by about 250 places.  I'm using the same truncated SVD combined with the most correlated variables for CITE, and just truncated SVD for Multiome.  The difference is that I've changed my NN architecture to include skip connections from the inputs and each hidden layer to the output layer.  </p>\n<p>More ensembling to do over the next 22 hours….</p>\n<p>Good luck, everyone!</p>\n<p><strong>FINAL UPDATE</strong><br>\nUsing the data I described above (CITE - tSVD + 100 most correlated + Day, Multi - tSVD + Day) and using NN with skip connections, I was able to move from ~600th place to 278th place by ensembling multiple models from the keras_tuner search results.</p>\n<p>Now we wait for the Private LB and I predict a big shake up.  We'll wee where we all land.</p>\n<p>Thanks to everyone that helped me along the way and best of luck to you all!  I look forward to reading discussions and notebooks of your best (and not so best) submissions.</p>",
  "messages": [
    {
      "id": 2009038,
      "postDate": "2022-10-29T15:19:56.253Z",
      "content": "<p>Last weekend I was able to break into the the top 1/3 of the leaderboard with a score of 0.810.  Looking at the leaderboard today, I see that scores of 0.812 have come to dominate the board, with 322 submissions at 0.812!!  Now, with only a couple of weeks left, I'm contemplating what to try next.</p>\n<p>Some of the things I've tried:</p>\n<ul>\n<li><p>My current best submission uses truncated SVD (seems this is the most common approach) and a Keras Tuner tuned MLP for CITE and Multiome using Negative Log Correlation Loss.  I do 3-fold cross-validation by donor, and combine the predictions from each fold for the test set.</p></li>\n<li><p>I've tried combining more than 1 model per fold with as many as 15 or 50 models per fold - this doesn't improve my score.</p></li>\n<li><p>XGBoost and LightGBM did not improve performance for me - 0.809 score.</p></li>\n<li><p>CV-tuned elastic net only got a score of 0.803.</p></li>\n<li><p>KNN came in with 0.802. </p></li>\n<li><p>I've also tried developing independent models for each cell type, but performance went down compared to general models.</p></li>\n<li><p>I have not tried considering the day as a factor.</p></li>\n</ul>\n<p>All of these are single model approaches, and my guess is that to see improvement on the leaderboard, I need to try stacking or blending.  I haven't tried those approaches yet.  I also have not used any public submissions as additions to my model.</p>\n<p>I'm curious to know what types of techniques have shown promise for others.  Is it really a stacking/blending world out there?  </p>\n<p><strong>UPDATE</strong><br>\nToday (Nov 4th) I see there are 120 submissions at 0.811 and 432 submissions at 0.812.  Seems everyone is moving up the boards into the 0.812 range!  </p>\n<p>Given I have the time, I'm hoping to follow advice from <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> to try adding most correlated features to the truncated SVD.  I've seen this elsewhere, too, and it seems the best option to try next.  I'll update here with the results once I have them…</p>\n<p><strong>UPDATE</strong><br>\nI <em>FINALLY</em> made it into the .811 range.  🙌</p>\n<p>Over the weekend I changed my CITE dataset to use the top 100 most correlated features for each output with no other dimensionality reduction.  This left me at 0.810, but yesterday I combined this dataset with the truncated SVD of the remaining features not in my correlated list.  Same model - MLP tuned with Keras-tuner.  </p>\n<p>I'm still nowhere near where I would like to be, but I did move up the LB by about 100.</p>\n<p>Less than a week left…more work to be done and more approaches to try….</p>\n<p><strong>UPDATE</strong><br>\nI just now made it to the .812 group!  Woot!  I'm testing a couple more things and will update here again later.  Very exciting!</p>\n<p>Since last night, my standing on the LB has improved by about 250 places.  I'm using the same truncated SVD combined with the most correlated variables for CITE, and just truncated SVD for Multiome.  The difference is that I've changed my NN architecture to include skip connections from the inputs and each hidden layer to the output layer.  </p>\n<p>More ensembling to do over the next 22 hours….</p>\n<p>Good luck, everyone!</p>\n<p><strong>FINAL UPDATE</strong><br>\nUsing the data I described above (CITE - tSVD + 100 most correlated + Day, Multi - tSVD + Day) and using NN with skip connections, I was able to move from ~600th place to 278th place by ensembling multiple models from the keras_tuner search results.</p>\n<p>Now we wait for the Private LB and I predict a big shake up.  We'll wee where we all land.</p>\n<p>Thanks to everyone that helped me along the way and best of luck to you all!  I look forward to reading discussions and notebooks of your best (and not so best) submissions.</p>",
      "rawMarkdown": "Last weekend I was able to break into the the top 1/3 of the leaderboard with a score of 0.810.  Looking at the leaderboard today, I see that scores of 0.812 have come to dominate the board, with 322 submissions at 0.812!!  Now, with only a couple of weeks left, I'm contemplating what to try next.\n\nSome of the things I've tried:\n\n- My current best submission uses truncated SVD (seems this is the most common approach) and a Keras Tuner tuned MLP for CITE and Multiome using Negative Log Correlation Loss.  I do 3-fold cross-validation by donor, and combine the predictions from each fold for the test set.\n\n- I've tried combining more than 1 model per fold with as many as 15 or 50 models per fold - this doesn't improve my score.\n\n- XGBoost and LightGBM did not improve performance for me - 0.809 score.\n\n- CV-tuned elastic net only got a score of 0.803.\n\n- KNN came in with 0.802. \n\n- I've also tried developing independent models for each cell type, but performance went down compared to general models.\n\n- I have not tried considering the day as a factor.\n\nAll of these are single model approaches, and my guess is that to see improvement on the leaderboard, I need to try stacking or blending.  I haven't tried those approaches yet.  I also have not used any public submissions as additions to my model.\n\nI'm curious to know what types of techniques have shown promise for others.  Is it really a stacking/blending world out there?  \n\n**UPDATE**\nToday (Nov 4th) I see there are 120 submissions at 0.811 and 432 submissions at 0.812.  Seems everyone is moving up the boards into the 0.812 range!  \n\nGiven I have the time, I'm hoping to follow advice from @ahmedelfazouan to try adding most correlated features to the truncated SVD.  I've seen this elsewhere, too, and it seems the best option to try next.  I'll update here with the results once I have them...\n\n**UPDATE**\nI _FINALLY_ made it into the .811 range.  🙌\n\nOver the weekend I changed my CITE dataset to use the top 100 most correlated features for each output with no other dimensionality reduction.  This left me at 0.810, but yesterday I combined this dataset with the truncated SVD of the remaining features not in my correlated list.  Same model - MLP tuned with Keras-tuner.  \n\nI'm still nowhere near where I would like to be, but I did move up the LB by about 100.\n\nLess than a week left...more work to be done and more approaches to try....\n\n**UPDATE**\nI just now made it to the .812 group!  Woot!  I'm testing a couple more things and will update here again later.  Very exciting!\n\nSince last night, my standing on the LB has improved by about 250 places.  I'm using the same truncated SVD combined with the most correlated variables for CITE, and just truncated SVD for Multiome.  The difference is that I've changed my NN architecture to include skip connections from the inputs and each hidden layer to the output layer.  \n\nMore ensembling to do over the next 22 hours....\n\nGood luck, everyone!\n\n**FINAL UPDATE**\nUsing the data I described above (CITE - tSVD + 100 most correlated + Day, Multi - tSVD + Day) and using NN with skip connections, I was able to move from ~600th place to 278th place by ensembling multiple models from the keras_tuner search results.\n\nNow we wait for the Private LB and I predict a big shake up.  We'll wee where we all land.\n\nThanks to everyone that helped me along the way and best of luck to you all!  I look forward to reading discussions and notebooks of your best (and not so best) submissions.\n",
      "votes": 19
    },
    {
      "id": 2010237,
      "postDate": "2022-10-30T16:02:57.997Z",
      "content": "<p>No it's not all about blending, one method that worked well is training the model using only columns that have the highest correlations with the target (for each target)</p>",
      "rawMarkdown": "No it's not all about blending, one method that worked well is training the model using only columns that have the highest correlations with the target (for each target)",
      "votes": 2,
      "replies": [
        {
          "id": 2010339,
          "postDate": "2022-10-30T17:47:32.107Z",
          "content": "<p>Yes, I saw a post of that idea but hadn't seen an actual submission.  I may have to give that a try.</p>",
          "rawMarkdown": "Yes, I saw a post of that idea but hadn't seen an actual submission.  I may have to give that a try.",
          "votes": 1
        },
        {
          "id": 2010416,
          "postDate": "2022-10-30T18:27:15.657Z",
          "content": "<p>I'm using it for it for citeseq, it is working so well, but it didn't work for multiome.</p>",
          "rawMarkdown": "I'm using it for it for citeseq, it is working so well, but it didn't work for multiome.",
          "votes": 1
        },
        {
          "id": 2010439,
          "postDate": "2022-10-30T18:51:23.047Z",
          "content": "<p>Do you develop a unique model for each output measurement, or did you pre-group the outputs into correlated groups?  I can imagine there are a lot of ways to go about the details of that.</p>",
          "rawMarkdown": "Do you develop a unique model for each output measurement, or did you pre-group the outputs into correlated groups?  I can imagine there are a lot of ways to go about the details of that.",
          "votes": 1
        },
        {
          "id": 2010466,
          "postDate": "2022-10-30T19:07:27.377Z",
          "content": "<p>One model for each target (so in total : 140 models for citeseq), trained with the most 800 correlated features, for correlation I'm taking the absolute value, because I think negative correlations could also have some information for the target.</p>",
          "rawMarkdown": "One model for each target (so in total : 140 models for citeseq), trained with the most 800 correlated features, for correlation I'm taking the absolute value, because I think negative correlations could also have some information for the target.",
          "votes": 4
        },
        {
          "id": 2010519,
          "postDate": "2022-10-30T20:04:47.267Z",
          "content": "<p>Interesting!  I may give that a try, if I have time.</p>",
          "rawMarkdown": "Interesting!  I may give that a try, if I have time.",
          "votes": 1
        },
        {
          "id": 2017255,
          "postDate": "2022-11-04T16:36:15.343Z",
          "content": "<p>May I ask which model you use in this setting, I've tried this before but the result is not good as the correlation loss + mlp</p>",
          "rawMarkdown": "May I ask which model you use in this setting, I've tried this before but the result is not good as the correlation loss + mlp"
        },
        {
          "id": 2018656,
          "postDate": "2022-11-05T21:12:09.877Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2018661,
          "postDate": "2022-11-05T21:17:30.480Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2018662,
          "postDate": "2022-11-05T21:22:27.483Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2018663,
          "postDate": "2022-11-05T21:24:37.530Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 2018668,
          "postDate": "2022-11-05T21:39:05.237Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 2019822,
          "postDate": "2022-11-07T02:24:19.100Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2032083,
          "postDate": "2022-11-16T12:20:04.610Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2016084,
      "postDate": "2022-11-03T19:06:03.083Z",
      "content": "<p>using some feature selection methods (e.g., <a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-021-02544-3\" target=\"_blank\">https://genomebiology.biomedcentral.com/articles/10.1186/s13059-021-02544-3</a>) might be useful as  it looks like major issues is with the amount of noise these reading carry.</p>",
      "rawMarkdown": "using some feature selection methods (e.g., https://genomebiology.biomedcentral.com/articles/10.1186/s13059-021-02544-3) might be useful as  it looks like major issues is with the amount of noise these reading carry.",
      "votes": 1,
      "replies": [
        {
          "id": 2016149,
          "postDate": "2022-11-03T20:11:05.273Z",
          "content": "<p>I was hopeful that truncSVD was doing some portion of the feature selection by reducing the dimensionality down to the maximum variance dimensions.  I have seen others add the 100 most important or most correlated genes to the truncSVD dimensions as well.  <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> mentioned the 800 most correlated features in his answers to my question.  Is there a notebook or dataset that actually lists those important features, rather than recoding the whole thing myself?  I haven't been able to find them if they exist.</p>\n<p>I've also used XGBoost and would have assumed decision trees natural feature selection process would be enough, but all my work with XGBoost has led to lower CVs and LB standings.  </p>",
          "rawMarkdown": "I was hopeful that truncSVD was doing some portion of the feature selection by reducing the dimensionality down to the maximum variance dimensions.  I have seen others add the 100 most important or most correlated genes to the truncSVD dimensions as well.  @ahmedelfazouan mentioned the 800 most correlated features in his answers to my question.  Is there a notebook or dataset that actually lists those important features, rather than recoding the whole thing myself?  I haven't been able to find them if they exist.\n\nI've also used XGBoost and would have assumed decision trees natural feature selection process would be enough, but all my work with XGBoost has led to lower CVs and LB standings.  ",
          "votes": 1
        },
        {
          "id": 2020620,
          "postDate": "2022-11-07T16:37:50.130Z",
          "content": "<p>Something like this?<br>\n<a href=\"https://www.kaggle.com/datasets/karlopintaric/citeseq-columns-ordered-by-correlation-with-target?select=orderedCols_TargetCorr.csv\" target=\"_blank\">https://www.kaggle.com/datasets/karlopintaric/citeseq-columns-ordered-by-correlation-with-target?select=orderedCols_TargetCorr.csv</a></p>",
          "rawMarkdown": "Something like this?\nhttps://www.kaggle.com/datasets/karlopintaric/citeseq-columns-ordered-by-correlation-with-target?select=orderedCols_TargetCorr.csv",
          "votes": 2
        },
        {
          "id": 2020813,
          "postDate": "2022-11-07T19:50:36.573Z",
          "content": "<p>That looks perfect!  I also found this:  <a href=\"https://www.kaggle.com/code/kaggledummie007/cite-targets-vs-features-correlation\" target=\"_blank\">https://www.kaggle.com/code/kaggledummie007/cite-targets-vs-features-correlation</a></p>",
          "rawMarkdown": "That looks perfect!  I also found this:  https://www.kaggle.com/code/kaggledummie007/cite-targets-vs-features-correlation",
          "votes": 2
        }
      ]
    },
    {
      "id": 2029115,
      "postDate": "2022-11-14T13:11:29.787Z",
      "content": "<p>Thanks for sharing. It was very helpful.</p>",
      "rawMarkdown": "Thanks for sharing. It was very helpful."
    },
    {
      "id": 2017929,
      "postDate": "2022-11-05T08:17:05.610Z",
      "content": "<p>You can get to 0.811 with a simple TabNet for both Citeseq and Multiome. </p>",
      "rawMarkdown": "You can get to 0.811 with a simple TabNet for both Citeseq and Multiome. ",
      "replies": [
        {
          "id": 2019494,
          "postDate": "2022-11-06T16:34:37.647Z",
          "content": "<p>I haven't tried TabNet.  My skills aren't quite up to implementing on my own, but I do see a few versions on Github and one on PyPi.  Do you have a preferred TabNet package?</p>",
          "rawMarkdown": "I haven't tried TabNet.  My skills aren't quite up to implementing on my own, but I do see a few versions on Github and one on PyPi.  Do you have a preferred TabNet package?"
        },
        {
          "id": 2020322,
          "postDate": "2022-11-07T11:01:12.807Z",
          "content": "<p><a href=\"https://www.kaggle.com/kirkdco\" target=\"_blank\">@kirkdco</a> I think pytorch-tabnet (<a href=\"https://github.com/dreamquark-ai/tabnet\" target=\"_blank\">https://github.com/dreamquark-ai/tabnet</a>) should be what you are looking for, very easy to use and install (pip) and probably the most used implementation of tabnet. I don't know about the competition's specificity though. </p>",
          "rawMarkdown": "@kirkdco I think pytorch-tabnet (https://github.com/dreamquark-ai/tabnet) should be what you are looking for, very easy to use and install (pip) and probably the most used implementation of tabnet. I don't know about the competition's specificity though. ",
          "votes": 1
        },
        {
          "id": 2021184,
          "postDate": "2022-11-08T04:28:02.867Z",
          "content": "<p>I will give that a try.  Thanks, <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> !</p>",
          "rawMarkdown": "I will give that a try.  Thanks, @optimo !"
        }
      ]
    },
    {
      "id": 2029078,
      "postDate": "2022-11-14T12:22:03.297Z",
      "content": "<p>Wonderful thank!!</p>",
      "rawMarkdown": "Wonderful thank!!"
    },
    {
      "id": 2028497,
      "postDate": "2022-11-14T02:32:04.913Z",
      "content": "<p>Thanks for sharing, great!</p>",
      "rawMarkdown": "Thanks for sharing, great!"
    }
  ],
  "comments": [
    {
      "id": 2010237,
      "author_name": "Ahmed El Fazouani",
      "author_url": "",
      "post_date": "2022-10-30T16:02:57.997000",
      "content": "<p>No it's not all about blending, one method that worked well is training the model using only columns that have the highest correlations with the target (for each target)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2010339,
          "author_name": "KirkDCO",
          "author_url": "",
          "post_date": "2022-10-30T17:47:32.107000",
          "content": "<p>Yes, I saw a post of that idea but hadn't seen an actual submission.  I may have to give that a try.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2010416,
          "author_name": "Ahmed El Fazouani",
          "author_url": "",
          "post_date": "2022-10-30T18:27:15.657000",
          "content": "<p>I'm using it for it for citeseq, it is working so well, but it didn't work for multiome.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2010439,
          "author_name": "KirkDCO",
          "author_url": "",
          "post_date": "2022-10-30T18:51:23.047000",
          "content": "<p>Do you develop a unique model for each output measurement, or did you pre-group the outputs into correlated groups?  I can imagine there are a lot of ways to go about the details of that.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2010466,
          "author_name": "Ahmed El Fazouani",
          "author_url": "",
          "post_date": "2022-10-30T19:07:27.377000",
          "content": "<p>One model for each target (so in total : 140 models for citeseq), trained with the most 800 correlated features, for correlation I'm taking the absolute value, because I think negative correlations could also have some information for the target.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2010519,
          "author_name": "KirkDCO",
          "author_url": "",
          "post_date": "2022-10-30T20:04:47.267000",
          "content": "<p>Interesting!  I may give that a try, if I have time.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2017255,
          "author_name": "no-magic",
          "author_url": "",
          "post_date": "2022-11-04T16:36:15.343000",
          "content": "<p>May I ask which model you use in this setting, I've tried this before but the result is not good as the correlation loss + mlp</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2018656,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-11-05T21:12:09.877000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2018661,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-11-05T21:17:30.480000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2018662,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-11-05T21:22:27.483000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2018663,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-11-05T21:24:37.530000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2018668,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-11-05T21:39:05.237000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2019822,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-11-07T02:24:19.100000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2032083,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-11-16T12:20:04.610000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2016084,
      "author_name": "Sandeep Kumar",
      "author_url": "",
      "post_date": "2022-11-03T19:06:03.083000",
      "content": "<p>using some feature selection methods (e.g., <a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-021-02544-3\" target=\"_blank\">https://genomebiology.biomedcentral.com/articles/10.1186/s13059-021-02544-3</a>) might be useful as  it looks like major issues is with the amount of noise these reading carry.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2016149,
          "author_name": "KirkDCO",
          "author_url": "",
          "post_date": "2022-11-03T20:11:05.273000",
          "content": "<p>I was hopeful that truncSVD was doing some portion of the feature selection by reducing the dimensionality down to the maximum variance dimensions.  I have seen others add the 100 most important or most correlated genes to the truncSVD dimensions as well.  <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> mentioned the 800 most correlated features in his answers to my question.  Is there a notebook or dataset that actually lists those important features, rather than recoding the whole thing myself?  I haven't been able to find them if they exist.</p>\n<p>I've also used XGBoost and would have assumed decision trees natural feature selection process would be enough, but all my work with XGBoost has led to lower CVs and LB standings.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2020620,
          "author_name": "Karlo Pintaric",
          "author_url": "",
          "post_date": "2022-11-07T16:37:50.130000",
          "content": "<p>Something like this?<br>\n<a href=\"https://www.kaggle.com/datasets/karlopintaric/citeseq-columns-ordered-by-correlation-with-target?select=orderedCols_TargetCorr.csv\" target=\"_blank\">https://www.kaggle.com/datasets/karlopintaric/citeseq-columns-ordered-by-correlation-with-target?select=orderedCols_TargetCorr.csv</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2020813,
          "author_name": "KirkDCO",
          "author_url": "",
          "post_date": "2022-11-07T19:50:36.573000",
          "content": "<p>That looks perfect!  I also found this:  <a href=\"https://www.kaggle.com/code/kaggledummie007/cite-targets-vs-features-correlation\" target=\"_blank\">https://www.kaggle.com/code/kaggledummie007/cite-targets-vs-features-correlation</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2029115,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-14T13:11:29.787000",
      "content": "<p>Thanks for sharing. It was very helpful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2017929,
      "author_name": "tarick.morty",
      "author_url": "",
      "post_date": "2022-11-05T08:17:05.610000",
      "content": "<p>You can get to 0.811 with a simple TabNet for both Citeseq and Multiome. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2019494,
          "author_name": "KirkDCO",
          "author_url": "",
          "post_date": "2022-11-06T16:34:37.647000",
          "content": "<p>I haven't tried TabNet.  My skills aren't quite up to implementing on my own, but I do see a few versions on Github and one on PyPi.  Do you have a preferred TabNet package?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2020322,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2022-11-07T11:01:12.807000",
          "content": "<p><a href=\"https://www.kaggle.com/kirkdco\" target=\"_blank\">@kirkdco</a> I think pytorch-tabnet (<a href=\"https://github.com/dreamquark-ai/tabnet\" target=\"_blank\">https://github.com/dreamquark-ai/tabnet</a>) should be what you are looking for, very easy to use and install (pip) and probably the most used implementation of tabnet. I don't know about the competition's specificity though. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2021184,
          "author_name": "KirkDCO",
          "author_url": "",
          "post_date": "2022-11-08T04:28:02.867000",
          "content": "<p>I will give that a try.  Thanks, <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2029078,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-14T12:22:03.297000",
      "content": "<p>Wonderful thank!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2028497,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-14T02:32:04.913000",
      "content": "<p>Thanks for sharing, great!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2009038": "Last weekend I was able to break into the the top 1/3 of the leaderboard with a score of 0.810.  Looking at the leaderboard today, I see that scores of 0.812 have come to dominate the board, with 322 submissions at 0.812!!  Now, with only a couple of weeks left, I'm contemplating what to try next.\n\nSome of the things I've tried:\n\n- My current best submission uses truncated SVD (seems this is the most common approach) and a Keras Tuner tuned MLP for CITE and Multiome using Negative Log Correlation Loss.  I do 3-fold cross-validation by donor, and combine the predictions from each fold for the test set.\n\n- I've tried combining more than 1 model per fold with as many as 15 or 50 models per fold - this doesn't improve my score.\n\n- XGBoost and LightGBM did not improve performance for me - 0.809 score.\n\n- CV-tuned elastic net only got a score of 0.803.\n\n- KNN came in with 0.802. \n\n- I've also tried developing independent models for each cell type, but performance went down compared to general models.\n\n- I have not tried considering the day as a factor.\n\nAll of these are single model approaches, and my guess is that to see improvement on the leaderboard, I need to try stacking or blending.  I haven't tried those approaches yet.  I also have not used any public submissions as additions to my model.\n\nI'm curious to know what types of techniques have shown promise for others.  Is it really a stacking/blending world out there?  \n\n**UPDATE**\nToday (Nov 4th) I see there are 120 submissions at 0.811 and 432 submissions at 0.812.  Seems everyone is moving up the boards into the 0.812 range!  \n\nGiven I have the time, I'm hoping to follow advice from @ahmedelfazouan to try adding most correlated features to the truncated SVD.  I've seen this elsewhere, too, and it seems the best option to try next.  I'll update here with the results once I have them...\n\n**UPDATE**\nI _FINALLY_ made it into the .811 range.  🙌\n\nOver the weekend I changed my CITE dataset to use the top 100 most correlated features for each output with no other dimensionality reduction.  This left me at 0.810, but yesterday I combined this dataset with the truncated SVD of the remaining features not in my correlated list.  Same model - MLP tuned with Keras-tuner.  \n\nI'm still nowhere near where I would like to be, but I did move up the LB by about 100.\n\nLess than a week left...more work to be done and more approaches to try....\n\n**UPDATE**\nI just now made it to the .812 group!  Woot!  I'm testing a couple more things and will update here again later.  Very exciting!\n\nSince last night, my standing on the LB has improved by about 250 places.  I'm using the same truncated SVD combined with the most correlated variables for CITE, and just truncated SVD for Multiome.  The difference is that I've changed my NN architecture to include skip connections from the inputs and each hidden layer to the output layer.  \n\nMore ensembling to do over the next 22 hours....\n\nGood luck, everyone!\n\n**FINAL UPDATE**\nUsing the data I described above (CITE - tSVD + 100 most correlated + Day, Multi - tSVD + Day) and using NN with skip connections, I was able to move from ~600th place to 278th place by ensembling multiple models from the keras_tuner search results.\n\nNow we wait for the Private LB and I predict a big shake up.  We'll wee where we all land.\n\nThanks to everyone that helped me along the way and best of luck to you all!  I look forward to reading discussions and notebooks of your best (and not so best) submissions.\n",
    "2010237": "No it's not all about blending, one method that worked well is training the model using only columns that have the highest correlations with the target (for each target)",
    "2016084": "using some feature selection methods (e.g., https://genomebiology.biomedcentral.com/articles/10.1186/s13059-021-02544-3) might be useful as  it looks like major issues is with the amount of noise these reading carry.",
    "2029115": "Thanks for sharing. It was very helpful.",
    "2017929": "You can get to 0.811 with a simple TabNet for both Citeseq and Multiome. ",
    "2029078": "Wonderful thank!!",
    "2028497": "Thanks for sharing, great!"
  }
}