{
  "id": 366504,
  "title": "13th place and how to",
  "url": "/competitions/open-problems-multimodal/writeups/q-13th-place-and-how-to",
  "author_name": "",
  "post_date": "2022-11-16T13:12:15.840Z",
  "votes": 46,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hey guys, it's been a while since my last actively participated competition and this is a good experience! Though I am not active any more, I frequently come back to kaggle to browse new ideas. One thing I notice is that it is not really straightforward to learn new things by just reading top solutions if you were not active in that competition… And one factor could be that we mostly share what finally worked (and/or what didn't work), but not the thought process of getting there. For someone like me, either a bit too lazy or a bit too busy to actively participate, or someone who might be a bit inexperienced, I think it would be more beneficial to share \"how I got here\" more than \"here I am\". So in that spirit, I would like to start sharing solutions this way…</p>\n<h5>Some background:</h5>\n<p>My job can be demanding so whenever I can delegate the work to computers, I do, whenever I cannot, I minimize the time needed for coding such that I can leverage fragamented time as much as possible. And I mainly participate in this competition to learn how feature extraction could work with high dimension inputs and outputs. And they shape what next experiments I decided to try, and mistakes I made along the way so I thought it is important to share.</p>\n<h5>My journey:</h5>\n<ol>\n<li><p>started by reading the Discussion to understand what the data is about, and walk through popular public kernels to understand what can be used as baselines. I noticed that mostly TruncatedSVD was used to reduce dimensions. </p></li>\n<li><p>I figured it might be a good idea to just train a mlp model with all features included. And I was too lazy to build a local validation pipeline so I just randomly sampled 10% as validation set. And the result wasn't so good on public LB compared to public kernels.</p></li>\n<li><p>Then I was thinking, ok, maybe it was because mlp was bad. But xgboost/lightgbm on cpu would take forever to train, and would run out of memory on gpu. So need a better neural network model. What about Tabnet (terrible). Ok. There was one paper I remember claiming similar performance to xgboost. OK. found it. <a href=\"https://arxiv.org/pdf/2112.02962.pdf\" target=\"_blank\">https://arxiv.org/pdf/2112.02962.pdf</a> (It is called DANets). Better but still not as good…</p></li>\n<li><p>Now back to public kernel as the baseline. Maybe instead of SVD, we can use autoencoder? OK, only linear autoencoder  performed ok-ish, any nonlinearity didn't work… no matter what tricks (e.g. swap noise augmentation) used. Hmmm…</p></li>\n<li><p>Back to TSVD + MLP as baseline again… Let's make this baseline better first. First swap MLP with DANets. And let's just focus on cite since it only have high dim inputs. Whatever works for cite should work for multi right? (TimeMachine: Nope!)</p></li>\n<li><p>Let's standardize the data since PCA likes it. OK. Slightly better. The explained variance seems quite low, but adding more components as features doesn't seem too helpful. So likely the inputs are quite noisy. Don't really know what to do. Well, we can always add different decomposition methods if no better ideas. Added NMF, FactorAnalysis, FastICA. Ok. All of them worked. </p></li>\n<li><p>Now let's also train some xgboost model since now we can train on gpu. Ok cool. Averaging xgboost with DANets improves results significantly.</p></li>\n<li><p>Maybe should try autoencoder again…. Read some papers. Ok. Still didn't work.</p></li>\n<li><p>Let's try some popular nonlinear dimension reduction techniques. UMAP, TriMAP, PaCMAP. Doesn't seem working.</p></li>\n<li><p>XGBoost or LightGBM train one model per target which seems wasteful and may not consider the correlation between outputs. Let's see if there is a better way out there. <a href=\"https://arxiv.org/abs/1909.04373\" target=\"_blank\">https://arxiv.org/abs/1909.04373</a> found this GBDT-MO. and it has code. Tried. Didn't work so well. </p></li>\n<li><p>Read in the Discussion that we can select features by matching input/output names for Cite. Tried and it worked. Neat.</p></li>\n<li><p>Also read in the Discussion that the 0s in inputs may not be actual 0s could also be missing. OK. Tried to calculate the mean/std by considering all 0s as missing and then calculate PCA. Adding to the features improved the model a bit.</p></li>\n<li><p>Realized that if 0s can be treated as missing then we can calculate PCA differently too according to this old paper (<a href=\"https://www.sciencedirect.com/science/article/abs/pii/S016974399600007X)\" target=\"_blank\">https://www.sciencedirect.com/science/article/abs/pii/S016974399600007X)</a>. Helped a bit.</p></li>\n<li><p>OK… stop the laziness and spend the weekend on building proper cross validation pipeline and retrain the models for cite. Nice. a huge jump from 0.812 to 0.814 on leaderboard.</p></li>\n<li><p>Saw some discussion about the supplymentary raw data. Well, I believe the host has done the best preprocessing and probably not helpful, also it is a pain to handle two datasets, so ignored. (TimeMachine: Big Mistake!)</p></li>\n<li><p>Approaching the last week of competition so rewire the cite pipeline for multi. Hoping to see a huge jump on score for multi as seen on cite. But that didn't happen. So… They are actually very different… Maybe I should have accepted some team merging invite earlier….</p></li>\n<li><p>stack a few DANets and GBDT models. </p></li>\n<li><p>picked the wrong submissions for final evaluation. But what can you do.</p></li>\n</ol>\n<h5>Model Summary</h5>\n<p>Cite:<br>\n(PCA + NMF + FA + ICA + NanPCA + Missing Value PCA (i.e. NIPALS) ) + (DANets + XGBoost)</p>\n<p>Multi:<br>\n(PCA + NMF) + (DANets + XGBoost with SVDed Output)</p>\n<h5>Looking Back, to improve the score further</h5>\n<ol>\n<li>Should have spent more time on data preprocessing.</li>\n<li>Should anticipate and prepare for a complex ensemble pipeline so save all the models and predictions properly along the way for multi layer stacking.</li>\n<li>Shouldn't have assumed Cite and Multi being similar and go for team up.</li>\n</ol>\n<p>Hopefully it is helpful (also to those who didn't participate!)</p>",
  "messages": [
    {
      "id": "2032146",
      "postDate": "11/16/2022 12:59:20",
      "content": "<p>Hey guys, it's been a while since my last actively participated competition and this is a good experience! Though I am not active any more, I frequently come back to kaggle to browse new ideas. One thing I notice is that it is not really straightforward to learn new things by just reading top solutions if you were not active in that competition… And one factor could be that we mostly share what finally worked (and/or what didn't work), but not the thought process of getting there. For someone like me, either a bit too lazy or a bit too busy to actively participate, or someone who might be a bit inexperienced, I think it would be more beneficial to share \"how I got here\" more than \"here I am\". So in that spirit, I would like to start sharing solutions this way…</p>\n<h5>Some background:</h5>\n<p>My job can be demanding so whenever I can delegate the work to computers, I do, whenever I cannot, I minimize the time needed for coding such that I can leverage fragamented time as much as possible. And I mainly participate in this competition to learn how feature extraction could work with high dimension inputs and outputs. And they shape what next experiments I decided to try, and mistakes I made along the way so I thought it is important to share.</p>\n<h5>My journey:</h5>\n<ol>\n<li><p>started by reading the Discussion to understand what the data is about, and walk through popular public kernels to understand what can be used as baselines. I noticed that mostly TruncatedSVD was used to reduce dimensions. </p></li>\n<li><p>I figured it might be a good idea to just train a mlp model with all features included. And I was too lazy to build a local validation pipeline so I just randomly sampled 10% as validation set. And the result wasn't so good on public LB compared to public kernels.</p></li>\n<li><p>Then I was thinking, ok, maybe it was because mlp was bad. But xgboost/lightgbm on cpu would take forever to train, and would run out of memory on gpu. So need a better neural network model. What about Tabnet (terrible). Ok. There was one paper I remember claiming similar performance to xgboost. OK. found it. <a href=\"https://arxiv.org/pdf/2112.02962.pdf\" target=\"_blank\">https://arxiv.org/pdf/2112.02962.pdf</a> (It is called DANets). Better but still not as good…</p></li>\n<li><p>Now back to public kernel as the baseline. Maybe instead of SVD, we can use autoencoder? OK, only linear autoencoder  performed ok-ish, any nonlinearity didn't work… no matter what tricks (e.g. swap noise augmentation) used. Hmmm…</p></li>\n<li><p>Back to TSVD + MLP as baseline again… Let's make this baseline better first. First swap MLP with DANets. And let's just focus on cite since it only have high dim inputs. Whatever works for cite should work for multi right? (TimeMachine: Nope!)</p></li>\n<li><p>Let's standardize the data since PCA likes it. OK. Slightly better. The explained variance seems quite low, but adding more components as features doesn't seem too helpful. So likely the inputs are quite noisy. Don't really know what to do. Well, we can always add different decomposition methods if no better ideas. Added NMF, FactorAnalysis, FastICA. Ok. All of them worked. </p></li>\n<li><p>Now let's also train some xgboost model since now we can train on gpu. Ok cool. Averaging xgboost with DANets improves results significantly.</p></li>\n<li><p>Maybe should try autoencoder again…. Read some papers. Ok. Still didn't work.</p></li>\n<li><p>Let's try some popular nonlinear dimension reduction techniques. UMAP, TriMAP, PaCMAP. Doesn't seem working.</p></li>\n<li><p>XGBoost or LightGBM train one model per target which seems wasteful and may not consider the correlation between outputs. Let's see if there is a better way out there. <a href=\"https://arxiv.org/abs/1909.04373\" target=\"_blank\">https://arxiv.org/abs/1909.04373</a> found this GBDT-MO. and it has code. Tried. Didn't work so well. </p></li>\n<li><p>Read in the Discussion that we can select features by matching input/output names for Cite. Tried and it worked. Neat.</p></li>\n<li><p>Also read in the Discussion that the 0s in inputs may not be actual 0s could also be missing. OK. Tried to calculate the mean/std by considering all 0s as missing and then calculate PCA. Adding to the features improved the model a bit.</p></li>\n<li><p>Realized that if 0s can be treated as missing then we can calculate PCA differently too according to this old paper (<a href=\"https://www.sciencedirect.com/science/article/abs/pii/S016974399600007X)\" target=\"_blank\">https://www.sciencedirect.com/science/article/abs/pii/S016974399600007X)</a>. Helped a bit.</p></li>\n<li><p>OK… stop the laziness and spend the weekend on building proper cross validation pipeline and retrain the models for cite. Nice. a huge jump from 0.812 to 0.814 on leaderboard.</p></li>\n<li><p>Saw some discussion about the supplymentary raw data. Well, I believe the host has done the best preprocessing and probably not helpful, also it is a pain to handle two datasets, so ignored. (TimeMachine: Big Mistake!)</p></li>\n<li><p>Approaching the last week of competition so rewire the cite pipeline for multi. Hoping to see a huge jump on score for multi as seen on cite. But that didn't happen. So… They are actually very different… Maybe I should have accepted some team merging invite earlier….</p></li>\n<li><p>stack a few DANets and GBDT models. </p></li>\n<li><p>picked the wrong submissions for final evaluation. But what can you do.</p></li>\n</ol>\n<h5>Model Summary</h5>\n<p>Cite:<br>\n(PCA + NMF + FA + ICA + NanPCA + Missing Value PCA (i.e. NIPALS) ) + (DANets + XGBoost)</p>\n<p>Multi:<br>\n(PCA + NMF) + (DANets + XGBoost with SVDed Output)</p>\n<h5>Looking Back, to improve the score further</h5>\n<ol>\n<li>Should have spent more time on data preprocessing.</li>\n<li>Should anticipate and prepare for a complex ensemble pipeline so save all the models and predictions properly along the way for multi layer stacking.</li>\n<li>Shouldn't have assumed Cite and Multi being similar and go for team up.</li>\n</ol>\n<p>Hopefully it is helpful (also to those who didn't participate!)</p>",
      "rawMarkdown": "Hey guys, it's been a while since my last actively participated competition and this is a good experience! Though I am not active any more, I frequently come back to kaggle to browse new ideas. One thing I notice is that it is not really straightforward to learn new things by just reading top solutions if you were not active in that competition... And one factor could be that we mostly share what finally worked (and/or what didn't work), but not the thought process of getting there. For someone like me, either a bit too lazy or a bit too busy to actively participate, or someone who might be a bit inexperienced, I think it would be more beneficial to share \"how I got here\" more than \"here I am\". So in that spirit, I would like to start sharing solutions this way...\n\n\n##### Some background: \nMy job can be demanding so whenever I can delegate the work to computers, I do, whenever I cannot, I minimize the time needed for coding such that I can leverage fragamented time as much as possible. And I mainly participate in this competition to learn how feature extraction could work with high dimension inputs and outputs. And they shape what next experiments I decided to try, and mistakes I made along the way so I thought it is important to share.\n\n\n##### My journey:\n\n1. started by reading the Discussion to understand what the data is about, and walk through popular public kernels to understand what can be used as baselines. I noticed that mostly TruncatedSVD was used to reduce dimensions. \n\n2. I figured it might be a good idea to just train a mlp model with all features included. And I was too lazy to build a local validation pipeline so I just randomly sampled 10% as validation set. And the result wasn't so good on public LB compared to public kernels.\n\n3. Then I was thinking, ok, maybe it was because mlp was bad. But xgboost/lightgbm on cpu would take forever to train, and would run out of memory on gpu. So need a better neural network model. What about Tabnet (terrible). Ok. There was one paper I remember claiming similar performance to xgboost. OK. found it. https://arxiv.org/pdf/2112.02962.pdf (It is called DANets). Better but still not as good...\n\n4. Now back to public kernel as the baseline. Maybe instead of SVD, we can use autoencoder? OK, only linear autoencoder  performed ok-ish, any nonlinearity didn't work... no matter what tricks (e.g. swap noise augmentation) used. Hmmm...\n\n5. Back to TSVD + MLP as baseline again... Let's make this baseline better first. First swap MLP with DANets. And let's just focus on cite since it only have high dim inputs. Whatever works for cite should work for multi right? (TimeMachine: Nope!)\n\n6. Let's standardize the data since PCA likes it. OK. Slightly better. The explained variance seems quite low, but adding more components as features doesn't seem too helpful. So likely the inputs are quite noisy. Don't really know what to do. Well, we can always add different decomposition methods if no better ideas. Added NMF, FactorAnalysis, FastICA. Ok. All of them worked. \n\n7. Now let's also train some xgboost model since now we can train on gpu. Ok cool. Averaging xgboost with DANets improves results significantly.\n\n8. Maybe should try autoencoder again.... Read some papers. Ok. Still didn't work.\n\n9. Let's try some popular nonlinear dimension reduction techniques. UMAP, TriMAP, PaCMAP. Doesn't seem working.\n\n10. XGBoost or LightGBM train one model per target which seems wasteful and may not consider the correlation between outputs. Let's see if there is a better way out there. https://arxiv.org/abs/1909.04373 found this GBDT-MO. and it has code. Tried. Didn't work so well. \n\n11. Read in the Discussion that we can select features by matching input/output names for Cite. Tried and it worked. Neat.\n\n12. Also read in the Discussion that the 0s in inputs may not be actual 0s could also be missing. OK. Tried to calculate the mean/std by considering all 0s as missing and then calculate PCA. Adding to the features improved the model a bit.\n\n13. Realized that if 0s can be treated as missing then we can calculate PCA differently too according to this old paper (https://www.sciencedirect.com/science/article/abs/pii/S016974399600007X). Helped a bit.\n\n14. OK... stop the laziness and spend the weekend on building proper cross validation pipeline and retrain the models for cite. Nice. a huge jump from 0.812 to 0.814 on leaderboard.\n\n15. Saw some discussion about the supplymentary raw data. Well, I believe the host has done the best preprocessing and probably not helpful, also it is a pain to handle two datasets, so ignored. (TimeMachine: Big Mistake!)\n\n16. Approaching the last week of competition so rewire the cite pipeline for multi. Hoping to see a huge jump on score for multi as seen on cite. But that didn't happen. So... They are actually very different... Maybe I should have accepted some team merging invite earlier....\n\n17. stack a few DANets and GBDT models. \n\n18. picked the wrong submissions for final evaluation. But what can you do.\n\n\n##### Model Summary\nCite:\n(PCA + NMF + FA + ICA + NanPCA + Missing Value PCA (i.e. NIPALS) ) + (DANets + XGBoost)\n\nMulti:\n(PCA + NMF) + (DANets + XGBoost with SVDed Output)\n\n\n##### Looking Back, to improve the score further\n1. Should have spent more time on data preprocessing.\n2. Should anticipate and prepare for a complex ensemble pipeline so save all the models and predictions properly along the way for multi layer stacking.\n3. Shouldn't have assumed Cite and Multi being similar and go for team up.\n\n\nHopefully it is helpful (also to those who didn't participate!)",
      "votes": null
    },
    {
      "id": "2032173",
      "postDate": "11/16/2022 13:08:57",
      "content": "<p>wow, amazing. Great to have you back and share your journey!</p>",
      "rawMarkdown": "wow, amazing. Great to have you back and share your journey!",
      "votes": null
    },
    {
      "id": "2032177",
      "postDate": "11/16/2022 13:13:46",
      "content": "<p>Appreciate the \"thought process\" details here. Thank you.</p>",
      "rawMarkdown": "Appreciate the \"thought process\" details here. Thank you.",
      "votes": null
    },
    {
      "id": "2032185",
      "postDate": "11/16/2022 13:19:51",
      "content": "<p>did you use the xgb multioutput mode? <a href=\"https://xgboost.readthedocs.io/en/stable/tutorials/multioutput.html\" target=\"_blank\">https://xgboost.readthedocs.io/en/stable/tutorials/multioutput.html</a> ? </p>",
      "rawMarkdown": "did you use the xgb multioutput mode? https://xgboost.readthedocs.io/en/stable/tutorials/multioutput.html ?",
      "votes": null
    },
    {
      "id": "2032204",
      "postDate": "11/16/2022 13:36:48",
      "content": "<p>Yes. And you can run it on GPU too! </p>",
      "rawMarkdown": "Yes. And you can run it on GPU too!",
      "votes": null
    },
    {
      "id": "2032761",
      "postDate": "11/16/2022 19:58:15",
      "content": "<p>Thank you very much. This full-story format is really awesome. even for those who did participate in the competition, it is most helpful for learning. </p>",
      "rawMarkdown": "Thank you very much. This full-story format is really awesome. even for those who did participate in the competition, it is most helpful for learning.",
      "votes": null
    },
    {
      "id": "2032938",
      "postDate": "11/16/2022 22:38:58",
      "content": "<p>Tons of learning here.<br>\nI usually has  the same issue with finding dimensionality reduction technique that works and get better results too. (like you mentioned, UMAP, TriMAP, PaCMAP. Doesn't seem working)</p>",
      "rawMarkdown": "Tons of learning here.\nI usually has  the same issue with finding dimensionality reduction technique that works and get better results too. (like you mentioned, UMAP, TriMAP, PaCMAP. Doesn't seem working)",
      "votes": null
    },
    {
      "id": "2033202",
      "postDate": "11/17/2022 04:55:25",
      "content": "<p><a href=\"https://www.kaggle.com/xiaozhouwang\" target=\"_blank\">@xiaozhouwang</a> Thank you for sharing, one starter question, what is MLP?</p>",
      "rawMarkdown": "xiaozhouwang Thank you for sharing, one starter question, what is MLP?",
      "votes": null
    },
    {
      "id": "2033734",
      "postDate": "11/17/2022 13:38:49",
      "content": "<p>it is multi layer perceptron, a.k.a fully connected neural network <a href=\"https://en.wikipedia.org/wiki/Multilayer_perceptron\" target=\"_blank\">https://en.wikipedia.org/wiki/Multilayer_perceptron</a></p>",
      "rawMarkdown": "it is multi layer perceptron, a.k.a fully connected neural network https://en.wikipedia.org/wiki/Multilayer_perceptron",
      "votes": null
    },
    {
      "id": "2034303",
      "postDate": "11/18/2022 02:59:21",
      "content": "<p>Thank you for sharing. This is very helpful for a newbie like me.<br>\nAbout DANets, I couldn't find its implementation. Could you share your code? </p>",
      "rawMarkdown": "Thank you for sharing. This is very helpful for a newbie like me.\nAbout DANets, I couldn't find its implementation. Could you share your code?",
      "votes": null
    },
    {
      "id": "2034314",
      "postDate": "11/18/2022 03:20:03",
      "content": "<p>haha yeah… I wished you accepted our merge request.</p>\n<p>I am impressed that you can do that many experiments within such a short time and with your startup work. I am also a startup guy, I am interested in how did you manage/track your experiment? I tried using excel or W&amp;B, but non of them truly sped up or organized my workflow.  Lately, I am considering setting on MLOps framework like kubeflow in my personal computer to see if it is helpful, what do you think?</p>",
      "rawMarkdown": "haha yeah... I wished you accepted our merge request.\n\nI am impressed that you can do that many experiments within such a short time and with your startup work. I am also a startup guy, I am interested in how did you manage/track your experiment? I tried using excel or W&B, but non of them truly sped up or organized my workflow.  Lately, I am considering setting on MLOps framework like kubeflow in my personal computer to see if it is helpful, what do you think?",
      "votes": null
    },
    {
      "id": "2035499",
      "postDate": "11/18/2022 23:43:14",
      "content": "<p>Thank you so much! This is extremely helpful, especially for illustrating the importance of cleaning data and creating a good validation set. Congrats on an amazing result!</p>",
      "rawMarkdown": "Thank you so much! This is extremely helpful, especially for illustrating the importance of cleaning data and creating a good validation set. Congrats on an amazing result!",
      "votes": null
    },
    {
      "id": "2035501",
      "postDate": "11/18/2022 23:47:10",
      "content": "<p>just a lot of messy code + W&amp;B runs… I don't know the best way to trade off code quality with the speed of experiments actually. The good thing is that you have to do things with quality at work. So I guess that balances things 😄</p>",
      "rawMarkdown": "just a lot of messy code + W&B runs... I don't know the best way to trade off code quality with the speed of experiments actually. The good thing is that you have to do things with quality at work. So I guess that balances things 😄",
      "votes": null
    },
    {
      "id": "2035503",
      "postDate": "11/18/2022 23:48:33",
      "content": "<p>if I got time to clean up the code will post here</p>",
      "rawMarkdown": "if I got time to clean up the code will post here",
      "votes": null
    },
    {
      "id": "2036695",
      "postDate": "11/20/2022 04:47:33",
      "content": "<p>Always great to hear your thoughts, LB. I learned a lot from you when first starting out on Kaggle.</p>",
      "rawMarkdown": "Always great to hear your thoughts, LB. I learned a lot from you when first starting out on Kaggle.",
      "votes": null
    },
    {
      "id": "2041642",
      "postDate": "11/24/2022 05:19:30",
      "content": "<p>Congratulations and thanks for posting your great solution Little Boat!</p>",
      "rawMarkdown": "Congratulations and thanks for posting your great solution Little Boat!",
      "votes": null
    },
    {
      "id": "2042233",
      "postDate": "11/24/2022 14:53:22",
      "content": "<p>Thank you very much for sharing your wonderful ideas ! </p>\n<p>May I kindly ask you: </p>\n<p>what Python packages you used for <br>\n1) NanPCA  <br>\n2) Missing Value PCA (i.e. NIPALS) </p>\n<p>?</p>\n<p>I cannot google anything like \"NanPCA\". For NIPALS there is package <a href=\"https://python-nipals.readthedocs.io/en/latest/usage.html\" target=\"_blank\">https://python-nipals.readthedocs.io/en/latest/usage.html</a><br>\nthere is not any single line of example how to use it. </p>\n<p>Or you are using \"R\" ?</p>\n<p>=======<br>\nPS<br>\nFor Python I found:<br>\n<a href=\"https://www.statsmodels.org/dev/generated/statsmodels.multivariate.pca.PCA.html\" target=\"_blank\">https://www.statsmodels.org/dev/generated/statsmodels.multivariate.pca.PCA.html</a><br>\nmissing = ‘fill-em’ - use EM algorithm to fill missing value. ncomp should be set to the number of factors required.</p>",
      "rawMarkdown": "Thank you very much for sharing your wonderful ideas ! \n\nMay I kindly ask you: \n\nwhat Python packages you used for \n1) NanPCA  \n2) Missing Value PCA (i.e. NIPALS) \n\n?\n\nI cannot google anything like \"NanPCA\". For NIPALS there is package https://python-nipals.readthedocs.io/en/latest/usage.html\nthere is not any single line of example how to use it. \n\nOr you are using \"R\" ?\n\n=======\nPS\nFor Python I found:\nhttps://www.statsmodels.org/dev/generated/statsmodels.multivariate.pca.PCA.html\nmissing = ‘fill-em’ - use EM algorithm to fill missing value. ncomp should be set to the number of factors required.",
      "votes": null
    },
    {
      "id": "2077131",
      "postDate": "12/27/2022 09:13:27",
      "content": "<p><a href=\"https://www.kaggle.com/xiaozhouwang\" target=\"_blank\">@xiaozhouwang</a> , thank you very much for sharing! Can you clarify which cross-validation pipeline did you use in the end?</p>",
      "rawMarkdown": "xiaozhouwang , thank you very much for sharing! Can you clarify which cross-validation pipeline did you use in the end?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2032173,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/16/2022 13:08:57",
      "content": "<p>wow, amazing. Great to have you back and share your journey!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2032177,
      "author_name": "lukeschiefelbein",
      "author_url": "",
      "post_date": "11/16/2022 13:13:46",
      "content": "<p>Appreciate the \"thought process\" details here. Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2032185,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/16/2022 13:19:51",
      "content": "<p>did you use the xgb multioutput mode? <a href=\"https://xgboost.readthedocs.io/en/stable/tutorials/multioutput.html\" target=\"_blank\">https://xgboost.readthedocs.io/en/stable/tutorials/multioutput.html</a> ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2032204,
          "author_name": "xiaozhouwang",
          "author_url": "",
          "post_date": "11/16/2022 13:36:48",
          "content": "<p>Yes. And you can run it on GPU too! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2032761,
      "author_name": "chrisrichardmiles",
      "author_url": "",
      "post_date": "11/16/2022 19:58:15",
      "content": "<p>Thank you very much. This full-story format is really awesome. even for those who did participate in the competition, it is most helpful for learning. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2032938,
      "author_name": "mahyararani",
      "author_url": "",
      "post_date": "11/16/2022 22:38:58",
      "content": "<p>Tons of learning here.<br>\nI usually has  the same issue with finding dimensionality reduction technique that works and get better results too. (like you mentioned, UMAP, TriMAP, PaCMAP. Doesn't seem working)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2033202,
      "author_name": "akmalmir",
      "author_url": "",
      "post_date": "11/17/2022 04:55:25",
      "content": "<p><a href=\"https://www.kaggle.com/xiaozhouwang\" target=\"_blank\">@xiaozhouwang</a> Thank you for sharing, one starter question, what is MLP?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2033734,
          "author_name": "xiaozhouwang",
          "author_url": "",
          "post_date": "11/17/2022 13:38:49",
          "content": "<p>it is multi layer perceptron, a.k.a fully connected neural network <a href=\"https://en.wikipedia.org/wiki/Multilayer_perceptron\" target=\"_blank\">https://en.wikipedia.org/wiki/Multilayer_perceptron</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2034303,
      "author_name": "ludditep",
      "author_url": "",
      "post_date": "11/18/2022 02:59:21",
      "content": "<p>Thank you for sharing. This is very helpful for a newbie like me.<br>\nAbout DANets, I couldn't find its implementation. Could you share your code? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2035503,
          "author_name": "xiaozhouwang",
          "author_url": "",
          "post_date": "11/18/2022 23:48:33",
          "content": "<p>if I got time to clean up the code will post here</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2034314,
      "author_name": "kingychiu",
      "author_url": "",
      "post_date": "11/18/2022 03:20:03",
      "content": "<p>haha yeah… I wished you accepted our merge request.</p>\n<p>I am impressed that you can do that many experiments within such a short time and with your startup work. I am also a startup guy, I am interested in how did you manage/track your experiment? I tried using excel or W&amp;B, but non of them truly sped up or organized my workflow.  Lately, I am considering setting on MLOps framework like kubeflow in my personal computer to see if it is helpful, what do you think?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2035501,
          "author_name": "xiaozhouwang",
          "author_url": "",
          "post_date": "11/18/2022 23:47:10",
          "content": "<p>just a lot of messy code + W&amp;B runs… I don't know the best way to trade off code quality with the speed of experiments actually. The good thing is that you have to do things with quality at work. So I guess that balances things 😄</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2035499,
      "author_name": "perrypineapple",
      "author_url": "",
      "post_date": "11/18/2022 23:43:14",
      "content": "<p>Thank you so much! This is extremely helpful, especially for illustrating the importance of cleaning data and creating a good validation set. Congrats on an amazing result!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2036695,
      "author_name": "jpmiller",
      "author_url": "",
      "post_date": "11/20/2022 04:47:33",
      "content": "<p>Always great to hear your thoughts, LB. I learned a lot from you when first starting out on Kaggle.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2041642,
      "author_name": "songqizhou",
      "author_url": "",
      "post_date": "11/24/2022 05:19:30",
      "content": "<p>Congratulations and thanks for posting your great solution Little Boat!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2042233,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "11/24/2022 14:53:22",
      "content": "<p>Thank you very much for sharing your wonderful ideas ! </p>\n<p>May I kindly ask you: </p>\n<p>what Python packages you used for <br>\n1) NanPCA  <br>\n2) Missing Value PCA (i.e. NIPALS) </p>\n<p>?</p>\n<p>I cannot google anything like \"NanPCA\". For NIPALS there is package <a href=\"https://python-nipals.readthedocs.io/en/latest/usage.html\" target=\"_blank\">https://python-nipals.readthedocs.io/en/latest/usage.html</a><br>\nthere is not any single line of example how to use it. </p>\n<p>Or you are using \"R\" ?</p>\n<p>=======<br>\nPS<br>\nFor Python I found:<br>\n<a href=\"https://www.statsmodels.org/dev/generated/statsmodels.multivariate.pca.PCA.html\" target=\"_blank\">https://www.statsmodels.org/dev/generated/statsmodels.multivariate.pca.PCA.html</a><br>\nmissing = ‘fill-em’ - use EM algorithm to fill missing value. ncomp should be set to the number of factors required.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2077131,
      "author_name": "shitovvladimir",
      "author_url": "",
      "post_date": "12/27/2022 09:13:27",
      "content": "<p><a href=\"https://www.kaggle.com/xiaozhouwang\" target=\"_blank\">@xiaozhouwang</a> , thank you very much for sharing! Can you clarify which cross-validation pipeline did you use in the end?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2032146": "Hey guys, it's been a while since my last actively participated competition and this is a good experience! Though I am not active any more, I frequently come back to kaggle to browse new ideas. One thing I notice is that it is not really straightforward to learn new things by just reading top solutions if you were not active in that competition... And one factor could be that we mostly share what finally worked (and/or what didn't work), but not the thought process of getting there. For someone like me, either a bit too lazy or a bit too busy to actively participate, or someone who might be a bit inexperienced, I think it would be more beneficial to share \"how I got here\" more than \"here I am\". So in that spirit, I would like to start sharing solutions this way...\n\n\n##### Some background: \nMy job can be demanding so whenever I can delegate the work to computers, I do, whenever I cannot, I minimize the time needed for coding such that I can leverage fragamented time as much as possible. And I mainly participate in this competition to learn how feature extraction could work with high dimension inputs and outputs. And they shape what next experiments I decided to try, and mistakes I made along the way so I thought it is important to share.\n\n\n##### My journey:\n\n1. started by reading the Discussion to understand what the data is about, and walk through popular public kernels to understand what can be used as baselines. I noticed that mostly TruncatedSVD was used to reduce dimensions. \n\n2. I figured it might be a good idea to just train a mlp model with all features included. And I was too lazy to build a local validation pipeline so I just randomly sampled 10% as validation set. And the result wasn't so good on public LB compared to public kernels.\n\n3. Then I was thinking, ok, maybe it was because mlp was bad. But xgboost/lightgbm on cpu would take forever to train, and would run out of memory on gpu. So need a better neural network model. What about Tabnet (terrible). Ok. There was one paper I remember claiming similar performance to xgboost. OK. found it. https://arxiv.org/pdf/2112.02962.pdf (It is called DANets). Better but still not as good...\n\n4. Now back to public kernel as the baseline. Maybe instead of SVD, we can use autoencoder? OK, only linear autoencoder  performed ok-ish, any nonlinearity didn't work... no matter what tricks (e.g. swap noise augmentation) used. Hmmm...\n\n5. Back to TSVD + MLP as baseline again... Let's make this baseline better first. First swap MLP with DANets. And let's just focus on cite since it only have high dim inputs. Whatever works for cite should work for multi right? (TimeMachine: Nope!)\n\n6. Let's standardize the data since PCA likes it. OK. Slightly better. The explained variance seems quite low, but adding more components as features doesn't seem too helpful. So likely the inputs are quite noisy. Don't really know what to do. Well, we can always add different decomposition methods if no better ideas. Added NMF, FactorAnalysis, FastICA. Ok. All of them worked. \n\n7. Now let's also train some xgboost model since now we can train on gpu. Ok cool. Averaging xgboost with DANets improves results significantly.\n\n8. Maybe should try autoencoder again.... Read some papers. Ok. Still didn't work.\n\n9. Let's try some popular nonlinear dimension reduction techniques. UMAP, TriMAP, PaCMAP. Doesn't seem working.\n\n10. XGBoost or LightGBM train one model per target which seems wasteful and may not consider the correlation between outputs. Let's see if there is a better way out there. https://arxiv.org/abs/1909.04373 found this GBDT-MO. and it has code. Tried. Didn't work so well. \n\n11. Read in the Discussion that we can select features by matching input/output names for Cite. Tried and it worked. Neat.\n\n12. Also read in the Discussion that the 0s in inputs may not be actual 0s could also be missing. OK. Tried to calculate the mean/std by considering all 0s as missing and then calculate PCA. Adding to the features improved the model a bit.\n\n13. Realized that if 0s can be treated as missing then we can calculate PCA differently too according to this old paper (https://www.sciencedirect.com/science/article/abs/pii/S016974399600007X). Helped a bit.\n\n14. OK... stop the laziness and spend the weekend on building proper cross validation pipeline and retrain the models for cite. Nice. a huge jump from 0.812 to 0.814 on leaderboard.\n\n15. Saw some discussion about the supplymentary raw data. Well, I believe the host has done the best preprocessing and probably not helpful, also it is a pain to handle two datasets, so ignored. (TimeMachine: Big Mistake!)\n\n16. Approaching the last week of competition so rewire the cite pipeline for multi. Hoping to see a huge jump on score for multi as seen on cite. But that didn't happen. So... They are actually very different... Maybe I should have accepted some team merging invite earlier....\n\n17. stack a few DANets and GBDT models. \n\n18. picked the wrong submissions for final evaluation. But what can you do.\n\n\n##### Model Summary\nCite:\n(PCA + NMF + FA + ICA + NanPCA + Missing Value PCA (i.e. NIPALS) ) + (DANets + XGBoost)\n\nMulti:\n(PCA + NMF) + (DANets + XGBoost with SVDed Output)\n\n\n##### Looking Back, to improve the score further\n1. Should have spent more time on data preprocessing.\n2. Should anticipate and prepare for a complex ensemble pipeline so save all the models and predictions properly along the way for multi layer stacking.\n3. Shouldn't have assumed Cite and Multi being similar and go for team up.\n\n\nHopefully it is helpful (also to those who didn't participate!)",
    "2032173": "wow, amazing. Great to have you back and share your journey!",
    "2032177": "Appreciate the \"thought process\" details here. Thank you.",
    "2032185": "did you use the xgb multioutput mode? https://xgboost.readthedocs.io/en/stable/tutorials/multioutput.html ?",
    "2032204": "Yes. And you can run it on GPU too!",
    "2032761": "Thank you very much. This full-story format is really awesome. even for those who did participate in the competition, it is most helpful for learning.",
    "2032938": "Tons of learning here.\nI usually has  the same issue with finding dimensionality reduction technique that works and get better results too. (like you mentioned, UMAP, TriMAP, PaCMAP. Doesn't seem working)",
    "2033202": "xiaozhouwang Thank you for sharing, one starter question, what is MLP?",
    "2033734": "it is multi layer perceptron, a.k.a fully connected neural network https://en.wikipedia.org/wiki/Multilayer_perceptron",
    "2034303": "Thank you for sharing. This is very helpful for a newbie like me.\nAbout DANets, I couldn't find its implementation. Could you share your code?",
    "2034314": "haha yeah... I wished you accepted our merge request.\n\nI am impressed that you can do that many experiments within such a short time and with your startup work. I am also a startup guy, I am interested in how did you manage/track your experiment? I tried using excel or W&B, but non of them truly sped up or organized my workflow.  Lately, I am considering setting on MLOps framework like kubeflow in my personal computer to see if it is helpful, what do you think?",
    "2035499": "Thank you so much! This is extremely helpful, especially for illustrating the importance of cleaning data and creating a good validation set. Congrats on an amazing result!",
    "2035501": "just a lot of messy code + W&B runs... I don't know the best way to trade off code quality with the speed of experiments actually. The good thing is that you have to do things with quality at work. So I guess that balances things 😄",
    "2035503": "if I got time to clean up the code will post here",
    "2036695": "Always great to hear your thoughts, LB. I learned a lot from you when first starting out on Kaggle.",
    "2041642": "Congratulations and thanks for posting your great solution Little Boat!",
    "2042233": "Thank you very much for sharing your wonderful ideas ! \n\nMay I kindly ask you: \n\nwhat Python packages you used for \n1) NanPCA  \n2) Missing Value PCA (i.e. NIPALS) \n\n?\n\nI cannot google anything like \"NanPCA\". For NIPALS there is package https://python-nipals.readthedocs.io/en/latest/usage.html\nthere is not any single line of example how to use it. \n\nOr you are using \"R\" ?\n\n=======\nPS\nFor Python I found:\nhttps://www.statsmodels.org/dev/generated/statsmodels.multivariate.pca.PCA.html\nmissing = ‘fill-em’ - use EM algorithm to fill missing value. ncomp should be set to the number of factors required.",
    "2077131": "xiaozhouwang , thank you very much for sharing! Can you clarify which cross-validation pipeline did you use in the end?"
  },
  "source": "meta"
}