{
  "id": 366417,
  "title": "[6th private - 3rd public] Summary of our solution",
  "url": "/competitions/open-problems-multimodal/writeups/risk-zalopay-aggressive-6th-private-3rd-public-sum",
  "author_name": "",
  "post_date": "2022-11-16T09:58:09.837Z",
  "votes": 40,
  "comment_count": 9,
  "views": 0,
  "content": "<p>First of all, thanks to the Kaggle team and the host for providing this cutting-edge technology dataset and hosting this great challenge. I learned a lot from the competition. Many thanks to my teammates <a href=\"https://www.kaggle.com/mathormad\" target=\"_blank\">@mathormad</a> <a href=\"https://www.kaggle.com/nguyenvlm\" target=\"_blank\">@nguyenvlm</a> for hard-working days. This is a summary of what we did to get 3rd place in LB and 6th place in private</p>\n<h1>Local CV (for both cite and multiome task)</h1>\n<p>We use <strong>stratified-5-Fold</strong> (stratify by day/donor/celltype) CV strategy. We do not use GroupKFold to avoid overfitting in public, and the private dataset day is also quite \"far\" from the training dataset, so it's a bit risky to use a CV split by day (we've tried it, but both CV and LB drop). In our opinion, <strong>this is the key to our stable placement in both public and private leaderboards. Our best submission in CV are also the best on Public and Private leaderboard.</strong></p>\n<h1>Feature engineering (cite)</h1>\n<ul>\n<li>Feature selection: remove all 0 cols group by metadata (day/donor/cell_type/train/test) --&gt; remove ~4,500 features</li>\n<li>Dimension reduction: we use different methods to increase the diversity for the final ensemble: 240 n_components SVD/quantiledPCA and denoise Autoencoder (256 latent dims). For SVD, increase n_iter params slightly improve CV by 2e-4 (but run longer) (ours is 50 iter, the default of sklearn library is only 5). The denoise autoencoder helps us improve CV by 1e-3 (comparing to SVD/PCA)</li>\n<li>Feature importance<ul>\n<li>Use name matching from this discussion <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349242\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349242</a></li>\n<li>Search using xgb feature importance: for cite, we fit multiple xgb for each target (full 22k feature) then choose the top 5  important features of each target. In total, we get around 500 important features for our model.</li></ul></li>\n<li>TargetEncoder for xgb: we apply target encode the cell_type for each target (each cell_type will be represented by a 140-dims vector)</li>\n</ul>\n<h1>Feature engineering (multiome)</h1>\n<ul>\n<li>Feature selection: remove chY features, which is not correlated to our target. We also remove all 0 cols group by metadata (day/donor/cell_type/train/test) --&gt; remove ~ 500 features</li>\n<li>Dimension reduction: we use 256 n_components SVD fitted with 200iter</li>\n</ul>\n<h1>Training process</h1>\n<ul>\n<li>Use custom loss (weighted correlation and MSE loss)</li>\n<li>Use c-mixup to increase the diversity for tabnet</li>\n<li>SWA when training MLP, Denoise Autocoder, Tabnet and 1D-CNN</li>\n<li>Adam optimizer with high learning rate (1e-2)</li>\n<li>(Multiome) XGB is trained with the PCA of target as label</li>\n</ul>\n<h1>External data</h1>\n<p>For cite task, we <strong>apply the whole training process above for the <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359355\" target=\"_blank\">raw count dataset</a></strong>, then ensemble with the original one. This significantly boosts the performance by 1e-3<br>\nFor multiome task, we do not use the raw count dataset because it's lacking some rows, we cannot match the OOF between the original and raw count so we do not have CV score to validate. I think it should work but we do not have enough time to rematch the OOF</p>\n<h1>Stacking feature &amp; Pseudo-labeling</h1>\n<p>We concatenate the prediction of 1 model with current features as input to another model (e.g. use output of xgb as input of MLP, and vice versa), which improves the performance of a single model around 5e-4. <br>\nWe also use pseudo labeling (not much improvement, about 2-3e-4)</p>\n<h1>Post process</h1>\n<ul>\n<li>(Cite) Apply standard scaler (axis 1) for the output of each fold before taking the average</li>\n<li>(Multiome) Some of the target is all 0, so we remove them from the training process and replace 0 later. This helps us increase our CV by 1e-3</li>\n</ul>\n<h1>Ensemble</h1>\n<ul>\n<li>(Cite) Our final submission is a blending of the following models:<ul>\n<li>MLP</li>\n<li>XGB</li>\n<li>Tabnet</li>\n<li>1D-CNN<br>\nEach model is trained on the original and raw count dataset -&gt; 8 models in total</li></ul></li>\n<li>(Multiome) Blending of:<ul>\n<li>MLP</li>\n<li>XGB</li>\n<li>Tabnet</li></ul></li>\n</ul>\n<h1>Tryhard</h1>\n<p>We worked 6 hours/day, from the very beginning of the competition to the last hour. Especially for this competition when we have 2 tasks with 2 datasets, a lot of work and experiments to do. The most important thing I learned during this competition is that the more time you spend, the higher place you will be.</p>\n<h1>What does not work</h1>\n<ul>\n<li>Encoder for multiome task</li>\n<li>Use important features in multiome task</li>\n<li>Some bio techniques such as Ivis (dimension reduction), Magic (denoising)</li>\n<li>Use raw label for training</li>\n<li>…</li>\n</ul>",
  "messages": [
    {
      "id": "2031508",
      "postDate": "11/16/2022 05:59:05",
      "content": "<p>First of all, thanks to the Kaggle team and the host for providing this cutting-edge technology dataset and hosting this great challenge. I learned a lot from the competition. Many thanks to my teammates <a href=\"https://www.kaggle.com/mathormad\" target=\"_blank\">@mathormad</a> <a href=\"https://www.kaggle.com/nguyenvlm\" target=\"_blank\">@nguyenvlm</a> for hard-working days. This is a summary of what we did to get 3rd place in LB and 6th place in private</p>\n<h1>Local CV (for both cite and multiome task)</h1>\n<p>We use <strong>stratified-5-Fold</strong> (stratify by day/donor/celltype) CV strategy. We do not use GroupKFold to avoid overfitting in public, and the private dataset day is also quite \"far\" from the training dataset, so it's a bit risky to use a CV split by day (we've tried it, but both CV and LB drop). In our opinion, <strong>this is the key to our stable placement in both public and private leaderboards. Our best submission in CV are also the best on Public and Private leaderboard.</strong></p>\n<h1>Feature engineering (cite)</h1>\n<ul>\n<li>Feature selection: remove all 0 cols group by metadata (day/donor/cell_type/train/test) --&gt; remove ~4,500 features</li>\n<li>Dimension reduction: we use different methods to increase the diversity for the final ensemble: 240 n_components SVD/quantiledPCA and denoise Autoencoder (256 latent dims). For SVD, increase n_iter params slightly improve CV by 2e-4 (but run longer) (ours is 50 iter, the default of sklearn library is only 5). The denoise autoencoder helps us improve CV by 1e-3 (comparing to SVD/PCA)</li>\n<li>Feature importance<ul>\n<li>Use name matching from this discussion <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349242\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349242</a></li>\n<li>Search using xgb feature importance: for cite, we fit multiple xgb for each target (full 22k feature) then choose the top 5  important features of each target. In total, we get around 500 important features for our model.</li></ul></li>\n<li>TargetEncoder for xgb: we apply target encode the cell_type for each target (each cell_type will be represented by a 140-dims vector)</li>\n</ul>\n<h1>Feature engineering (multiome)</h1>\n<ul>\n<li>Feature selection: remove chY features, which is not correlated to our target. We also remove all 0 cols group by metadata (day/donor/cell_type/train/test) --&gt; remove ~ 500 features</li>\n<li>Dimension reduction: we use 256 n_components SVD fitted with 200iter</li>\n</ul>\n<h1>Training process</h1>\n<ul>\n<li>Use custom loss (weighted correlation and MSE loss)</li>\n<li>Use c-mixup to increase the diversity for tabnet</li>\n<li>SWA when training MLP, Denoise Autocoder, Tabnet and 1D-CNN</li>\n<li>Adam optimizer with high learning rate (1e-2)</li>\n<li>(Multiome) XGB is trained with the PCA of target as label</li>\n</ul>\n<h1>External data</h1>\n<p>For cite task, we <strong>apply the whole training process above for the <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359355\" target=\"_blank\">raw count dataset</a></strong>, then ensemble with the original one. This significantly boosts the performance by 1e-3<br>\nFor multiome task, we do not use the raw count dataset because it's lacking some rows, we cannot match the OOF between the original and raw count so we do not have CV score to validate. I think it should work but we do not have enough time to rematch the OOF</p>\n<h1>Stacking feature &amp; Pseudo-labeling</h1>\n<p>We concatenate the prediction of 1 model with current features as input to another model (e.g. use output of xgb as input of MLP, and vice versa), which improves the performance of a single model around 5e-4. <br>\nWe also use pseudo labeling (not much improvement, about 2-3e-4)</p>\n<h1>Post process</h1>\n<ul>\n<li>(Cite) Apply standard scaler (axis 1) for the output of each fold before taking the average</li>\n<li>(Multiome) Some of the target is all 0, so we remove them from the training process and replace 0 later. This helps us increase our CV by 1e-3</li>\n</ul>\n<h1>Ensemble</h1>\n<ul>\n<li>(Cite) Our final submission is a blending of the following models:<ul>\n<li>MLP</li>\n<li>XGB</li>\n<li>Tabnet</li>\n<li>1D-CNN<br>\nEach model is trained on the original and raw count dataset -&gt; 8 models in total</li></ul></li>\n<li>(Multiome) Blending of:<ul>\n<li>MLP</li>\n<li>XGB</li>\n<li>Tabnet</li></ul></li>\n</ul>\n<h1>Tryhard</h1>\n<p>We worked 6 hours/day, from the very beginning of the competition to the last hour. Especially for this competition when we have 2 tasks with 2 datasets, a lot of work and experiments to do. The most important thing I learned during this competition is that the more time you spend, the higher place you will be.</p>\n<h1>What does not work</h1>\n<ul>\n<li>Encoder for multiome task</li>\n<li>Use important features in multiome task</li>\n<li>Some bio techniques such as Ivis (dimension reduction), Magic (denoising)</li>\n<li>Use raw label for training</li>\n<li>…</li>\n</ul>",
      "rawMarkdown": "First of all, thanks to the Kaggle team and the host for providing this cutting-edge technology dataset and hosting this great challenge. I learned a lot from the competition. Many thanks to my teammates @mathormad @nguyenvlm for hard-working days. This is a summary of what we did to get 3rd place in LB and 6th place in private\n\n# Local CV (for both cite and multiome task)\nWe use **stratified-5-Fold** (stratify by day/donor/celltype) CV strategy. We do not use GroupKFold to avoid overfitting in public, and the private dataset day is also quite \"far\" from the training dataset, so it's a bit risky to use a CV split by day (we've tried it, but both CV and LB drop). In our opinion, **this is the key to our stable placement in both public and private leaderboards. Our best submission in CV are also the best on Public and Private leaderboard.**\n\n\n# Feature engineering (cite)\n- Feature selection: remove all 0 cols group by metadata (day/donor/cell_type/train/test) --> remove ~4,500 features\n- Dimension reduction: we use different methods to increase the diversity for the final ensemble: 240 n_components SVD/quantiledPCA and denoise Autoencoder (256 latent dims). For SVD, increase n_iter params slightly improve CV by 2e-4 (but run longer) (ours is 50 iter, the default of sklearn library is only 5). The denoise autoencoder helps us improve CV by 1e-3 (comparing to SVD/PCA)\n- Feature importance\n    - Use name matching from this discussion https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349242\n    - Search using xgb feature importance: for cite, we fit multiple xgb for each target (full 22k feature) then choose the top 5  important features of each target. In total, we get around 500 important features for our model.\n- TargetEncoder for xgb: we apply target encode the cell_type for each target (each cell_type will be represented by a 140-dims vector)\n\n\n# Feature engineering (multiome)\n- Feature selection: remove chY features, which is not correlated to our target. We also remove all 0 cols group by metadata (day/donor/cell_type/train/test) --> remove ~ 500 features\n- Dimension reduction: we use 256 n_components SVD fitted with 200iter\n\n# Training process\n- Use custom loss (weighted correlation and MSE loss)\n- Use c-mixup to increase the diversity for tabnet\n- SWA when training MLP, Denoise Autocoder, Tabnet and 1D-CNN\n- Adam optimizer with high learning rate (1e-2)\n- (Multiome) XGB is trained with the PCA of target as label\n\n# External data\nFor cite task, we **apply the whole training process above for the [raw count dataset](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359355)**, then ensemble with the original one. This significantly boosts the performance by 1e-3\nFor multiome task, we do not use the raw count dataset because it's lacking some rows, we cannot match the OOF between the original and raw count so we do not have CV score to validate. I think it should work but we do not have enough time to rematch the OOF\n\n# Stacking feature & Pseudo-labeling\nWe concatenate the prediction of 1 model with current features as input to another model (e.g. use output of xgb as input of MLP, and vice versa), which improves the performance of a single model around 5e-4. \nWe also use pseudo labeling (not much improvement, about 2-3e-4)\n\n# Post process\n- (Cite) Apply standard scaler (axis 1) for the output of each fold before taking the average\n- (Multiome) Some of the target is all 0, so we remove them from the training process and replace 0 later. This helps us increase our CV by 1e-3\n\n# Ensemble\n- (Cite) Our final submission is a blending of the following models:\n  - MLP\n  - XGB\n  - Tabnet\n  - 1D-CNN\nEach model is trained on the original and raw count dataset -> 8 models in total\n- (Multiome) Blending of:\n    - MLP\n    - XGB\n    - Tabnet\n\n\n# Tryhard\nWe worked 6 hours/day, from the very beginning of the competition to the last hour. Especially for this competition when we have 2 tasks with 2 datasets, a lot of work and experiments to do. The most important thing I learned during this competition is that the more time you spend, the higher place you will be.\n\n# What does not work\n- Encoder for multiome task\n- Use important features in multiome task\n- Some bio techniques such as Ivis (dimension reduction), Magic (denoising)\n- Use raw label for training\n- ...",
      "votes": null
    },
    {
      "id": "2031549",
      "postDate": "11/16/2022 06:45:15",
      "content": "<p><a href=\"https://www.kaggle.com/minhtu123\" target=\"_blank\">@minhtu123</a> Thank you for sharing, one starter question, which is first, feature selection and then dimension reduction or vice versa?</p>",
      "rawMarkdown": "minhtu123 Thank you for sharing, one starter question, which is first, feature selection and then dimension reduction or vice versa?",
      "votes": null
    },
    {
      "id": "2031554",
      "postDate": "11/16/2022 06:48:30",
      "content": "<p>Feature selection first, then the reduction. We want the embedding feature learned from clean data</p>",
      "rawMarkdown": "Feature selection first, then the reduction. We want the embedding feature learned from clean data",
      "votes": null
    },
    {
      "id": "2031609",
      "postDate": "11/16/2022 07:17:56",
      "content": "<p>Thank you for sharing, may I ask what is the difference between raw count dataset and original dataset?  Does raw count dataset means dataset composed by 500 important features?</p>",
      "rawMarkdown": "Thank you for sharing, may I ask what is the difference between raw count dataset and original dataset?  Does raw count dataset means dataset composed by 500 important features?",
      "votes": null
    },
    {
      "id": "2031624",
      "postDate": "11/16/2022 07:25:14",
      "content": "<p>The raw count dataset is this one: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359355\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359355</a></p>",
      "rawMarkdown": "The raw count dataset is this one: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359355",
      "votes": null
    },
    {
      "id": "2031631",
      "postDate": "11/16/2022 07:31:22",
      "content": "<p>Thank you for your reply, I understand now, thank  you very much!</p>",
      "rawMarkdown": "Thank you for your reply, I understand now, thank  you very much!",
      "votes": null
    },
    {
      "id": "2031807",
      "postDate": "11/16/2022 08:50:43",
      "content": "<p>Important features were working for me in both multiome and siteseq. But I used random forest and another method for selection (thresholding). Also I use correlation features for siteseq (feature-target correlation for whole pairs of it)</p>",
      "rawMarkdown": "Important features were working for me in both multiome and siteseq. But I used random forest and another method for selection (thresholding). Also I use correlation features for siteseq (feature-target correlation for whole pairs of it)",
      "votes": null
    },
    {
      "id": "2032869",
      "postDate": "11/16/2022 21:00:01",
      "content": "<p>Hearty congratulations for the result! Hats off to all of you for the dedication and perseverance! 6 hours per day for the length of the competition is a commendable perseverance!! </p>\n<p>All the best <a href=\"https://www.kaggle.com/minhtu123\" target=\"_blank\">@minhtu123</a> and team!!</p>",
      "rawMarkdown": "Hearty congratulations for the result! Hats off to all of you for the dedication and perseverance! 6 hours per day for the length of the competition is a commendable perseverance!! \n\nAll the best @minhtu123 and team!!",
      "votes": null
    },
    {
      "id": "2033195",
      "postDate": "11/17/2022 04:49:09",
      "content": "<p>To my mind, any important features searching strategy works as well. I tried some of them (correlation features, shap, random selection, etc) and they all lead to a common subset of features (let's say all strategies result overlap more than 80%)</p>",
      "rawMarkdown": "To my mind, any important features searching strategy works as well. I tried some of them (correlation features, shap, random selection, etc) and they all lead to a common subset of features (let's say all strategies result overlap more than 80%)",
      "votes": null
    },
    {
      "id": "2047199",
      "postDate": "11/28/2022 17:14:16",
      "content": "<p><a href=\"https://www.kaggle.com/minhtu123\" target=\"_blank\">@minhtu123</a> </p>\n<p>Thanks for sharing and congratulations with the gold medal !</p>\n<p>May I ask you about the following, you write:<br>\n\"Search using xgb feature importance: for cite, we fit multiple xgb for each target (full 22k feature) then choose the top 5 important features of each target. In total, we get around 500 important features for our model.\"</p>\n<p>Is is possible to share these results ?</p>\n<p>Thanks in advance ! </p>",
      "rawMarkdown": "minhtu123 \n\nThanks for sharing and congratulations with the gold medal !\n\nMay I ask you about the following, you write:\n\"Search using xgb feature importance: for cite, we fit multiple xgb for each target (full 22k feature) then choose the top 5 important features of each target. In total, we get around 500 important features for our model.\"\n\nIs is possible to share these results ?\n\nThanks in advance !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2031549,
      "author_name": "akmalmir",
      "author_url": "",
      "post_date": "11/16/2022 06:45:15",
      "content": "<p><a href=\"https://www.kaggle.com/minhtu123\" target=\"_blank\">@minhtu123</a> Thank you for sharing, one starter question, which is first, feature selection and then dimension reduction or vice versa?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2031554,
          "author_name": "minhtu123",
          "author_url": "",
          "post_date": "11/16/2022 06:48:30",
          "content": "<p>Feature selection first, then the reduction. We want the embedding feature learned from clean data</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2031609,
      "author_name": "learningandchanging",
      "author_url": "",
      "post_date": "11/16/2022 07:17:56",
      "content": "<p>Thank you for sharing, may I ask what is the difference between raw count dataset and original dataset?  Does raw count dataset means dataset composed by 500 important features?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2031624,
          "author_name": "minhtu123",
          "author_url": "",
          "post_date": "11/16/2022 07:25:14",
          "content": "<p>The raw count dataset is this one: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359355\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359355</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2031631,
          "author_name": "learningandchanging",
          "author_url": "",
          "post_date": "11/16/2022 07:31:22",
          "content": "<p>Thank you for your reply, I understand now, thank  you very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2031807,
      "author_name": "bejeweled",
      "author_url": "",
      "post_date": "11/16/2022 08:50:43",
      "content": "<p>Important features were working for me in both multiome and siteseq. But I used random forest and another method for selection (thresholding). Also I use correlation features for siteseq (feature-target correlation for whole pairs of it)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2033195,
          "author_name": "minhtu123",
          "author_url": "",
          "post_date": "11/17/2022 04:49:09",
          "content": "<p>To my mind, any important features searching strategy works as well. I tried some of them (correlation features, shap, random selection, etc) and they all lead to a common subset of features (let's say all strategies result overlap more than 80%)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2032869,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "11/16/2022 21:00:01",
      "content": "<p>Hearty congratulations for the result! Hats off to all of you for the dedication and perseverance! 6 hours per day for the length of the competition is a commendable perseverance!! </p>\n<p>All the best <a href=\"https://www.kaggle.com/minhtu123\" target=\"_blank\">@minhtu123</a> and team!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2047199,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "11/28/2022 17:14:16",
      "content": "<p><a href=\"https://www.kaggle.com/minhtu123\" target=\"_blank\">@minhtu123</a> </p>\n<p>Thanks for sharing and congratulations with the gold medal !</p>\n<p>May I ask you about the following, you write:<br>\n\"Search using xgb feature importance: for cite, we fit multiple xgb for each target (full 22k feature) then choose the top 5 important features of each target. In total, we get around 500 important features for our model.\"</p>\n<p>Is is possible to share these results ?</p>\n<p>Thanks in advance ! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2031508": "First of all, thanks to the Kaggle team and the host for providing this cutting-edge technology dataset and hosting this great challenge. I learned a lot from the competition. Many thanks to my teammates @mathormad @nguyenvlm for hard-working days. This is a summary of what we did to get 3rd place in LB and 6th place in private\n\n# Local CV (for both cite and multiome task)\nWe use **stratified-5-Fold** (stratify by day/donor/celltype) CV strategy. We do not use GroupKFold to avoid overfitting in public, and the private dataset day is also quite \"far\" from the training dataset, so it's a bit risky to use a CV split by day (we've tried it, but both CV and LB drop). In our opinion, **this is the key to our stable placement in both public and private leaderboards. Our best submission in CV are also the best on Public and Private leaderboard.**\n\n\n# Feature engineering (cite)\n- Feature selection: remove all 0 cols group by metadata (day/donor/cell_type/train/test) --> remove ~4,500 features\n- Dimension reduction: we use different methods to increase the diversity for the final ensemble: 240 n_components SVD/quantiledPCA and denoise Autoencoder (256 latent dims). For SVD, increase n_iter params slightly improve CV by 2e-4 (but run longer) (ours is 50 iter, the default of sklearn library is only 5). The denoise autoencoder helps us improve CV by 1e-3 (comparing to SVD/PCA)\n- Feature importance\n    - Use name matching from this discussion https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349242\n    - Search using xgb feature importance: for cite, we fit multiple xgb for each target (full 22k feature) then choose the top 5  important features of each target. In total, we get around 500 important features for our model.\n- TargetEncoder for xgb: we apply target encode the cell_type for each target (each cell_type will be represented by a 140-dims vector)\n\n\n# Feature engineering (multiome)\n- Feature selection: remove chY features, which is not correlated to our target. We also remove all 0 cols group by metadata (day/donor/cell_type/train/test) --> remove ~ 500 features\n- Dimension reduction: we use 256 n_components SVD fitted with 200iter\n\n# Training process\n- Use custom loss (weighted correlation and MSE loss)\n- Use c-mixup to increase the diversity for tabnet\n- SWA when training MLP, Denoise Autocoder, Tabnet and 1D-CNN\n- Adam optimizer with high learning rate (1e-2)\n- (Multiome) XGB is trained with the PCA of target as label\n\n# External data\nFor cite task, we **apply the whole training process above for the [raw count dataset](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359355)**, then ensemble with the original one. This significantly boosts the performance by 1e-3\nFor multiome task, we do not use the raw count dataset because it's lacking some rows, we cannot match the OOF between the original and raw count so we do not have CV score to validate. I think it should work but we do not have enough time to rematch the OOF\n\n# Stacking feature & Pseudo-labeling\nWe concatenate the prediction of 1 model with current features as input to another model (e.g. use output of xgb as input of MLP, and vice versa), which improves the performance of a single model around 5e-4. \nWe also use pseudo labeling (not much improvement, about 2-3e-4)\n\n# Post process\n- (Cite) Apply standard scaler (axis 1) for the output of each fold before taking the average\n- (Multiome) Some of the target is all 0, so we remove them from the training process and replace 0 later. This helps us increase our CV by 1e-3\n\n# Ensemble\n- (Cite) Our final submission is a blending of the following models:\n  - MLP\n  - XGB\n  - Tabnet\n  - 1D-CNN\nEach model is trained on the original and raw count dataset -> 8 models in total\n- (Multiome) Blending of:\n    - MLP\n    - XGB\n    - Tabnet\n\n\n# Tryhard\nWe worked 6 hours/day, from the very beginning of the competition to the last hour. Especially for this competition when we have 2 tasks with 2 datasets, a lot of work and experiments to do. The most important thing I learned during this competition is that the more time you spend, the higher place you will be.\n\n# What does not work\n- Encoder for multiome task\n- Use important features in multiome task\n- Some bio techniques such as Ivis (dimension reduction), Magic (denoising)\n- Use raw label for training\n- ...",
    "2031549": "minhtu123 Thank you for sharing, one starter question, which is first, feature selection and then dimension reduction or vice versa?",
    "2031554": "Feature selection first, then the reduction. We want the embedding feature learned from clean data",
    "2031609": "Thank you for sharing, may I ask what is the difference between raw count dataset and original dataset?  Does raw count dataset means dataset composed by 500 important features?",
    "2031624": "The raw count dataset is this one: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359355",
    "2031631": "Thank you for your reply, I understand now, thank  you very much!",
    "2031807": "Important features were working for me in both multiome and siteseq. But I used random forest and another method for selection (thresholding). Also I use correlation features for siteseq (feature-target correlation for whole pairs of it)",
    "2032869": "Hearty congratulations for the result! Hats off to all of you for the dedication and perseverance! 6 hours per day for the length of the competition is a commendable perseverance!! \n\nAll the best @minhtu123 and team!!",
    "2033195": "To my mind, any important features searching strategy works as well. I tried some of them (correlation features, shap, random selection, etc) and they all lead to a common subset of features (let's say all strategies result overlap more than 80%)",
    "2047199": "minhtu123 \n\nThanks for sharing and congratulations with the gold medal !\n\nMay I ask you about the following, you write:\n\"Search using xgb feature importance: for cite, we fit multiple xgb for each target (full 22k feature) then choose the top 5 important features of each target. In total, we get around 500 important features for our model.\"\n\nIs is possible to share these results ?\n\nThanks in advance !"
  },
  "source": "meta"
}