{
  "id": 368052,
  "title": "16th Place Solution Summary",
  "url": "/competitions/open-problems-multimodal/writeups/aypy-16th-place-solution-summary",
  "author_name": "",
  "post_date": "2022-11-23T22:32:43.723Z",
  "votes": 19,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks to the competition host, Kaggle team, Saturn cloud team and congrats to all the winners!</p>\n<p>Although many great solutions have already been posted and my solution may not contain new approaches, I would like to leave my efforts over the past two months. In this term I learned a lot. Thanks to the all competitors for a great game.</p>\n<p>Please forgive me if it is difficult to read this post or see the schematic diagram below, as my command of English is not very good and my educational background was different from computer science or machine learning area.</p>\n<hr>\n<h1>Overview:</h1>\n<p>The machine learning algorithms used were as follows:</p>\n<ul>\n<li>Two MLPs (4 and 9 hidden layers, the variants of <a href=\"https://www.kaggle.com/code/pourchot/all-in-one-citeseq-multiome-with-keras?scriptVersionId=108466116\" target=\"_blank\">Laurent Pourchot's model</a>)</li>\n<li>Conv1d (almost all the same as the <a href=\"https://www.kaggle.com/c/lish-moa/discussion/202256\" target=\"_blank\">tmp’s 1D-CNN model for tabular data on MoA competition</a>)</li>\n<li>LGBM</li>\n</ul>\n<p>In my case, stacking scheme boosted the score. The outputs of level 1 models were concatenated and then used as input for level 2. It was effective to apply dimensionality reduction to concatenated level 1 outputs after standardization. When the outputs were just concatenated without dimensionality reduction, the score of level 2 was rather lower than that of level 1.</p>\n<p>Also, ensemble worked well. I created several models with slight difference (different dimensionality reduction algorithms, feature extraction methods and loss functions) and blended them. In addition, each learning process was performed on 15 random-seeds and results were averaged.<br>\n<br></p>\n<h1>CV scheme:</h1>\n<p>I used simple KFold (k = 5). Fortunately, I resulted in shakeup in private LB.<br>\n<br></p>\n<h1>Citeseq:</h1>\n<p>The diagram of my Citeseq stacking scheme is as follow.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11204962%2F65d09e1ed96e210c63bab1d627b8497b%2Fsolution%20scheme-vstack.jpg?generation=1669194022322983&amp;alt=media\" alt=\"\"></p>\n<h4>Preprocess</h4>\n<p>Before dimensionality reduction or feature extraction in level 1, the set of all features which are constant in the train or test were eliminated according to the <a href=\"https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart\" target=\"_blank\">AmbrosM’s Code</a>.</p>\n<ul>\n<li>Dimensionality reduction: tSVD and PCA (n_components = 64) were used separately and inference results were finally blended.</li>\n<li>Feature extraction: According to the <a href=\"https://www.kaggle.com/code/fabiencrom/msci-correlations-eda-citeseq/notebook\" target=\"_blank\">Fabien Crom's Code</a>, I picked features with high Pearson correlation coefficients for the targets. In order to gain diversity as much as possible, I changed the picking query for each models. For example:<br>\n・Extract the RNAs in order of highest <strong>average</strong> of correlation to 140 proteins<br>\n・Extract the RNAs that have a high correlation value for a <strong>single</strong> protein, <strong>not the average</strong><br>\n・Change how many of the top RNAs are extracted<br>\n・With or without dimensionality reduction after extracted<br>\n<br></li>\n</ul>\n<h1>Multiome:</h1>\n<p>The scheme is almost the same as that of Citeseq. The differences are as follows:</p>\n<ul>\n<li>Extracted features were not used. In the Multiome case, the score was deteriorated when they were concatenated with 64 dims compressed features.</li>\n<li>For NN algorisms, only MSE was used as loss function.</li>\n<li>The target vectors were compressed to 512 (for NN) or 128 (for LGBM) dims by tSVD and used in training. The inferred vectors were decompressed to 23418 dims by inverse SVD.</li>\n</ul>",
  "messages": [
    {
      "id": "2040678",
      "postDate": "11/23/2022 09:33:23",
      "content": "<p>Thanks to the competition host, Kaggle team, Saturn cloud team and congrats to all the winners!</p>\n<p>Although many great solutions have already been posted and my solution may not contain new approaches, I would like to leave my efforts over the past two months. In this term I learned a lot. Thanks to the all competitors for a great game.</p>\n<p>Please forgive me if it is difficult to read this post or see the schematic diagram below, as my command of English is not very good and my educational background was different from computer science or machine learning area.</p>\n<hr>\n<h1>Overview:</h1>\n<p>The machine learning algorithms used were as follows:</p>\n<ul>\n<li>Two MLPs (4 and 9 hidden layers, the variants of <a href=\"https://www.kaggle.com/code/pourchot/all-in-one-citeseq-multiome-with-keras?scriptVersionId=108466116\" target=\"_blank\">Laurent Pourchot's model</a>)</li>\n<li>Conv1d (almost all the same as the <a href=\"https://www.kaggle.com/c/lish-moa/discussion/202256\" target=\"_blank\">tmp’s 1D-CNN model for tabular data on MoA competition</a>)</li>\n<li>LGBM</li>\n</ul>\n<p>In my case, stacking scheme boosted the score. The outputs of level 1 models were concatenated and then used as input for level 2. It was effective to apply dimensionality reduction to concatenated level 1 outputs after standardization. When the outputs were just concatenated without dimensionality reduction, the score of level 2 was rather lower than that of level 1.</p>\n<p>Also, ensemble worked well. I created several models with slight difference (different dimensionality reduction algorithms, feature extraction methods and loss functions) and blended them. In addition, each learning process was performed on 15 random-seeds and results were averaged.<br>\n<br></p>\n<h1>CV scheme:</h1>\n<p>I used simple KFold (k = 5). Fortunately, I resulted in shakeup in private LB.<br>\n<br></p>\n<h1>Citeseq:</h1>\n<p>The diagram of my Citeseq stacking scheme is as follow.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11204962%2F65d09e1ed96e210c63bab1d627b8497b%2Fsolution%20scheme-vstack.jpg?generation=1669194022322983&amp;alt=media\" alt=\"\"></p>\n<h4>Preprocess</h4>\n<p>Before dimensionality reduction or feature extraction in level 1, the set of all features which are constant in the train or test were eliminated according to the <a href=\"https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart\" target=\"_blank\">AmbrosM’s Code</a>.</p>\n<ul>\n<li>Dimensionality reduction: tSVD and PCA (n_components = 64) were used separately and inference results were finally blended.</li>\n<li>Feature extraction: According to the <a href=\"https://www.kaggle.com/code/fabiencrom/msci-correlations-eda-citeseq/notebook\" target=\"_blank\">Fabien Crom's Code</a>, I picked features with high Pearson correlation coefficients for the targets. In order to gain diversity as much as possible, I changed the picking query for each models. For example:<br>\n・Extract the RNAs in order of highest <strong>average</strong> of correlation to 140 proteins<br>\n・Extract the RNAs that have a high correlation value for a <strong>single</strong> protein, <strong>not the average</strong><br>\n・Change how many of the top RNAs are extracted<br>\n・With or without dimensionality reduction after extracted<br>\n<br></li>\n</ul>\n<h1>Multiome:</h1>\n<p>The scheme is almost the same as that of Citeseq. The differences are as follows:</p>\n<ul>\n<li>Extracted features were not used. In the Multiome case, the score was deteriorated when they were concatenated with 64 dims compressed features.</li>\n<li>For NN algorisms, only MSE was used as loss function.</li>\n<li>The target vectors were compressed to 512 (for NN) or 128 (for LGBM) dims by tSVD and used in training. The inferred vectors were decompressed to 23418 dims by inverse SVD.</li>\n</ul>",
      "rawMarkdown": "Thanks to the competition host, Kaggle team, Saturn cloud team and congrats to all the winners!\n\nAlthough many great solutions have already been posted and my solution may not contain new approaches, I would like to leave my efforts over the past two months. In this term I learned a lot. Thanks to the all competitors for a great game.\n\nPlease forgive me if it is difficult to read this post or see the schematic diagram below, as my command of English is not very good and my educational background was different from computer science or machine learning area.\n***\n# Overview:\nThe machine learning algorithms used were as follows:\n- Two MLPs (4 and 9 hidden layers, the variants of [Laurent Pourchot's model](https://www.kaggle.com/code/pourchot/all-in-one-citeseq-multiome-with-keras?scriptVersionId=108466116))\n- Conv1d (almost all the same as the [tmp’s 1D-CNN model for tabular data on MoA competition](https://www.kaggle.com/c/lish-moa/discussion/202256))\n- LGBM\n\nIn my case, stacking scheme boosted the score. The outputs of level 1 models were concatenated and then used as input for level 2. It was effective to apply dimensionality reduction to concatenated level 1 outputs after standardization. When the outputs were just concatenated without dimensionality reduction, the score of level 2 was rather lower than that of level 1.\n\nAlso, ensemble worked well. I created several models with slight difference (different dimensionality reduction algorithms, feature extraction methods and loss functions) and blended them. In addition, each learning process was performed on 15 random-seeds and results were averaged.\n<br>\n# CV scheme:\nI used simple KFold (k = 5). Fortunately, I resulted in shakeup in private LB.\n<br>\n# Citeseq:\nThe diagram of my Citeseq stacking scheme is as follow.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11204962%2F65d09e1ed96e210c63bab1d627b8497b%2Fsolution%20scheme-vstack.jpg?generation=1669194022322983&alt=media)\n\n#### Preprocess\nBefore dimensionality reduction or feature extraction in level 1, the set of all features which are constant in the train or test were eliminated according to the [AmbrosM’s Code](https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart).\n- Dimensionality reduction: tSVD and PCA (n_components = 64) were used separately and inference results were finally blended.\n- Feature extraction: According to the [Fabien Crom's Code](https://www.kaggle.com/code/fabiencrom/msci-correlations-eda-citeseq/notebook), I picked features with high Pearson correlation coefficients for the targets. In order to gain diversity as much as possible, I changed the picking query for each models. For example:\n・Extract the RNAs in order of highest **average** of correlation to 140 proteins\n・Extract the RNAs that have a high correlation value for a **single** protein, **not the average**\n・Change how many of the top RNAs are extracted\n・With or without dimensionality reduction after extracted\n<br>\n# Multiome:\nThe scheme is almost the same as that of Citeseq. The differences are as follows:\n- Extracted features were not used. In the Multiome case, the score was deteriorated when they were concatenated with 64 dims compressed features.\n- For NN algorisms, only MSE was used as loss function.\n- The target vectors were compressed to 512 (for NN) or 128 (for LGBM) dims by tSVD and used in training. The inferred vectors were decompressed to 23418 dims by inverse SVD.",
      "votes": null
    },
    {
      "id": "2041635",
      "postDate": "11/24/2022 05:16:30",
      "content": "<p>Congratulations and thanks for posting your solution AyPy!</p>",
      "rawMarkdown": "Congratulations and thanks for posting your solution AyPy!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2041635,
      "author_name": "songqizhou",
      "author_url": "",
      "post_date": "11/24/2022 05:16:30",
      "content": "<p>Congratulations and thanks for posting your solution AyPy!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2040678": "Thanks to the competition host, Kaggle team, Saturn cloud team and congrats to all the winners!\n\nAlthough many great solutions have already been posted and my solution may not contain new approaches, I would like to leave my efforts over the past two months. In this term I learned a lot. Thanks to the all competitors for a great game.\n\nPlease forgive me if it is difficult to read this post or see the schematic diagram below, as my command of English is not very good and my educational background was different from computer science or machine learning area.\n***\n# Overview:\nThe machine learning algorithms used were as follows:\n- Two MLPs (4 and 9 hidden layers, the variants of [Laurent Pourchot's model](https://www.kaggle.com/code/pourchot/all-in-one-citeseq-multiome-with-keras?scriptVersionId=108466116))\n- Conv1d (almost all the same as the [tmp’s 1D-CNN model for tabular data on MoA competition](https://www.kaggle.com/c/lish-moa/discussion/202256))\n- LGBM\n\nIn my case, stacking scheme boosted the score. The outputs of level 1 models were concatenated and then used as input for level 2. It was effective to apply dimensionality reduction to concatenated level 1 outputs after standardization. When the outputs were just concatenated without dimensionality reduction, the score of level 2 was rather lower than that of level 1.\n\nAlso, ensemble worked well. I created several models with slight difference (different dimensionality reduction algorithms, feature extraction methods and loss functions) and blended them. In addition, each learning process was performed on 15 random-seeds and results were averaged.\n<br>\n# CV scheme:\nI used simple KFold (k = 5). Fortunately, I resulted in shakeup in private LB.\n<br>\n# Citeseq:\nThe diagram of my Citeseq stacking scheme is as follow.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11204962%2F65d09e1ed96e210c63bab1d627b8497b%2Fsolution%20scheme-vstack.jpg?generation=1669194022322983&alt=media)\n\n#### Preprocess\nBefore dimensionality reduction or feature extraction in level 1, the set of all features which are constant in the train or test were eliminated according to the [AmbrosM’s Code](https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart).\n- Dimensionality reduction: tSVD and PCA (n_components = 64) were used separately and inference results were finally blended.\n- Feature extraction: According to the [Fabien Crom's Code](https://www.kaggle.com/code/fabiencrom/msci-correlations-eda-citeseq/notebook), I picked features with high Pearson correlation coefficients for the targets. In order to gain diversity as much as possible, I changed the picking query for each models. For example:\n・Extract the RNAs in order of highest **average** of correlation to 140 proteins\n・Extract the RNAs that have a high correlation value for a **single** protein, **not the average**\n・Change how many of the top RNAs are extracted\n・With or without dimensionality reduction after extracted\n<br>\n# Multiome:\nThe scheme is almost the same as that of Citeseq. The differences are as follows:\n- Extracted features were not used. In the Multiome case, the score was deteriorated when they were concatenated with 64 dims compressed features.\n- For NN algorisms, only MSE was used as loss function.\n- The target vectors were compressed to 512 (for NN) or 128 (for LGBM) dims by tSVD and used in training. The inferred vectors were decompressed to 23418 dims by inverse SVD.",
    "2041635": "Congratulations and thanks for posting your solution AyPy!"
  },
  "source": "meta"
}