{
  "id": 368421,
  "title": "43rd Place : Summary and What Worked Well",
  "url": "/competitions/open-problems-multimodal/discussion/368421",
  "author_name": "tarick.morty",
  "post_date": "2022-11-25T09:37:06.987000",
  "votes": 15,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Many thanks to the organizers for this interesting multimodal challenge and giving us unique multi-target time series dataset to ideate upon.</p>\n<h3><strong>Data Preparation and Feature Pipeline</strong></h3>\n<ul>\n<li>TruncatedSVD of both cite and multi inputs - 100 components</li>\n<li>Binarized data TruncatedSVD of both cite and multi inputs - 100 components</li>\n<li>TruncatedSVD of multi targets - 256 components</li>\n<li>PCA of both cite and multi inputs - 40 components - used only in some of the models for additional features</li>\n<li>Most correlated raw features for individual cite targets</li>\n<li>Usage of '<em>Day</em>' as a feature</li>\n</ul>\n<h3><strong>CV Scheme</strong></h3>\n<ul>\n<li>GroupKFold by donor for both cite and multi - used with higher weightage in the final pipeline</li>\n<li>KFold for both cite and multi - since it was also correlated, kept it in the final pipeline with low weightage</li>\n</ul>\n<h3><strong>Modeling Pipeline</strong></h3>\n<ul>\n<li>MLPs with varied number of layers without binary components for both cite and multi (0.813 on public)</li>\n<li>MLPs with a mixture of both binary and non binary components for both cite and multi (0.813 on public)</li>\n<li>TabNet and LGBM model with dimensionality reduced cite data and multi (0.812 on public)</li>\n<li>Individual models with <em>highly correlated important features</em> per target with LGBM, XGB and CB for cite data (jump to 0.8142 on public, resulted in best ensemble)</li>\n</ul>\n<h4><strong>TakeAways</strong></h4>\n<ul>\n<li>pyBoost</li>\n<li>day similarity analysis (as <a href=\"https://www.kaggle.com/l0glikelihood\" target=\"_blank\">@l0glikelihood</a> trained only on day 7 for multiome)</li>\n</ul>\n<p>Would have expected to go upward with the shakeup, realized that other teams really did great and many congratulations to them. It was a wonderful competition, one of it's kind. </p>\n<p>Cheers!</p>",
  "messages": [
    {
      "id": 2043092,
      "postDate": "2022-11-25T09:37:06.987Z",
      "content": "<p>Many thanks to the organizers for this interesting multimodal challenge and giving us unique multi-target time series dataset to ideate upon.</p>\n<h3><strong>Data Preparation and Feature Pipeline</strong></h3>\n<ul>\n<li>TruncatedSVD of both cite and multi inputs - 100 components</li>\n<li>Binarized data TruncatedSVD of both cite and multi inputs - 100 components</li>\n<li>TruncatedSVD of multi targets - 256 components</li>\n<li>PCA of both cite and multi inputs - 40 components - used only in some of the models for additional features</li>\n<li>Most correlated raw features for individual cite targets</li>\n<li>Usage of '<em>Day</em>' as a feature</li>\n</ul>\n<h3><strong>CV Scheme</strong></h3>\n<ul>\n<li>GroupKFold by donor for both cite and multi - used with higher weightage in the final pipeline</li>\n<li>KFold for both cite and multi - since it was also correlated, kept it in the final pipeline with low weightage</li>\n</ul>\n<h3><strong>Modeling Pipeline</strong></h3>\n<ul>\n<li>MLPs with varied number of layers without binary components for both cite and multi (0.813 on public)</li>\n<li>MLPs with a mixture of both binary and non binary components for both cite and multi (0.813 on public)</li>\n<li>TabNet and LGBM model with dimensionality reduced cite data and multi (0.812 on public)</li>\n<li>Individual models with <em>highly correlated important features</em> per target with LGBM, XGB and CB for cite data (jump to 0.8142 on public, resulted in best ensemble)</li>\n</ul>\n<h4><strong>TakeAways</strong></h4>\n<ul>\n<li>pyBoost</li>\n<li>day similarity analysis (as <a href=\"https://www.kaggle.com/l0glikelihood\" target=\"_blank\">@l0glikelihood</a> trained only on day 7 for multiome)</li>\n</ul>\n<p>Would have expected to go upward with the shakeup, realized that other teams really did great and many congratulations to them. It was a wonderful competition, one of it's kind. </p>\n<p>Cheers!</p>",
      "rawMarkdown": "Many thanks to the organizers for this interesting multimodal challenge and giving us unique multi-target time series dataset to ideate upon.\n\n### **Data Preparation and Feature Pipeline**\n\n- TruncatedSVD of both cite and multi inputs - 100 components\n- Binarized data TruncatedSVD of both cite and multi inputs - 100 components\n- TruncatedSVD of multi targets - 256 components\n- PCA of both cite and multi inputs - 40 components - used only in some of the models for additional features\n- Most correlated raw features for individual cite targets\n- Usage of '*Day*' as a feature\n\n### **CV Scheme**\n\n- GroupKFold by donor for both cite and multi - used with higher weightage in the final pipeline\n- KFold for both cite and multi - since it was also correlated, kept it in the final pipeline with low weightage\n\n### **Modeling Pipeline**\n\n- MLPs with varied number of layers without binary components for both cite and multi (0.813 on public)\n- MLPs with a mixture of both binary and non binary components for both cite and multi (0.813 on public)\n- TabNet and LGBM model with dimensionality reduced cite data and multi (0.812 on public)\n- Individual models with *highly correlated important features* per target with LGBM, XGB and CB for cite data (jump to 0.8142 on public, resulted in best ensemble)\n\n#### **TakeAways**\n\n- pyBoost\n- day similarity analysis (as @l0glikelihood trained only on day 7 for multiome)\n\nWould have expected to go upward with the shakeup, realized that other teams really did great and many congratulations to them. It was a wonderful competition, one of it's kind. \n\nCheers!",
      "votes": 15
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2043092": "Many thanks to the organizers for this interesting multimodal challenge and giving us unique multi-target time series dataset to ideate upon.\n\n### **Data Preparation and Feature Pipeline**\n\n- TruncatedSVD of both cite and multi inputs - 100 components\n- Binarized data TruncatedSVD of both cite and multi inputs - 100 components\n- TruncatedSVD of multi targets - 256 components\n- PCA of both cite and multi inputs - 40 components - used only in some of the models for additional features\n- Most correlated raw features for individual cite targets\n- Usage of '*Day*' as a feature\n\n### **CV Scheme**\n\n- GroupKFold by donor for both cite and multi - used with higher weightage in the final pipeline\n- KFold for both cite and multi - since it was also correlated, kept it in the final pipeline with low weightage\n\n### **Modeling Pipeline**\n\n- MLPs with varied number of layers without binary components for both cite and multi (0.813 on public)\n- MLPs with a mixture of both binary and non binary components for both cite and multi (0.813 on public)\n- TabNet and LGBM model with dimensionality reduced cite data and multi (0.812 on public)\n- Individual models with *highly correlated important features* per target with LGBM, XGB and CB for cite data (jump to 0.8142 on public, resulted in best ensemble)\n\n#### **TakeAways**\n\n- pyBoost\n- day similarity analysis (as @l0glikelihood trained only on day 7 for multiome)\n\nWould have expected to go upward with the shakeup, realized that other teams really did great and many congratulations to them. It was a wonderful competition, one of it's kind. \n\nCheers!"
  }
}