{
  "id": 359683,
  "title": "Feature shop for you - lots of engineered features by dimensional reduction and clustering ",
  "url": "/competitions/open-problems-multimodal/discussion/359683",
  "author_name": "Alexander Chervov",
  "post_date": "2022-10-13T05:57:41.556000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>You may try to concatenate those features to your models, not guarantee it will improve score, but might be curious to try.</p>\n<p>Take a look on <a href=\"https://www.kaggle.com/datasets/alexandervc/feature-shop-for-multimodal-singlecell-competition\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/feature-shop-for-multimodal-singlecell-competition</a><br>\nVarious dimensional reductions algorithms (PCA, UMAP, ICA …) with various parameters were used.<br>\nSo the outcome - kind of engineered features - which can be concatenated to current features and might improve score (or might not). </p>\n<p>Since these algorithms are not fast on huge datasets, as a first step fast PCA (and TruncatedSVD) were used to reduce dimensions to 200 and 500. And only after that these algorithms were launched - so execution time is reasonable. <br>\nReduction by PCA/tSVD were done on Saturn cloud where we have 132 G RAM.  And files were downloaded to Kaggle. </p>\n<p>Files contain from 2 to 100 features - each file from different method+params. The name of file informs on method and params.  First train samples placed, later test samples </p>\n<p>Also clustering algorithm can used as feature engineering. <br>\nSome attempts seems show that  Kmeans can work, other fail crashing RAM or shows unreasonable results (according to my limited experience)</p>\n<p>PS<br>\nUse of dimensional reductions for visualization can be found here:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-eda-targets-citeseq-02#PCA-dimensional-reduction-and-visualizations\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-eda-targets-citeseq-02#PCA-dimensional-reduction-and-visualizations</a><br>\nand here: <a href=\"https://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo-targets-citeseq-01?scriptVersionId=106428078&amp;cellId=25\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo-targets-citeseq-01?scriptVersionId=106428078&amp;cellId=25</a></p>",
  "messages": [
    {
      "id": 1985133,
      "postDate": "2022-10-13T05:57:41.557Z",
      "content": "<p>You may try to concatenate those features to your models, not guarantee it will improve score, but might be curious to try.</p>\n<p>Take a look on <a href=\"https://www.kaggle.com/datasets/alexandervc/feature-shop-for-multimodal-singlecell-competition\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/feature-shop-for-multimodal-singlecell-competition</a><br>\nVarious dimensional reductions algorithms (PCA, UMAP, ICA …) with various parameters were used.<br>\nSo the outcome - kind of engineered features - which can be concatenated to current features and might improve score (or might not). </p>\n<p>Since these algorithms are not fast on huge datasets, as a first step fast PCA (and TruncatedSVD) were used to reduce dimensions to 200 and 500. And only after that these algorithms were launched - so execution time is reasonable. <br>\nReduction by PCA/tSVD were done on Saturn cloud where we have 132 G RAM.  And files were downloaded to Kaggle. </p>\n<p>Files contain from 2 to 100 features - each file from different method+params. The name of file informs on method and params.  First train samples placed, later test samples </p>\n<p>Also clustering algorithm can used as feature engineering. <br>\nSome attempts seems show that  Kmeans can work, other fail crashing RAM or shows unreasonable results (according to my limited experience)</p>\n<p>PS<br>\nUse of dimensional reductions for visualization can be found here:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-eda-targets-citeseq-02#PCA-dimensional-reduction-and-visualizations\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-eda-targets-citeseq-02#PCA-dimensional-reduction-and-visualizations</a><br>\nand here: <a href=\"https://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo-targets-citeseq-01?scriptVersionId=106428078&amp;cellId=25\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo-targets-citeseq-01?scriptVersionId=106428078&amp;cellId=25</a></p>",
      "rawMarkdown": "You may try to concatenate those features to your models, not guarantee it will improve score, but might be curious to try.\n\nTake a look on https://www.kaggle.com/datasets/alexandervc/feature-shop-for-multimodal-singlecell-competition\nVarious dimensional reductions algorithms (PCA, UMAP, ICA ...) with various parameters were used.\nSo the outcome - kind of engineered features - which can be concatenated to current features and might improve score (or might not). \n\nSince these algorithms are not fast on huge datasets, as a first step fast PCA (and TruncatedSVD) were used to reduce dimensions to 200 and 500. And only after that these algorithms were launched - so execution time is reasonable. \nReduction by PCA/tSVD were done on Saturn cloud where we have 132 G RAM.  And files were downloaded to Kaggle. \n\nFiles contain from 2 to 100 features - each file from different method+params. The name of file informs on method and params.  First train samples placed, later test samples \n\nAlso clustering algorithm can used as feature engineering. \nSome attempts seems show that  Kmeans can work, other fail crashing RAM or shows unreasonable results (according to my limited experience)\n\nPS\nUse of dimensional reductions for visualization can be found here:\nhttps://www.kaggle.com/code/alexandervc/mmscel-eda-targets-citeseq-02#PCA-dimensional-reduction-and-visualizations\nand here: https://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo-targets-citeseq-01?scriptVersionId=106428078&cellId=25",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1985133": "You may try to concatenate those features to your models, not guarantee it will improve score, but might be curious to try.\n\nTake a look on https://www.kaggle.com/datasets/alexandervc/feature-shop-for-multimodal-singlecell-competition\nVarious dimensional reductions algorithms (PCA, UMAP, ICA ...) with various parameters were used.\nSo the outcome - kind of engineered features - which can be concatenated to current features and might improve score (or might not). \n\nSince these algorithms are not fast on huge datasets, as a first step fast PCA (and TruncatedSVD) were used to reduce dimensions to 200 and 500. And only after that these algorithms were launched - so execution time is reasonable. \nReduction by PCA/tSVD were done on Saturn cloud where we have 132 G RAM.  And files were downloaded to Kaggle. \n\nFiles contain from 2 to 100 features - each file from different method+params. The name of file informs on method and params.  First train samples placed, later test samples \n\nAlso clustering algorithm can used as feature engineering. \nSome attempts seems show that  Kmeans can work, other fail crashing RAM or shows unreasonable results (according to my limited experience)\n\nPS\nUse of dimensional reductions for visualization can be found here:\nhttps://www.kaggle.com/code/alexandervc/mmscel-eda-targets-citeseq-02#PCA-dimensional-reduction-and-visualizations\nand here: https://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo-targets-citeseq-01?scriptVersionId=106428078&cellId=25"
  }
}