{
  "id": 338905,
  "title": "Dimensionality Reduction?",
  "url": "/competitions/amex-default-prediction/discussion/338905",
  "author_name": "",
  "post_date": "2022-07-22T14:05:41.227707Z",
  "votes": 5,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Good day everyone,</p>\n<p>As for the size of the dataset, the basic feature selection algorithms and model building take too much time to execute. I am new to ml applications, and I wonder if there is any good strategy to reduce the dimensionality of the dataset or to boost the speed.</p>\n<p>Any advice is welcomed :)</p>",
  "messages": [
    {
      "id": "1866485",
      "postDate": "07/22/2022 14:05:41",
      "content": "<p>Good day everyone,</p>\n<p>As for the size of the dataset, the basic feature selection algorithms and model building take too much time to execute. I am new to ml applications, and I wonder if there is any good strategy to reduce the dimensionality of the dataset or to boost the speed.</p>\n<p>Any advice is welcomed :)</p>",
      "rawMarkdown": "Good day everyone,\n\nAs for the size of the dataset, the basic feature selection algorithms and model building take too much time to execute. I am new to ml applications, and I wonder if there is any good strategy to reduce the dimensionality of the dataset or to boost the speed.\n\nAny advice is welcomed :)",
      "votes": null
    },
    {
      "id": "1866503",
      "postDate": "07/22/2022 14:31:29",
      "content": "<p>This technique is often required in industry, in Kaggle competitions I feel that aiming to get high scores, all features should be kept, except those that negatively affect the result. As for your question, I generally use PCA and remove all features those are highly correlated with the others. Of course, there are many other methods, which needs to be selected depands on the specific data set.</p>",
      "rawMarkdown": "This technique is often required in industry, in Kaggle competitions I feel that aiming to get high scores, all features should be kept, except those that negatively affect the result. As for your question, I generally use PCA and remove all features those are highly correlated with the others. Of course, there are many other methods, which needs to be selected depands on the specific data set.",
      "votes": null
    },
    {
      "id": "1866579",
      "postDate": "07/22/2022 15:37:55",
      "content": "<p>Usually one could use PCA for this purpose. This is a very important and a crucial step for the associated model development. </p>",
      "rawMarkdown": "Usually one could use PCA for this purpose. This is a very important and a crucial step for the associated model development.",
      "votes": null
    },
    {
      "id": "1866583",
      "postDate": "07/22/2022 15:38:47",
      "content": "<p>Thanks so much for your advice! I will go with what you suggest. So, don't algorithms like Boruta add much value to the competition? </p>",
      "rawMarkdown": "Thanks so much for your advice! I will go with what you suggest. So, don't algorithms like Boruta add much value to the competition?",
      "votes": null
    },
    {
      "id": "1866586",
      "postDate": "07/22/2022 15:39:15",
      "content": "<p>Thanks so much for your advice! I will try this.</p>",
      "rawMarkdown": "Thanks so much for your advice! I will try this.",
      "votes": null
    },
    {
      "id": "1866808",
      "postDate": "07/22/2022 18:36:29",
      "content": "<p>I prefer using a denoising autoencoder for dimensionality reduction</p>",
      "rawMarkdown": "I prefer using a denoising autoencoder for dimensionality reduction",
      "votes": null
    },
    {
      "id": "1866829",
      "postDate": "07/22/2022 19:05:22",
      "content": "<p>That is a good idea, have you tried autoencoder in this competition yet? I wanted to do it with the raw data not the data after FE. I cannot however find a effective way.</p>",
      "rawMarkdown": "That is a good idea, have you tried autoencoder in this competition yet? I wanted to do it with the raw data not the data after FE. I cannot however find a effective way.",
      "votes": null
    },
    {
      "id": "1866909",
      "postDate": "07/22/2022 20:18:55",
      "content": "<p>my intuition: no, it won't help. I had some expriances that feature selection (by doing EDA) helps to improve the final scores in some kaggle competetions. Well, maybe you should have a try either use some algorithms or just through data analysis.</p>",
      "rawMarkdown": "my intuition: no, it won't help. I had some expriances that feature selection (by doing EDA) helps to improve the final scores in some kaggle competetions. Well, maybe you should have a try either use some algorithms or just through data analysis.",
      "votes": null
    },
    {
      "id": "1868042",
      "postDate": "07/23/2022 16:40:40",
      "content": "<p><a href=\"https://www.kaggle.com/meli19\" target=\"_blank\">@meli19</a> What threshold did u use to remove highly correlated features ? </p>",
      "rawMarkdown": "meli19 What threshold did u use to remove highly correlated features ?",
      "votes": null
    },
    {
      "id": "1869451",
      "postDate": "07/24/2022 18:48:06",
      "content": "<p>No I have not tried it yet. But dimensionality reduction usually improves the training/inference speed as well as decreases memory. But it does not often lead to a performance boost unless it removes noisy features. Definitely necessary on limited resources especially when computing on the edge.</p>",
      "rawMarkdown": "No I have not tried it yet. But dimensionality reduction usually improves the training/inference speed as well as decreases memory. But it does not often lead to a performance boost unless it removes noisy features. Definitely necessary on limited resources especially when computing on the edge.",
      "votes": null
    },
    {
      "id": "1870003",
      "postDate": "07/25/2022 07:32:57",
      "content": "<p>There is no fixed value, and it needs to be selected according to the specific data set. Sometimes 0.85 is a good value, sometimes 0.95+ is needed so you can make sure that the prediction accuracy remains.</p>",
      "rawMarkdown": "There is no fixed value, and it needs to be selected according to the specific data set. Sometimes 0.85 is a good value, sometimes 0.95+ is needed so you can make sure that the prediction accuracy remains.",
      "votes": null
    },
    {
      "id": "1870094",
      "postDate": "07/25/2022 09:31:17",
      "content": "<p><a href=\"https://www.kaggle.com/meli19\" target=\"_blank\">@meli19</a> Have you tried autoencoders yet? What did you experience?</p>",
      "rawMarkdown": "meli19 Have you tried autoencoders yet? What did you experience?",
      "votes": null
    },
    {
      "id": "1879558",
      "postDate": "08/01/2022 05:59:54",
      "content": "<p>Have you used PCA on this competition? I've tried to do so in the whole training set (using IncrementalPCA from cuml, RAPIDS) but, when the total number of features (after FE) is more than ~200, the PCA Solver breaks because of memory issues.</p>",
      "rawMarkdown": "Have you used PCA on this competition? I've tried to do so in the whole training set (using IncrementalPCA from cuml, RAPIDS) but, when the total number of features (after FE) is more than ~200, the PCA Solver breaks because of memory issues.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1866503,
      "author_name": "meli19",
      "author_url": "",
      "post_date": "07/22/2022 14:31:29",
      "content": "<p>This technique is often required in industry, in Kaggle competitions I feel that aiming to get high scores, all features should be kept, except those that negatively affect the result. As for your question, I generally use PCA and remove all features those are highly correlated with the others. Of course, there are many other methods, which needs to be selected depands on the specific data set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1866583,
          "author_name": "anaraghazada",
          "author_url": "",
          "post_date": "07/22/2022 15:38:47",
          "content": "<p>Thanks so much for your advice! I will go with what you suggest. So, don't algorithms like Boruta add much value to the competition? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1866909,
          "author_name": "meli19",
          "author_url": "",
          "post_date": "07/22/2022 20:18:55",
          "content": "<p>my intuition: no, it won't help. I had some expriances that feature selection (by doing EDA) helps to improve the final scores in some kaggle competetions. Well, maybe you should have a try either use some algorithms or just through data analysis.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1868042,
          "author_name": "kushal1506",
          "author_url": "",
          "post_date": "07/23/2022 16:40:40",
          "content": "<p><a href=\"https://www.kaggle.com/meli19\" target=\"_blank\">@meli19</a> What threshold did u use to remove highly correlated features ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1870003,
          "author_name": "meli19",
          "author_url": "",
          "post_date": "07/25/2022 07:32:57",
          "content": "<p>There is no fixed value, and it needs to be selected according to the specific data set. Sometimes 0.85 is a good value, sometimes 0.95+ is needed so you can make sure that the prediction accuracy remains.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1866579,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "07/22/2022 15:37:55",
      "content": "<p>Usually one could use PCA for this purpose. This is a very important and a crucial step for the associated model development. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1866586,
          "author_name": "anaraghazada",
          "author_url": "",
          "post_date": "07/22/2022 15:39:15",
          "content": "<p>Thanks so much for your advice! I will try this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1879558,
          "author_name": "carlosasdesouza",
          "author_url": "",
          "post_date": "08/01/2022 05:59:54",
          "content": "<p>Have you used PCA on this competition? I've tried to do so in the whole training set (using IncrementalPCA from cuml, RAPIDS) but, when the total number of features (after FE) is more than ~200, the PCA Solver breaks because of memory issues.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1866808,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "07/22/2022 18:36:29",
      "content": "<p>I prefer using a denoising autoencoder for dimensionality reduction</p>",
      "votes": null,
      "replies": [
        {
          "id": 1866829,
          "author_name": "meli19",
          "author_url": "",
          "post_date": "07/22/2022 19:05:22",
          "content": "<p>That is a good idea, have you tried autoencoder in this competition yet? I wanted to do it with the raw data not the data after FE. I cannot however find a effective way.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1869451,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "07/24/2022 18:48:06",
          "content": "<p>No I have not tried it yet. But dimensionality reduction usually improves the training/inference speed as well as decreases memory. But it does not often lead to a performance boost unless it removes noisy features. Definitely necessary on limited resources especially when computing on the edge.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1870094,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "07/25/2022 09:31:17",
          "content": "<p><a href=\"https://www.kaggle.com/meli19\" target=\"_blank\">@meli19</a> Have you tried autoencoders yet? What did you experience?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1866485": "Good day everyone,\n\nAs for the size of the dataset, the basic feature selection algorithms and model building take too much time to execute. I am new to ml applications, and I wonder if there is any good strategy to reduce the dimensionality of the dataset or to boost the speed.\n\nAny advice is welcomed :)",
    "1866503": "This technique is often required in industry, in Kaggle competitions I feel that aiming to get high scores, all features should be kept, except those that negatively affect the result. As for your question, I generally use PCA and remove all features those are highly correlated with the others. Of course, there are many other methods, which needs to be selected depands on the specific data set.",
    "1866579": "Usually one could use PCA for this purpose. This is a very important and a crucial step for the associated model development.",
    "1866583": "Thanks so much for your advice! I will go with what you suggest. So, don't algorithms like Boruta add much value to the competition?",
    "1866586": "Thanks so much for your advice! I will try this.",
    "1866808": "I prefer using a denoising autoencoder for dimensionality reduction",
    "1866829": "That is a good idea, have you tried autoencoder in this competition yet? I wanted to do it with the raw data not the data after FE. I cannot however find a effective way.",
    "1866909": "my intuition: no, it won't help. I had some expriances that feature selection (by doing EDA) helps to improve the final scores in some kaggle competetions. Well, maybe you should have a try either use some algorithms or just through data analysis.",
    "1868042": "meli19 What threshold did u use to remove highly correlated features ?",
    "1869451": "No I have not tried it yet. But dimensionality reduction usually improves the training/inference speed as well as decreases memory. But it does not often lead to a performance boost unless it removes noisy features. Definitely necessary on limited resources especially when computing on the edge.",
    "1870003": "There is no fixed value, and it needs to be selected according to the specific data set. Sometimes 0.85 is a good value, sometimes 0.95+ is needed so you can make sure that the prediction accuracy remains.",
    "1870094": "meli19 Have you tried autoencoders yet? What did you experience?",
    "1879558": "Have you used PCA on this competition? I've tried to do so in the whole training set (using IncrementalPCA from cuml, RAPIDS) but, when the total number of features (after FE) is more than ~200, the PCA Solver breaks because of memory issues."
  },
  "source": "meta"
}