{
  "id": 357681,
  "title": "Strategies for reducing data size",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/357681",
  "author_name": "",
  "post_date": "2022-10-05T08:43:39.433033100Z",
  "votes": 9,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hallo</p>\n<p>I wish to share my way of reducing data size.</p>\n<p>I run the training code on my computer where I have 32Gb RAM. Nevertheless, it is not enough to parse the data using Pandas and Scikit Learn.</p>\n<p><strong>Dealing with missing values</strong></p>\n<p>For the missing values, I used either mean or mode, depending on the feature distribution (mode for skewed data and mean for the relatively normal distribution).<br>\nI run one imputer model for each table and saved it. For imputing, I loaded each model and run each of them on a table (train and test), then parsed a mean of the result. I saved the resulting tables in csv files.</p>\n<p><strong>Reducing dimensionality and size</strong></p>\n<p>First of all I run PCA and noticed that only 20 features are important. In fact, 20 features cover 90% of cumulative variance. So I run a PCA with 20 components on maximum of tables I could merge without getting an OOM error  -- that means, in my case, 6 tables (train from 0 to 5), and I saved the model with Pickle.</p>\n<p>I run this model on the preprocessed tables. After PCA, the values in features were between -1 and 1. Therefore, I changed the type of the features from <strong>float64</strong> to <strong>float16</strong>. I also changed the type of the targets to <strong>int8</strong>. The size of the tables was threefold less than before. I saved again the resulting files into csv files, and used them afterwards for classification.</p>",
  "messages": [
    {
      "id": "1972610",
      "postDate": "10/05/2022 08:43:39",
      "content": "<p>Hallo</p>\n<p>I wish to share my way of reducing data size.</p>\n<p>I run the training code on my computer where I have 32Gb RAM. Nevertheless, it is not enough to parse the data using Pandas and Scikit Learn.</p>\n<p><strong>Dealing with missing values</strong></p>\n<p>For the missing values, I used either mean or mode, depending on the feature distribution (mode for skewed data and mean for the relatively normal distribution).<br>\nI run one imputer model for each table and saved it. For imputing, I loaded each model and run each of them on a table (train and test), then parsed a mean of the result. I saved the resulting tables in csv files.</p>\n<p><strong>Reducing dimensionality and size</strong></p>\n<p>First of all I run PCA and noticed that only 20 features are important. In fact, 20 features cover 90% of cumulative variance. So I run a PCA with 20 components on maximum of tables I could merge without getting an OOM error  -- that means, in my case, 6 tables (train from 0 to 5), and I saved the model with Pickle.</p>\n<p>I run this model on the preprocessed tables. After PCA, the values in features were between -1 and 1. Therefore, I changed the type of the features from <strong>float64</strong> to <strong>float16</strong>. I also changed the type of the targets to <strong>int8</strong>. The size of the tables was threefold less than before. I saved again the resulting files into csv files, and used them afterwards for classification.</p>",
      "rawMarkdown": "Hallo\n\nI wish to share my way of reducing data size.\n\nI run the training code on my computer where I have 32Gb RAM. Nevertheless, it is not enough to parse the data using Pandas and Scikit Learn.\n\n**Dealing with missing values**\n\nFor the missing values, I used either mean or mode, depending on the feature distribution (mode for skewed data and mean for the relatively normal distribution).\nI run one imputer model for each table and saved it. For imputing, I loaded each model and run each of them on a table (train and test), then parsed a mean of the result. I saved the resulting tables in csv files.\n\n**Reducing dimensionality and size**\n\nFirst of all I run PCA and noticed that only 20 features are important. In fact, 20 features cover 90% of cumulative variance. So I run a PCA with 20 components on maximum of tables I could merge without getting an OOM error  -- that means, in my case, 6 tables (train from 0 to 5), and I saved the model with Pickle.\n\nI run this model on the preprocessed tables. After PCA, the values in features were between -1 and 1. Therefore, I changed the type of the features from **float64** to **float16**. I also changed the type of the targets to **int8**. The size of the tables was threefold less than before. I saved again the resulting files into csv files, and used them afterwards for classification.",
      "votes": null
    },
    {
      "id": "1973876",
      "postDate": "10/05/2022 23:18:31",
      "content": "<p>is it okay to run PCA on time dependent features?</p>",
      "rawMarkdown": "is it okay to run PCA on time dependent features?",
      "votes": null
    },
    {
      "id": "1973910",
      "postDate": "10/06/2022 00:09:11",
      "content": "<p>I considered them as tabular (check the title of the competition!) and run a ML classification on it, so I run PCA. I shall try also sequential algorithms, and treat them as sequences.</p>",
      "rawMarkdown": "I considered them as tabular (check the title of the competition!) and run a ML classification on it, so I run PCA. I shall try also sequential algorithms, and treat them as sequences.",
      "votes": null
    },
    {
      "id": "1978255",
      "postDate": "10/08/2022 16:31:43",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/catadanna\" target=\"_blank\">@catadanna</a>, thanks for sharing, it's  explained in the dataset descriptions that NaNs are due to players not being on the part of the game due to their vehicle being destroyed, so replacing them with mean or average values will probably decrease the performance of your model</p>",
      "rawMarkdown": "Hello @catadanna, thanks for sharing, it's  explained in the dataset descriptions that NaNs are due to players not being on the part of the game due to their vehicle being destroyed, so replacing them with mean or average values will probably decrease the performance of your model",
      "votes": null
    },
    {
      "id": "1978493",
      "postDate": "10/08/2022 18:26:12",
      "content": "<p>Thank you,, indeed, you are right. Replacing them with zeros is not relevant neither. I used a big negative value instead.</p>",
      "rawMarkdown": "Thank you,, indeed, you are right. Replacing them with zeros is not relevant neither. I used a big negative value instead.",
      "votes": null
    },
    {
      "id": "1978609",
      "postDate": "10/08/2022 19:54:18",
      "content": "<p>Try to build feature around them I been counting the total amount of NaNs as a feature</p>",
      "rawMarkdown": "Try to build feature around them I been counting the total amount of NaNs as a feature",
      "votes": null
    },
    {
      "id": "1978635",
      "postDate": "10/08/2022 20:45:59",
      "content": "<p>Indeed, that works well. I already have a function which does that, somewhere, I had to find it … </p>",
      "rawMarkdown": "Indeed, that works well. I already have a function which does that, somewhere, I had to find it ...",
      "votes": null
    },
    {
      "id": "1978677",
      "postDate": "10/08/2022 22:53:03",
      "content": "<p>def missing_values_feat(df):<br>\n    team_a = ['p0_pos_x','p1_pos_x','p2_pos_x']<br>\n    team_b = ['p3_pos_x','p4_pos_x','p5_pos_x']</p>\n<pre><code>df[\"missing_team_A_players\"] = df[team_a].isnull().sum(axis = 1)\ndf[\"missing_team_B_players\"] = df[team_b].isnull().sum(axis = 1)\n\nreturn df\n</code></pre>\n<p>trn_data = missing_values_feat(trn_data)<br>\ntst_data = missing_values_feat(tst_data)</p>",
      "rawMarkdown": "def missing_values_feat(df):\n    team_a = ['p0_pos_x','p1_pos_x','p2_pos_x']\n    team_b = ['p3_pos_x','p4_pos_x','p5_pos_x']\n    \n    df[\"missing_team_A_players\"] = df[team_a].isnull().sum(axis = 1)\n    df[\"missing_team_B_players\"] = df[team_b].isnull().sum(axis = 1)\n\n    return df\n\ntrn_data = missing_values_feat(trn_data)\ntst_data = missing_values_feat(tst_data)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1973876,
      "author_name": "peace2022",
      "author_url": "",
      "post_date": "10/05/2022 23:18:31",
      "content": "<p>is it okay to run PCA on time dependent features?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1973910,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "10/06/2022 00:09:11",
          "content": "<p>I considered them as tabular (check the title of the competition!) and run a ML classification on it, so I run PCA. I shall try also sequential algorithms, and treat them as sequences.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1978255,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "10/08/2022 16:31:43",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/catadanna\" target=\"_blank\">@catadanna</a>, thanks for sharing, it's  explained in the dataset descriptions that NaNs are due to players not being on the part of the game due to their vehicle being destroyed, so replacing them with mean or average values will probably decrease the performance of your model</p>",
      "votes": null,
      "replies": [
        {
          "id": 1978493,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "10/08/2022 18:26:12",
          "content": "<p>Thank you,, indeed, you are right. Replacing them with zeros is not relevant neither. I used a big negative value instead.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1978609,
          "author_name": "cv13j0",
          "author_url": "",
          "post_date": "10/08/2022 19:54:18",
          "content": "<p>Try to build feature around them I been counting the total amount of NaNs as a feature</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1978635,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "10/08/2022 20:45:59",
          "content": "<p>Indeed, that works well. I already have a function which does that, somewhere, I had to find it … </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1978677,
          "author_name": "cv13j0",
          "author_url": "",
          "post_date": "10/08/2022 22:53:03",
          "content": "<p>def missing_values_feat(df):<br>\n    team_a = ['p0_pos_x','p1_pos_x','p2_pos_x']<br>\n    team_b = ['p3_pos_x','p4_pos_x','p5_pos_x']</p>\n<pre><code>df[\"missing_team_A_players\"] = df[team_a].isnull().sum(axis = 1)\ndf[\"missing_team_B_players\"] = df[team_b].isnull().sum(axis = 1)\n\nreturn df\n</code></pre>\n<p>trn_data = missing_values_feat(trn_data)<br>\ntst_data = missing_values_feat(tst_data)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1972610": "Hallo\n\nI wish to share my way of reducing data size.\n\nI run the training code on my computer where I have 32Gb RAM. Nevertheless, it is not enough to parse the data using Pandas and Scikit Learn.\n\n**Dealing with missing values**\n\nFor the missing values, I used either mean or mode, depending on the feature distribution (mode for skewed data and mean for the relatively normal distribution).\nI run one imputer model for each table and saved it. For imputing, I loaded each model and run each of them on a table (train and test), then parsed a mean of the result. I saved the resulting tables in csv files.\n\n**Reducing dimensionality and size**\n\nFirst of all I run PCA and noticed that only 20 features are important. In fact, 20 features cover 90% of cumulative variance. So I run a PCA with 20 components on maximum of tables I could merge without getting an OOM error  -- that means, in my case, 6 tables (train from 0 to 5), and I saved the model with Pickle.\n\nI run this model on the preprocessed tables. After PCA, the values in features were between -1 and 1. Therefore, I changed the type of the features from **float64** to **float16**. I also changed the type of the targets to **int8**. The size of the tables was threefold less than before. I saved again the resulting files into csv files, and used them afterwards for classification.",
    "1973876": "is it okay to run PCA on time dependent features?",
    "1973910": "I considered them as tabular (check the title of the competition!) and run a ML classification on it, so I run PCA. I shall try also sequential algorithms, and treat them as sequences.",
    "1978255": "Hello @catadanna, thanks for sharing, it's  explained in the dataset descriptions that NaNs are due to players not being on the part of the game due to their vehicle being destroyed, so replacing them with mean or average values will probably decrease the performance of your model",
    "1978493": "Thank you,, indeed, you are right. Replacing them with zeros is not relevant neither. I used a big negative value instead.",
    "1978609": "Try to build feature around them I been counting the total amount of NaNs as a feature",
    "1978635": "Indeed, that works well. I already have a function which does that, somewhere, I had to find it ...",
    "1978677": "def missing_values_feat(df):\n    team_a = ['p0_pos_x','p1_pos_x','p2_pos_x']\n    team_b = ['p3_pos_x','p4_pos_x','p5_pos_x']\n    \n    df[\"missing_team_A_players\"] = df[team_a].isnull().sum(axis = 1)\n    df[\"missing_team_B_players\"] = df[team_b].isnull().sum(axis = 1)\n\n    return df\n\ntrn_data = missing_values_feat(trn_data)\ntst_data = missing_values_feat(tst_data)"
  },
  "source": "meta"
}