{
  "id": 327268,
  "title": "Compressed Dataset with targets (~10x compression) & Notebook",
  "url": "/competitions/amex-default-prediction/discussion/327268",
  "author_name": "Munum",
  "post_date": "2022-05-26T12:59:41.505000",
  "votes": 20,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I have seen a few compressed datasets be made so far but I have found them to be incomplete. Either they are too slow to load in or do not contain the training targets. To solve this I have made my own version of the dataset <a href=\"https://www.kaggle.com/datasets/munumbutt/amexfeather\" target=\"_blank\">here</a></p>\n<p>I have converted numerical columns to <code>float16</code>, categorical columns to <code>category</code>, the timestamp to <code>pd.datetime</code> and appended the target to the training data to make a single dataframe, ready for modelling.</p>\n<p>The data is in feather format which is extremely fast to load in. The train files are now 1.73GB and testing 3.55GB.</p>\n<p>Here is an end-to-end <a href=\"https://www.kaggle.com/code/munumbutt/simple-lgbm-starter\" target=\"_blank\">notebook </a> for using this data on Kaggle, without OOM error, using LightGBM</p>",
  "messages": [
    {
      "id": 1802105,
      "postDate": "2022-05-26T12:59:41.507Z",
      "content": "<p>I have seen a few compressed datasets be made so far but I have found them to be incomplete. Either they are too slow to load in or do not contain the training targets. To solve this I have made my own version of the dataset <a href=\"https://www.kaggle.com/datasets/munumbutt/amexfeather\" target=\"_blank\">here</a></p>\n<p>I have converted numerical columns to <code>float16</code>, categorical columns to <code>category</code>, the timestamp to <code>pd.datetime</code> and appended the target to the training data to make a single dataframe, ready for modelling.</p>\n<p>The data is in feather format which is extremely fast to load in. The train files are now 1.73GB and testing 3.55GB.</p>\n<p>Here is an end-to-end <a href=\"https://www.kaggle.com/code/munumbutt/simple-lgbm-starter\" target=\"_blank\">notebook </a> for using this data on Kaggle, without OOM error, using LightGBM</p>",
      "rawMarkdown": "I have seen a few compressed datasets be made so far but I have found them to be incomplete. Either they are too slow to load in or do not contain the training targets. To solve this I have made my own version of the dataset [here](https://www.kaggle.com/datasets/munumbutt/amexfeather)\n\nI have converted numerical columns to `float16`, categorical columns to `category`, the timestamp to `pd.datetime` and appended the target to the training data to make a single dataframe, ready for modelling.\n\nThe data is in feather format which is extremely fast to load in. The train files are now 1.73GB and testing 3.55GB.\n\nHere is an end-to-end [notebook ](https://www.kaggle.com/code/munumbutt/simple-lgbm-starter) for using this data on Kaggle, without OOM error, using LightGBM",
      "votes": 19
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1802105": "I have seen a few compressed datasets be made so far but I have found them to be incomplete. Either they are too slow to load in or do not contain the training targets. To solve this I have made my own version of the dataset [here](https://www.kaggle.com/datasets/munumbutt/amexfeather)\n\nI have converted numerical columns to `float16`, categorical columns to `category`, the timestamp to `pd.datetime` and appended the target to the training data to make a single dataframe, ready for modelling.\n\nThe data is in feather format which is extremely fast to load in. The train files are now 1.73GB and testing 3.55GB.\n\nHere is an end-to-end [notebook ](https://www.kaggle.com/code/munumbutt/simple-lgbm-starter) for using this data on Kaggle, without OOM error, using LightGBM"
  }
}