{
  "id": 327908,
  "title": "American Express Default Prediction Snappy Parquet Files (50 GB -> 10 GB)",
  "url": "/competitions/amex-default-prediction/discussion/327908",
  "author_name": "",
  "post_date": "2022-05-29T21:32:05.437789400Z",
  "votes": 9,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi everyone, this is my first dataset submission, so I hope you'll provide some feedback on my methodology! </p>\n<p>I have created snappy parquet versions of:</p>\n<ul>\n<li>train_data.csv</li>\n<li>test_data.csv</li>\n<li>train_labels.csv</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>File Name</th>\n<th>CSV Disk Size</th>\n<th>Snappy Parquet Disk Size</th>\n<th>Pandas In-Memory Size</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>test_data</code></td>\n<td>32GB</td>\n<td>6.5GB</td>\n<td>8.5GB</td>\n</tr>\n<tr>\n<td><code>train_data</code></td>\n<td>15GB</td>\n<td>3.2GB</td>\n<td>4.1GB</td>\n</tr>\n<tr>\n<td><code>train_labels</code></td>\n<td>29MB</td>\n<td>27MB</td>\n<td>7.0MB</td>\n</tr>\n</tbody>\n</table>\n<p>Please see the link to my dataset <a href=\"https://www.kaggle.com/datasets/josephcorrado76/american-express-default-prediction-snappy-parquet\" target=\"_blank\">here</a>.</p>\n<p>My plan is to continue to update and improve the dataset over time. I will peruse the EDA and featurizations other people are performing, and add those as featurized versions over time. <br>\nI also plan to include notebooks showing interesting/useful EDA that other people are doing under the dataset. </p>\n<p>Please note that, for now, all I've done is:</p>\n<ul>\n<li>convert <code>customer_ID</code> to a proper <code>string</code></li>\n<li>converted <code>S_2</code> to a datetime type</li>\n<li>converted all categorical columns provided on the competition data documentation page to <code>category</code></li>\n<li>converted all <code>float64</code> columns to <code>float32</code> to get the size down and processing speed up. I am still debating if there is sufficient justification to convert to <code>float16</code></li>\n</ul>\n<p>I am looking for feedback, suggestions on new features to include into the dataset and suggestions for EDA/visualization to publish along with the data. </p>\n<p>Thank you!</p>",
  "messages": [
    {
      "id": "1805174",
      "postDate": "05/29/2022 21:32:05",
      "content": "<p>Hi everyone, this is my first dataset submission, so I hope you'll provide some feedback on my methodology! </p>\n<p>I have created snappy parquet versions of:</p>\n<ul>\n<li>train_data.csv</li>\n<li>test_data.csv</li>\n<li>train_labels.csv</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>File Name</th>\n<th>CSV Disk Size</th>\n<th>Snappy Parquet Disk Size</th>\n<th>Pandas In-Memory Size</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>test_data</code></td>\n<td>32GB</td>\n<td>6.5GB</td>\n<td>8.5GB</td>\n</tr>\n<tr>\n<td><code>train_data</code></td>\n<td>15GB</td>\n<td>3.2GB</td>\n<td>4.1GB</td>\n</tr>\n<tr>\n<td><code>train_labels</code></td>\n<td>29MB</td>\n<td>27MB</td>\n<td>7.0MB</td>\n</tr>\n</tbody>\n</table>\n<p>Please see the link to my dataset <a href=\"https://www.kaggle.com/datasets/josephcorrado76/american-express-default-prediction-snappy-parquet\" target=\"_blank\">here</a>.</p>\n<p>My plan is to continue to update and improve the dataset over time. I will peruse the EDA and featurizations other people are performing, and add those as featurized versions over time. <br>\nI also plan to include notebooks showing interesting/useful EDA that other people are doing under the dataset. </p>\n<p>Please note that, for now, all I've done is:</p>\n<ul>\n<li>convert <code>customer_ID</code> to a proper <code>string</code></li>\n<li>converted <code>S_2</code> to a datetime type</li>\n<li>converted all categorical columns provided on the competition data documentation page to <code>category</code></li>\n<li>converted all <code>float64</code> columns to <code>float32</code> to get the size down and processing speed up. I am still debating if there is sufficient justification to convert to <code>float16</code></li>\n</ul>\n<p>I am looking for feedback, suggestions on new features to include into the dataset and suggestions for EDA/visualization to publish along with the data. </p>\n<p>Thank you!</p>",
      "rawMarkdown": "Hi everyone, this is my first dataset submission, so I hope you'll provide some feedback on my methodology! \n\nI have created snappy parquet versions of:\n* train_data.csv\n* test_data.csv\n* train_labels.csv\n\n|File Name|CSV Disk Size|Snappy Parquet Disk Size|Pandas In-Memory Size|\n|---|---|---|---|\n|`test_data`|32GB|6.5GB|8.5GB|\n|`train_data`|15GB|3.2GB|4.1GB|\n|`train_labels`|29MB|27MB|7.0MB|\n\n\nPlease see the link to my dataset [here](https://www.kaggle.com/datasets/josephcorrado76/american-express-default-prediction-snappy-parquet).\n\nMy plan is to continue to update and improve the dataset over time. I will peruse the EDA and featurizations other people are performing, and add those as featurized versions over time. \nI also plan to include notebooks showing interesting/useful EDA that other people are doing under the dataset. \n\nPlease note that, for now, all I've done is:\n* convert `customer_ID` to a proper `string`\n* converted `S_2` to a datetime type\n* converted all categorical columns provided on the competition data documentation page to `category`\n* converted all `float64` columns to `float32` to get the size down and processing speed up. I am still debating if there is sufficient justification to convert to `float16`\n\nI am looking for feedback, suggestions on new features to include into the dataset and suggestions for EDA/visualization to publish along with the data. \n\nThank you!",
      "votes": null
    },
    {
      "id": "1805186",
      "postDate": "05/29/2022 22:11:45",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/josephcorrado76\" target=\"_blank\">@josephcorrado76</a>, thanks for the detailed explanation of the transformation used. I'm planning to run my own local models and I will take some ideas to improve memory efficiencty</p>",
      "rawMarkdown": "Hello @josephcorrado76, thanks for the detailed explanation of the transformation used. I'm planning to run my own local models and I will take some ideas to improve memory efficiencty",
      "votes": null
    },
    {
      "id": "1805785",
      "postDate": "05/30/2022 13:49:28",
      "content": "<p>Hey Carl! No problem! For memory efficiency, without dropping columns and knowing anything additional about the features, the transformations I've done here have reduced the in memory footprint as a pandas dataframe to about 12 GB (which is still too large for the kaggle notebooks). I think you still would need to only have the training or testing data in memory at one time. If you just load my data files as is, you should get similar optimizations. </p>",
      "rawMarkdown": "Hey Carl! No problem! For memory efficiency, without dropping columns and knowing anything additional about the features, the transformations I've done here have reduced the in memory footprint as a pandas dataframe to about 12 GB (which is still too large for the kaggle notebooks). I think you still would need to only have the training or testing data in memory at one time. If you just load my data files as is, you should get similar optimizations.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1805186,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "05/29/2022 22:11:45",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/josephcorrado76\" target=\"_blank\">@josephcorrado76</a>, thanks for the detailed explanation of the transformation used. I'm planning to run my own local models and I will take some ideas to improve memory efficiencty</p>",
      "votes": null,
      "replies": [
        {
          "id": 1805785,
          "author_name": "josephcorrado76",
          "author_url": "",
          "post_date": "05/30/2022 13:49:28",
          "content": "<p>Hey Carl! No problem! For memory efficiency, without dropping columns and knowing anything additional about the features, the transformations I've done here have reduced the in memory footprint as a pandas dataframe to about 12 GB (which is still too large for the kaggle notebooks). I think you still would need to only have the training or testing data in memory at one time. If you just load my data files as is, you should get similar optimizations. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1805174": "Hi everyone, this is my first dataset submission, so I hope you'll provide some feedback on my methodology! \n\nI have created snappy parquet versions of:\n* train_data.csv\n* test_data.csv\n* train_labels.csv\n\n|File Name|CSV Disk Size|Snappy Parquet Disk Size|Pandas In-Memory Size|\n|---|---|---|---|\n|`test_data`|32GB|6.5GB|8.5GB|\n|`train_data`|15GB|3.2GB|4.1GB|\n|`train_labels`|29MB|27MB|7.0MB|\n\n\nPlease see the link to my dataset [here](https://www.kaggle.com/datasets/josephcorrado76/american-express-default-prediction-snappy-parquet).\n\nMy plan is to continue to update and improve the dataset over time. I will peruse the EDA and featurizations other people are performing, and add those as featurized versions over time. \nI also plan to include notebooks showing interesting/useful EDA that other people are doing under the dataset. \n\nPlease note that, for now, all I've done is:\n* convert `customer_ID` to a proper `string`\n* converted `S_2` to a datetime type\n* converted all categorical columns provided on the competition data documentation page to `category`\n* converted all `float64` columns to `float32` to get the size down and processing speed up. I am still debating if there is sufficient justification to convert to `float16`\n\nI am looking for feedback, suggestions on new features to include into the dataset and suggestions for EDA/visualization to publish along with the data. \n\nThank you!",
    "1805186": "Hello @josephcorrado76, thanks for the detailed explanation of the transformation used. I'm planning to run my own local models and I will take some ideas to improve memory efficiencty",
    "1805785": "Hey Carl! No problem! For memory efficiency, without dropping columns and knowing anything additional about the features, the transformations I've done here have reduced the in memory footprint as a pandas dataframe to about 12 GB (which is still too large for the kaggle notebooks). I think you still would need to only have the training or testing data in memory at one time. If you just load my data files as is, you should get similar optimizations."
  },
  "source": "meta"
}