{
  "id": 330432,
  "title": "Memory friendly dataset in Parquet",
  "url": "/competitions/amex-default-prediction/discussion/330432",
  "author_name": "Bartosz Mikulski",
  "post_date": "2022-06-12T09:59:12.019000",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I converted the dataset from CSV to parquet files, you can find it here <a href=\"https://www.kaggle.com/datasets/bartmiki/amex-parquet\" target=\"_blank\">https://www.kaggle.com/datasets/bartmiki/amex-parquet</a>.</p>\n<p>On 2022-06-22 I updated the dataset. The following information is for the newest version:</p>\n<p>Changes done between <code>.csv</code> and <code>.parquet</code> files:</p>\n<ul>\n<li>all float64 converted to float32</li>\n<li>all boolean converted to int8 (if there was a missing value I used -127 - minimal int8 value, as int8 does not support NAs - Int8 does, but it takes more space)</li>\n<li>converted all of the categorical columns (based on competition description) to int8:<ul>\n<li>D_63 and D_64 have string values hence I encoded them to numbers and provided mapping in <code>.json</code> files. In this case, missing values were given a custom category: M (as missing).</li>\n<li>other categorical that were numbers were converted to int8, with missing values filled with -127 (see above for a reason)</li></ul></li>\n<li>multiple float columns were converted to the uint8 as there were 0s and 1s with noise (see this excellent thread for explanation: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514)\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514)</a>. However, I didn't convert all columns mentioned in this thread. I only converted \"disguised\" boolean values not all integers in float format. Why? I need to dig more about these conversions and I don't want to share something that I do not understand. The logic can be found in a script attached to the dataset: <code>refine_dataset.disguised_boolean</code>.</li>\n<li>index is set to <code>customer_ID</code></li>\n<li>training set contains <code>target</code> as int8</li>\n<li><code>S_2</code> was converted to <code>datetime</code></li>\n</ul>\n<p>Training data info:</p>\n<pre><code>Index: 5531451 entries, 0000099d6bd597052cdcda90ffabf56573fe9d7c79be5fbac11a8ed792feb62a to fffff1d38b785cef84adeace64f8f83db3a0c31e8d92eaba8b115f71cab04681\nColumns: 190 entries, S_2 to target\ndtypes: datetime64[ns](1), float32(134), int8(55)\nmemory usage: 3.1+ GB\n</code></pre>\n<p>Testing data info:</p>\n<pre><code>Index: 11363762 entries, 00000469ba478561f23a92a868bd366de6f6527a684c9a2e78fb826dcac3b9b7 to fffffa7cf7e453e1acc6a1426475d5cb9400859f82ff61cceb803ea8ec37634d\nColumns: 189 entries, S_2 to D_145\ndtypes: datetime64[ns](1), float32(134), int8(54)\nmemory usage: 6.4+ GB\n</code></pre>",
  "messages": [
    {
      "id": 1818128,
      "postDate": "2022-06-12T09:59:12.020Z",
      "content": "<p>I converted the dataset from CSV to parquet files, you can find it here <a href=\"https://www.kaggle.com/datasets/bartmiki/amex-parquet\" target=\"_blank\">https://www.kaggle.com/datasets/bartmiki/amex-parquet</a>.</p>\n<p>On 2022-06-22 I updated the dataset. The following information is for the newest version:</p>\n<p>Changes done between <code>.csv</code> and <code>.parquet</code> files:</p>\n<ul>\n<li>all float64 converted to float32</li>\n<li>all boolean converted to int8 (if there was a missing value I used -127 - minimal int8 value, as int8 does not support NAs - Int8 does, but it takes more space)</li>\n<li>converted all of the categorical columns (based on competition description) to int8:<ul>\n<li>D_63 and D_64 have string values hence I encoded them to numbers and provided mapping in <code>.json</code> files. In this case, missing values were given a custom category: M (as missing).</li>\n<li>other categorical that were numbers were converted to int8, with missing values filled with -127 (see above for a reason)</li></ul></li>\n<li>multiple float columns were converted to the uint8 as there were 0s and 1s with noise (see this excellent thread for explanation: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514)\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514)</a>. However, I didn't convert all columns mentioned in this thread. I only converted \"disguised\" boolean values not all integers in float format. Why? I need to dig more about these conversions and I don't want to share something that I do not understand. The logic can be found in a script attached to the dataset: <code>refine_dataset.disguised_boolean</code>.</li>\n<li>index is set to <code>customer_ID</code></li>\n<li>training set contains <code>target</code> as int8</li>\n<li><code>S_2</code> was converted to <code>datetime</code></li>\n</ul>\n<p>Training data info:</p>\n<pre><code>Index: 5531451 entries, 0000099d6bd597052cdcda90ffabf56573fe9d7c79be5fbac11a8ed792feb62a to fffff1d38b785cef84adeace64f8f83db3a0c31e8d92eaba8b115f71cab04681\nColumns: 190 entries, S_2 to target\ndtypes: datetime64[ns](1), float32(134), int8(55)\nmemory usage: 3.1+ GB\n</code></pre>\n<p>Testing data info:</p>\n<pre><code>Index: 11363762 entries, 00000469ba478561f23a92a868bd366de6f6527a684c9a2e78fb826dcac3b9b7 to fffffa7cf7e453e1acc6a1426475d5cb9400859f82ff61cceb803ea8ec37634d\nColumns: 189 entries, S_2 to D_145\ndtypes: datetime64[ns](1), float32(134), int8(54)\nmemory usage: 6.4+ GB\n</code></pre>",
      "rawMarkdown": "I converted the dataset from CSV to parquet files, you can find it here https://www.kaggle.com/datasets/bartmiki/amex-parquet.\n\nOn 2022-06-22 I updated the dataset. The following information is for the newest version:\n\nChanges done between `.csv` and `.parquet` files:\n* all float64 converted to float32\n* all boolean converted to int8 (if there was a missing value I used -127 - minimal int8 value, as int8 does not support NAs - Int8 does, but it takes more space)\n* converted all of the categorical columns (based on competition description) to int8:\n  * D_63 and D_64 have string values hence I encoded them to numbers and provided mapping in `.json` files. In this case, missing values were given a custom category: M (as missing).\n  * other categorical that were numbers were converted to int8, with missing values filled with -127 (see above for a reason)\n* multiple float columns were converted to the uint8 as there were 0s and 1s with noise (see this excellent thread for explanation: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514). However, I didn't convert all columns mentioned in this thread. I only converted \"disguised\" boolean values not all integers in float format. Why? I need to dig more about these conversions and I don't want to share something that I do not understand. The logic can be found in a script attached to the dataset: `refine_dataset.disguised_boolean`.\n* index is set to `customer_ID`\n* training set contains `target` as int8\n* `S_2` was converted to `datetime`\n\nTraining data info:\n```\nIndex: 5531451 entries, 0000099d6bd597052cdcda90ffabf56573fe9d7c79be5fbac11a8ed792feb62a to fffff1d38b785cef84adeace64f8f83db3a0c31e8d92eaba8b115f71cab04681\nColumns: 190 entries, S_2 to target\ndtypes: datetime64[ns](1), float32(134), int8(55)\nmemory usage: 3.1+ GB\n```\n\nTesting data info:\n```\nIndex: 11363762 entries, 00000469ba478561f23a92a868bd366de6f6527a684c9a2e78fb826dcac3b9b7 to fffffa7cf7e453e1acc6a1426475d5cb9400859f82ff61cceb803ea8ec37634d\nColumns: 189 entries, S_2 to D_145\ndtypes: datetime64[ns](1), float32(134), int8(54)\nmemory usage: 6.4+ GB\n```",
      "votes": 4
    },
    {
      "id": 1818913,
      "postDate": "2022-06-13T09:20:48.063Z",
      "content": "<p>Very good initiative! Thanks for sharing</p>",
      "rawMarkdown": "Very good initiative! Thanks for sharing",
      "votes": 1
    },
    {
      "id": 1821647,
      "postDate": "2022-06-15T19:10:47.363Z",
      "content": "<p>Many thanks!</p>\n<p>For the next time…how did you convert the file from CSV to parquet, and save it? There isn't much info online…</p>",
      "rawMarkdown": "Many thanks!\n\nFor the next time...how did you convert the file from CSV to parquet, and save it? There isn't much info online...",
      "replies": [
        {
          "id": 1829654,
          "postDate": "2022-06-22T19:28:24.190Z",
          "content": "<p>I'll attach the required scripts to the dataset. However, I recommend having 32 GB of RAM (and no additional apps running in the background), especially for the test set (merging partial files to a single one). Or to have a bigger swap file. I should update the dataset in a moment.</p>",
          "rawMarkdown": "I'll attach the required scripts to the dataset. However, I recommend having 32 GB of RAM (and no additional apps running in the background), especially for the test set (merging partial files to a single one). Or to have a bigger swap file. I should update the dataset in a moment."
        }
      ]
    },
    {
      "id": 1821939,
      "postDate": "2022-06-16T00:02:38.563Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1829650,
          "postDate": "2022-06-22T19:24:23.910Z",
          "content": "<p>I'll address that in the future dataset version. About to drop a new version</p>",
          "rawMarkdown": "I'll address that in the future dataset version. About to drop a new version"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1818913,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-06-13T09:20:48.063000",
      "content": "<p>Very good initiative! Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1821647,
      "author_name": "Federico Trotta",
      "author_url": "",
      "post_date": "2022-06-15T19:10:47.363000",
      "content": "<p>Many thanks!</p>\n<p>For the next time…how did you convert the file from CSV to parquet, and save it? There isn't much info online…</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1829654,
          "author_name": "Bartosz Mikulski",
          "author_url": "",
          "post_date": "2022-06-22T19:28:24.190000",
          "content": "<p>I'll attach the required scripts to the dataset. However, I recommend having 32 GB of RAM (and no additional apps running in the background), especially for the test set (merging partial files to a single one). Or to have a bigger swap file. I should update the dataset in a moment.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1821939,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-16T00:02:38.563000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1829650,
          "author_name": "Bartosz Mikulski",
          "author_url": "",
          "post_date": "2022-06-22T19:24:23.910000",
          "content": "<p>I'll address that in the future dataset version. About to drop a new version</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1818128": "I converted the dataset from CSV to parquet files, you can find it here https://www.kaggle.com/datasets/bartmiki/amex-parquet.\n\nOn 2022-06-22 I updated the dataset. The following information is for the newest version:\n\nChanges done between `.csv` and `.parquet` files:\n* all float64 converted to float32\n* all boolean converted to int8 (if there was a missing value I used -127 - minimal int8 value, as int8 does not support NAs - Int8 does, but it takes more space)\n* converted all of the categorical columns (based on competition description) to int8:\n  * D_63 and D_64 have string values hence I encoded them to numbers and provided mapping in `.json` files. In this case, missing values were given a custom category: M (as missing).\n  * other categorical that were numbers were converted to int8, with missing values filled with -127 (see above for a reason)\n* multiple float columns were converted to the uint8 as there were 0s and 1s with noise (see this excellent thread for explanation: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514). However, I didn't convert all columns mentioned in this thread. I only converted \"disguised\" boolean values not all integers in float format. Why? I need to dig more about these conversions and I don't want to share something that I do not understand. The logic can be found in a script attached to the dataset: `refine_dataset.disguised_boolean`.\n* index is set to `customer_ID`\n* training set contains `target` as int8\n* `S_2` was converted to `datetime`\n\nTraining data info:\n```\nIndex: 5531451 entries, 0000099d6bd597052cdcda90ffabf56573fe9d7c79be5fbac11a8ed792feb62a to fffff1d38b785cef84adeace64f8f83db3a0c31e8d92eaba8b115f71cab04681\nColumns: 190 entries, S_2 to target\ndtypes: datetime64[ns](1), float32(134), int8(55)\nmemory usage: 3.1+ GB\n```\n\nTesting data info:\n```\nIndex: 11363762 entries, 00000469ba478561f23a92a868bd366de6f6527a684c9a2e78fb826dcac3b9b7 to fffffa7cf7e453e1acc6a1426475d5cb9400859f82ff61cceb803ea8ec37634d\nColumns: 189 entries, S_2 to D_145\ndtypes: datetime64[ns](1), float32(134), int8(54)\nmemory usage: 6.4+ GB\n```",
    "1818913": "Very good initiative! Thanks for sharing",
    "1821647": "Many thanks!\n\nFor the next time...how did you convert the file from CSV to parquet, and save it? There isn't much info online...",
    "1821939": ""
  }
}