{
  "id": 327612,
  "title": "Which format of data best ! Used in large amount of data?",
  "url": "/competitions/amex-default-prediction/discussion/327612",
  "author_name": "",
  "post_date": "2022-05-28T05:31:49.635769200Z",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>This competition data is very large in my experience of approach!</p>\n<blockquote>\n  <p><strong>which one is best! parquet or pickle?</strong></p>\n</blockquote>\n<p>Every time memory is full and restarts my notebook!</p>\n<blockquote>\n  <p><strong>Any buddy suggest! How to effectively approach</strong></p>\n</blockquote>",
  "messages": [
    {
      "id": "1803739",
      "postDate": "05/28/2022 05:31:49",
      "content": "<p>This competition data is very large in my experience of approach!</p>\n<blockquote>\n  <p><strong>which one is best! parquet or pickle?</strong></p>\n</blockquote>\n<p>Every time memory is full and restarts my notebook!</p>\n<blockquote>\n  <p><strong>Any buddy suggest! How to effectively approach</strong></p>\n</blockquote>",
      "rawMarkdown": "This competition data is very large in my experience of approach!\n\n> **which one is best! parquet or pickle?**\n\nEvery time memory is full and restarts my notebook!\n\n> **Any buddy suggest! How to effectively approach**",
      "votes": null
    },
    {
      "id": "1803756",
      "postDate": "05/28/2022 06:28:38",
      "content": "<p>Hi Venkatkumar,</p>\n<p>If you try using Feather format, it works just fine, both train and test together without memory issues.<br>\nAnd to make it easier, someone has already uploaded the dataset in Feather format.</p>\n<p>Regards,<br>\nAnurag</p>",
      "rawMarkdown": "Hi Venkatkumar,\n\nIf you try using Feather format, it works just fine, both train and test together without memory issues.\nAnd to make it easier, someone has already uploaded the dataset in Feather format.\n\nRegards,\nAnurag",
      "votes": null
    },
    {
      "id": "1803767",
      "postDate": "05/28/2022 06:48:31",
      "content": "<p>I went with parquet and feather both work fines and takes less space. And I actually used both for train data -&gt; parquet and for test data -&gt; feather, because by using both parquet files I was getting OOM. And in feather file that I used didn't include \"target\" column so went with both formats.</p>\n<p>Do tell me if there's some problem with using different formats.</p>",
      "rawMarkdown": "I went with parquet and feather both work fines and takes less space. And I actually used both for train data -> parquet and for test data -> feather, because by using both parquet files I was getting OOM. And in feather file that I used didn't include \"target\" column so went with both formats.\n\nDo tell me if there's some problem with using different formats.",
      "votes": null
    },
    {
      "id": "1803783",
      "postDate": "05/28/2022 07:21:47",
      "content": "<p>Yes, buddy! I acceptable<br>\nI didn't try .far format,  I will try</p>\n<p>I was trying the parquet format! it is good but occupy the more memory using different model approach</p>\n<p>Thanks for sharing your experience buddy <a href=\"https://www.kaggle.com/devkhant24\" target=\"_blank\">@devkhant24</a> </p>",
      "rawMarkdown": "Yes, buddy! I acceptable\nI didn't try .far format,  I will try\n\nI was trying the parquet format! it is good but occupy the more memory using different model approach\n\nThanks for sharing your experience buddy @devkhant24",
      "votes": null
    },
    {
      "id": "1803784",
      "postDate": "05/28/2022 07:22:40",
      "content": "<p>Yes! i definitely try this method<br>\nThanks for sharing your experience</p>",
      "rawMarkdown": "Yes! i definitely try this method\nThanks for sharing your experience",
      "votes": null
    },
    {
      "id": "1803859",
      "postDate": "05/28/2022 08:54:17",
      "content": "<p>An awesome <a href=\"https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets/notebook\" target=\"_blank\">Tutorial on reading large datasets</a> by <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>  !</p>",
      "rawMarkdown": "An awesome [Tutorial on reading large datasets](https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets/notebook) by @rohanrao  !",
      "votes": null
    },
    {
      "id": "1803931",
      "postDate": "05/28/2022 10:42:50",
      "content": "<p>In addition to all the file type suggestions, note that you can also save the data as multiple files. For example, there are <code>458913</code> unique customers in train. For example, you can create 10 files where each file has <code>50_000</code> customers. Then when you do feature engineering, you can load one file at a time and engineer then save the result. Then load another etc. </p>\n<p>Using small chunks of data will allow you to do your feature engineering on GPU which is much faster than CPU. And reading and writing to disk from GPU is often faster than reading and writing from CPU. When you are ready to train, most models allow training from chunked data. Load chunk (one or more to fill memory), shuffle chunk(s), create batches within and train. Load another chunk (one or more), etc.</p>\n<p>Of course whatever file format or files you use, you should reduce the memory of each column. For example, use <code>float32</code> and <code>int32</code> instead of <code>float64</code> and <code>int64</code>. And convert the <code>customer_ID</code> column into <code>int32</code> or <code>int64</code> after label encoding. Also the date column can be compressed into three columns of <code>int8</code> for year, month, day.</p>",
      "rawMarkdown": "In addition to all the file type suggestions, note that you can also save the data as multiple files. For example, there are `458913` unique customers in train. For example, you can create 10 files where each file has `50_000` customers. Then when you do feature engineering, you can load one file at a time and engineer then save the result. Then load another etc. \n\nUsing small chunks of data will allow you to do your feature engineering on GPU which is much faster than CPU. And reading and writing to disk from GPU is often faster than reading and writing from CPU. When you are ready to train, most models allow training from chunked data. Load chunk (one or more to fill memory), shuffle chunk(s), create batches within and train. Load another chunk (one or more), etc.\n\nOf course whatever file format or files you use, you should reduce the memory of each column. For example, use `float32` and `int32` instead of `float64` and `int64`. And convert the `customer_ID` column into `int32` or `int64` after label encoding. Also the date column can be compressed into three columns of `int8` for year, month, day.",
      "votes": null
    },
    {
      "id": "1804008",
      "postDate": "05/28/2022 12:35:49",
      "content": "<p>Thankyou buddy! <a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> 👍</p>",
      "rawMarkdown": "Thankyou buddy! @steubk 👍",
      "votes": null
    },
    {
      "id": "1804010",
      "postDate": "05/28/2022 12:38:17",
      "content": "<p>Thanks, buddy <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<p>This is totally new approach for me! I definitely try this approach! Once again thanks for sharing your experience </p>",
      "rawMarkdown": "Thanks, buddy @cdeotte \n\nThis is totally new approach for me! I definitely try this approach! Once again thanks for sharing your experience",
      "votes": null
    },
    {
      "id": "1805581",
      "postDate": "05/30/2022 10:23:58",
      "content": "<p>Thanks, very informative to deal with large datasets! </p>",
      "rawMarkdown": "Thanks, very informative to deal with large datasets!",
      "votes": null
    },
    {
      "id": "1805583",
      "postDate": "05/30/2022 10:28:17",
      "content": "<p>Maybe take a look at dask_cudf and cudf libraries for manipulating data</p>",
      "rawMarkdown": "Maybe take a look at dask_cudf and cudf libraries for manipulating data",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1803756,
      "author_name": "anuragmuk",
      "author_url": "",
      "post_date": "05/28/2022 06:28:38",
      "content": "<p>Hi Venkatkumar,</p>\n<p>If you try using Feather format, it works just fine, both train and test together without memory issues.<br>\nAnd to make it easier, someone has already uploaded the dataset in Feather format.</p>\n<p>Regards,<br>\nAnurag</p>",
      "votes": null,
      "replies": [
        {
          "id": 1803784,
          "author_name": "venkatkumar001",
          "author_url": "",
          "post_date": "05/28/2022 07:22:40",
          "content": "<p>Yes! i definitely try this method<br>\nThanks for sharing your experience</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1803767,
      "author_name": "devkhant24",
      "author_url": "",
      "post_date": "05/28/2022 06:48:31",
      "content": "<p>I went with parquet and feather both work fines and takes less space. And I actually used both for train data -&gt; parquet and for test data -&gt; feather, because by using both parquet files I was getting OOM. And in feather file that I used didn't include \"target\" column so went with both formats.</p>\n<p>Do tell me if there's some problem with using different formats.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1803783,
          "author_name": "venkatkumar001",
          "author_url": "",
          "post_date": "05/28/2022 07:21:47",
          "content": "<p>Yes, buddy! I acceptable<br>\nI didn't try .far format,  I will try</p>\n<p>I was trying the parquet format! it is good but occupy the more memory using different model approach</p>\n<p>Thanks for sharing your experience buddy <a href=\"https://www.kaggle.com/devkhant24\" target=\"_blank\">@devkhant24</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1803859,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "05/28/2022 08:54:17",
      "content": "<p>An awesome <a href=\"https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets/notebook\" target=\"_blank\">Tutorial on reading large datasets</a> by <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>  !</p>",
      "votes": null,
      "replies": [
        {
          "id": 1804008,
          "author_name": "venkatkumar001",
          "author_url": "",
          "post_date": "05/28/2022 12:35:49",
          "content": "<p>Thankyou buddy! <a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> 👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1803931,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "05/28/2022 10:42:50",
      "content": "<p>In addition to all the file type suggestions, note that you can also save the data as multiple files. For example, there are <code>458913</code> unique customers in train. For example, you can create 10 files where each file has <code>50_000</code> customers. Then when you do feature engineering, you can load one file at a time and engineer then save the result. Then load another etc. </p>\n<p>Using small chunks of data will allow you to do your feature engineering on GPU which is much faster than CPU. And reading and writing to disk from GPU is often faster than reading and writing from CPU. When you are ready to train, most models allow training from chunked data. Load chunk (one or more to fill memory), shuffle chunk(s), create batches within and train. Load another chunk (one or more), etc.</p>\n<p>Of course whatever file format or files you use, you should reduce the memory of each column. For example, use <code>float32</code> and <code>int32</code> instead of <code>float64</code> and <code>int64</code>. And convert the <code>customer_ID</code> column into <code>int32</code> or <code>int64</code> after label encoding. Also the date column can be compressed into three columns of <code>int8</code> for year, month, day.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1804010,
          "author_name": "venkatkumar001",
          "author_url": "",
          "post_date": "05/28/2022 12:38:17",
          "content": "<p>Thanks, buddy <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<p>This is totally new approach for me! I definitely try this approach! Once again thanks for sharing your experience </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1805581,
          "author_name": "paulojunqueira",
          "author_url": "",
          "post_date": "05/30/2022 10:23:58",
          "content": "<p>Thanks, very informative to deal with large datasets! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1805583,
      "author_name": "paulojunqueira",
      "author_url": "",
      "post_date": "05/30/2022 10:28:17",
      "content": "<p>Maybe take a look at dask_cudf and cudf libraries for manipulating data</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1803739": "This competition data is very large in my experience of approach!\n\n> **which one is best! parquet or pickle?**\n\nEvery time memory is full and restarts my notebook!\n\n> **Any buddy suggest! How to effectively approach**",
    "1803756": "Hi Venkatkumar,\n\nIf you try using Feather format, it works just fine, both train and test together without memory issues.\nAnd to make it easier, someone has already uploaded the dataset in Feather format.\n\nRegards,\nAnurag",
    "1803767": "I went with parquet and feather both work fines and takes less space. And I actually used both for train data -> parquet and for test data -> feather, because by using both parquet files I was getting OOM. And in feather file that I used didn't include \"target\" column so went with both formats.\n\nDo tell me if there's some problem with using different formats.",
    "1803783": "Yes, buddy! I acceptable\nI didn't try .far format,  I will try\n\nI was trying the parquet format! it is good but occupy the more memory using different model approach\n\nThanks for sharing your experience buddy @devkhant24",
    "1803784": "Yes! i definitely try this method\nThanks for sharing your experience",
    "1803859": "An awesome [Tutorial on reading large datasets](https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets/notebook) by @rohanrao  !",
    "1803931": "In addition to all the file type suggestions, note that you can also save the data as multiple files. For example, there are `458913` unique customers in train. For example, you can create 10 files where each file has `50_000` customers. Then when you do feature engineering, you can load one file at a time and engineer then save the result. Then load another etc. \n\nUsing small chunks of data will allow you to do your feature engineering on GPU which is much faster than CPU. And reading and writing to disk from GPU is often faster than reading and writing from CPU. When you are ready to train, most models allow training from chunked data. Load chunk (one or more to fill memory), shuffle chunk(s), create batches within and train. Load another chunk (one or more), etc.\n\nOf course whatever file format or files you use, you should reduce the memory of each column. For example, use `float32` and `int32` instead of `float64` and `int64`. And convert the `customer_ID` column into `int32` or `int64` after label encoding. Also the date column can be compressed into three columns of `int8` for year, month, day.",
    "1804008": "Thankyou buddy! @steubk 👍",
    "1804010": "Thanks, buddy @cdeotte \n\nThis is totally new approach for me! I definitely try this approach! Once again thanks for sharing your experience",
    "1805581": "Thanks, very informative to deal with large datasets!",
    "1805583": "Maybe take a look at dask_cudf and cudf libraries for manipulating data"
  },
  "source": "meta"
}