{
  "id": 341538,
  "title": "How To Compress Humongous Dataset",
  "url": "/competitions/amex-default-prediction/discussion/341538",
  "author_name": "",
  "post_date": "2022-08-03T10:02:29.399386600Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello Kagglers, could you please share information on how one can compress the size of the dataset for this competition from the 50+GB size to a few MB?</p>",
  "messages": [
    {
      "id": "1882502",
      "postDate": "08/03/2022 10:02:29",
      "content": "<p>Hello Kagglers, could you please share information on how one can compress the size of the dataset for this competition from the 50+GB size to a few MB?</p>",
      "rawMarkdown": "Hello Kagglers, could you please share information on how one can compress the size of the dataset for this competition from the 50+GB size to a few MB?",
      "votes": null
    },
    {
      "id": "1882571",
      "postDate": "08/03/2022 10:43:25",
      "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> has created a compressed dataset for this competition, I highly recommend to use it (I also am using it). I hope this solves the issue. </p>",
      "rawMarkdown": "raddar has created a compressed dataset for this competition, I highly recommend to use it (I also am using it). I hope this solves the issue.",
      "votes": null
    },
    {
      "id": "1882667",
      "postDate": "08/03/2022 11:19:40",
      "content": "<p>Well, I do not think there is a compression algorithm to do that. What you need to do for that is to change features' data type to another type that requires less memory like float64 to float32. Of course, this operation would also result in losing some information but most of the time, that loss would be negligible. With this, you can make size much more smaller than the original size.</p>",
      "rawMarkdown": "Well, I do not think there is a compression algorithm to do that. What you need to do for that is to change features' data type to another type that requires less memory like float64 to float32. Of course, this operation would also result in losing some information but most of the time, that loss would be negligible. With this, you can make size much more smaller than the original size.",
      "votes": null
    },
    {
      "id": "1882887",
      "postDate": "08/03/2022 13:17:58",
      "content": "<p>Thank you for the information </p>",
      "rawMarkdown": "Thank you for the information",
      "votes": null
    },
    {
      "id": "1882889",
      "postDate": "08/03/2022 13:18:21",
      "content": "<p>Thank you for your help </p>",
      "rawMarkdown": "Thank you for your help",
      "votes": null
    },
    {
      "id": "1883566",
      "postDate": "08/03/2022 23:02:42",
      "content": "<p>Yes, Raddar's train dataset is only 1.6GB</p>",
      "rawMarkdown": "Yes, Raddar's train dataset is only 1.6GB",
      "votes": null
    },
    {
      "id": "1892221",
      "postDate": "08/10/2022 00:49:00",
      "content": "<p>Hello,</p>\n<p>If you are having trouble reading the test_data.csv file ( 30+ GB ), try using a batch read algorithm - Here is a brief example…</p>\n<h1>#</h1>\n<h1>read data in chunks of 1 million rows at a time</h1>\n<p>#<br>\ntestDataSet = pd.DataFrame()<br>\ntotal = 0<br>\nfor chunk in pd.read_csv(\"./input/amex-default-prediction/test_data.csv\", iterator=True, chunksize=1000000):</p>\n<pre><code>testDataSet = pd.concat([testDataSet, chunk])\ntotal += len(chunk)\nprint(\"Read CSV chunk: \", len(chunk), type(chunk), total)\n</code></pre>\n<p>print(\"Read csv with chunks: total rows ==  \", total )</p>\n<p>Hope it helps!<br>\nMichael</p>",
      "rawMarkdown": "Hello,\n\nIf you are having trouble reading the test_data.csv file ( 30+ GB ), try using a batch read algorithm - Here is a brief example...\n\n# ##################################################\n# read data in chunks of 1 million rows at a time\n#\ntestDataSet = pd.DataFrame()\ntotal = 0\nfor chunk in pd.read_csv(\"./input/amex-default-prediction/test_data.csv\", iterator=True, chunksize=1000000):\n\n    testDataSet = pd.concat([testDataSet, chunk])\n    total += len(chunk)\n    print(\"Read CSV chunk: \", len(chunk), type(chunk), total)\n\nprint(\"Read csv with chunks: total rows ==  \", total )\n\nHope it helps!\nMichael",
      "votes": null
    },
    {
      "id": "1894844",
      "postDate": "08/11/2022 18:39:56",
      "content": "<p>Thank you for your help. It surely will </p>",
      "rawMarkdown": "Thank you for your help. It surely will",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1882571,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "08/03/2022 10:43:25",
      "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> has created a compressed dataset for this competition, I highly recommend to use it (I also am using it). I hope this solves the issue. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1882887,
          "author_name": "lordxerxes",
          "author_url": "",
          "post_date": "08/03/2022 13:17:58",
          "content": "<p>Thank you for the information </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1883566,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/03/2022 23:02:42",
          "content": "<p>Yes, Raddar's train dataset is only 1.6GB</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1882667,
      "author_name": "memocan64",
      "author_url": "",
      "post_date": "08/03/2022 11:19:40",
      "content": "<p>Well, I do not think there is a compression algorithm to do that. What you need to do for that is to change features' data type to another type that requires less memory like float64 to float32. Of course, this operation would also result in losing some information but most of the time, that loss would be negligible. With this, you can make size much more smaller than the original size.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1882889,
          "author_name": "lordxerxes",
          "author_url": "",
          "post_date": "08/03/2022 13:18:21",
          "content": "<p>Thank you for your help </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1892221,
      "author_name": "michaelfromoldwick",
      "author_url": "",
      "post_date": "08/10/2022 00:49:00",
      "content": "<p>Hello,</p>\n<p>If you are having trouble reading the test_data.csv file ( 30+ GB ), try using a batch read algorithm - Here is a brief example…</p>\n<h1>#</h1>\n<h1>read data in chunks of 1 million rows at a time</h1>\n<p>#<br>\ntestDataSet = pd.DataFrame()<br>\ntotal = 0<br>\nfor chunk in pd.read_csv(\"./input/amex-default-prediction/test_data.csv\", iterator=True, chunksize=1000000):</p>\n<pre><code>testDataSet = pd.concat([testDataSet, chunk])\ntotal += len(chunk)\nprint(\"Read CSV chunk: \", len(chunk), type(chunk), total)\n</code></pre>\n<p>print(\"Read csv with chunks: total rows ==  \", total )</p>\n<p>Hope it helps!<br>\nMichael</p>",
      "votes": null,
      "replies": [
        {
          "id": 1894844,
          "author_name": "lordxerxes",
          "author_url": "",
          "post_date": "08/11/2022 18:39:56",
          "content": "<p>Thank you for your help. It surely will </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1882502": "Hello Kagglers, could you please share information on how one can compress the size of the dataset for this competition from the 50+GB size to a few MB?",
    "1882571": "raddar has created a compressed dataset for this competition, I highly recommend to use it (I also am using it). I hope this solves the issue.",
    "1882667": "Well, I do not think there is a compression algorithm to do that. What you need to do for that is to change features' data type to another type that requires less memory like float64 to float32. Of course, this operation would also result in losing some information but most of the time, that loss would be negligible. With this, you can make size much more smaller than the original size.",
    "1882887": "Thank you for the information",
    "1882889": "Thank you for your help",
    "1883566": "Yes, Raddar's train dataset is only 1.6GB",
    "1892221": "Hello,\n\nIf you are having trouble reading the test_data.csv file ( 30+ GB ), try using a batch read algorithm - Here is a brief example...\n\n# ##################################################\n# read data in chunks of 1 million rows at a time\n#\ntestDataSet = pd.DataFrame()\ntotal = 0\nfor chunk in pd.read_csv(\"./input/amex-default-prediction/test_data.csv\", iterator=True, chunksize=1000000):\n\n    testDataSet = pd.concat([testDataSet, chunk])\n    total += len(chunk)\n    print(\"Read CSV chunk: \", len(chunk), type(chunk), total)\n\nprint(\"Read csv with chunks: total rows ==  \", total )\n\nHope it helps!\nMichael",
    "1894844": "Thank you for your help. It surely will"
  },
  "source": "meta"
}