{
  "id": 75196,
  "title": "For Future use: Reduce memory footprint of testset by factor 2-3",
  "url": "/competitions/PLAsTiCC-2018/discussion/75196",
  "author_name": "Marc",
  "post_date": "2018-12-19T10:07:30.756000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Congrats to the winners and medalists here first of all.</p>\n\n<p>I just wanted to share a quick tip, which may be obvious to the pros here, but was absent from the kernels as far as I could see:</p>\n\n<p><strong>If you specify the datatypes when loading the csv with pandas, you can reduce the memory footprint of the testset from 20.3 GB to 9.3 GB</strong> without information loss. If you are willing to reduce precision a bit, you can get it to 6.7 GB. This is useful, even if you have more than enough RAM and even speeds up processing a bit.</p>\n\n<p>I have created a quick kernel to show how this is used and also a little benchmark of different pandas storage formats for quick reference (all using the test dataset of this competition)\n<a href=\"https://www.kaggle.com/marcmuc/large-csv-datasets-with-pandas-use-less-memory\">https://www.kaggle.com/marcmuc/large-csv-datasets-with-pandas-use-less-memory</a></p>",
  "messages": [
    {
      "id": 441977,
      "postDate": "2018-12-19T10:07:30.757Z",
      "content": "<p>Congrats to the winners and medalists here first of all.</p>\n\n<p>I just wanted to share a quick tip, which may be obvious to the pros here, but was absent from the kernels as far as I could see:</p>\n\n<p><strong>If you specify the datatypes when loading the csv with pandas, you can reduce the memory footprint of the testset from 20.3 GB to 9.3 GB</strong> without information loss. If you are willing to reduce precision a bit, you can get it to 6.7 GB. This is useful, even if you have more than enough RAM and even speeds up processing a bit.</p>\n\n<p>I have created a quick kernel to show how this is used and also a little benchmark of different pandas storage formats for quick reference (all using the test dataset of this competition)\n<a href=\"https://www.kaggle.com/marcmuc/large-csv-datasets-with-pandas-use-less-memory\">https://www.kaggle.com/marcmuc/large-csv-datasets-with-pandas-use-less-memory</a></p>",
      "rawMarkdown": "Congrats to the winners and medalists here first of all.\n\nI just wanted to share a quick tip, which may be obvious to the pros here, but was absent from the kernels as far as I could see:\n\n**If you specify the datatypes when loading the csv with pandas, you can reduce the memory footprint of the testset from 20.3 GB to 9.3 GB** without information loss. If you are willing to reduce precision a bit, you can get it to 6.7 GB. This is useful, even if you have more than enough RAM and even speeds up processing a bit.\n\nI have created a quick kernel to show how this is used and also a little benchmark of different pandas storage formats for quick reference (all using the test dataset of this competition)\nhttps://www.kaggle.com/marcmuc/large-csv-datasets-with-pandas-use-less-memory",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "441977": "Congrats to the winners and medalists here first of all.\n\nI just wanted to share a quick tip, which may be obvious to the pros here, but was absent from the kernels as far as I could see:\n\n**If you specify the datatypes when loading the csv with pandas, you can reduce the memory footprint of the testset from 20.3 GB to 9.3 GB** without information loss. If you are willing to reduce precision a bit, you can get it to 6.7 GB. This is useful, even if you have more than enough RAM and even speeds up processing a bit.\n\nI have created a quick kernel to show how this is used and also a little benchmark of different pandas storage formats for quick reference (all using the test dataset of this competition)\nhttps://www.kaggle.com/marcmuc/large-csv-datasets-with-pandas-use-less-memory"
  }
}