{
  "id": 76457,
  "title": "Full of memory problem when reading train.parquet & test.parquet",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/76457",
  "author_name": "",
  "post_date": "2019-01-03T03:28:58.639468Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>In kaggle's kernel, or in my computer with 16GB RAM, it is impossible to read train.parquet and test parquet because of full of memory (RAM).\nHow could I deal with this problem??</p>",
  "messages": [
    {
      "id": "449372",
      "postDate": "01/03/2019 03:28:58",
      "content": "<p>In kaggle's kernel, or in my computer with 16GB RAM, it is impossible to read train.parquet and test parquet because of full of memory (RAM).\nHow could I deal with this problem??</p>",
      "rawMarkdown": "In kaggle's kernel, or in my computer with 16GB RAM, it is impossible to read train.parquet and test parquet because of full of memory (RAM).\nHow could I deal with this problem??",
      "votes": null
    },
    {
      "id": "449597",
      "postDate": "01/03/2019 12:14:27",
      "content": "<p>You can process the parquet file in chunks, defining chunks by a list of columns. After reading the test meta data through</p>\n\n<p>df_test = pd.read_csv('../input/metadata_test.csv')</p>\n\n<p>idx = df_test.signal_id.values</p>\n\n<p>idx defines the list of colums. With chunk_start and chunk_size defining the start and size of a chunk you can read the columns through</p>\n\n<p>chunk_columns = [str(j) for j in idx[chunk_start:chunk_start + chunk_size]]</p>\n\n<p>signals = pq.read_pandas('../input/test.parquet', columns=chunk_columns).to_pandas().values</p>",
      "rawMarkdown": "You can process the parquet file in chunks, defining chunks by a list of columns. After reading the test meta data through\n\ndf\\_test = pd.read\\_csv('../input/metadata\\_test.csv')\n\nidx = df\\_test.signal_id.values\n\nidx defines the list of colums. With chunk_start and chunk\\_size defining the start and size of a chunk you can read the columns through\n\nchunk\\_columns = [str(j) for j in idx[chunk\\_start:chunk\\_start + chunk\\_size]]\n\nsignals = pq.read\\_pandas('../input/test.parquet', columns=chunk\\_columns).to\\_pandas().values",
      "votes": null
    },
    {
      "id": "450440",
      "postDate": "01/05/2019 00:32:57",
      "content": "<p>Consider training the model first before even trying to pull in the test data set.  The test data set is considerably bigger than the training set.  I suppose one could look at the test data set to see if the features have the same distribution as the training set, but that feels kinda dirty.</p>",
      "rawMarkdown": "Consider training the model first before even trying to pull in the test data set.  The test data set is considerably bigger than the training set.  I suppose one could look at the test data set to see if the features have the same distribution as the training set, but that feels kinda dirty.",
      "votes": null
    },
    {
      "id": "450540",
      "postDate": "01/05/2019 07:55:14",
      "content": "<p>Thanks for your kind advice!!</p>",
      "rawMarkdown": "Thanks for your kind advice!!",
      "votes": null
    },
    {
      "id": "450542",
      "postDate": "01/05/2019 07:56:29",
      "content": "<p>Yes, I definitely agree with what you mentioned! Test dataset should not be used, when I extract features and train my model! Thanks~</p>",
      "rawMarkdown": "Yes, I definitely agree with what you mentioned! Test dataset should not be used, when I extract features and train my model! Thanks~",
      "votes": null
    },
    {
      "id": "452471",
      "postDate": "01/08/2019 19:13:57",
      "content": "<p>If the \"left-over\" chunks are still taking up too much memory, then you could do something that is less common in Python: <code>del variable</code> and run the garbage collector directly afterwards (it helps to run in manually to free a few more MB). In general, this dataset is a good way to remember how messed up the Python ecosystem is in terms of efficiency and overhead.</p>",
      "rawMarkdown": "If the \"left-over\" chunks are still taking up too much memory, then you could do something that is less common in Python: ```del variable``` and run the garbage collector directly afterwards (it helps to run in manually to free a few more MB). In general, this dataset is a good way to remember how messed up the Python ecosystem is in terms of efficiency and overhead.",
      "votes": null
    },
    {
      "id": "452491",
      "postDate": "01/08/2019 20:08:12",
      "content": "<p>If converting to a pandas dataframe casting  to int8 will reduce memory usage. </p>",
      "rawMarkdown": "If converting to a pandas dataframe casting  to int8 will reduce memory usage.",
      "votes": null
    },
    {
      "id": "455848",
      "postDate": "01/14/2019 17:31:57",
      "content": "<p>For anyone experiencing the same RAM limitations, I highly recommend reviewing this kernel where it details how to load in subsets of the parquet file.</p>\n\n<p><a href=\"https://www.kaggle.com/sohier/reading-the-data-with-python\">https://www.kaggle.com/sohier/reading-the-data-with-python</a></p>",
      "rawMarkdown": "For anyone experiencing the same RAM limitations, I highly recommend reviewing this kernel where it details how to load in subsets of the parquet file.\n\nhttps://www.kaggle.com/sohier/reading-the-data-with-python",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 449597,
      "author_name": "helgith",
      "author_url": "",
      "post_date": "01/03/2019 12:14:27",
      "content": "<p>You can process the parquet file in chunks, defining chunks by a list of columns. After reading the test meta data through</p>\n\n<p>df_test = pd.read_csv('../input/metadata_test.csv')</p>\n\n<p>idx = df_test.signal_id.values</p>\n\n<p>idx defines the list of colums. With chunk_start and chunk_size defining the start and size of a chunk you can read the columns through</p>\n\n<p>chunk_columns = [str(j) for j in idx[chunk_start:chunk_start + chunk_size]]</p>\n\n<p>signals = pq.read_pandas('../input/test.parquet', columns=chunk_columns).to_pandas().values</p>",
      "votes": null,
      "replies": [
        {
          "id": 450540,
          "author_name": "mykim89",
          "author_url": "",
          "post_date": "01/05/2019 07:55:14",
          "content": "<p>Thanks for your kind advice!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452471,
          "author_name": "simonwenkel",
          "author_url": "",
          "post_date": "01/08/2019 19:13:57",
          "content": "<p>If the \"left-over\" chunks are still taking up too much memory, then you could do something that is less common in Python: <code>del variable</code> and run the garbage collector directly afterwards (it helps to run in manually to free a few more MB). In general, this dataset is a good way to remember how messed up the Python ecosystem is in terms of efficiency and overhead.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 450440,
      "author_name": "mnight08",
      "author_url": "",
      "post_date": "01/05/2019 00:32:57",
      "content": "<p>Consider training the model first before even trying to pull in the test data set.  The test data set is considerably bigger than the training set.  I suppose one could look at the test data set to see if the features have the same distribution as the training set, but that feels kinda dirty.</p>",
      "votes": null,
      "replies": [
        {
          "id": 450542,
          "author_name": "mykim89",
          "author_url": "",
          "post_date": "01/05/2019 07:56:29",
          "content": "<p>Yes, I definitely agree with what you mentioned! Test dataset should not be used, when I extract features and train my model! Thanks~</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 452491,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "01/08/2019 20:08:12",
      "content": "<p>If converting to a pandas dataframe casting  to int8 will reduce memory usage. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 455848,
      "author_name": "jeffreyegan",
      "author_url": "",
      "post_date": "01/14/2019 17:31:57",
      "content": "<p>For anyone experiencing the same RAM limitations, I highly recommend reviewing this kernel where it details how to load in subsets of the parquet file.</p>\n\n<p><a href=\"https://www.kaggle.com/sohier/reading-the-data-with-python\">https://www.kaggle.com/sohier/reading-the-data-with-python</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "449372": "In kaggle's kernel, or in my computer with 16GB RAM, it is impossible to read train.parquet and test parquet because of full of memory (RAM).\nHow could I deal with this problem??",
    "449597": "You can process the parquet file in chunks, defining chunks by a list of columns. After reading the test meta data through\n\ndf\\_test = pd.read\\_csv('../input/metadata\\_test.csv')\n\nidx = df\\_test.signal_id.values\n\nidx defines the list of colums. With chunk_start and chunk\\_size defining the start and size of a chunk you can read the columns through\n\nchunk\\_columns = [str(j) for j in idx[chunk\\_start:chunk\\_start + chunk\\_size]]\n\nsignals = pq.read\\_pandas('../input/test.parquet', columns=chunk\\_columns).to\\_pandas().values",
    "450440": "Consider training the model first before even trying to pull in the test data set.  The test data set is considerably bigger than the training set.  I suppose one could look at the test data set to see if the features have the same distribution as the training set, but that feels kinda dirty.",
    "450540": "Thanks for your kind advice!!",
    "450542": "Yes, I definitely agree with what you mentioned! Test dataset should not be used, when I extract features and train my model! Thanks~",
    "452471": "If the \"left-over\" chunks are still taking up too much memory, then you could do something that is less common in Python: ```del variable``` and run the garbage collector directly afterwards (it helps to run in manually to free a few more MB). In general, this dataset is a good way to remember how messed up the Python ecosystem is in terms of efficiency and overhead.",
    "452491": "If converting to a pandas dataframe casting  to int8 will reduce memory usage.",
    "455848": "For anyone experiencing the same RAM limitations, I highly recommend reviewing this kernel where it details how to load in subsets of the parquet file.\n\nhttps://www.kaggle.com/sohier/reading-the-data-with-python"
  },
  "source": "meta"
}