{
  "id": 342547,
  "title": "RAM out of memory in colab ",
  "url": "/competitions/amex-default-prediction/discussion/342547",
  "author_name": "Nava Bharath Myneni",
  "post_date": "2022-08-07T19:15:42.410000",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi,I am using colab free version to participate in this competition.colab provides 12.8GB ram and when i try to load test parquet compressed dataset(6 GB size) for predictions.I am running out of memory in colab. Can somebody suggest best approach to handle test predictions and what is the recommended hardware configuration to handle this test set?</p>",
  "messages": [
    {
      "id": 1888780,
      "postDate": "2022-08-07T19:15:42.410Z",
      "content": "<p>Hi,I am using colab free version to participate in this competition.colab provides 12.8GB ram and when i try to load test parquet compressed dataset(6 GB size) for predictions.I am running out of memory in colab. Can somebody suggest best approach to handle test predictions and what is the recommended hardware configuration to handle this test set?</p>",
      "rawMarkdown": "Hi,I am using colab free version to participate in this competition.colab provides 12.8GB ram and when i try to load test parquet compressed dataset(6 GB size) for predictions.I am running out of memory in colab. Can somebody suggest best approach to handle test predictions and what is the recommended hardware configuration to handle this test set?",
      "votes": 3
    },
    {
      "id": 1894915,
      "postDate": "2022-08-11T19:41:51.010Z",
      "content": "<p>You can try with pandas read_feather() for loading the entire volume of data to a python variable. Or you can try to run it on paperspace as it  gives more RAM compared to Google Colab. Google Colab session often goes off even if we are working on a very lesser volume of data. Also another method can be, reading the csv data in chunksize.  </p>",
      "rawMarkdown": "You can try with pandas read_feather() for loading the entire volume of data to a python variable. Or you can try to run it on paperspace as it  gives more RAM compared to Google Colab. Google Colab session often goes off even if we are working on a very lesser volume of data. Also another method can be, reading the csv data in chunksize.  ",
      "votes": 1
    },
    {
      "id": 1890241,
      "postDate": "2022-08-08T16:13:13.787Z",
      "content": "<p>Use parquet data set by <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> . I am able to load it without any issue using Kaggle Notebook</p>",
      "rawMarkdown": "Use parquet data set by @raddar . I am able to load it without any issue using Kaggle Notebook",
      "votes": 1
    },
    {
      "id": 1889265,
      "postDate": "2022-08-08T05:34:15.673Z",
      "content": "<p>Using kaggle's GPUs is better at this point. Dataset is ready to use so hopefully, you could probably use it without an error. </p>",
      "rawMarkdown": "Using kaggle's GPUs is better at this point. Dataset is ready to use so hopefully, you could probably use it without an error. ",
      "votes": 2,
      "replies": [
        {
          "id": 1890130,
          "postDate": "2022-08-08T15:04:14.350Z",
          "content": "<p>No problem in downloading test dataset to hard disk of colab. problem arises when i load the dataset as dataframe into memory.even kaggle GPU provides 15.9 GB ram, are you able to do prediction using kaggle GPU…if yes,what library you are using to load test file ?</p>",
          "rawMarkdown": "No problem in downloading test dataset to hard disk of colab. problem arises when i load the dataset as dataframe into memory.even kaggle GPU provides 15.9 GB ram, are you able to do prediction using kaggle GPU...if yes,what library you are using to load test file ?",
          "votes": 1
        },
        {
          "id": 1890237,
          "postDate": "2022-08-08T16:10:51.780Z",
          "content": "<p>Yes, we are able to do prediction and pandas are being used in reading the data but I assume that did not help you.<br>\nI checked the internet and I found these two. Hope it works,</p>\n<p>df = pd.read_csv(myfile,sep='\\t') # didn't work, memory error<br>\ndf = pd.read_csv(myfile,sep='\\t',low_memory=False) # worked fine and in less than 30 seconds</p>\n<p>If you are not using 32bit python in windows but are looking to improve your memory efficiency while reading csv files, there is a trick. The pandas.read_csv function takes an option called dtype. This lets pandas know what types exist inside your csv data.</p>",
          "rawMarkdown": "Yes, we are able to do prediction and pandas are being used in reading the data but I assume that did not help you.\nI checked the internet and I found these two. Hope it works,\n\ndf = pd.read_csv(myfile,sep='\\t') # didn't work, memory error\ndf = pd.read_csv(myfile,sep='\\t',low_memory=False) # worked fine and in less than 30 seconds\n\nIf you are not using 32bit python in windows but are looking to improve your memory efficiency while reading csv files, there is a trick. The pandas.read_csv function takes an option called dtype. This lets pandas know what types exist inside your csv data.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1889100,
      "postDate": "2022-08-08T01:22:08.830Z",
      "content": "<p>I had the same problem, my approach to solve it was to select just the columns that are really useful for my model and then select that columns in the read_parquet, like this<br>\n<code>df = pd.read_parquet('../input/amex-parquet/test_data.parquet', columns=['customer_ID', 'S_2', 'D_39', 'B_1'......])</code><br>\nI hope it helps</p>",
      "rawMarkdown": "I had the same problem, my approach to solve it was to select just the columns that are really useful for my model and then select that columns in the read_parquet, like this\n`df = pd.read_parquet('../input/amex-parquet/test_data.parquet', columns=['customer_ID', 'S_2', 'D_39', 'B_1'......])`\nI hope it helps",
      "votes": 2
    },
    {
      "id": 1889260,
      "postDate": "2022-08-08T05:28:25.060Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1894915,
      "author_name": "Mahalakshmi Thimmappa",
      "author_url": "",
      "post_date": "2022-08-11T19:41:51.010000",
      "content": "<p>You can try with pandas read_feather() for loading the entire volume of data to a python variable. Or you can try to run it on paperspace as it  gives more RAM compared to Google Colab. Google Colab session often goes off even if we are working on a very lesser volume of data. Also another method can be, reading the csv data in chunksize.  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1890241,
      "author_name": "Dileep Kumar",
      "author_url": "",
      "post_date": "2022-08-08T16:13:13.787000",
      "content": "<p>Use parquet data set by <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> . I am able to load it without any issue using Kaggle Notebook</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1889265,
      "author_name": "uraninjo",
      "author_url": "",
      "post_date": "2022-08-08T05:34:15.673000",
      "content": "<p>Using kaggle's GPUs is better at this point. Dataset is ready to use so hopefully, you could probably use it without an error. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1890130,
          "author_name": "Nava Bharath Myneni",
          "author_url": "",
          "post_date": "2022-08-08T15:04:14.350000",
          "content": "<p>No problem in downloading test dataset to hard disk of colab. problem arises when i load the dataset as dataframe into memory.even kaggle GPU provides 15.9 GB ram, are you able to do prediction using kaggle GPU…if yes,what library you are using to load test file ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1890237,
          "author_name": "uraninjo",
          "author_url": "",
          "post_date": "2022-08-08T16:10:51.780000",
          "content": "<p>Yes, we are able to do prediction and pandas are being used in reading the data but I assume that did not help you.<br>\nI checked the internet and I found these two. Hope it works,</p>\n<p>df = pd.read_csv(myfile,sep='\\t') # didn't work, memory error<br>\ndf = pd.read_csv(myfile,sep='\\t',low_memory=False) # worked fine and in less than 30 seconds</p>\n<p>If you are not using 32bit python in windows but are looking to improve your memory efficiency while reading csv files, there is a trick. The pandas.read_csv function takes an option called dtype. This lets pandas know what types exist inside your csv data.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1889100,
      "author_name": "RAUL ERNESTO GUILLEN",
      "author_url": "",
      "post_date": "2022-08-08T01:22:08.830000",
      "content": "<p>I had the same problem, my approach to solve it was to select just the columns that are really useful for my model and then select that columns in the read_parquet, like this<br>\n<code>df = pd.read_parquet('../input/amex-parquet/test_data.parquet', columns=['customer_ID', 'S_2', 'D_39', 'B_1'......])</code><br>\nI hope it helps</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1889260,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-08T05:28:25.060000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1888780": "Hi,I am using colab free version to participate in this competition.colab provides 12.8GB ram and when i try to load test parquet compressed dataset(6 GB size) for predictions.I am running out of memory in colab. Can somebody suggest best approach to handle test predictions and what is the recommended hardware configuration to handle this test set?",
    "1894915": "You can try with pandas read_feather() for loading the entire volume of data to a python variable. Or you can try to run it on paperspace as it  gives more RAM compared to Google Colab. Google Colab session often goes off even if we are working on a very lesser volume of data. Also another method can be, reading the csv data in chunksize.  ",
    "1890241": "Use parquet data set by @raddar . I am able to load it without any issue using Kaggle Notebook",
    "1889265": "Using kaggle's GPUs is better at this point. Dataset is ready to use so hopefully, you could probably use it without an error. ",
    "1889100": "I had the same problem, my approach to solve it was to select just the columns that are really useful for my model and then select that columns in the read_parquet, like this\n`df = pd.read_parquet('../input/amex-parquet/test_data.parquet', columns=['customer_ID', 'S_2', 'D_39', 'B_1'......])`\nI hope it helps",
    "1889260": ""
  }
}