{
  "id": 338877,
  "title": "For Colab Users, I have a question about memory overflow problem.",
  "url": "/competitions/amex-default-prediction/discussion/338877",
  "author_name": "M_Murata",
  "post_date": "2022-07-22T11:50:05.852000",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I used kaggle notebook, but it often occur memory overflow problem.<br>\nTherefore, I tried to switch GoogleColab. But it also occur same problem. I registered Colab Pro.</p>\n<p>If you use Colab Pro+, does it occur same problem? if you know how to avoid this problem, please tell us.</p>",
  "messages": [
    {
      "id": 1866317,
      "postDate": "2022-07-22T11:50:05.853Z",
      "content": "<p>I used kaggle notebook, but it often occur memory overflow problem.<br>\nTherefore, I tried to switch GoogleColab. But it also occur same problem. I registered Colab Pro.</p>\n<p>If you use Colab Pro+, does it occur same problem? if you know how to avoid this problem, please tell us.</p>",
      "rawMarkdown": "I used kaggle notebook, but it often occur memory overflow problem.\nTherefore, I tried to switch GoogleColab. But it also occur same problem. I registered Colab Pro.\n\nIf you use Colab Pro+, does it occur same problem? if you know how to avoid this problem, please tell us.",
      "votes": 1
    },
    {
      "id": 1866374,
      "postDate": "2022-07-22T12:37:25.963Z",
      "content": "<p>If you switch to a high ram environment you should be able to run almost everything, if I remember correctly colab pro offers 25gb of ram. In case it doesn't work for you, there is an interesting platform that I use, it's called paparspace gradient, it offers you 30 gb of ram even in its free version.</p>",
      "rawMarkdown": "If you switch to a high ram environment you should be able to run almost everything, if I remember correctly colab pro offers 25gb of ram. In case it doesn't work for you, there is an interesting platform that I use, it's called paparspace gradient, it offers you 30 gb of ram even in its free version.",
      "votes": 2,
      "replies": [
        {
          "id": 1866426,
          "postDate": "2022-07-22T13:24:52.043Z",
          "content": "<p>Yes, it's good, but it's not private first, the timeout is at a maximum of 6 hours, and version free is always out of capacity </p>",
          "rawMarkdown": "Yes, it's good, but it's not private first, the timeout is at a maximum of 6 hours, and version free is always out of capacity ",
          "votes": 2
        },
        {
          "id": 1866432,
          "postDate": "2022-07-22T13:30:02.580Z",
          "content": "<p>Kaggle notebook is much better and, I think it's enough for this competition if we use it with the optimum method.</p>",
          "rawMarkdown": "Kaggle notebook is much better and, I think it's enough for this competition if we use it with the optimum method.",
          "votes": 2
        },
        {
          "id": 1866443,
          "postDate": "2022-07-22T13:35:26.677Z",
          "content": "<p>Yes, it's true, it has its cons, it would be like the last resort to ram problems. I've used the pro version for a while now, which also has 30 GB ram and I have had no problems in that regard, the only annoying thing is the 6-hour limit.</p>",
          "rawMarkdown": "Yes, it's true, it has its cons, it would be like the last resort to ram problems. I've used the pro version for a while now, which also has 30 GB ram and I have had no problems in that regard, the only annoying thing is the 6-hour limit.",
          "votes": 2
        },
        {
          "id": 1866978,
          "postDate": "2022-07-22T23:03:07.113Z",
          "content": "<p><a href=\"https://www.kaggle.com/maxdiazbattan\" target=\"_blank\">@maxdiazbattan</a> <a href=\"https://www.kaggle.com/youneseloiarm\" target=\"_blank\">@youneseloiarm</a><br>\nThank you for your polite reply! :D<br>\nI'll try paparspace gradient.</p>\n<p><a href=\"https://www.kaggle.com/youneseloiarm\" target=\"_blank\">@youneseloiarm</a><br>\nThank you for reply. is it also enough to preprocessing? (ex. aggregation, creating features)</p>",
          "rawMarkdown": "@maxdiazbattan @youneseloiarm\nThank you for your polite reply! :D\nI'll try paparspace gradient.\n\n@youneseloiarm\nThank you for reply. is it also enough to preprocessing? (ex. aggregation, creating features)",
          "votes": 1
        },
        {
          "id": 1866988,
          "postDate": "2022-07-22T23:37:04.397Z",
          "content": "<p>You're welcome <a href=\"https://www.kaggle.com/masashimurata\" target=\"_blank\">@masashimurata</a>. Sorry to butt in, but in my case, the key to overcoming ram problems was to work in batches, you can try something like this:</p>\n<pre><code>import pyarrow.parquet as pq\n\ndef feature_engineering(path, batch_size, cols_to_group, cols_to_agg, agregation):\n\n    batch_size = batch_size \n    batches = pq.ParquetFile(path).iter_batches(batch_size, use_pandas_metadata=True) \n\n    agg_list = []\n\n    for df in batches:\n        df = df.to_pandas()\n\n        df['customer_ID_hash'] = df['customer_ID'].apply(lambda x: int(x[-16:],16) ).astype('int64')\n\n        agg_df = df.groupby(cols_to_group)[cols_to_agg].agg(agregation).astype('float32')    \n        agg_list.append(agg_df)\n\n        del df, agg_df\n        gc.collect()\n\n    concat_df = pd.concat(agg_list)\n    concat_df.columns = concat_df.columns.map(lambda x: '|'.join([str(i) for i in x]))\n\n    concat_df = concat_df.reset_index().drop_duplicates(subset=['customer_ID_hash'], keep='last')\n    gc.collect()\n\n    return concat_df\n\nagg_train_df = feature_engineering('../input/amex-data-integer-dtypes-parquet-format/train.parquet', 500000, 'customer_ID_hash', cat_features, ['count', 'last', 'nunique'])\n</code></pre>\n<p>I hope something like this helps you, greetings!</p>",
          "rawMarkdown": "You're welcome @masashimurata. Sorry to butt in, but in my case, the key to overcoming ram problems was to work in batches, you can try something like this:\n\n```\nimport pyarrow.parquet as pq\n\ndef feature_engineering(path, batch_size, cols_to_group, cols_to_agg, agregation):\n    \n    batch_size = batch_size \n    batches = pq.ParquetFile(path).iter_batches(batch_size, use_pandas_metadata=True) \n    \n    agg_list = []\n    \n    for df in batches:\n        df = df.to_pandas()\n\n        df['customer_ID_hash'] = df['customer_ID'].apply(lambda x: int(x[-16:],16) ).astype('int64')\n   \n        agg_df = df.groupby(cols_to_group)[cols_to_agg].agg(agregation).astype('float32')    \n        agg_list.append(agg_df)\n        \n        del df, agg_df\n        gc.collect()\n        \n    concat_df = pd.concat(agg_list)\n    concat_df.columns = concat_df.columns.map(lambda x: '|'.join([str(i) for i in x]))\n    \n    concat_df = concat_df.reset_index().drop_duplicates(subset=['customer_ID_hash'], keep='last')\n    gc.collect()\n        \n    return concat_df\n\nagg_train_df = feature_engineering('../input/amex-data-integer-dtypes-parquet-format/train.parquet', 500000, 'customer_ID_hash', cat_features, ['count', 'last', 'nunique'])\n\n\n```\n\nI hope something like this helps you, greetings!",
          "votes": 1
        },
        {
          "id": 1875299,
          "postDate": "2022-07-28T21:52:31.307Z",
          "content": "<p>Currently, I don't have any problems in the training step, but in the final step, there are a lot of problems with the huge size of the data test, my solution is to split the data test into 10 parts, and concrete them after finishing each part, So, you can use Kaggle Notebook without any problem from training to submission steps.</p>",
          "rawMarkdown": "Currently, I don't have any problems in the training step, but in the final step, there are a lot of problems with the huge size of the data test, my solution is to split the data test into 10 parts, and concrete them after finishing each part, So, you can use Kaggle Notebook without any problem from training to submission steps."
        }
      ]
    },
    {
      "id": 1866977,
      "postDate": "2022-07-22T23:02:39.350Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1866374,
      "author_name": "Maximiliano Diaz Battan",
      "author_url": "",
      "post_date": "2022-07-22T12:37:25.963000",
      "content": "<p>If you switch to a high ram environment you should be able to run almost everything, if I remember correctly colab pro offers 25gb of ram. In case it doesn't work for you, there is an interesting platform that I use, it's called paparspace gradient, it offers you 30 gb of ram even in its free version.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1866426,
          "author_name": "EL Younes",
          "author_url": "",
          "post_date": "2022-07-22T13:24:52.043000",
          "content": "<p>Yes, it's good, but it's not private first, the timeout is at a maximum of 6 hours, and version free is always out of capacity </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1866432,
          "author_name": "EL Younes",
          "author_url": "",
          "post_date": "2022-07-22T13:30:02.580000",
          "content": "<p>Kaggle notebook is much better and, I think it's enough for this competition if we use it with the optimum method.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1866443,
          "author_name": "Maximiliano Diaz Battan",
          "author_url": "",
          "post_date": "2022-07-22T13:35:26.677000",
          "content": "<p>Yes, it's true, it has its cons, it would be like the last resort to ram problems. I've used the pro version for a while now, which also has 30 GB ram and I have had no problems in that regard, the only annoying thing is the 6-hour limit.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1866978,
          "author_name": "M_Murata",
          "author_url": "",
          "post_date": "2022-07-22T23:03:07.113000",
          "content": "<p><a href=\"https://www.kaggle.com/maxdiazbattan\" target=\"_blank\">@maxdiazbattan</a> <a href=\"https://www.kaggle.com/youneseloiarm\" target=\"_blank\">@youneseloiarm</a><br>\nThank you for your polite reply! :D<br>\nI'll try paparspace gradient.</p>\n<p><a href=\"https://www.kaggle.com/youneseloiarm\" target=\"_blank\">@youneseloiarm</a><br>\nThank you for reply. is it also enough to preprocessing? (ex. aggregation, creating features)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1866988,
          "author_name": "Maximiliano Diaz Battan",
          "author_url": "",
          "post_date": "2022-07-22T23:37:04.397000",
          "content": "<p>You're welcome <a href=\"https://www.kaggle.com/masashimurata\" target=\"_blank\">@masashimurata</a>. Sorry to butt in, but in my case, the key to overcoming ram problems was to work in batches, you can try something like this:</p>\n<pre><code>import pyarrow.parquet as pq\n\ndef feature_engineering(path, batch_size, cols_to_group, cols_to_agg, agregation):\n\n    batch_size = batch_size \n    batches = pq.ParquetFile(path).iter_batches(batch_size, use_pandas_metadata=True) \n\n    agg_list = []\n\n    for df in batches:\n        df = df.to_pandas()\n\n        df['customer_ID_hash'] = df['customer_ID'].apply(lambda x: int(x[-16:],16) ).astype('int64')\n\n        agg_df = df.groupby(cols_to_group)[cols_to_agg].agg(agregation).astype('float32')    \n        agg_list.append(agg_df)\n\n        del df, agg_df\n        gc.collect()\n\n    concat_df = pd.concat(agg_list)\n    concat_df.columns = concat_df.columns.map(lambda x: '|'.join([str(i) for i in x]))\n\n    concat_df = concat_df.reset_index().drop_duplicates(subset=['customer_ID_hash'], keep='last')\n    gc.collect()\n\n    return concat_df\n\nagg_train_df = feature_engineering('../input/amex-data-integer-dtypes-parquet-format/train.parquet', 500000, 'customer_ID_hash', cat_features, ['count', 'last', 'nunique'])\n</code></pre>\n<p>I hope something like this helps you, greetings!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1875299,
          "author_name": "EL Younes",
          "author_url": "",
          "post_date": "2022-07-28T21:52:31.307000",
          "content": "<p>Currently, I don't have any problems in the training step, but in the final step, there are a lot of problems with the huge size of the data test, my solution is to split the data test into 10 parts, and concrete them after finishing each part, So, you can use Kaggle Notebook without any problem from training to submission steps.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1866977,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-22T23:02:39.350000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1866317": "I used kaggle notebook, but it often occur memory overflow problem.\nTherefore, I tried to switch GoogleColab. But it also occur same problem. I registered Colab Pro.\n\nIf you use Colab Pro+, does it occur same problem? if you know how to avoid this problem, please tell us.",
    "1866374": "If you switch to a high ram environment you should be able to run almost everything, if I remember correctly colab pro offers 25gb of ram. In case it doesn't work for you, there is an interesting platform that I use, it's called paparspace gradient, it offers you 30 gb of ram even in its free version.",
    "1866977": ""
  }
}