{
  "id": 347126,
  "title": "16 GB RAM for LGBM Instead of 13 GB",
  "url": "/competitions/amex-default-prediction/discussion/347126",
  "author_name": "Mohamed Eltayeb",
  "post_date": "2022-08-23T01:09:41.346000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello everyone,<br>\nAs that LGBM uses the RAM, and the preprocessing and feature engineering uses the GPU, I think it is a good idea to have two notebooks. One utilizes the GPU, while the other one use the CPU. </p>\n<p>After Completeing the preprocessing in the GPU kernel use to_parquet() to save it, then use the following code to get the link of the file:</p>\n<p><code>from IPython.display import FileLink</code><br>\n<code>FileLink(r'train_df.parquet')</code></p>\n<p>And just use the link to import the data to the other kernel using read_parquet().</p>\n<p>When exporting the data, If the kernel crashed due to memory limit try to save the training data in parts:</p>\n<p><code>train_df[:int(len(train_df.index)/2)].to_parquet('train_df_1.parquet')</code><br>\n<code>train_df[int(len(train_df.index)/2):].to_parquet('train_df_2.parquet')</code></p>\n<p>Then just concatnate them when importing.</p>\n<p>NOTE: It is important to use to_pandas() before exporting the data, as that LGBM will not recognize data types of cudf library.</p>",
  "messages": [
    {
      "id": 1909820,
      "postDate": "2022-08-23T01:09:41.347Z",
      "content": "<p>Hello everyone,<br>\nAs that LGBM uses the RAM, and the preprocessing and feature engineering uses the GPU, I think it is a good idea to have two notebooks. One utilizes the GPU, while the other one use the CPU. </p>\n<p>After Completeing the preprocessing in the GPU kernel use to_parquet() to save it, then use the following code to get the link of the file:</p>\n<p><code>from IPython.display import FileLink</code><br>\n<code>FileLink(r'train_df.parquet')</code></p>\n<p>And just use the link to import the data to the other kernel using read_parquet().</p>\n<p>When exporting the data, If the kernel crashed due to memory limit try to save the training data in parts:</p>\n<p><code>train_df[:int(len(train_df.index)/2)].to_parquet('train_df_1.parquet')</code><br>\n<code>train_df[int(len(train_df.index)/2):].to_parquet('train_df_2.parquet')</code></p>\n<p>Then just concatnate them when importing.</p>\n<p>NOTE: It is important to use to_pandas() before exporting the data, as that LGBM will not recognize data types of cudf library.</p>",
      "rawMarkdown": "Hello everyone,\nAs that LGBM uses the RAM, and the preprocessing and feature engineering uses the GPU, I think it is a good idea to have two notebooks. One utilizes the GPU, while the other one use the CPU. \n\nAfter Completeing the preprocessing in the GPU kernel use to_parquet() to save it, then use the following code to get the link of the file:\n\n`from IPython.display import FileLink`\n`FileLink(r'train_df.parquet')`\n\nAnd just use the link to import the data to the other kernel using read_parquet().\n\nWhen exporting the data, If the kernel crashed due to memory limit try to save the training data in parts:\n\n`train_df[:int(len(train_df.index)/2)].to_parquet('train_df_1.parquet')`\n`train_df[int(len(train_df.index)/2):].to_parquet('train_df_2.parquet')`\n\nThen just concatnate them when importing.\n\nNOTE: It is important to use to_pandas() before exporting the data, as that LGBM will not recognize data types of cudf library.",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1909820": "Hello everyone,\nAs that LGBM uses the RAM, and the preprocessing and feature engineering uses the GPU, I think it is a good idea to have two notebooks. One utilizes the GPU, while the other one use the CPU. \n\nAfter Completeing the preprocessing in the GPU kernel use to_parquet() to save it, then use the following code to get the link of the file:\n\n`from IPython.display import FileLink`\n`FileLink(r'train_df.parquet')`\n\nAnd just use the link to import the data to the other kernel using read_parquet().\n\nWhen exporting the data, If the kernel crashed due to memory limit try to save the training data in parts:\n\n`train_df[:int(len(train_df.index)/2)].to_parquet('train_df_1.parquet')`\n`train_df[int(len(train_df.index)/2):].to_parquet('train_df_2.parquet')`\n\nThen just concatnate them when importing.\n\nNOTE: It is important to use to_pandas() before exporting the data, as that LGBM will not recognize data types of cudf library."
  }
}