{
  "id": 391166,
  "title": "How can you train many batches (>15) in Kaggle?",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/391166",
  "author_name": "",
  "post_date": "2023-02-28T15:59:32.580823600Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I've been following the graphnet-example uploaded by <a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a>. I have (unsuccessfully) been trying to train on larger samples of data since, from empirical experience, more data = more performance. However, I've reached a plateau where I am restricted by memory constraints. Has anyone managed to train Graphnet (on Kaggle) with large batches of data? </p>\n<p>Note: for each batch (parquet train file) the corresponding SQLite db is ~2GB</p>",
  "messages": [
    {
      "id": "2163136",
      "postDate": "02/28/2023 15:59:32",
      "content": "<p>I've been following the graphnet-example uploaded by <a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a>. I have (unsuccessfully) been trying to train on larger samples of data since, from empirical experience, more data = more performance. However, I've reached a plateau where I am restricted by memory constraints. Has anyone managed to train Graphnet (on Kaggle) with large batches of data? </p>\n<p>Note: for each batch (parquet train file) the corresponding SQLite db is ~2GB</p>",
      "rawMarkdown": "I've been following the graphnet-example uploaded by @rasmusrse. I have (unsuccessfully) been trying to train on larger samples of data since, from empirical experience, more data = more performance. However, I've reached a plateau where I am restricted by memory constraints. Has anyone managed to train Graphnet (on Kaggle) with large batches of data? \n\nNote: for each batch (parquet train file) the corresponding SQLite db is ~2GB",
      "votes": null
    },
    {
      "id": "2163155",
      "postDate": "02/28/2023 16:18:08",
      "content": "<p>You may want to rewrite the dataloader for it to upload the new db and erase the old one from memory each time when previous db is completely processed by batches</p>\n<p>But again, as Rasmus has pointed out, that may be a walkaround for some different problem</p>",
      "rawMarkdown": "You may want to rewrite the dataloader for it to upload the new db and erase the old one from memory each time when previous db is completely processed by batches\n\nBut again, as Rasmus has pointed out, that may be a walkaround for some different problem",
      "votes": null
    },
    {
      "id": "2163163",
      "postDate": "02/28/2023 16:23:10",
      "content": "<p>15 batches in SQLite format should take up around 30 gb of disk space - not RAM. Using the example, you should never have more than batch_size*(1 + prefetch_factor) events in memory (because of SQLite) - so I'm surprised that you're running out of memory. Could you elaborate on what you're doing?</p>",
      "rawMarkdown": "15 batches in SQLite format should take up around 30 gb of disk space - not RAM. Using the example, you should never have more than batch_size*(1 + prefetch_factor) events in memory (because of SQLite) - so I'm surprised that you're running out of memory. Could you elaborate on what you're doing?",
      "votes": null
    },
    {
      "id": "2164196",
      "postDate": "03/01/2023 11:49:54",
      "content": "<p>Hello! I have same problem with <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a>, when I try to convert multiple batches into a single database by adjusting batch_id in kaggle env. Even if I convert just one batch, as the code indicates, a memory(RAM) overflow will occur. <br>\n<code>convert_to_sqlite(meta_data_path,\n                  database_path=database_path,\n                  input_data_folder=input_data_folder,\n                  batch_ids = [51])</code></p>",
      "rawMarkdown": "Hello! I have same problem with @alejopaullier, when I try to convert multiple batches into a single database by adjusting batch_id in kaggle env. Even if I convert just one batch, as the code indicates, a memory(RAM) overflow will occur. \n`convert_to_sqlite(meta_data_path,\n                  database_path=database_path,\n                  input_data_folder=input_data_folder,\n                  batch_ids = [51])`",
      "votes": null
    },
    {
      "id": "2164289",
      "postDate": "03/01/2023 13:06:39",
      "content": "<p>Thanks Rasmus and sorry for not being so clear. When I create a database from many batches the output is larger than what Kaggle supports in the output directory (<code>/kaggle/working/</code>), which is around ~20GB, which limits the size of the database I can create.</p>",
      "rawMarkdown": "Thanks Rasmus and sorry for not being so clear. When I create a database from many batches the output is larger than what Kaggle supports in the output directory (`/kaggle/working/`), which is around ~20GB, which limits the size of the database I can create.",
      "votes": null
    },
    {
      "id": "2166105",
      "postDate": "03/02/2023 16:08:10",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> <a href=\"https://www.kaggle.com/lfz5cv\" target=\"_blank\">@lfz5cv</a> please try to put them in <code>/tmp</code>. For example,</p>\n<pre><code>SAVE_FOLDER = \"/tmp/output/\"\nos.makedirs(SAVE_FOLDER, exist_ok=True)\n</code></pre>\n<p>Then you can save up to 70GB data on disk.</p>",
      "rawMarkdown": "Hi @alejopaullier @lfz5cv please try to put them in `/tmp`. For example,\n\n```\nSAVE_FOLDER = \"/tmp/output/\"\nos.makedirs(SAVE_FOLDER, exist_ok=True)\n```\n\nThen you can save up to 70GB data on disk.",
      "votes": null
    },
    {
      "id": "2187301",
      "postDate": "03/18/2023 15:55:04",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a>, <a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a><br>\nI have similar problems.</p>\n<p>By using the sample code, we can generate the SQLite table separately like:</p>\n<pre><code>batch_1.db\nbatch_2.db\n...\nbatch_50.db\n</code></pre>\n<p>, which table can be saved at the storage and each db is ~2GB.</p>\n<p>However, when we train by using these all batch (1-50), I think we have to merge the db like <code>batch_1_50.db</code>.<br>\nIn my understanding, we can set the config file like:</p>\n<pre><code>config = {\n        \"path\": f'{INPUT_DIR}/batch_1_50.db',\n        \"inference_database_path\": f'{INPUT_DIR}/batch_51.db',\n        ...\n        'train_selection': './train_selection_max_200_pulses_1_50_batch.csv',\n        'validate_selection': './validate_selection_max_200_pulses_51_batch.csv',\n        ...\n}\n</code></pre>\n<p>.<br>\nTherefore, we have to load <code>batch_1_50.db</code> to the memory which is huge size(2GB * 50 ~ 100GB).<br>\nI'm wondering we need to prepare more than 100GB size memory.<br>\nI'm not sure this idea is correct or not.<br>\nI would appreciate if you or anyone could answer my question.</p>",
      "rawMarkdown": "Hello, @alejopaullier, @rasmusrse\nI have similar problems.\n\nBy using the sample code, we can generate the SQLite table separately like:\n```\nbatch_1.db\nbatch_2.db\n...\nbatch_50.db\n```\n, which table can be saved at the storage and each db is ~2GB.\n\nHowever, when we train by using these all batch (1-50), I think we have to merge the db like `batch_1_50.db`.\nIn my understanding, we can set the config file like:\n\n```\nconfig = {\n        \"path\": f'{INPUT_DIR}/batch_1_50.db',\n        \"inference_database_path\": f'{INPUT_DIR}/batch_51.db',\n        ...\n        'train_selection': './train_selection_max_200_pulses_1_50_batch.csv',\n        'validate_selection': './validate_selection_max_200_pulses_51_batch.csv',\n        ...\n}\n```\n.\nTherefore, we have to load `batch_1_50.db` to the memory which is huge size(2GB * 50 ~ 100GB).\nI'm wondering we need to prepare more than 100GB size memory.\nI'm not sure this idea is correct or not.\nI would appreciate if you or anyone could answer my question.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2163155,
      "author_name": "avtobusbratiev",
      "author_url": "",
      "post_date": "02/28/2023 16:18:08",
      "content": "<p>You may want to rewrite the dataloader for it to upload the new db and erase the old one from memory each time when previous db is completely processed by batches</p>\n<p>But again, as Rasmus has pointed out, that may be a walkaround for some different problem</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2163163,
      "author_name": "rasmusrse",
      "author_url": "",
      "post_date": "02/28/2023 16:23:10",
      "content": "<p>15 batches in SQLite format should take up around 30 gb of disk space - not RAM. Using the example, you should never have more than batch_size*(1 + prefetch_factor) events in memory (because of SQLite) - so I'm surprised that you're running out of memory. Could you elaborate on what you're doing?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2164196,
          "author_name": "lfz5cv",
          "author_url": "",
          "post_date": "03/01/2023 11:49:54",
          "content": "<p>Hello! I have same problem with <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a>, when I try to convert multiple batches into a single database by adjusting batch_id in kaggle env. Even if I convert just one batch, as the code indicates, a memory(RAM) overflow will occur. <br>\n<code>convert_to_sqlite(meta_data_path,\n                  database_path=database_path,\n                  input_data_folder=input_data_folder,\n                  batch_ids = [51])</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2164289,
          "author_name": "alejopaullier",
          "author_url": "",
          "post_date": "03/01/2023 13:06:39",
          "content": "<p>Thanks Rasmus and sorry for not being so clear. When I create a database from many batches the output is larger than what Kaggle supports in the output directory (<code>/kaggle/working/</code>), which is around ~20GB, which limits the size of the database I can create.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2166105,
              "author_name": "forcewithme",
              "author_url": "",
              "post_date": "03/02/2023 16:08:10",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> <a href=\"https://www.kaggle.com/lfz5cv\" target=\"_blank\">@lfz5cv</a> please try to put them in <code>/tmp</code>. For example,</p>\n<pre><code>SAVE_FOLDER = \"/tmp/output/\"\nos.makedirs(SAVE_FOLDER, exist_ok=True)\n</code></pre>\n<p>Then you can save up to 70GB data on disk.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2187301,
      "author_name": "tetsuro731",
      "author_url": "",
      "post_date": "03/18/2023 15:55:04",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a>, <a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a><br>\nI have similar problems.</p>\n<p>By using the sample code, we can generate the SQLite table separately like:</p>\n<pre><code>batch_1.db\nbatch_2.db\n...\nbatch_50.db\n</code></pre>\n<p>, which table can be saved at the storage and each db is ~2GB.</p>\n<p>However, when we train by using these all batch (1-50), I think we have to merge the db like <code>batch_1_50.db</code>.<br>\nIn my understanding, we can set the config file like:</p>\n<pre><code>config = {\n        \"path\": f'{INPUT_DIR}/batch_1_50.db',\n        \"inference_database_path\": f'{INPUT_DIR}/batch_51.db',\n        ...\n        'train_selection': './train_selection_max_200_pulses_1_50_batch.csv',\n        'validate_selection': './validate_selection_max_200_pulses_51_batch.csv',\n        ...\n}\n</code></pre>\n<p>.<br>\nTherefore, we have to load <code>batch_1_50.db</code> to the memory which is huge size(2GB * 50 ~ 100GB).<br>\nI'm wondering we need to prepare more than 100GB size memory.<br>\nI'm not sure this idea is correct or not.<br>\nI would appreciate if you or anyone could answer my question.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2163136": "I've been following the graphnet-example uploaded by @rasmusrse. I have (unsuccessfully) been trying to train on larger samples of data since, from empirical experience, more data = more performance. However, I've reached a plateau where I am restricted by memory constraints. Has anyone managed to train Graphnet (on Kaggle) with large batches of data? \n\nNote: for each batch (parquet train file) the corresponding SQLite db is ~2GB",
    "2163155": "You may want to rewrite the dataloader for it to upload the new db and erase the old one from memory each time when previous db is completely processed by batches\n\nBut again, as Rasmus has pointed out, that may be a walkaround for some different problem",
    "2163163": "15 batches in SQLite format should take up around 30 gb of disk space - not RAM. Using the example, you should never have more than batch_size*(1 + prefetch_factor) events in memory (because of SQLite) - so I'm surprised that you're running out of memory. Could you elaborate on what you're doing?",
    "2164196": "Hello! I have same problem with @alejopaullier, when I try to convert multiple batches into a single database by adjusting batch_id in kaggle env. Even if I convert just one batch, as the code indicates, a memory(RAM) overflow will occur. \n`convert_to_sqlite(meta_data_path,\n                  database_path=database_path,\n                  input_data_folder=input_data_folder,\n                  batch_ids = [51])`",
    "2164289": "Thanks Rasmus and sorry for not being so clear. When I create a database from many batches the output is larger than what Kaggle supports in the output directory (`/kaggle/working/`), which is around ~20GB, which limits the size of the database I can create.",
    "2166105": "Hi @alejopaullier @lfz5cv please try to put them in `/tmp`. For example,\n\n```\nSAVE_FOLDER = \"/tmp/output/\"\nos.makedirs(SAVE_FOLDER, exist_ok=True)\n```\n\nThen you can save up to 70GB data on disk.",
    "2187301": "Hello, @alejopaullier, @rasmusrse\nI have similar problems.\n\nBy using the sample code, we can generate the SQLite table separately like:\n```\nbatch_1.db\nbatch_2.db\n...\nbatch_50.db\n```\n, which table can be saved at the storage and each db is ~2GB.\n\nHowever, when we train by using these all batch (1-50), I think we have to merge the db like `batch_1_50.db`.\nIn my understanding, we can set the config file like:\n\n```\nconfig = {\n        \"path\": f'{INPUT_DIR}/batch_1_50.db',\n        \"inference_database_path\": f'{INPUT_DIR}/batch_51.db',\n        ...\n        'train_selection': './train_selection_max_200_pulses_1_50_batch.csv',\n        'validate_selection': './validate_selection_max_200_pulses_51_batch.csv',\n        ...\n}\n```\n.\nTherefore, we have to load `batch_1_50.db` to the memory which is huge size(2GB * 50 ~ 100GB).\nI'm wondering we need to prepare more than 100GB size memory.\nI'm not sure this idea is correct or not.\nI would appreciate if you or anyone could answer my question."
  },
  "source": "meta"
}