{
  "id": 401353,
  "title": "Processing of large train data",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/401353",
  "author_name": "",
  "post_date": "2023-04-12T21:11:18.272931500Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello dear data scientists, I could see that the training files for this competition are really huge and that they quickly run out of memory, what important advice can you give me?</p>",
  "messages": [
    {
      "id": "2219769",
      "postDate": "04/12/2023 21:11:18",
      "content": "<p>Hello dear data scientists, I could see that the training files for this competition are really huge and that they quickly run out of memory, what important advice can you give me?</p>",
      "rawMarkdown": "Hello dear data scientists, I could see that the training files for this competition are really huge and that they quickly run out of memory, what important advice can you give me?",
      "votes": null
    },
    {
      "id": "2219845",
      "postDate": "04/12/2023 23:59:39",
      "content": "<p>I haven't worked on a dataset that big yet but here's what I would think of doing:</p>\n<ol>\n<li>Chunk the data. So cut it into smaller parts then preprocess the data that way - this way you know everything you need to do and don't have to wait a ton of time for it to run.</li>\n<li>Another option is to use the GPU's Google Collab offers - I believe they are stronger than what Kaggle currently has.</li>\n<li>Store the data in a relational database</li>\n</ol>\n<p>The first two options are probably the easiest out of the three.</p>",
      "rawMarkdown": "I haven't worked on a dataset that big yet but here's what I would think of doing:\n\n1. Chunk the data. So cut it into smaller parts then preprocess the data that way - this way you know everything you need to do and don't have to wait a ton of time for it to run.\n2.  Another option is to use the GPU's Google Collab offers - I believe they are stronger than what Kaggle currently has.\n3. Store the data in a relational database\n\nThe first two options are probably the easiest out of the three.",
      "votes": null
    },
    {
      "id": "2219877",
      "postDate": "04/13/2023 00:53:43",
      "content": "<p>There are many excellent notebook to solve your problem, like <a href=\"https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching\" target=\"_blank\">this</a> in which writer use chunk and sampler to avoid runing out of memory, and <a href=\"https://www.kaggle.com/code/roger92/can-we-speed-up-dataloader-io-ver2\" target=\"_blank\">this</a> in which writer use sqlite to avoid runing out of memory</p>",
      "rawMarkdown": "There are many excellent notebook to solve your problem, like [this](https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching) in which writer use chunk and sampler to avoid runing out of memory, and [this](https://www.kaggle.com/code/roger92/can-we-speed-up-dataloader-io-ver2) in which writer use sqlite to avoid runing out of memory",
      "votes": null
    },
    {
      "id": "2219892",
      "postDate": "04/13/2023 01:25:53",
      "content": "<p>Perfect, thank you Adrian…</p>",
      "rawMarkdown": "Perfect, thank you Adrian...",
      "votes": null
    },
    {
      "id": "2219894",
      "postDate": "04/13/2023 01:28:26",
      "content": "<p>Thank you, Roger… I checked your notebook…</p>",
      "rawMarkdown": "Thank you, Roger... I checked your notebook...",
      "votes": null
    },
    {
      "id": "2221302",
      "postDate": "04/14/2023 05:58:44",
      "content": "<p>Use efficient data formats such as Parquet or Feather that reduce the file size and speed up the loading time.</p>",
      "rawMarkdown": "Use efficient data formats such as Parquet or Feather that reduce the file size and speed up the loading time.",
      "votes": null
    },
    {
      "id": "2221641",
      "postDate": "04/14/2023 12:30:25",
      "content": "<p>Thank you, Yue Sun in this time I am using Vaex with great result… Thank you.</p>",
      "rawMarkdown": "Thank you, Yue Sun in this time I am using Vaex with great result... Thank you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2219845,
      "author_name": "adriandiazny",
      "author_url": "",
      "post_date": "04/12/2023 23:59:39",
      "content": "<p>I haven't worked on a dataset that big yet but here's what I would think of doing:</p>\n<ol>\n<li>Chunk the data. So cut it into smaller parts then preprocess the data that way - this way you know everything you need to do and don't have to wait a ton of time for it to run.</li>\n<li>Another option is to use the GPU's Google Collab offers - I believe they are stronger than what Kaggle currently has.</li>\n<li>Store the data in a relational database</li>\n</ol>\n<p>The first two options are probably the easiest out of the three.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2219892,
          "author_name": "henryjavier",
          "author_url": "",
          "post_date": "04/13/2023 01:25:53",
          "content": "<p>Perfect, thank you Adrian…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2219877,
      "author_name": "roger92",
      "author_url": "",
      "post_date": "04/13/2023 00:53:43",
      "content": "<p>There are many excellent notebook to solve your problem, like <a href=\"https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching\" target=\"_blank\">this</a> in which writer use chunk and sampler to avoid runing out of memory, and <a href=\"https://www.kaggle.com/code/roger92/can-we-speed-up-dataloader-io-ver2\" target=\"_blank\">this</a> in which writer use sqlite to avoid runing out of memory</p>",
      "votes": null,
      "replies": [
        {
          "id": 2219894,
          "author_name": "henryjavier",
          "author_url": "",
          "post_date": "04/13/2023 01:28:26",
          "content": "<p>Thank you, Roger… I checked your notebook…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2221302,
      "author_name": "yus002",
      "author_url": "",
      "post_date": "04/14/2023 05:58:44",
      "content": "<p>Use efficient data formats such as Parquet or Feather that reduce the file size and speed up the loading time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2221641,
          "author_name": "henryjavier",
          "author_url": "",
          "post_date": "04/14/2023 12:30:25",
          "content": "<p>Thank you, Yue Sun in this time I am using Vaex with great result… Thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2219769": "Hello dear data scientists, I could see that the training files for this competition are really huge and that they quickly run out of memory, what important advice can you give me?",
    "2219845": "I haven't worked on a dataset that big yet but here's what I would think of doing:\n\n1. Chunk the data. So cut it into smaller parts then preprocess the data that way - this way you know everything you need to do and don't have to wait a ton of time for it to run.\n2.  Another option is to use the GPU's Google Collab offers - I believe they are stronger than what Kaggle currently has.\n3. Store the data in a relational database\n\nThe first two options are probably the easiest out of the three.",
    "2219877": "There are many excellent notebook to solve your problem, like [this](https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching) in which writer use chunk and sampler to avoid runing out of memory, and [this](https://www.kaggle.com/code/roger92/can-we-speed-up-dataloader-io-ver2) in which writer use sqlite to avoid runing out of memory",
    "2219892": "Perfect, thank you Adrian...",
    "2219894": "Thank you, Roger... I checked your notebook...",
    "2221302": "Use efficient data formats such as Parquet or Feather that reduce the file size and speed up the loading time.",
    "2221641": "Thank you, Yue Sun in this time I am using Vaex with great result... Thank you."
  },
  "source": "meta"
}