{
  "id": 397688,
  "title": "The training dataset into small chunks",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/397688",
  "author_name": "",
  "post_date": "2023-03-26T18:40:39.033309800Z",
  "votes": 6,
  "comment_count": 11,
  "views": 0,
  "content": "<p>When using statistics such as the mean obtained by groupby <code>session_id</code> and <code>level_group</code> as features, it is not necessary to store the entire time series dataset in memory. Therefore, the policy is to split train.csv into smaller chunks and calculate statistics for each of these chunks. However, it must be handled in such a way that session_id is not split into multiple chunks.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2473952%2F1688baae242a52befb659da01ffbbb40%2Fchunk_image.png?generation=1679855369596237&amp;alt=media\" alt=\"chunks_image\"></p>\n<p>The specific process is shown in <a href=\"https://www.kaggle.com/kentatanaka1data/the-training-dataset-into-small-chunks\" target=\"_blank\">my notebook</a>. If you know of any good libraries, etc. that better optimize the process of splitting into chunks as it is done in my notebook, please let me know.</p>",
  "messages": [
    {
      "id": "2198194",
      "postDate": "03/26/2023 18:40:39",
      "content": "<p>When using statistics such as the mean obtained by groupby <code>session_id</code> and <code>level_group</code> as features, it is not necessary to store the entire time series dataset in memory. Therefore, the policy is to split train.csv into smaller chunks and calculate statistics for each of these chunks. However, it must be handled in such a way that session_id is not split into multiple chunks.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2473952%2F1688baae242a52befb659da01ffbbb40%2Fchunk_image.png?generation=1679855369596237&amp;alt=media\" alt=\"chunks_image\"></p>\n<p>The specific process is shown in <a href=\"https://www.kaggle.com/kentatanaka1data/the-training-dataset-into-small-chunks\" target=\"_blank\">my notebook</a>. If you know of any good libraries, etc. that better optimize the process of splitting into chunks as it is done in my notebook, please let me know.</p>",
      "rawMarkdown": "When using statistics such as the mean obtained by groupby `session_id` and `level_group` as features, it is not necessary to store the entire time series dataset in memory. Therefore, the policy is to split train.csv into smaller chunks and calculate statistics for each of these chunks. However, it must be handled in such a way that session_id is not split into multiple chunks.\n\n![chunks_image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2473952%2F1688baae242a52befb659da01ffbbb40%2Fchunk_image.png?generation=1679855369596237&alt=media)\n\nThe specific process is shown in [my notebook](https://www.kaggle.com/kentatanaka1data/the-training-dataset-into-small-chunks). If you know of any good libraries, etc. that better optimize the process of splitting into chunks as it is done in my notebook, please let me know.",
      "votes": null
    },
    {
      "id": "2198232",
      "postDate": "03/26/2023 19:22:52",
      "content": "<p>Thanks for sharing. I also read train data in chunks in my notebook <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680\" target=\"_blank\">here</a>. The technique i use is to first read the entire <code>session_id</code> column and then determine where rows can safely be split (in code cell #2). Afterward (in code cell #7), i use <code>nrows</code> and <code>skip_rows</code> and read chunks manually.</p>",
      "rawMarkdown": "Thanks for sharing. I also read train data in chunks in my notebook [here][1]. The technique i use is to first read the entire `session_id` column and then determine where rows can safely be split (in code cell #2). Afterward (in code cell #7), i use `nrows` and `skip_rows` and read chunks manually.\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680",
      "votes": null
    },
    {
      "id": "2198245",
      "postDate": "03/26/2023 19:51:30",
      "content": "<p>Thank you for your comment. I have been referring to your notebook for a long time, but I did not know that you had added a new process to separate the chunks. I thought it is simple and looked good, but I didn't know where you defined <code>reads</code> and <code>skips</code>. I will read the code a little more closely.</p>",
      "rawMarkdown": "Thank you for your comment. I have been referring to your notebook for a long time, but I did not know that you had added a new process to separate the chunks. I thought it is simple and looked good, but I didn't know where you defined `reads` and `skips`. I will read the code a little more closely.",
      "votes": null
    },
    {
      "id": "2198283",
      "postDate": "03/26/2023 21:02:20",
      "content": "<p>My original notebook did not use chunks because the original competition was small enough. After Kaggle doubled the train data, i needed to add chunks to avoid memory error. The variables <code>reads</code> and <code>skips</code> are defined in the hidden <code>code cell #2</code>, so you'll need to click \"Show hidden code\" to see it.</p>",
      "rawMarkdown": "My original notebook did not use chunks because the original competition was small enough. After Kaggle doubled the train data, i needed to add chunks to avoid memory error. The variables `reads` and `skips` are defined in the hidden `code cell #2`, so you'll need to click \"Show hidden code\" to see it.",
      "votes": null
    },
    {
      "id": "2198455",
      "postDate": "03/27/2023 03:33:37",
      "content": "<p>I overlooked the hidden code in cell#2. Thanks for the additional explanation.</p>",
      "rawMarkdown": "I overlooked the hidden code in cell#2. Thanks for the additional explanation.",
      "votes": null
    },
    {
      "id": "2198610",
      "postDate": "03/27/2023 07:20:38",
      "content": "<p>I recommend using <code>polars</code> instead of <code>pandas</code>. It's much more memory efficient, and you <a href=\"https://www.kaggle.com/code/demche/polars-memory-usage-optimization\" target=\"_blank\">don't even need</a> to read the entire data from a CSV file.</p>",
      "rawMarkdown": "I recommend using `polars` instead of `pandas`. It's much more memory efficient, and you [don't even need](https://www.kaggle.com/code/demche/polars-memory-usage-optimization) to read the entire data from a CSV file.",
      "votes": null
    },
    {
      "id": "2198681",
      "postDate": "03/27/2023 08:39:47",
      "content": "<p>Thank you for sharing the information, I will refer to the lazy load method to narrow down the target dataset using <code>scan_csv</code>, which is not possible with pandas.</p>",
      "rawMarkdown": "Thank you for sharing the information, I will refer to the lazy load method to narrow down the target dataset using `scan_csv`, which is not possible with pandas.",
      "votes": null
    },
    {
      "id": "2198904",
      "postDate": "03/27/2023 11:11:43",
      "content": "<p>Thanks for sharing your fabulous insights.</p>",
      "rawMarkdown": "Thanks for sharing your fabulous insights.",
      "votes": null
    },
    {
      "id": "2199456",
      "postDate": "03/27/2023 18:06:41",
      "content": "<p>The entire train can be downloaded in a 953 MB dataframe <a href=\"https://www.kaggle.com/code/vadimkamaev/reading-data-953-mb\" target=\"_blank\">https://www.kaggle.com/code/vadimkamaev/reading-data-953-mb</a></p>",
      "rawMarkdown": "The entire train can be downloaded in a 953 MB dataframe https://www.kaggle.com/code/vadimkamaev/reading-data-953-mb",
      "votes": null
    },
    {
      "id": "2199477",
      "postDate": "03/27/2023 18:21:50",
      "content": "<p>Thanks for sharing the information. It is an important perspective to use memory efficiently by specifying the dtype of each column.<br>\nBy the way, the purpose of this issue is that I want to split up the training data. That is because I think it is useful when we have more training data or when we need more memory for feature engineering.</p>",
      "rawMarkdown": "Thanks for sharing the information. It is an important perspective to use memory efficiently by specifying the dtype of each column.\nBy the way, the purpose of this issue is that I want to split up the training data. That is because I think it is useful when we have more training data or when we need more memory for feature engineering.",
      "votes": null
    },
    {
      "id": "2226955",
      "postDate": "04/19/2023 12:07:04",
      "content": "<p>Thanks for sharing!!　It is a very attractive code.　I will refer to it.　</p>",
      "rawMarkdown": "Thanks for sharing!!　It is a very attractive code.　I will refer to it.",
      "votes": null
    },
    {
      "id": "2226976",
      "postDate": "04/19/2023 12:27:54",
      "content": "<p>Thank you very much. However, please note that this code requires that the data in the resource file to be in the order of session_id, as in  this competition's dataset. So it may be better to use other code that others have commented on.</p>",
      "rawMarkdown": "Thank you very much. However, please note that this code requires that the data in the resource file to be in the order of session_id, as in  this competition's dataset. So it may be better to use other code that others have commented on.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2198232,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/26/2023 19:22:52",
      "content": "<p>Thanks for sharing. I also read train data in chunks in my notebook <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680\" target=\"_blank\">here</a>. The technique i use is to first read the entire <code>session_id</code> column and then determine where rows can safely be split (in code cell #2). Afterward (in code cell #7), i use <code>nrows</code> and <code>skip_rows</code> and read chunks manually.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2198245,
          "author_name": "kentatanaka1data",
          "author_url": "",
          "post_date": "03/26/2023 19:51:30",
          "content": "<p>Thank you for your comment. I have been referring to your notebook for a long time, but I did not know that you had added a new process to separate the chunks. I thought it is simple and looked good, but I didn't know where you defined <code>reads</code> and <code>skips</code>. I will read the code a little more closely.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2198283,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "03/26/2023 21:02:20",
              "content": "<p>My original notebook did not use chunks because the original competition was small enough. After Kaggle doubled the train data, i needed to add chunks to avoid memory error. The variables <code>reads</code> and <code>skips</code> are defined in the hidden <code>code cell #2</code>, so you'll need to click \"Show hidden code\" to see it.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2198455,
                  "author_name": "kentatanaka1data",
                  "author_url": "",
                  "post_date": "03/27/2023 03:33:37",
                  "content": "<p>I overlooked the hidden code in cell#2. Thanks for the additional explanation.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2198610,
      "author_name": "demche",
      "author_url": "",
      "post_date": "03/27/2023 07:20:38",
      "content": "<p>I recommend using <code>polars</code> instead of <code>pandas</code>. It's much more memory efficient, and you <a href=\"https://www.kaggle.com/code/demche/polars-memory-usage-optimization\" target=\"_blank\">don't even need</a> to read the entire data from a CSV file.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2198681,
          "author_name": "kentatanaka1data",
          "author_url": "",
          "post_date": "03/27/2023 08:39:47",
          "content": "<p>Thank you for sharing the information, I will refer to the lazy load method to narrow down the target dataset using <code>scan_csv</code>, which is not possible with pandas.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2198904,
      "author_name": "tariqbashir",
      "author_url": "",
      "post_date": "03/27/2023 11:11:43",
      "content": "<p>Thanks for sharing your fabulous insights.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2199456,
      "author_name": "vadimkamaev",
      "author_url": "",
      "post_date": "03/27/2023 18:06:41",
      "content": "<p>The entire train can be downloaded in a 953 MB dataframe <a href=\"https://www.kaggle.com/code/vadimkamaev/reading-data-953-mb\" target=\"_blank\">https://www.kaggle.com/code/vadimkamaev/reading-data-953-mb</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2199477,
          "author_name": "kentatanaka1data",
          "author_url": "",
          "post_date": "03/27/2023 18:21:50",
          "content": "<p>Thanks for sharing the information. It is an important perspective to use memory efficiently by specifying the dtype of each column.<br>\nBy the way, the purpose of this issue is that I want to split up the training data. That is because I think it is useful when we have more training data or when we need more memory for feature engineering.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2226955,
      "author_name": "risakashiwabara",
      "author_url": "",
      "post_date": "04/19/2023 12:07:04",
      "content": "<p>Thanks for sharing!!　It is a very attractive code.　I will refer to it.　</p>",
      "votes": null,
      "replies": [
        {
          "id": 2226976,
          "author_name": "kentatanaka1data",
          "author_url": "",
          "post_date": "04/19/2023 12:27:54",
          "content": "<p>Thank you very much. However, please note that this code requires that the data in the resource file to be in the order of session_id, as in  this competition's dataset. So it may be better to use other code that others have commented on.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2198194": "When using statistics such as the mean obtained by groupby `session_id` and `level_group` as features, it is not necessary to store the entire time series dataset in memory. Therefore, the policy is to split train.csv into smaller chunks and calculate statistics for each of these chunks. However, it must be handled in such a way that session_id is not split into multiple chunks.\n\n![chunks_image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2473952%2F1688baae242a52befb659da01ffbbb40%2Fchunk_image.png?generation=1679855369596237&alt=media)\n\nThe specific process is shown in [my notebook](https://www.kaggle.com/kentatanaka1data/the-training-dataset-into-small-chunks). If you know of any good libraries, etc. that better optimize the process of splitting into chunks as it is done in my notebook, please let me know.",
    "2198232": "Thanks for sharing. I also read train data in chunks in my notebook [here][1]. The technique i use is to first read the entire `session_id` column and then determine where rows can safely be split (in code cell #2). Afterward (in code cell #7), i use `nrows` and `skip_rows` and read chunks manually.\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680",
    "2198245": "Thank you for your comment. I have been referring to your notebook for a long time, but I did not know that you had added a new process to separate the chunks. I thought it is simple and looked good, but I didn't know where you defined `reads` and `skips`. I will read the code a little more closely.",
    "2198283": "My original notebook did not use chunks because the original competition was small enough. After Kaggle doubled the train data, i needed to add chunks to avoid memory error. The variables `reads` and `skips` are defined in the hidden `code cell #2`, so you'll need to click \"Show hidden code\" to see it.",
    "2198455": "I overlooked the hidden code in cell#2. Thanks for the additional explanation.",
    "2198610": "I recommend using `polars` instead of `pandas`. It's much more memory efficient, and you [don't even need](https://www.kaggle.com/code/demche/polars-memory-usage-optimization) to read the entire data from a CSV file.",
    "2198681": "Thank you for sharing the information, I will refer to the lazy load method to narrow down the target dataset using `scan_csv`, which is not possible with pandas.",
    "2198904": "Thanks for sharing your fabulous insights.",
    "2199456": "The entire train can be downloaded in a 953 MB dataframe https://www.kaggle.com/code/vadimkamaev/reading-data-953-mb",
    "2199477": "Thanks for sharing the information. It is an important perspective to use memory efficiently by specifying the dtype of each column.\nBy the way, the purpose of this issue is that I want to split up the training data. That is because I think it is useful when we have more training data or when we need more memory for feature engineering.",
    "2226955": "Thanks for sharing!!　It is a very attractive code.　I will refer to it.",
    "2226976": "Thank you very much. However, please note that this code requires that the data in the resource file to be in the order of session_id, as in  this competition's dataset. So it may be better to use other code that others have commented on."
  },
  "source": "meta"
}