{
  "id": 497296,
  "title": "Strategies to train a model with a large data set ??",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/497296",
  "author_name": "C.O",
  "post_date": "2024-04-24T07:17:59.217000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>This is the first time I'm going to train a model on a large dataset. I'm having memory problems and wanted to know if you had any tips or resources (videos, articles, etc.) to give me on how to train a model with a large dataset. </p>\n<p>The first idea I had was to take a sample with the .sample pandas method of around 60_000 rows and increase this sample until the performance no longer increases or simply take a fixed sample without increasing it but I don't know if this is a good idea or a good practice. </p>",
  "messages": [
    {
      "id": 2771363,
      "postDate": "2024-04-24T07:59:31.837Z",
      "content": "<p>There are two ways to solve memory problems: reducing data memory or increasing the amount of memory that can be accommodated.</p>\n<p>Reducing the memory of data is the function: reduce_memory_usage,example:<a href=\"https://www.kaggle.com/code/yunsuxiaozi/home-credit-linearregression-is-all-you-need\" target=\"_blank\">https://www.kaggle.com/code/yunsuxiaozi/home-credit-linearregression-is-all-you-need</a> .</p>\n<p>Increasing the amount of memory that can be accommodated is equivalent to using polars instead of pandas.</p>",
      "rawMarkdown": "There are two ways to solve memory problems: reducing data memory or increasing the amount of memory that can be accommodated.\n\nReducing the memory of data is the function: reduce_memory_usage,example:https://www.kaggle.com/code/yunsuxiaozi/home-credit-linearregression-is-all-you-need .\n\nIncreasing the amount of memory that can be accommodated is equivalent to using polars instead of pandas.",
      "votes": 1,
      "replies": [
        {
          "id": 2771387,
          "postDate": "2024-04-24T08:17:24.537Z",
          "content": "<p>Thanks for your answer, I already use this function reduce_memory_usage but when I want to train for example an XGBoost on all the data it's impossible that's why I'm wondering if taking a sample is a good practice when dealing with large datasets and you can't increase the amount of memory.</p>",
          "rawMarkdown": "Thanks for your answer, I already use this function reduce_memory_usage but when I want to train for example an XGBoost on all the data it's impossible that's why I'm wondering if taking a sample is a good practice when dealing with large datasets and you can't increase the amount of memory."
        }
      ]
    },
    {
      "id": 2771290,
      "postDate": "2024-04-24T07:17:59.217Z",
      "content": "<p>This is the first time I'm going to train a model on a large dataset. I'm having memory problems and wanted to know if you had any tips or resources (videos, articles, etc.) to give me on how to train a model with a large dataset. </p>\n<p>The first idea I had was to take a sample with the .sample pandas method of around 60_000 rows and increase this sample until the performance no longer increases or simply take a fixed sample without increasing it but I don't know if this is a good idea or a good practice. </p>",
      "rawMarkdown": "This is the first time I'm going to train a model on a large dataset. I'm having memory problems and wanted to know if you had any tips or resources (videos, articles, etc.) to give me on how to train a model with a large dataset. \n\nThe first idea I had was to take a sample with the .sample pandas method of around 60_000 rows and increase this sample until the performance no longer increases or simply take a fixed sample without increasing it but I don't know if this is a good idea or a good practice. ",
      "votes": 1
    },
    {
      "id": 2772081,
      "postDate": "2024-04-24T14:17:20.417Z",
      "content": "<p>You can try to train offline, and then only infer online.</p>",
      "rawMarkdown": "You can try to train offline, and then only infer online.",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2771363,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "2024-04-24T07:59:31.837000",
      "content": "<p>There are two ways to solve memory problems: reducing data memory or increasing the amount of memory that can be accommodated.</p>\n<p>Reducing the memory of data is the function: reduce_memory_usage,example:<a href=\"https://www.kaggle.com/code/yunsuxiaozi/home-credit-linearregression-is-all-you-need\" target=\"_blank\">https://www.kaggle.com/code/yunsuxiaozi/home-credit-linearregression-is-all-you-need</a> .</p>\n<p>Increasing the amount of memory that can be accommodated is equivalent to using polars instead of pandas.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2771387,
          "author_name": "C.O",
          "author_url": "",
          "post_date": "2024-04-24T08:17:24.537000",
          "content": "<p>Thanks for your answer, I already use this function reduce_memory_usage but when I want to train for example an XGBoost on all the data it's impossible that's why I'm wondering if taking a sample is a good practice when dealing with large datasets and you can't increase the amount of memory.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2772081,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-24T14:17:20.417000",
      "content": "<p>You can try to train offline, and then only infer online.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2771363": "There are two ways to solve memory problems: reducing data memory or increasing the amount of memory that can be accommodated.\n\nReducing the memory of data is the function: reduce_memory_usage,example:https://www.kaggle.com/code/yunsuxiaozi/home-credit-linearregression-is-all-you-need .\n\nIncreasing the amount of memory that can be accommodated is equivalent to using polars instead of pandas.",
    "2771290": "This is the first time I'm going to train a model on a large dataset. I'm having memory problems and wanted to know if you had any tips or resources (videos, articles, etc.) to give me on how to train a model with a large dataset. \n\nThe first idea I had was to take a sample with the .sample pandas method of around 60_000 rows and increase this sample until the performance no longer increases or simply take a fixed sample without increasing it but I don't know if this is a good idea or a good practice. ",
    "2772081": "You can try to train offline, and then only infer online."
  }
}