{
  "id": 541242,
  "title": "how much memory needed for train all data online?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/541242",
  "author_name": "",
  "post_date": "2024-10-18T09:35:46.003511700Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>“Your notebook tried to allocate more memory than is available.”</p>\n<p>I forked one public training notebook, above error is reported.</p>",
  "messages": [
    {
      "id": "3021235",
      "postDate": "10/18/2024 09:35:46",
      "content": "<p>“Your notebook tried to allocate more memory than is available.”</p>\n<p>I forked one public training notebook, above error is reported.</p>",
      "rawMarkdown": "“Your notebook tried to allocate more memory than is available.”\n\nI forked one public training notebook, above error is reported.",
      "votes": null
    },
    {
      "id": "3021243",
      "postDate": "10/18/2024 09:53:23",
      "content": "<p>I recommend you to avoid kaggle for training models in this assignment due to the training data size <a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">@dragonzhang</a> </p>",
      "rawMarkdown": "I recommend you to avoid kaggle for training models in this assignment due to the training data size @dragonzhang",
      "votes": null
    },
    {
      "id": "3021387",
      "postDate": "10/18/2024 12:57:30",
      "content": "<p>thx for your reply.  However, I have no better resource than Kaggle.    no money for computer power.<br>\nI have managed to train partial data.</p>",
      "rawMarkdown": "thx for your reply.  However, I have no better resource than Kaggle.    no money for computer power.\nI have managed to train partial data.",
      "votes": null
    },
    {
      "id": "3024783",
      "postDate": "10/22/2024 02:01:57",
      "content": "<p>maybe you can try this:</p>\n<pre><code>import polars as pl\nimport pandas as pd\nfrom tqdm import tqdm\nimport gc\n\nlazy_df = pl()\n\nunique_date_ids = lazy_df(pl()())()()()\n\ndf = pd()\n\nbatch_size = \n   ((, (unique_date_ids), batch_size), desc=):\n    batch_date_ids = unique_date_ids\n    df_chunk = lazy_df(pl()(batch_date_ids))()()\n    df = pd(, ignore_index=True)\n     df_chunk\n    gc() \n</code></pre>\n<p>but it may take you sometime to load the data.</p>",
      "rawMarkdown": "maybe you can try this:\n\n```\nimport polars as pl\nimport pandas as pd\nfrom tqdm import tqdm\nimport gc\n\nlazy_df = pl.scan_parquet('/kaggle/input/jane-street-realtime-marketdata-forecasting/train.parquet')\n\nunique_date_ids = lazy_df.select(pl.col(\"date_id\").unique()).collect().to_series().to_list()\n\ndf = pd.DataFrame()\n\nbatch_size = 10\nfor i in tqdm(range(0, len(unique_date_ids[-500:]), batch_size), desc=\"Processing batches\"):\n    batch_date_ids = unique_date_ids[-500:][i:i + batch_size]\n    df_chunk = lazy_df.filter(pl.col(\"date_id\").is_in(batch_date_ids)).collect().to_pandas()\n    df = pd.concat([df, df_chunk], ignore_index=True)\n    del df_chunk\n    gc.collect() \n```\n\nbut it may take you sometime to load the data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3021243,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "10/18/2024 09:53:23",
      "content": "<p>I recommend you to avoid kaggle for training models in this assignment due to the training data size <a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">@dragonzhang</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 3021387,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "10/18/2024 12:57:30",
          "content": "<p>thx for your reply.  However, I have no better resource than Kaggle.    no money for computer power.<br>\nI have managed to train partial data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3024783,
      "author_name": "carrotwait",
      "author_url": "",
      "post_date": "10/22/2024 02:01:57",
      "content": "<p>maybe you can try this:</p>\n<pre><code>import polars as pl\nimport pandas as pd\nfrom tqdm import tqdm\nimport gc\n\nlazy_df = pl()\n\nunique_date_ids = lazy_df(pl()())()()()\n\ndf = pd()\n\nbatch_size = \n   ((, (unique_date_ids), batch_size), desc=):\n    batch_date_ids = unique_date_ids\n    df_chunk = lazy_df(pl()(batch_date_ids))()()\n    df = pd(, ignore_index=True)\n     df_chunk\n    gc() \n</code></pre>\n<p>but it may take you sometime to load the data.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3021235": "“Your notebook tried to allocate more memory than is available.”\n\nI forked one public training notebook, above error is reported.",
    "3021243": "I recommend you to avoid kaggle for training models in this assignment due to the training data size @dragonzhang",
    "3021387": "thx for your reply.  However, I have no better resource than Kaggle.    no money for computer power.\nI have managed to train partial data.",
    "3024783": "maybe you can try this:\n\n```\nimport polars as pl\nimport pandas as pd\nfrom tqdm import tqdm\nimport gc\n\nlazy_df = pl.scan_parquet('/kaggle/input/jane-street-realtime-marketdata-forecasting/train.parquet')\n\nunique_date_ids = lazy_df.select(pl.col(\"date_id\").unique()).collect().to_series().to_list()\n\ndf = pd.DataFrame()\n\nbatch_size = 10\nfor i in tqdm(range(0, len(unique_date_ids[-500:]), batch_size), desc=\"Processing batches\"):\n    batch_date_ids = unique_date_ids[-500:][i:i + batch_size]\n    df_chunk = lazy_df.filter(pl.col(\"date_id\").is_in(batch_date_ids)).collect().to_pandas()\n    df = pd.concat([df, df_chunk], ignore_index=True)\n    del df_chunk\n    gc.collect() \n```\n\nbut it may take you sometime to load the data."
  },
  "source": "meta"
}