{
  "id": 541057,
  "title": "Suggestion Needed for the Platform for Training the Model",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/541057",
  "author_name": "",
  "post_date": "2024-10-17T08:41:05.331803800Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello Everyone, I am getting out of memory error on Kaggle for this competition, where are you guys training this models? Can you please suggest which platform you are using for training the model?</p>",
  "messages": [
    {
      "id": "3020184",
      "postDate": "10/17/2024 08:41:05",
      "content": "<p>Hello Everyone, I am getting out of memory error on Kaggle for this competition, where are you guys training this models? Can you please suggest which platform you are using for training the model?</p>",
      "rawMarkdown": "Hello Everyone, I am getting out of memory error on Kaggle for this competition, where are you guys training this models? Can you please suggest which platform you are using for training the model?",
      "votes": null
    },
    {
      "id": "3020186",
      "postDate": "10/17/2024 08:43:47",
      "content": "<p>Try to load a small partition like (part-0.parquet) for training.</p>",
      "rawMarkdown": "Try to load a small partition like (part-0.parquet) for training.",
      "votes": null
    },
    {
      "id": "3020188",
      "postDate": "10/17/2024 08:44:06",
      "content": "<p>Use <a href=\"runpod.io\" target=\"_blank\">runpod.io</a> - Kaggle kernels will create OOM issues with such a large dataset <a href=\"https://www.kaggle.com/salman1127\" target=\"_blank\">@salman1127</a> \nTheir A6000 GPU is very good</p>",
      "rawMarkdown": "Use [runpod.io](runpod.io) - Kaggle kernels will create OOM issues with such a large dataset @salman1127 \nTheir A6000 GPU is very good",
      "votes": null
    },
    {
      "id": "3036627",
      "postDate": "11/04/2024 18:40:57",
      "content": "<p>How do I use the model here on kaggle then?</p>",
      "rawMarkdown": "How do I use the model here on kaggle then?",
      "votes": null
    },
    {
      "id": "3036668",
      "postDate": "11/04/2024 19:34:29",
      "content": "<p>Save the model using joblib and infer here on Kaggle <a href=\"https://www.kaggle.com/koushiksahu\" target=\"_blank\">@koushiksahu</a> </p>",
      "rawMarkdown": "Save the model using joblib and infer here on Kaggle @koushiksahu",
      "votes": null
    },
    {
      "id": "3062167",
      "postDate": "12/03/2024 11:00:05",
      "content": "<p>thanks for tthe suggestion!</p>",
      "rawMarkdown": "thanks for tthe suggestion!",
      "votes": null
    },
    {
      "id": "3087051",
      "postDate": "01/03/2025 03:28:50",
      "content": "<p>import dask.dataframe as dd</p>\n<p>def load_data_dask(file_paths):<br>\n    ddf = dd.read_parquet(file_paths, engine='pyarrow')<br>\n    return ddf.compute()</p>\n<p>train_files = [f'{input_path}/train.parquet/partition_id={i}/part-0.parquet' for i in range(10)]<br>\ndf = load_data_dask(train_files[0:10])</p>",
      "rawMarkdown": "import dask.dataframe as dd\n\ndef load_data_dask(file_paths):\n    ddf = dd.read_parquet(file_paths, engine='pyarrow')\n    return ddf.compute()\n    \ntrain_files = [f'{input_path}/train.parquet/partition_id={i}/part-0.parquet' for i in range(10)]\ndf = load_data_dask(train_files[0:10])",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3020186,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "10/17/2024 08:43:47",
      "content": "<p>Try to load a small partition like (part-0.parquet) for training.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3020188,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "10/17/2024 08:44:06",
      "content": "<p>Use <a href=\"runpod.io\" target=\"_blank\">runpod.io</a> - Kaggle kernels will create OOM issues with such a large dataset <a href=\"https://www.kaggle.com/salman1127\" target=\"_blank\">@salman1127</a> \nTheir A6000 GPU is very good</p>",
      "votes": null,
      "replies": [
        {
          "id": 3036627,
          "author_name": "koushiksahu",
          "author_url": "",
          "post_date": "11/04/2024 18:40:57",
          "content": "<p>How do I use the model here on kaggle then?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3036668,
              "author_name": "ravi20076",
              "author_url": "",
              "post_date": "11/04/2024 19:34:29",
              "content": "<p>Save the model using joblib and infer here on Kaggle <a href=\"https://www.kaggle.com/koushiksahu\" target=\"_blank\">@koushiksahu</a> </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3062167,
          "author_name": "junhaochan",
          "author_url": "",
          "post_date": "12/03/2024 11:00:05",
          "content": "<p>thanks for tthe suggestion!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3087051,
      "author_name": "namjos",
      "author_url": "",
      "post_date": "01/03/2025 03:28:50",
      "content": "<p>import dask.dataframe as dd</p>\n<p>def load_data_dask(file_paths):<br>\n    ddf = dd.read_parquet(file_paths, engine='pyarrow')<br>\n    return ddf.compute()</p>\n<p>train_files = [f'{input_path}/train.parquet/partition_id={i}/part-0.parquet' for i in range(10)]<br>\ndf = load_data_dask(train_files[0:10])</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3020184": "Hello Everyone, I am getting out of memory error on Kaggle for this competition, where are you guys training this models? Can you please suggest which platform you are using for training the model?",
    "3020186": "Try to load a small partition like (part-0.parquet) for training.",
    "3020188": "Use [runpod.io](runpod.io) - Kaggle kernels will create OOM issues with such a large dataset @salman1127 \nTheir A6000 GPU is very good",
    "3036627": "How do I use the model here on kaggle then?",
    "3036668": "Save the model using joblib and infer here on Kaggle @koushiksahu",
    "3062167": "thanks for tthe suggestion!",
    "3087051": "import dask.dataframe as dd\n\ndef load_data_dask(file_paths):\n    ddf = dd.read_parquet(file_paths, engine='pyarrow')\n    return ddf.compute()\n    \ntrain_files = [f'{input_path}/train.parquet/partition_id={i}/part-0.parquet' for i in range(10)]\ndf = load_data_dask(train_files[0:10])"
  },
  "source": "meta"
}