{
  "id": 393936,
  "title": "How should we train 1 to 50 batches by GraphNeT?",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/393936",
  "author_name": "",
  "post_date": "2023-03-11T11:36:57.028747500Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>According to the following notebook,<br>\nwe’ve already know how to build the models by using GraphNet:<br>\n<a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-example/notebook\" target=\"_blank\">https://www.kaggle.com/code/rasmusrse/graphnet-example/notebook</a></p>\n<p>In this notebook, we can train by using batch-1 data and predict by using batch-51 data.</p>\n<p>On the other hand, the current baseline notebook is based on the pre-trained model which is based from 1 to 50 batch training data:</p>\n<p><a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission/notebook\" target=\"_blank\">https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission/notebook</a></p>\n<p>I want to reproduce this pre-trained model at the first step, but I think it’s <br>\ndifficult to generate the same model by myself due to the resource problems.</p>\n<p>For example, SQLite data from batch-1 is about 2 GB.<br>\nTherefore, we have to use 2 * 50 = 100 GB for the preparation.<br>\nIn addition, calculation time is very long if we try to train all 1 to 50 batch data.<br>\nTherefore, it’s difficult to train all data at Kaggle environment.</p>\n<p>I’m wondering what is the best method to reproduce or update the model.<br>\nIn my opinion there are several strategies to do:</p>\n<ol>\n<li>Prepare the different environment with better resource like Google Colab, GCP or local PC.</li>\n<li>Use parquet file directly instead of SQLite formula</li>\n<li>Try different method</li>\n</ol>\n<p>According to the following link, the pre-trained model was generated by the external machine and we are recommended to try to use the parquet file directly.</p>\n<p><a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-example/comments#2173476\" target=\"_blank\">https://www.kaggle.com/code/rasmusrse/graphnet-example/comments#2173476</a></p>\n<p>I’m not sure what is the best way.<br>\nI would be happy if anyone could share the ideas/suggestions.</p>",
  "messages": [
    {
      "id": "2177328",
      "postDate": "03/11/2023 11:36:57",
      "content": "<p>According to the following notebook,<br>\nwe’ve already know how to build the models by using GraphNet:<br>\n<a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-example/notebook\" target=\"_blank\">https://www.kaggle.com/code/rasmusrse/graphnet-example/notebook</a></p>\n<p>In this notebook, we can train by using batch-1 data and predict by using batch-51 data.</p>\n<p>On the other hand, the current baseline notebook is based on the pre-trained model which is based from 1 to 50 batch training data:</p>\n<p><a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission/notebook\" target=\"_blank\">https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission/notebook</a></p>\n<p>I want to reproduce this pre-trained model at the first step, but I think it’s <br>\ndifficult to generate the same model by myself due to the resource problems.</p>\n<p>For example, SQLite data from batch-1 is about 2 GB.<br>\nTherefore, we have to use 2 * 50 = 100 GB for the preparation.<br>\nIn addition, calculation time is very long if we try to train all 1 to 50 batch data.<br>\nTherefore, it’s difficult to train all data at Kaggle environment.</p>\n<p>I’m wondering what is the best method to reproduce or update the model.<br>\nIn my opinion there are several strategies to do:</p>\n<ol>\n<li>Prepare the different environment with better resource like Google Colab, GCP or local PC.</li>\n<li>Use parquet file directly instead of SQLite formula</li>\n<li>Try different method</li>\n</ol>\n<p>According to the following link, the pre-trained model was generated by the external machine and we are recommended to try to use the parquet file directly.</p>\n<p><a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-example/comments#2173476\" target=\"_blank\">https://www.kaggle.com/code/rasmusrse/graphnet-example/comments#2173476</a></p>\n<p>I’m not sure what is the best way.<br>\nI would be happy if anyone could share the ideas/suggestions.</p>",
      "rawMarkdown": "According to the following notebook,\nwe’ve already know how to build the models by using GraphNet:\nhttps://www.kaggle.com/code/rasmusrse/graphnet-example/notebook\n\nIn this notebook, we can train by using batch-1 data and predict by using batch-51 data.\n\nOn the other hand, the current baseline notebook is based on the pre-trained model which is based from 1 to 50 batch training data:\n\nhttps://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission/notebook\n\nI want to reproduce this pre-trained model at the first step, but I think it’s \ndifficult to generate the same model by myself due to the resource problems.\n\nFor example, SQLite data from batch-1 is about 2 GB.\nTherefore, we have to use 2 * 50 = 100 GB for the preparation.\nIn addition, calculation time is very long if we try to train all 1 to 50 batch data.\nTherefore, it’s difficult to train all data at Kaggle environment.\n\nI’m wondering what is the best method to reproduce or update the model.\nIn my opinion there are several strategies to do:\n\n1. Prepare the different environment with better resource like Google Colab, GCP or local PC.\n2. Use parquet file directly instead of SQLite formula\n3. Try different method\n\nAccording to the following link, the pre-trained model was generated by the external machine and we are recommended to try to use the parquet file directly.\n\nhttps://www.kaggle.com/code/rasmusrse/graphnet-example/comments#2173476\n\nI’m not sure what is the best way.\nI would be happy if anyone could share the ideas/suggestions.",
      "votes": null
    },
    {
      "id": "2177473",
      "postDate": "03/11/2023 14:03:10",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/393851#2176951\" target=\"_blank\">https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/393851#2176951</a></p>",
      "rawMarkdown": "https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/393851#2176951",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2177473,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "03/11/2023 14:03:10",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/393851#2176951\" target=\"_blank\">https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/393851#2176951</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2177328": "According to the following notebook,\nwe’ve already know how to build the models by using GraphNet:\nhttps://www.kaggle.com/code/rasmusrse/graphnet-example/notebook\n\nIn this notebook, we can train by using batch-1 data and predict by using batch-51 data.\n\nOn the other hand, the current baseline notebook is based on the pre-trained model which is based from 1 to 50 batch training data:\n\nhttps://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission/notebook\n\nI want to reproduce this pre-trained model at the first step, but I think it’s \ndifficult to generate the same model by myself due to the resource problems.\n\nFor example, SQLite data from batch-1 is about 2 GB.\nTherefore, we have to use 2 * 50 = 100 GB for the preparation.\nIn addition, calculation time is very long if we try to train all 1 to 50 batch data.\nTherefore, it’s difficult to train all data at Kaggle environment.\n\nI’m wondering what is the best method to reproduce or update the model.\nIn my opinion there are several strategies to do:\n\n1. Prepare the different environment with better resource like Google Colab, GCP or local PC.\n2. Use parquet file directly instead of SQLite formula\n3. Try different method\n\nAccording to the following link, the pre-trained model was generated by the external machine and we are recommended to try to use the parquet file directly.\n\nhttps://www.kaggle.com/code/rasmusrse/graphnet-example/comments#2173476\n\nI’m not sure what is the best way.\nI would be happy if anyone could share the ideas/suggestions.",
    "2177473": "https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/393851#2176951"
  },
  "source": "meta"
}