{
  "id": 388729,
  "title": "GraphNet Training is Slow. Is It Reasonable?",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/388729",
  "author_name": "",
  "post_date": "2023-02-19T07:38:21.569172300Z",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I tried to train the graph-net provided by the host on my own.<br>\nI've copied all the required codes from the git, and now I can train and predict some angles.<br>\nBut I found that the training is much slower than my LSTM.</p>\n<p>To give you some numbers, training the graph-net through 50 batches of data, with batch_size of 32, took about 5 hours while training my LSTM through 50 batches with batch_size of 128 took about 3 hours, with the same machine.</p>\n<p>Well, the experiment is not controlled at all.<br>\nThis is the first time using pytorch and GNN, so I don't have any sense of time…</p>\n<p>Is this long time consumption expectable? Is GNN generally slower than small LSTM?<br>\nOr does it indicate some potential inefficiency in my code?</p>",
  "messages": [
    {
      "id": "2150371",
      "postDate": "02/19/2023 07:38:21",
      "content": "<p>I tried to train the graph-net provided by the host on my own.<br>\nI've copied all the required codes from the git, and now I can train and predict some angles.<br>\nBut I found that the training is much slower than my LSTM.</p>\n<p>To give you some numbers, training the graph-net through 50 batches of data, with batch_size of 32, took about 5 hours while training my LSTM through 50 batches with batch_size of 128 took about 3 hours, with the same machine.</p>\n<p>Well, the experiment is not controlled at all.<br>\nThis is the first time using pytorch and GNN, so I don't have any sense of time…</p>\n<p>Is this long time consumption expectable? Is GNN generally slower than small LSTM?<br>\nOr does it indicate some potential inefficiency in my code?</p>",
      "rawMarkdown": "I tried to train the graph-net provided by the host on my own.\nI've copied all the required codes from the git, and now I can train and predict some angles.\nBut I found that the training is much slower than my LSTM.\n\nTo give you some numbers, training the graph-net through 50 batches of data, with batch_size of 32, took about 5 hours while training my LSTM through 50 batches with batch_size of 128 took about 3 hours, with the same machine.\n\nWell, the experiment is not controlled at all.\nThis is the first time using pytorch and GNN, so I don't have any sense of time...\n\nIs this long time consumption expectable? Is GNN generally slower than small LSTM?\nOr does it indicate some potential inefficiency in my code?",
      "votes": null
    },
    {
      "id": "2150506",
      "postDate": "02/19/2023 10:15:36",
      "content": "<p>At first glance, it seems weird that an LSTM (so a sequential model) outperforms a fully parallelized model in speed.<br>\nWhat are your machine ressources ? My implementation of graphnet takes several hours to train too.</p>",
      "rawMarkdown": "At first glance, it seems weird that an LSTM (so a sequential model) outperforms a fully parallelized model in speed.\nWhat are your machine ressources ? My implementation of graphnet takes several hours to train too.",
      "votes": null
    },
    {
      "id": "2150766",
      "postDate": "02/19/2023 15:07:09",
      "content": "<p>I'm using a machine with RTX2080.<br>\nSo, several hours is not a big deal, right?</p>",
      "rawMarkdown": "I'm using a machine with RTX2080.\nSo, several hours is not a big deal, right?",
      "votes": null
    },
    {
      "id": "2151442",
      "postDate": "02/20/2023 05:23:44",
      "content": "<p>\"training the graph-net through 50 batches of data, with batch_size of 32, took about 5 hours\"</p>\n<p>When you say 50 batches, do you mean the batches in the /train folder in the input data, or do you mean batches of 32 samples?</p>",
      "rawMarkdown": "\"training the graph-net through 50 batches of data, with batch_size of 32, took about 5 hours\"\n\nWhen you say 50 batches, do you mean the batches in the /train folder in the input data, or do you mean batches of 32 samples?",
      "votes": null
    },
    {
      "id": "2151860",
      "postDate": "02/20/2023 12:35:32",
      "content": "<p>The earlier one. 50 batch files, each containing the 200,000 events I meant.</p>",
      "rawMarkdown": "The earlier one. 50 batch files, each containing the 200,000 events I meant.",
      "votes": null
    },
    {
      "id": "2153497",
      "postDate": "02/21/2023 13:03:53",
      "content": "<p><a href=\"https://www.kaggle.com/seungmoklee\" target=\"_blank\">@seungmoklee</a> how did you train GraphNet with 50 batch files? At least for me that would mean a ~100GB file that would definitely lead to OOM</p>",
      "rawMarkdown": "seungmoklee how did you train GraphNet with 50 batch files? At least for me that would mean a ~100GB file that would definitely lead to OOM",
      "votes": null
    },
    {
      "id": "2153554",
      "postDate": "02/21/2023 13:44:23",
      "content": "<p>Using pytorch dataloader, I could read one batch, train it, discard the batch from memory, read the next batch, train it, …</p>",
      "rawMarkdown": "Using pytorch dataloader, I could read one batch, train it, discard the batch from memory, read the next batch, train it, ...",
      "votes": null
    },
    {
      "id": "2153557",
      "postDate": "02/21/2023 13:45:32",
      "content": "<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> You can find an example in the graph-net git.</p>",
      "rawMarkdown": "alejopaullier You can find an example in the graph-net git.",
      "votes": null
    },
    {
      "id": "2153580",
      "postDate": "02/21/2023 14:06:51",
      "content": "<p>Thanks a lot. I went for the SQLite approach mentioned in the graphnet_example notebook which involves creating a huge database</p>",
      "rawMarkdown": "Thanks a lot. I went for the SQLite approach mentioned in the graphnet_example notebook which involves creating a huge database",
      "votes": null
    },
    {
      "id": "2159340",
      "postDate": "02/25/2023 17:26:32",
      "content": "<p>Are the models of comparable size? I think your shared LSTM model was way smaller than graphnet.</p>",
      "rawMarkdown": "Are the models of comparable size? I think your shared LSTM model was way smaller than graphnet.",
      "votes": null
    },
    {
      "id": "2163149",
      "postDate": "02/28/2023 16:13:02",
      "content": "<p>You can with relatively little effort switch from the SQLite backend to the parquet backend in graphnet. We have a ParquetDataset class already - with a few light modifications it'll work with the competition data. </p>",
      "rawMarkdown": "You can with relatively little effort switch from the SQLite backend to the parquet backend in graphnet. We have a ParquetDataset class already - with a few light modifications it'll work with the competition data.",
      "votes": null
    },
    {
      "id": "2166215",
      "postDate": "03/02/2023 17:25:45",
      "content": "<p><a href=\"https://www.kaggle.com/seungmoklee\" target=\"_blank\">@seungmoklee</a> would you mind sharing how to load the parquet files. I've been struggling with how to use the parquet files for this competition</p>",
      "rawMarkdown": "seungmoklee would you mind sharing how to load the parquet files. I've been struggling with how to use the parquet files for this competition",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2150506,
      "author_name": "louisstefanuto",
      "author_url": "",
      "post_date": "02/19/2023 10:15:36",
      "content": "<p>At first glance, it seems weird that an LSTM (so a sequential model) outperforms a fully parallelized model in speed.<br>\nWhat are your machine ressources ? My implementation of graphnet takes several hours to train too.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2150766,
          "author_name": "seungmoklee",
          "author_url": "",
          "post_date": "02/19/2023 15:07:09",
          "content": "<p>I'm using a machine with RTX2080.<br>\nSo, several hours is not a big deal, right?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2151442,
      "author_name": "megaray",
      "author_url": "",
      "post_date": "02/20/2023 05:23:44",
      "content": "<p>\"training the graph-net through 50 batches of data, with batch_size of 32, took about 5 hours\"</p>\n<p>When you say 50 batches, do you mean the batches in the /train folder in the input data, or do you mean batches of 32 samples?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2151860,
          "author_name": "seungmoklee",
          "author_url": "",
          "post_date": "02/20/2023 12:35:32",
          "content": "<p>The earlier one. 50 batch files, each containing the 200,000 events I meant.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2153497,
              "author_name": "alejopaullier",
              "author_url": "",
              "post_date": "02/21/2023 13:03:53",
              "content": "<p><a href=\"https://www.kaggle.com/seungmoklee\" target=\"_blank\">@seungmoklee</a> how did you train GraphNet with 50 batch files? At least for me that would mean a ~100GB file that would definitely lead to OOM</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2153554,
                  "author_name": "seungmoklee",
                  "author_url": "",
                  "post_date": "02/21/2023 13:44:23",
                  "content": "<p>Using pytorch dataloader, I could read one batch, train it, discard the batch from memory, read the next batch, train it, …</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2153557,
                      "author_name": "seungmoklee",
                      "author_url": "",
                      "post_date": "02/21/2023 13:45:32",
                      "content": "<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> You can find an example in the graph-net git.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2153580,
                          "author_name": "alejopaullier",
                          "author_url": "",
                          "post_date": "02/21/2023 14:06:51",
                          "content": "<p>Thanks a lot. I went for the SQLite approach mentioned in the graphnet_example notebook which involves creating a huge database</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2163149,
                              "author_name": "rasmusrse",
                              "author_url": "",
                              "post_date": "02/28/2023 16:13:02",
                              "content": "<p>You can with relatively little effort switch from the SQLite backend to the parquet backend in graphnet. We have a ParquetDataset class already - with a few light modifications it'll work with the competition data. </p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            },
            {
              "id": 2166215,
              "author_name": "aneighbourhood",
              "author_url": "",
              "post_date": "03/02/2023 17:25:45",
              "content": "<p><a href=\"https://www.kaggle.com/seungmoklee\" target=\"_blank\">@seungmoklee</a> would you mind sharing how to load the parquet files. I've been struggling with how to use the parquet files for this competition</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2159340,
      "author_name": "dipamc77",
      "author_url": "",
      "post_date": "02/25/2023 17:26:32",
      "content": "<p>Are the models of comparable size? I think your shared LSTM model was way smaller than graphnet.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2150371": "I tried to train the graph-net provided by the host on my own.\nI've copied all the required codes from the git, and now I can train and predict some angles.\nBut I found that the training is much slower than my LSTM.\n\nTo give you some numbers, training the graph-net through 50 batches of data, with batch_size of 32, took about 5 hours while training my LSTM through 50 batches with batch_size of 128 took about 3 hours, with the same machine.\n\nWell, the experiment is not controlled at all.\nThis is the first time using pytorch and GNN, so I don't have any sense of time...\n\nIs this long time consumption expectable? Is GNN generally slower than small LSTM?\nOr does it indicate some potential inefficiency in my code?",
    "2150506": "At first glance, it seems weird that an LSTM (so a sequential model) outperforms a fully parallelized model in speed.\nWhat are your machine ressources ? My implementation of graphnet takes several hours to train too.",
    "2150766": "I'm using a machine with RTX2080.\nSo, several hours is not a big deal, right?",
    "2151442": "\"training the graph-net through 50 batches of data, with batch_size of 32, took about 5 hours\"\n\nWhen you say 50 batches, do you mean the batches in the /train folder in the input data, or do you mean batches of 32 samples?",
    "2151860": "The earlier one. 50 batch files, each containing the 200,000 events I meant.",
    "2153497": "seungmoklee how did you train GraphNet with 50 batch files? At least for me that would mean a ~100GB file that would definitely lead to OOM",
    "2153554": "Using pytorch dataloader, I could read one batch, train it, discard the batch from memory, read the next batch, train it, ...",
    "2153557": "alejopaullier You can find an example in the graph-net git.",
    "2153580": "Thanks a lot. I went for the SQLite approach mentioned in the graphnet_example notebook which involves creating a huge database",
    "2159340": "Are the models of comparable size? I think your shared LSTM model was way smaller than graphnet.",
    "2163149": "You can with relatively little effort switch from the SQLite backend to the parquet backend in graphnet. We have a ParquetDataset class already - with a few light modifications it'll work with the competition data.",
    "2166215": "seungmoklee would you mind sharing how to load the parquet files. I've been struggling with how to use the parquet files for this competition"
  },
  "source": "meta"
}