{
  "id": 383524,
  "title": "GraphNeT: DynEdge Baseline & Example",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/383524",
  "author_name": "",
  "post_date": "2023-02-04T06:38:48.120316500Z",
  "votes": 66,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hello everyone!</p>\n<p><a href=\"https://github.com/graphnet-team/graphnet\" target=\"_blank\">GraphNeT</a> is an open-source Python framework aimed at providing high quality, user friendly, end-to-end functionality to perform reconstruction tasks at neutrino telescopes using graph neural networks (GNNs). In a <a href=\"https://iopscience.iop.org/article/10.1088/1748-0221/17/11/P11003\" target=\"_blank\">recent paper</a> we showcased how a GNN implemented in GraphNeT performs on a series of reconstruction tasks in IceCube. </p>\n<p>We've created an example notebook that contains everything you need to install GraphNeT and train/apply this model to the competition data. Since the code is public, you should have every opportunity to tinker with the library, the GNN or to implement a completely different solution using GraphNeT. The notebook contains a pre-trained version of the GNN shown in the paper with a similar training configuration, which has been trained on batches 1 to 50. The notebook is available <a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-example\" target=\"_blank\">here</a>. </p>\n<p>We've submitted this pre-trained model to the leaderboard using <a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission\" target=\"_blank\">this notebook</a> as \"GraphNeT Baseline\". <strong>This submission is not a prize contender</strong>.</p>",
  "messages": [
    {
      "id": "2128870",
      "postDate": "02/04/2023 06:38:48",
      "content": "<p>Hello everyone!</p>\n<p><a href=\"https://github.com/graphnet-team/graphnet\" target=\"_blank\">GraphNeT</a> is an open-source Python framework aimed at providing high quality, user friendly, end-to-end functionality to perform reconstruction tasks at neutrino telescopes using graph neural networks (GNNs). In a <a href=\"https://iopscience.iop.org/article/10.1088/1748-0221/17/11/P11003\" target=\"_blank\">recent paper</a> we showcased how a GNN implemented in GraphNeT performs on a series of reconstruction tasks in IceCube. </p>\n<p>We've created an example notebook that contains everything you need to install GraphNeT and train/apply this model to the competition data. Since the code is public, you should have every opportunity to tinker with the library, the GNN or to implement a completely different solution using GraphNeT. The notebook contains a pre-trained version of the GNN shown in the paper with a similar training configuration, which has been trained on batches 1 to 50. The notebook is available <a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-example\" target=\"_blank\">here</a>. </p>\n<p>We've submitted this pre-trained model to the leaderboard using <a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission\" target=\"_blank\">this notebook</a> as \"GraphNeT Baseline\". <strong>This submission is not a prize contender</strong>.</p>",
      "rawMarkdown": "Hello everyone!\n\n[GraphNeT](https://github.com/graphnet-team/graphnet) is an open-source Python framework aimed at providing high quality, user friendly, end-to-end functionality to perform reconstruction tasks at neutrino telescopes using graph neural networks (GNNs). In a [recent paper](https://iopscience.iop.org/article/10.1088/1748-0221/17/11/P11003) we showcased how a GNN implemented in GraphNeT performs on a series of reconstruction tasks in IceCube. \n\nWe've created an example notebook that contains everything you need to install GraphNeT and train/apply this model to the competition data. Since the code is public, you should have every opportunity to tinker with the library, the GNN or to implement a completely different solution using GraphNeT. The notebook contains a pre-trained version of the GNN shown in the paper with a similar training configuration, which has been trained on batches 1 to 50. The notebook is available [here](https://www.kaggle.com/code/rasmusrse/graphnet-example). \n\nWe've submitted this pre-trained model to the leaderboard using [this notebook](https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission) as \"GraphNeT Baseline\". **This submission is not a prize contender**.",
      "votes": null
    },
    {
      "id": "2128904",
      "postDate": "02/04/2023 07:03:56",
      "content": "<p>Thank you for providing the baseline example. I believe it's a great approach that will enhance the competition and eventually improve the accuracy of neutrino reconstruction :) The performance of DybeEdge right out of the box is really impressive. How long does it take to train the model on 50 batches?</p>",
      "rawMarkdown": "Thank you for providing the baseline example. I believe it's a great approach that will enhance the competition and eventually improve the accuracy of neutrino reconstruction :) The performance of DybeEdge right out of the box is really impressive. How long does it take to train the model on 50 batches?",
      "votes": null
    },
    {
      "id": "2128927",
      "postDate": "02/04/2023 07:26:48",
      "content": "<p>Thank you! The training itself should conclude in well under the 12h Kaggle limit if on GPU. However, the conversion of parquet files to sqlite for 50 batches would probably take close to 12 hours. As I point out in the end of the example notebook, one could just modify the dataset class to read the competition parquet files directly instead.</p>",
      "rawMarkdown": "Thank you! The training itself should conclude in well under the 12h Kaggle limit if on GPU. However, the conversion of parquet files to sqlite for 50 batches would probably take close to 12 hours. As I point out in the end of the example notebook, one could just modify the dataset class to read the competition parquet files directly instead.",
      "votes": null
    },
    {
      "id": "2128983",
      "postDate": "02/04/2023 08:33:15",
      "content": "<p>Having baseline is great, early sharing prize was also a great idea imo. idk why <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> was having second thoughts about sharing or not his solution, it will for sure be overruned with only 2 weeks into competition. But 2 first weeks probably is not enough time to build a solid alternatives to existing solutions. Most probable ESP winner just first to semi-implement this same baseline shared by <a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a></p>",
      "rawMarkdown": "Having baseline is great, early sharing prize was also a great idea imo. idk why @anjum48 was having second thoughts about sharing or not his solution, it will for sure be overruned with only 2 weeks into competition. But 2 first weeks probably is not enough time to build a solid alternatives to existing solutions. Most probable ESP winner just first to semi-implement this same baseline shared by @rasmusrse",
      "votes": null
    },
    {
      "id": "2128995",
      "postDate": "02/04/2023 08:37:22",
      "content": "<p>You're totally right. The DynNet baseline is pretty solid IMO too - I have a long list of things that I tried that didn't work. It'll be interesting to see what people come up with</p>",
      "rawMarkdown": "You're totally right. The DynNet baseline is pretty solid IMO too - I have a long list of things that I tried that didn't work. It'll be interesting to see what people come up with",
      "votes": null
    },
    {
      "id": "2128998",
      "postDate": "02/04/2023 08:40:15",
      "content": "<p>just curious, what did you decide on sharing before you notice 1.018 graphnet baseline on LB? And how many graphnet solutions you managed to train over what, a week? given other things that didnt work</p>",
      "rawMarkdown": "just curious, what did you decide on sharing before you notice 1.018 graphnet baseline on LB? And how many graphnet solutions you managed to train over what, a week? given other things that didnt work",
      "votes": null
    },
    {
      "id": "2129004",
      "postDate": "02/04/2023 08:47:17",
      "content": "<p>I knew the Graphnet baseline was going to appear anyway. I saw that commits were being made in the Graphnet repo specific to this competition over the last few days.</p>\n<p>Regarding other solutions, I have a tonne of other GNNs, a few of which are now starting to surpass DynEdge. But in terms of training speed &amp; simplicity, DynEdge is very strong</p>",
      "rawMarkdown": "I knew the Graphnet baseline was going to appear anyway. I saw that commits were being made in the Graphnet repo specific to this competition over the last few days.\n\nRegarding other solutions, I have a tonne of other GNNs, a few of which are now starting to surpass DynEdge. But in terms of training speed & simplicity, DynEdge is very strong",
      "votes": null
    },
    {
      "id": "2129015",
      "postDate": "02/04/2023 08:56:25",
      "content": "<p>Thank you for sharing, i hope you ll get ESP, well deserved</p>",
      "rawMarkdown": "Thank you for sharing, i hope you ll get ESP, well deserved",
      "votes": null
    },
    {
      "id": "2129066",
      "postDate": "02/04/2023 09:38:03",
      "content": "<p>great sharing</p>",
      "rawMarkdown": "great sharing",
      "votes": null
    },
    {
      "id": "2129385",
      "postDate": "02/04/2023 15:34:48",
      "content": "<p>Roughly how long does a LB submission take with this baseline notebook?</p>",
      "rawMarkdown": "Roughly how long does a LB submission take with this baseline notebook?",
      "votes": null
    },
    {
      "id": "2129410",
      "postDate": "02/04/2023 15:59:39",
      "content": "<p>It takes less than 3 hours for the submission notebook to get scored. </p>",
      "rawMarkdown": "It takes less than 3 hours for the submission notebook to get scored.",
      "votes": null
    },
    {
      "id": "2129623",
      "postDate": "02/04/2023 18:56:07",
      "content": "<p><a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a> I get out of memory error on kaggle with GPU when trying to add in the code for converting to sqlite format below.</p>\n<pre><code>input_data_folder = \ngeometry_table = pd.read_csv()\nmeta_data_path = \n\n\ndatabase_path = \nconvert_to_sqlite(meta_data_path,\n                  database_path=database_path,\n                  input_data_folder=input_data_folder,\n                  batch_ids = [])\n</code></pre>\n<p>Note that if trying CPU run, I get this Import error (condensed):</p>\n<pre><code>OSError                                   Traceback (most recent call last)\n---&gt;   graphnet.data.sqlite.sqlite_utilities  create_table\n----&gt;       .sqlite_dataset  SQLiteDataset\n----&gt;   graphnet.data.dataset  Dataset, ColumnMissingException\n---&gt;   torch_geometric.data  Data\n----&gt;   torch_geometric.data\n----&gt;   .data  Data\n----&gt;   torch_sparse  SparseTensor\n/opt/conda/lib/python3/site-packages/torch_sparse/__init__.py  &lt;module&gt;\n          spec = cuda_spec  cpu_spec\n           spec   :\n---&gt;          torch.ops.load_library(spec.origin)\n/opt/conda/lib/python3/site-packages/torch/_ops.py  load_library(self, path)\n                 \n                 \n--&gt;              ctypes.CDLL(path)\n/opt/conda/lib/python3/ctypes/__init__.py  __init__(self, name, mode, handle, use_errno, use_last_error)\n              handle  :\n--&gt;              self._handle = _dlopen(self._name, mode)\n\nOSError: libcusparse.so: cannot  shared  file: No such file  directory\n</code></pre>\n<p>Any suggestions on fixing it?</p>\n<p>I suppose CPU run for this preprocessing step would be much better on kaggle anyways, if there's a way to fix that second error. Or a streamlined import with fewer pip installs or manual addition of just create_table might be even better?</p>\n<p>I would assume that the 32GB memory of CPU notebooks would be more than enough to eliminate the memory error side of things.</p>",
      "rawMarkdown": "rasmusrse I get out of memory error on kaggle with GPU when trying to add in the code for converting to sqlite format below.\n\n```python\ninput_data_folder = '/kaggle/input/icecube-neutrinos-in-deep-ice/train'\ngeometry_table = pd.read_csv('/kaggle/input/icecube-neutrinos-in-deep-ice/sensor_geometry.csv')\nmeta_data_path = '/kaggle/input/icecube-neutrinos-in-deep-ice/train_meta.parquet'\n\n#batch_51\ndatabase_path = '/kaggle/working/batch_51'\nconvert_to_sqlite(meta_data_path,\n                  database_path=database_path,\n                  input_data_folder=input_data_folder,\n                  batch_ids = [51])\n```\n\n\nNote that if trying CPU run, I get this Import error (condensed):\n\n```python\nOSError                                   Traceback (most recent call last)\n---> 10 from graphnet.data.sqlite.sqlite_utilities import create_table\n----> 9     from .sqlite_dataset import SQLiteDataset\n----> 7 from graphnet.data.dataset import Dataset, ColumnMissingException\n---> 10 from torch_geometric.data import Data\n----> 4 import torch_geometric.data\n----> 1 from .data import Data\n----> 9 from torch_sparse import SparseTensor\n/opt/conda/lib/python3.7/site-packages/torch_sparse/__init__.py in <module>\n     17     spec = cuda_spec or cpu_spec\n     18     if spec is not None:\n---> 19         torch.ops.load_library(spec.origin)\n/opt/conda/lib/python3.7/site-packages/torch/_ops.py in load_library(self, path)\n    218             # static (global) initialization code in order to register custom\n    219             # operators with the JIT.\n--> 220             ctypes.CDLL(path)\n/opt/conda/lib/python3.7/ctypes/__init__.py in __init__(self, name, mode, handle, use_errno, use_last_error)\n    363         if handle is None:\n--> 364             self._handle = _dlopen(self._name, mode)\n\nOSError: libcusparse.so.11: cannot open shared object file: No such file or directory\n```\n\nAny suggestions on fixing it?\n\nI suppose CPU run for this preprocessing step would be much better on kaggle anyways, if there's a way to fix that second error. Or a streamlined import with fewer pip installs or manual addition of just create_table might be even better?\n\nI would assume that the 32GB memory of CPU notebooks would be more than enough to eliminate the memory error side of things.",
      "votes": null
    },
    {
      "id": "2129688",
      "postDate": "02/04/2023 19:54:27",
      "content": "<p>Update: To create a CPU notebook that just does this task, I did the following:</p>\n<ul>\n<li>removed graphnet-related pip install and imports</li>\n<li>import sqlite3</li>\n<li>manually copied into my notebook three functions from graphnet's <a href=\"https://github.com/graphnet-team/graphnet/blob/57fd0f37b65cfc8f2171b800353e5a4c40e752cb/src/graphnet/data/sqlite/sqlite_utilities.py\" target=\"_blank\">sqlite_utilities.py</a> file:<ul>\n<li>run_sql_code</li>\n<li>attach_index</li>\n<li>create_table</li></ul></li>\n</ul>\n<p>Then it ran successfully in 10 minutes (599 seconds) for a single batch [52].</p>",
      "rawMarkdown": "Update: To create a CPU notebook that just does this task, I did the following:\n\n* removed graphnet-related pip install and imports\n* import sqlite3\n* manually copied into my notebook three functions from graphnet's [sqlite_utilities.py](https://github.com/graphnet-team/graphnet/blob/57fd0f37b65cfc8f2171b800353e5a4c40e752cb/src/graphnet/data/sqlite/sqlite_utilities.py) file:\n  * run_sql_code\n  * attach_index\n  * create_table\n\nThen it ran successfully in 10 minutes (599 seconds) for a single batch [52].",
      "votes": null
    },
    {
      "id": "2129987",
      "postDate": "02/05/2023 04:32:54",
      "content": "<p><a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a> Man, having published paper AND code, you are golden! Thanks</p>",
      "rawMarkdown": "rasmusrse Man, having published paper AND code, you are golden! Thanks",
      "votes": null
    },
    {
      "id": "2130177",
      "postDate": "02/05/2023 08:50:21",
      "content": "<p>Hey Robert - glad to hear you got it working. For reference, the error message you show here typically indicates that graphnet has been installed with GPU on a system without one. </p>",
      "rawMarkdown": "Hey Robert - glad to hear you got it working. For reference, the error message you show here typically indicates that graphnet has been installed with GPU on a system without one.",
      "votes": null
    },
    {
      "id": "2135094",
      "postDate": "02/08/2023 12:56:58",
      "content": "<p>If I want to download the dependecies on my on-prem environment, the way to do it is getting them <br>\nby using:<br>\nkaggle kernels output rasmusrse/graphnet-baseline-submission -p /path/to/dest</p>\n<p>Sorry for stupid question I am very new to kaggle and this is my first competition :)<br>\nThanks,</p>",
      "rawMarkdown": "If I want to download the dependecies on my on-prem environment, the way to do it is getting them \nby using:\nkaggle kernels output rasmusrse/graphnet-baseline-submission -p /path/to/dest\n\nSorry for stupid question I am very new to kaggle and this is my first competition :)\nThanks,",
      "votes": null
    },
    {
      "id": "2160316",
      "postDate": "02/26/2023 15:11:05",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a> thank you a lot for your great work!! Is there a reason why you use the direction task and not the combination of the azimuth/zenith tasks ? Do you think one approach is better than the other?</p>",
      "rawMarkdown": "Hi @rasmusrse thank you a lot for your great work!! Is there a reason why you use the direction task and not the combination of the azimuth/zenith tasks ? Do you think one approach is better than the other?",
      "votes": null
    },
    {
      "id": "2164986",
      "postDate": "03/01/2023 22:46:57",
      "content": "<p>Thank you for publishing the baseline!</p>",
      "rawMarkdown": "Thank you for publishing the baseline!",
      "votes": null
    },
    {
      "id": "2171502",
      "postDate": "03/06/2023 20:53:54",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a>  I'm getting the following error from the baseline notebook in the cell where the model, build, inference, and data functions are defined:<br>\n<code>OSError: /opt/conda/lib/python3.7/site-packages/torch_sparse/_version_cuda.so: undefined symbol: _ZN5torch3jit17parseSchemaOrNameERKSs</code></p>\n<p>Anyone have any idea what's going on?</p>",
      "rawMarkdown": "Hi @rasmusrse  I'm getting the following error from the baseline notebook in the cell where the model, build, inference, and data functions are defined:\n`OSError: /opt/conda/lib/python3.7/site-packages/torch_sparse/_version_cuda.so: undefined symbol: _ZN5torch3jit17parseSchemaOrNameERKSs`\n\nAnyone have any idea what's going on?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2128904,
      "author_name": "inartimiryasov",
      "author_url": "",
      "post_date": "02/04/2023 07:03:56",
      "content": "<p>Thank you for providing the baseline example. I believe it's a great approach that will enhance the competition and eventually improve the accuracy of neutrino reconstruction :) The performance of DybeEdge right out of the box is really impressive. How long does it take to train the model on 50 batches?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2128927,
          "author_name": "rasmusrse",
          "author_url": "",
          "post_date": "02/04/2023 07:26:48",
          "content": "<p>Thank you! The training itself should conclude in well under the 12h Kaggle limit if on GPU. However, the conversion of parquet files to sqlite for 50 batches would probably take close to 12 hours. As I point out in the end of the example notebook, one could just modify the dataset class to read the competition parquet files directly instead.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2128983,
      "author_name": "bakeryproducts",
      "author_url": "",
      "post_date": "02/04/2023 08:33:15",
      "content": "<p>Having baseline is great, early sharing prize was also a great idea imo. idk why <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> was having second thoughts about sharing or not his solution, it will for sure be overruned with only 2 weeks into competition. But 2 first weeks probably is not enough time to build a solid alternatives to existing solutions. Most probable ESP winner just first to semi-implement this same baseline shared by <a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2128995,
          "author_name": "anjum48",
          "author_url": "",
          "post_date": "02/04/2023 08:37:22",
          "content": "<p>You're totally right. The DynNet baseline is pretty solid IMO too - I have a long list of things that I tried that didn't work. It'll be interesting to see what people come up with</p>",
          "votes": null,
          "replies": [
            {
              "id": 2128998,
              "author_name": "bakeryproducts",
              "author_url": "",
              "post_date": "02/04/2023 08:40:15",
              "content": "<p>just curious, what did you decide on sharing before you notice 1.018 graphnet baseline on LB? And how many graphnet solutions you managed to train over what, a week? given other things that didnt work</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2129004,
                  "author_name": "anjum48",
                  "author_url": "",
                  "post_date": "02/04/2023 08:47:17",
                  "content": "<p>I knew the Graphnet baseline was going to appear anyway. I saw that commits were being made in the Graphnet repo specific to this competition over the last few days.</p>\n<p>Regarding other solutions, I have a tonne of other GNNs, a few of which are now starting to surpass DynEdge. But in terms of training speed &amp; simplicity, DynEdge is very strong</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2129015,
                      "author_name": "bakeryproducts",
                      "author_url": "",
                      "post_date": "02/04/2023 08:56:25",
                      "content": "<p>Thank you for sharing, i hope you ll get ESP, well deserved</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2129066,
      "author_name": "moro146",
      "author_url": "",
      "post_date": "02/04/2023 09:38:03",
      "content": "<p>great sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2129385,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "02/04/2023 15:34:48",
      "content": "<p>Roughly how long does a LB submission take with this baseline notebook?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2129410,
          "author_name": "rasmusrse",
          "author_url": "",
          "post_date": "02/04/2023 15:59:39",
          "content": "<p>It takes less than 3 hours for the submission notebook to get scored. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2129623,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "02/04/2023 18:56:07",
      "content": "<p><a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a> I get out of memory error on kaggle with GPU when trying to add in the code for converting to sqlite format below.</p>\n<pre><code>input_data_folder = \ngeometry_table = pd.read_csv()\nmeta_data_path = \n\n\ndatabase_path = \nconvert_to_sqlite(meta_data_path,\n                  database_path=database_path,\n                  input_data_folder=input_data_folder,\n                  batch_ids = [])\n</code></pre>\n<p>Note that if trying CPU run, I get this Import error (condensed):</p>\n<pre><code>OSError                                   Traceback (most recent call last)\n---&gt;   graphnet.data.sqlite.sqlite_utilities  create_table\n----&gt;       .sqlite_dataset  SQLiteDataset\n----&gt;   graphnet.data.dataset  Dataset, ColumnMissingException\n---&gt;   torch_geometric.data  Data\n----&gt;   torch_geometric.data\n----&gt;   .data  Data\n----&gt;   torch_sparse  SparseTensor\n/opt/conda/lib/python3/site-packages/torch_sparse/__init__.py  &lt;module&gt;\n          spec = cuda_spec  cpu_spec\n           spec   :\n---&gt;          torch.ops.load_library(spec.origin)\n/opt/conda/lib/python3/site-packages/torch/_ops.py  load_library(self, path)\n                 \n                 \n--&gt;              ctypes.CDLL(path)\n/opt/conda/lib/python3/ctypes/__init__.py  __init__(self, name, mode, handle, use_errno, use_last_error)\n              handle  :\n--&gt;              self._handle = _dlopen(self._name, mode)\n\nOSError: libcusparse.so: cannot  shared  file: No such file  directory\n</code></pre>\n<p>Any suggestions on fixing it?</p>\n<p>I suppose CPU run for this preprocessing step would be much better on kaggle anyways, if there's a way to fix that second error. Or a streamlined import with fewer pip installs or manual addition of just create_table might be even better?</p>\n<p>I would assume that the 32GB memory of CPU notebooks would be more than enough to eliminate the memory error side of things.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2129688,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "02/04/2023 19:54:27",
          "content": "<p>Update: To create a CPU notebook that just does this task, I did the following:</p>\n<ul>\n<li>removed graphnet-related pip install and imports</li>\n<li>import sqlite3</li>\n<li>manually copied into my notebook three functions from graphnet's <a href=\"https://github.com/graphnet-team/graphnet/blob/57fd0f37b65cfc8f2171b800353e5a4c40e752cb/src/graphnet/data/sqlite/sqlite_utilities.py\" target=\"_blank\">sqlite_utilities.py</a> file:<ul>\n<li>run_sql_code</li>\n<li>attach_index</li>\n<li>create_table</li></ul></li>\n</ul>\n<p>Then it ran successfully in 10 minutes (599 seconds) for a single batch [52].</p>",
          "votes": null,
          "replies": [
            {
              "id": 2130177,
              "author_name": "rasmusrse",
              "author_url": "",
              "post_date": "02/05/2023 08:50:21",
              "content": "<p>Hey Robert - glad to hear you got it working. For reference, the error message you show here typically indicates that graphnet has been installed with GPU on a system without one. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2129987,
      "author_name": "minhbtnguyen",
      "author_url": "",
      "post_date": "02/05/2023 04:32:54",
      "content": "<p><a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a> Man, having published paper AND code, you are golden! Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2135094,
      "author_name": "riccardogatti",
      "author_url": "",
      "post_date": "02/08/2023 12:56:58",
      "content": "<p>If I want to download the dependecies on my on-prem environment, the way to do it is getting them <br>\nby using:<br>\nkaggle kernels output rasmusrse/graphnet-baseline-submission -p /path/to/dest</p>\n<p>Sorry for stupid question I am very new to kaggle and this is my first competition :)<br>\nThanks,</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2160316,
      "author_name": "matthiasanderer",
      "author_url": "",
      "post_date": "02/26/2023 15:11:05",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a> thank you a lot for your great work!! Is there a reason why you use the direction task and not the combination of the azimuth/zenith tasks ? Do you think one approach is better than the other?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2164986,
      "author_name": "yassinetoughrai",
      "author_url": "",
      "post_date": "03/01/2023 22:46:57",
      "content": "<p>Thank you for publishing the baseline!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2171502,
      "author_name": "samuelzxu",
      "author_url": "",
      "post_date": "03/06/2023 20:53:54",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a>  I'm getting the following error from the baseline notebook in the cell where the model, build, inference, and data functions are defined:<br>\n<code>OSError: /opt/conda/lib/python3.7/site-packages/torch_sparse/_version_cuda.so: undefined symbol: _ZN5torch3jit17parseSchemaOrNameERKSs</code></p>\n<p>Anyone have any idea what's going on?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2128870": "Hello everyone!\n\n[GraphNeT](https://github.com/graphnet-team/graphnet) is an open-source Python framework aimed at providing high quality, user friendly, end-to-end functionality to perform reconstruction tasks at neutrino telescopes using graph neural networks (GNNs). In a [recent paper](https://iopscience.iop.org/article/10.1088/1748-0221/17/11/P11003) we showcased how a GNN implemented in GraphNeT performs on a series of reconstruction tasks in IceCube. \n\nWe've created an example notebook that contains everything you need to install GraphNeT and train/apply this model to the competition data. Since the code is public, you should have every opportunity to tinker with the library, the GNN or to implement a completely different solution using GraphNeT. The notebook contains a pre-trained version of the GNN shown in the paper with a similar training configuration, which has been trained on batches 1 to 50. The notebook is available [here](https://www.kaggle.com/code/rasmusrse/graphnet-example). \n\nWe've submitted this pre-trained model to the leaderboard using [this notebook](https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission) as \"GraphNeT Baseline\". **This submission is not a prize contender**.",
    "2128904": "Thank you for providing the baseline example. I believe it's a great approach that will enhance the competition and eventually improve the accuracy of neutrino reconstruction :) The performance of DybeEdge right out of the box is really impressive. How long does it take to train the model on 50 batches?",
    "2128927": "Thank you! The training itself should conclude in well under the 12h Kaggle limit if on GPU. However, the conversion of parquet files to sqlite for 50 batches would probably take close to 12 hours. As I point out in the end of the example notebook, one could just modify the dataset class to read the competition parquet files directly instead.",
    "2128983": "Having baseline is great, early sharing prize was also a great idea imo. idk why @anjum48 was having second thoughts about sharing or not his solution, it will for sure be overruned with only 2 weeks into competition. But 2 first weeks probably is not enough time to build a solid alternatives to existing solutions. Most probable ESP winner just first to semi-implement this same baseline shared by @rasmusrse",
    "2128995": "You're totally right. The DynNet baseline is pretty solid IMO too - I have a long list of things that I tried that didn't work. It'll be interesting to see what people come up with",
    "2128998": "just curious, what did you decide on sharing before you notice 1.018 graphnet baseline on LB? And how many graphnet solutions you managed to train over what, a week? given other things that didnt work",
    "2129004": "I knew the Graphnet baseline was going to appear anyway. I saw that commits were being made in the Graphnet repo specific to this competition over the last few days.\n\nRegarding other solutions, I have a tonne of other GNNs, a few of which are now starting to surpass DynEdge. But in terms of training speed & simplicity, DynEdge is very strong",
    "2129015": "Thank you for sharing, i hope you ll get ESP, well deserved",
    "2129066": "great sharing",
    "2129385": "Roughly how long does a LB submission take with this baseline notebook?",
    "2129410": "It takes less than 3 hours for the submission notebook to get scored.",
    "2129623": "rasmusrse I get out of memory error on kaggle with GPU when trying to add in the code for converting to sqlite format below.\n\n```python\ninput_data_folder = '/kaggle/input/icecube-neutrinos-in-deep-ice/train'\ngeometry_table = pd.read_csv('/kaggle/input/icecube-neutrinos-in-deep-ice/sensor_geometry.csv')\nmeta_data_path = '/kaggle/input/icecube-neutrinos-in-deep-ice/train_meta.parquet'\n\n#batch_51\ndatabase_path = '/kaggle/working/batch_51'\nconvert_to_sqlite(meta_data_path,\n                  database_path=database_path,\n                  input_data_folder=input_data_folder,\n                  batch_ids = [51])\n```\n\n\nNote that if trying CPU run, I get this Import error (condensed):\n\n```python\nOSError                                   Traceback (most recent call last)\n---> 10 from graphnet.data.sqlite.sqlite_utilities import create_table\n----> 9     from .sqlite_dataset import SQLiteDataset\n----> 7 from graphnet.data.dataset import Dataset, ColumnMissingException\n---> 10 from torch_geometric.data import Data\n----> 4 import torch_geometric.data\n----> 1 from .data import Data\n----> 9 from torch_sparse import SparseTensor\n/opt/conda/lib/python3.7/site-packages/torch_sparse/__init__.py in <module>\n     17     spec = cuda_spec or cpu_spec\n     18     if spec is not None:\n---> 19         torch.ops.load_library(spec.origin)\n/opt/conda/lib/python3.7/site-packages/torch/_ops.py in load_library(self, path)\n    218             # static (global) initialization code in order to register custom\n    219             # operators with the JIT.\n--> 220             ctypes.CDLL(path)\n/opt/conda/lib/python3.7/ctypes/__init__.py in __init__(self, name, mode, handle, use_errno, use_last_error)\n    363         if handle is None:\n--> 364             self._handle = _dlopen(self._name, mode)\n\nOSError: libcusparse.so.11: cannot open shared object file: No such file or directory\n```\n\nAny suggestions on fixing it?\n\nI suppose CPU run for this preprocessing step would be much better on kaggle anyways, if there's a way to fix that second error. Or a streamlined import with fewer pip installs or manual addition of just create_table might be even better?\n\nI would assume that the 32GB memory of CPU notebooks would be more than enough to eliminate the memory error side of things.",
    "2129688": "Update: To create a CPU notebook that just does this task, I did the following:\n\n* removed graphnet-related pip install and imports\n* import sqlite3\n* manually copied into my notebook three functions from graphnet's [sqlite_utilities.py](https://github.com/graphnet-team/graphnet/blob/57fd0f37b65cfc8f2171b800353e5a4c40e752cb/src/graphnet/data/sqlite/sqlite_utilities.py) file:\n  * run_sql_code\n  * attach_index\n  * create_table\n\nThen it ran successfully in 10 minutes (599 seconds) for a single batch [52].",
    "2129987": "rasmusrse Man, having published paper AND code, you are golden! Thanks",
    "2130177": "Hey Robert - glad to hear you got it working. For reference, the error message you show here typically indicates that graphnet has been installed with GPU on a system without one.",
    "2135094": "If I want to download the dependecies on my on-prem environment, the way to do it is getting them \nby using:\nkaggle kernels output rasmusrse/graphnet-baseline-submission -p /path/to/dest\n\nSorry for stupid question I am very new to kaggle and this is my first competition :)\nThanks,",
    "2160316": "Hi @rasmusrse thank you a lot for your great work!! Is there a reason why you use the direction task and not the combination of the azimuth/zenith tasks ? Do you think one approach is better than the other?",
    "2164986": "Thank you for publishing the baseline!",
    "2171502": "Hi @rasmusrse  I'm getting the following error from the baseline notebook in the cell where the model, build, inference, and data functions are defined:\n`OSError: /opt/conda/lib/python3.7/site-packages/torch_sparse/_version_cuda.so: undefined symbol: _ZN5torch3jit17parseSchemaOrNameERKSs`\n\nAnyone have any idea what's going on?"
  },
  "source": "meta"
}