{
  "id": 389358,
  "title": "Submission error ❌ - Limit your nb of pulses per event",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/389358",
  "author_name": "",
  "post_date": "2023-02-21T15:31:09.398788500Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello folks ! </p>\n<p>I am facing weird submission issues with my own Graphnet implementation. <br>\nTo describe the pb, the notebook makes predictions but the csv file is declined (seems to be bc the nb of rows doesn't match with the expected size).</p>\n<p>Yet that is weird bc my pipeline works with another simpler network. So the issue comes from the network itself (using sageconv convolutions instead of DynEdgeConv gives a valid graded submission). I added an assertion to check if the predictions are empty, but they are in fact always filled with values. I precise that when I run the notebook on a test batch, no NaN or inf etc. is returned.</p>\n<p>So far I tried : fill nan inf values + add assertion tests at various locations and wait for the notebook to crash, but debugging is almost impossible, bc of the Kaggle's submission encapsulating.</p>\n<p><strong>Has anyone faced similar issues in this challenge, or in past competitions ?</strong> This is getting frustrating as f***.</p>\n<p><strong>[EDIT]</strong> After extra experiments, it seems that my errors come from mispredictions on <strong>big events</strong>. Maybe when they are too many pulses in one event, the knn graph builder fails ? or returns a graph that creates a bug ? Any ideas ?</p>\n<p><strong>[EDIT2]</strong> Maybe, when the graph has many nodes, KNN creates multiple disjoint sub graphs, and it causes an error for an unknown reason ?</p>\n<p><strong>[EDIT3]</strong> I finally conclude that you should limit the nb of pulses per event when creating your graphs, first to reduce memory usage variance, second to \"remove\" huge outlier graphs (in practice we don't drop them, we only keep a subset of nodes). But I still don't understand why it lead to inference errors :) Hope it will help people facing similar issues. A really similar explanation can be found in the graphnet baseline example to justify the 200 pulse threshold.</p>",
  "messages": [
    {
      "id": "2153705",
      "postDate": "02/21/2023 15:31:09",
      "content": "<p>Hello folks ! </p>\n<p>I am facing weird submission issues with my own Graphnet implementation. <br>\nTo describe the pb, the notebook makes predictions but the csv file is declined (seems to be bc the nb of rows doesn't match with the expected size).</p>\n<p>Yet that is weird bc my pipeline works with another simpler network. So the issue comes from the network itself (using sageconv convolutions instead of DynEdgeConv gives a valid graded submission). I added an assertion to check if the predictions are empty, but they are in fact always filled with values. I precise that when I run the notebook on a test batch, no NaN or inf etc. is returned.</p>\n<p>So far I tried : fill nan inf values + add assertion tests at various locations and wait for the notebook to crash, but debugging is almost impossible, bc of the Kaggle's submission encapsulating.</p>\n<p><strong>Has anyone faced similar issues in this challenge, or in past competitions ?</strong> This is getting frustrating as f***.</p>\n<p><strong>[EDIT]</strong> After extra experiments, it seems that my errors come from mispredictions on <strong>big events</strong>. Maybe when they are too many pulses in one event, the knn graph builder fails ? or returns a graph that creates a bug ? Any ideas ?</p>\n<p><strong>[EDIT2]</strong> Maybe, when the graph has many nodes, KNN creates multiple disjoint sub graphs, and it causes an error for an unknown reason ?</p>\n<p><strong>[EDIT3]</strong> I finally conclude that you should limit the nb of pulses per event when creating your graphs, first to reduce memory usage variance, second to \"remove\" huge outlier graphs (in practice we don't drop them, we only keep a subset of nodes). But I still don't understand why it lead to inference errors :) Hope it will help people facing similar issues. A really similar explanation can be found in the graphnet baseline example to justify the 200 pulse threshold.</p>",
      "rawMarkdown": "Hello folks ! \n\nI am facing weird submission issues with my own Graphnet implementation. \nTo describe the pb, the notebook makes predictions but the csv file is declined (seems to be bc the nb of rows doesn't match with the expected size).\n\nYet that is weird bc my pipeline works with another simpler network. So the issue comes from the network itself (using sageconv convolutions instead of DynEdgeConv gives a valid graded submission). I added an assertion to check if the predictions are empty, but they are in fact always filled with values. I precise that when I run the notebook on a test batch, no NaN or inf etc. is returned.\n\nSo far I tried : fill nan inf values + add assertion tests at various locations and wait for the notebook to crash, but debugging is almost impossible, bc of the Kaggle's submission encapsulating.\n\n**Has anyone faced similar issues in this challenge, or in past competitions ?** This is getting frustrating as f***.\n\n**[EDIT]** After extra experiments, it seems that my errors come from mispredictions on **big events**. Maybe when they are too many pulses in one event, the knn graph builder fails ? or returns a graph that creates a bug ? Any ideas ?\n\n**[EDIT2]** Maybe, when the graph has many nodes, KNN creates multiple disjoint sub graphs, and it causes an error for an unknown reason ?\n\n**[EDIT3]** I finally conclude that you should limit the nb of pulses per event when creating your graphs, first to reduce memory usage variance, second to \"remove\" huge outlier graphs (in practice we don't drop them, we only keep a subset of nodes). But I still don't understand why it lead to inference errors :) Hope it will help people facing similar issues. A really similar explanation can be found in the graphnet baseline example to justify the 200 pulse threshold.",
      "votes": null
    },
    {
      "id": "2153722",
      "postDate": "02/21/2023 15:43:36",
      "content": "<p>It's a pain but debugging of these type errors needs to occur by running a couple of the train batches as if they are test.  So you need to make up your own test_meta.parquet file, sample_submission and a folder with a couple of train files.</p>\n<p>Kaggle could really do us a favor by having test stuff with a couple of batches and a lot more than 3 rows.  Than we would be able to debug many of the errors without the hassle of building our own fake test data of decent size.</p>\n<p>In previous competitions Kaggle staff has indicted that we need to be coders who think about errors and trap them in the code.  While I agree with that idea - having a test file with only 3 events is a JOKE.</p>",
      "rawMarkdown": "It's a pain but debugging of these type errors needs to occur by running a couple of the train batches as if they are test.  So you need to make up your own test_meta.parquet file, sample_submission and a folder with a couple of train files.\n\nKaggle could really do us a favor by having test stuff with a couple of batches and a lot more than 3 rows.  Than we would be able to debug many of the errors without the hassle of building our own fake test data of decent size.\n\nIn previous competitions Kaggle staff has indicted that we need to be coders who think about errors and trap them in the code.  While I agree with that idea - having a test file with only 3 events is a JOKE.",
      "votes": null
    },
    {
      "id": "2153760",
      "postDate": "02/21/2023 16:11:46",
      "content": "<p>I totally agree with you on that point: the test set is absurdly tiny.</p>\n<p>I tried my approach on multiple test batches, on my device, and I don't face such issues … so I am totally lost, I have been tracking every possible error case for 5 days, but no results so far</p>",
      "rawMarkdown": "I totally agree with you on that point: the test set is absurdly tiny.\n\nI tried my approach on multiple test batches, on my device, and I don't face such issues ... so I am totally lost, I have been tracking every possible error case for 5 days, but no results so far",
      "votes": null
    },
    {
      "id": "2154244",
      "postDate": "02/21/2023 22:55:56",
      "content": "<p>Make sure that the values for azimuth and zenith are within allowed limits. It might have to do with the precision of float values. For example 3.1415927 might cause an error for the zenith, whereas 3.1415926 does not.</p>",
      "rawMarkdown": "Make sure that the values for azimuth and zenith are within allowed limits. It might have to do with the precision of float values. For example 3.1415927 might cause an error for the zenith, whereas 3.1415926 does not.",
      "votes": null
    },
    {
      "id": "2154290",
      "postDate": "02/21/2023 23:33:00",
      "content": "<p>try re running notebook with limiting to <code>64</code> observations per event … see if you can do sucesfull sub</p>",
      "rawMarkdown": "try re running notebook with limiting to `64` observations per event ... see if you can do sucesfull sub",
      "votes": null
    },
    {
      "id": "2154357",
      "postDate": "02/22/2023 00:51:11",
      "content": "<p>When things run well on my local machine but fail on kaggle - the kaggle memory limit has been my most frequent root cause.  Last year I went so far as to remove the 64GB ram chips and replace to have only 16GB on one of my machines to find my OOM point.   </p>",
      "rawMarkdown": "When things run well on my local machine but fail on kaggle - the kaggle memory limit has been my most frequent root cause.  Last year I went so far as to remove the 64GB ram chips and replace to have only 16GB on one of my machines to find my OOM point.",
      "votes": null
    },
    {
      "id": "2155303",
      "postDate": "02/22/2023 14:56:39",
      "content": "<p>My current fix is to select only a subset of pulses, as you propose ! But really, I still can't figure why it fails on big graphs :)</p>",
      "rawMarkdown": "My current fix is to select only a subset of pulses, as you propose ! But really, I still can't figure why it fails on big graphs :)",
      "votes": null
    },
    {
      "id": "2155306",
      "postDate": "02/22/2023 14:58:30",
      "content": "<p>I predict XYZ coords and then convert them back to angles using arccos and arctan. So they are in the expected intervals. <br>\nBut definitively, it could have been an issue :)</p>",
      "rawMarkdown": "I predict XYZ coords and then convert them back to angles using arccos and arctan. So they are in the expected intervals. \nBut definitively, it could have been an issue :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2153722,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "02/21/2023 15:43:36",
      "content": "<p>It's a pain but debugging of these type errors needs to occur by running a couple of the train batches as if they are test.  So you need to make up your own test_meta.parquet file, sample_submission and a folder with a couple of train files.</p>\n<p>Kaggle could really do us a favor by having test stuff with a couple of batches and a lot more than 3 rows.  Than we would be able to debug many of the errors without the hassle of building our own fake test data of decent size.</p>\n<p>In previous competitions Kaggle staff has indicted that we need to be coders who think about errors and trap them in the code.  While I agree with that idea - having a test file with only 3 events is a JOKE.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2153760,
          "author_name": "louisstefanuto",
          "author_url": "",
          "post_date": "02/21/2023 16:11:46",
          "content": "<p>I totally agree with you on that point: the test set is absurdly tiny.</p>\n<p>I tried my approach on multiple test batches, on my device, and I don't face such issues … so I am totally lost, I have been tracking every possible error case for 5 days, but no results so far</p>",
          "votes": null,
          "replies": [
            {
              "id": 2154357,
              "author_name": "pcjimmmy",
              "author_url": "",
              "post_date": "02/22/2023 00:51:11",
              "content": "<p>When things run well on my local machine but fail on kaggle - the kaggle memory limit has been my most frequent root cause.  Last year I went so far as to remove the 64GB ram chips and replace to have only 16GB on one of my machines to find my OOM point.   </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2154244,
      "author_name": "taqseorangpun",
      "author_url": "",
      "post_date": "02/21/2023 22:55:56",
      "content": "<p>Make sure that the values for azimuth and zenith are within allowed limits. It might have to do with the precision of float values. For example 3.1415927 might cause an error for the zenith, whereas 3.1415926 does not.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2155306,
          "author_name": "louisstefanuto",
          "author_url": "",
          "post_date": "02/22/2023 14:58:30",
          "content": "<p>I predict XYZ coords and then convert them back to angles using arccos and arctan. So they are in the expected intervals. <br>\nBut definitively, it could have been an issue :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2154290,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "02/21/2023 23:33:00",
      "content": "<p>try re running notebook with limiting to <code>64</code> observations per event … see if you can do sucesfull sub</p>",
      "votes": null,
      "replies": [
        {
          "id": 2155303,
          "author_name": "louisstefanuto",
          "author_url": "",
          "post_date": "02/22/2023 14:56:39",
          "content": "<p>My current fix is to select only a subset of pulses, as you propose ! But really, I still can't figure why it fails on big graphs :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2153705": "Hello folks ! \n\nI am facing weird submission issues with my own Graphnet implementation. \nTo describe the pb, the notebook makes predictions but the csv file is declined (seems to be bc the nb of rows doesn't match with the expected size).\n\nYet that is weird bc my pipeline works with another simpler network. So the issue comes from the network itself (using sageconv convolutions instead of DynEdgeConv gives a valid graded submission). I added an assertion to check if the predictions are empty, but they are in fact always filled with values. I precise that when I run the notebook on a test batch, no NaN or inf etc. is returned.\n\nSo far I tried : fill nan inf values + add assertion tests at various locations and wait for the notebook to crash, but debugging is almost impossible, bc of the Kaggle's submission encapsulating.\n\n**Has anyone faced similar issues in this challenge, or in past competitions ?** This is getting frustrating as f***.\n\n**[EDIT]** After extra experiments, it seems that my errors come from mispredictions on **big events**. Maybe when they are too many pulses in one event, the knn graph builder fails ? or returns a graph that creates a bug ? Any ideas ?\n\n**[EDIT2]** Maybe, when the graph has many nodes, KNN creates multiple disjoint sub graphs, and it causes an error for an unknown reason ?\n\n**[EDIT3]** I finally conclude that you should limit the nb of pulses per event when creating your graphs, first to reduce memory usage variance, second to \"remove\" huge outlier graphs (in practice we don't drop them, we only keep a subset of nodes). But I still don't understand why it lead to inference errors :) Hope it will help people facing similar issues. A really similar explanation can be found in the graphnet baseline example to justify the 200 pulse threshold.",
    "2153722": "It's a pain but debugging of these type errors needs to occur by running a couple of the train batches as if they are test.  So you need to make up your own test_meta.parquet file, sample_submission and a folder with a couple of train files.\n\nKaggle could really do us a favor by having test stuff with a couple of batches and a lot more than 3 rows.  Than we would be able to debug many of the errors without the hassle of building our own fake test data of decent size.\n\nIn previous competitions Kaggle staff has indicted that we need to be coders who think about errors and trap them in the code.  While I agree with that idea - having a test file with only 3 events is a JOKE.",
    "2153760": "I totally agree with you on that point: the test set is absurdly tiny.\n\nI tried my approach on multiple test batches, on my device, and I don't face such issues ... so I am totally lost, I have been tracking every possible error case for 5 days, but no results so far",
    "2154244": "Make sure that the values for azimuth and zenith are within allowed limits. It might have to do with the precision of float values. For example 3.1415927 might cause an error for the zenith, whereas 3.1415926 does not.",
    "2154290": "try re running notebook with limiting to `64` observations per event ... see if you can do sucesfull sub",
    "2154357": "When things run well on my local machine but fail on kaggle - the kaggle memory limit has been my most frequent root cause.  Last year I went so far as to remove the 64GB ram chips and replace to have only 16GB on one of my machines to find my OOM point.",
    "2155303": "My current fix is to select only a subset of pulses, as you propose ! But really, I still can't figure why it fails on big graphs :)",
    "2155306": "I predict XYZ coords and then convert them back to angles using arccos and arctan. So they are in the expected intervals. \nBut definitively, it could have been an issue :)"
  },
  "source": "meta"
}