{
  "id": 393851,
  "title": "Pointnet model performance",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/393851",
  "author_name": "",
  "post_date": "2023-03-11T03:03:54.740872100Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/code/emanuelruzak/neutrinos3\" target=\"_blank\">https://www.kaggle.com/code/emanuelruzak/neutrinos3</a></p>\n<p>Pointnet is a model where each pulse vector (x,y,z,time,charge,…) is passed through the same MLP (1D convolution), then it takes the maximum of all the output vectors (maxpooling), then it passes through another MLP. which gives XYZ coordinates that can then be turned into azimuth/zenith, I used mean(arccos(cosinesimilarity)) as the error function since it is the same as Mean Angular Error.</p>\n<p>This model has permutation symmetry.</p>\n<p>I tried using vanilla pointnet (it didn't work with tnet layers, I think its because the problem is not symmetric under affine transformations)</p>\n<p>The model is overfitted, it gives a validation error of around 1.32, but a training error of 1.17<br>\nI used only 4 batches, if I use more, it exceeds RAM limit.<br>\nCan pointnet perform better, or is it the most it can get?</p>",
  "messages": [
    {
      "id": "2176904",
      "postDate": "03/11/2023 03:03:54",
      "content": "<p><a href=\"https://www.kaggle.com/code/emanuelruzak/neutrinos3\" target=\"_blank\">https://www.kaggle.com/code/emanuelruzak/neutrinos3</a></p>\n<p>Pointnet is a model where each pulse vector (x,y,z,time,charge,…) is passed through the same MLP (1D convolution), then it takes the maximum of all the output vectors (maxpooling), then it passes through another MLP. which gives XYZ coordinates that can then be turned into azimuth/zenith, I used mean(arccos(cosinesimilarity)) as the error function since it is the same as Mean Angular Error.</p>\n<p>This model has permutation symmetry.</p>\n<p>I tried using vanilla pointnet (it didn't work with tnet layers, I think its because the problem is not symmetric under affine transformations)</p>\n<p>The model is overfitted, it gives a validation error of around 1.32, but a training error of 1.17<br>\nI used only 4 batches, if I use more, it exceeds RAM limit.<br>\nCan pointnet perform better, or is it the most it can get?</p>",
      "rawMarkdown": "https://www.kaggle.com/code/emanuelruzak/neutrinos3\n\nPointnet is a model where each pulse vector (x,y,z,time,charge,...) is passed through the same MLP (1D convolution), then it takes the maximum of all the output vectors (maxpooling), then it passes through another MLP. which gives XYZ coordinates that can then be turned into azimuth/zenith, I used mean(arccos(cosinesimilarity)) as the error function since it is the same as Mean Angular Error.\n\nThis model has permutation symmetry.\n\nI tried using vanilla pointnet (it didn't work with tnet layers, I think its because the problem is not symmetric under affine transformations)\n\nThe model is overfitted, it gives a validation error of around 1.32, but a training error of 1.17\nI used only 4 batches, if I use more, it exceeds RAM limit.\nCan pointnet perform better, or is it the most it can get?",
      "votes": null
    },
    {
      "id": "2176951",
      "postDate": "03/11/2023 04:36:39",
      "content": "<p>Try training on more batches by loading data partially. Train on 1-2 batches, clear the memory, load next 2 batches and continue training. Applicable to training any model.</p>",
      "rawMarkdown": "Try training on more batches by loading data partially. Train on 1-2 batches, clear the memory, load next 2 batches and continue training. Applicable to training any model.",
      "votes": null
    },
    {
      "id": "2177779",
      "postDate": "03/11/2023 18:54:17",
      "content": "<p>That's an interesting model, thanks for sharing it!  I looked at your notebook and cannot make out how this model handles time and charge of the pulses - could you please explain it?</p>\n<p>A note about the permutation symmetry. If that means that the model always treats the points as interchangeable, I'm inclined to think that this is a weakness for this competition. One of the reasons it the pulses here are ordered in time, and this order very much influences the solution we are looking for.</p>",
      "rawMarkdown": "That's an interesting model, thanks for sharing it!  I looked at your notebook and cannot make out how this model handles time and charge of the pulses - could you please explain it?\n\nA note about the permutation symmetry. If that means that the model always treats the points as interchangeable, I'm inclined to think that this is a weakness for this competition. One of the reasons it the pulses here are ordered in time, and this order very much influences the solution we are looking for.",
      "votes": null
    },
    {
      "id": "2177824",
      "postDate": "03/11/2023 20:19:26",
      "content": "<p>Have you looked at PointNet++, an improvement by the same authors?  It uses neighbors of each point rather than the entire graph, similar to graphnet</p>",
      "rawMarkdown": "Have you looked at PointNet++, an improvement by the same authors?  It uses neighbors of each point rather than the entire graph, similar to graphnet",
      "votes": null
    },
    {
      "id": "2177825",
      "postDate": "03/11/2023 20:20:43",
      "content": "<p>I would definitely use more batches.  Why are you running of memory?  You only have to load a single parquet file at a time</p>",
      "rawMarkdown": "I would definitely use more batches.  Why are you running of memory?  You only have to load a single parquet file at a time",
      "votes": null
    },
    {
      "id": "2177888",
      "postDate": "03/11/2023 22:10:07",
      "content": "<p>The input is 128x9, it uses ZhaunBooty's dataset, and it includes time, charge, aux, x, y, z, rank, x_err, and z_err, it is made for training LSTM, my model treats every variable equally. It performed better than a simple MLP which didn't get under 1.5 at least in my case.<br>\nI think that a limitation of LSTM is that there might be many particles at the same time in different places and the pulses from each particle alternate in time randomly. But ignoring the order of time might be even worse. I tried using mean pooling, so the output changes if there are many similar pulses instead of one, but it still has around 1.3 validation accuracy</p>",
      "rawMarkdown": "The input is 128x9, it uses ZhaunBooty's dataset, and it includes time, charge, aux, x, y, z, rank, x_err, and z_err, it is made for training LSTM, my model treats every variable equally. It performed better than a simple MLP which didn't get under 1.5 at least in my case.\nI think that a limitation of LSTM is that there might be many particles at the same time in different places and the pulses from each particle alternate in time randomly. But ignoring the order of time might be even worse. I tried using mean pooling, so the output changes if there are many similar pulses instead of one, but it still has around 1.3 validation accuracy",
      "votes": null
    },
    {
      "id": "2177901",
      "postDate": "03/11/2023 22:17:18",
      "content": "<p>Thanks for the suggestion, I haven't tried it</p>",
      "rawMarkdown": "Thanks for the suggestion, I haven't tried it",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2176951,
      "author_name": "dipamc77",
      "author_url": "",
      "post_date": "03/11/2023 04:36:39",
      "content": "<p>Try training on more batches by loading data partially. Train on 1-2 batches, clear the memory, load next 2 batches and continue training. Applicable to training any model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2177779,
      "author_name": "alexz0",
      "author_url": "",
      "post_date": "03/11/2023 18:54:17",
      "content": "<p>That's an interesting model, thanks for sharing it!  I looked at your notebook and cannot make out how this model handles time and charge of the pulses - could you please explain it?</p>\n<p>A note about the permutation symmetry. If that means that the model always treats the points as interchangeable, I'm inclined to think that this is a weakness for this competition. One of the reasons it the pulses here are ordered in time, and this order very much influences the solution we are looking for.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2177888,
          "author_name": "emanuelruzak",
          "author_url": "",
          "post_date": "03/11/2023 22:10:07",
          "content": "<p>The input is 128x9, it uses ZhaunBooty's dataset, and it includes time, charge, aux, x, y, z, rank, x_err, and z_err, it is made for training LSTM, my model treats every variable equally. It performed better than a simple MLP which didn't get under 1.5 at least in my case.<br>\nI think that a limitation of LSTM is that there might be many particles at the same time in different places and the pulses from each particle alternate in time randomly. But ignoring the order of time might be even worse. I tried using mean pooling, so the output changes if there are many similar pulses instead of one, but it still has around 1.3 validation accuracy</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2177824,
      "author_name": "solverworld",
      "author_url": "",
      "post_date": "03/11/2023 20:19:26",
      "content": "<p>Have you looked at PointNet++, an improvement by the same authors?  It uses neighbors of each point rather than the entire graph, similar to graphnet</p>",
      "votes": null,
      "replies": [
        {
          "id": 2177901,
          "author_name": "emanuelruzak",
          "author_url": "",
          "post_date": "03/11/2023 22:17:18",
          "content": "<p>Thanks for the suggestion, I haven't tried it</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2177825,
      "author_name": "solverworld",
      "author_url": "",
      "post_date": "03/11/2023 20:20:43",
      "content": "<p>I would definitely use more batches.  Why are you running of memory?  You only have to load a single parquet file at a time</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2176904": "https://www.kaggle.com/code/emanuelruzak/neutrinos3\n\nPointnet is a model where each pulse vector (x,y,z,time,charge,...) is passed through the same MLP (1D convolution), then it takes the maximum of all the output vectors (maxpooling), then it passes through another MLP. which gives XYZ coordinates that can then be turned into azimuth/zenith, I used mean(arccos(cosinesimilarity)) as the error function since it is the same as Mean Angular Error.\n\nThis model has permutation symmetry.\n\nI tried using vanilla pointnet (it didn't work with tnet layers, I think its because the problem is not symmetric under affine transformations)\n\nThe model is overfitted, it gives a validation error of around 1.32, but a training error of 1.17\nI used only 4 batches, if I use more, it exceeds RAM limit.\nCan pointnet perform better, or is it the most it can get?",
    "2176951": "Try training on more batches by loading data partially. Train on 1-2 batches, clear the memory, load next 2 batches and continue training. Applicable to training any model.",
    "2177779": "That's an interesting model, thanks for sharing it!  I looked at your notebook and cannot make out how this model handles time and charge of the pulses - could you please explain it?\n\nA note about the permutation symmetry. If that means that the model always treats the points as interchangeable, I'm inclined to think that this is a weakness for this competition. One of the reasons it the pulses here are ordered in time, and this order very much influences the solution we are looking for.",
    "2177824": "Have you looked at PointNet++, an improvement by the same authors?  It uses neighbors of each point rather than the entire graph, similar to graphnet",
    "2177825": "I would definitely use more batches.  Why are you running of memory?  You only have to load a single parquet file at a time",
    "2177888": "The input is 128x9, it uses ZhaunBooty's dataset, and it includes time, charge, aux, x, y, z, rank, x_err, and z_err, it is made for training LSTM, my model treats every variable equally. It performed better than a simple MLP which didn't get under 1.5 at least in my case.\nI think that a limitation of LSTM is that there might be many particles at the same time in different places and the pulses from each particle alternate in time randomly. But ignoring the order of time might be even worse. I tried using mean pooling, so the output changes if there are many similar pulses instead of one, but it still has around 1.3 validation accuracy",
    "2177901": "Thanks for the suggestion, I haven't tried it"
  },
  "source": "meta"
}