{
  "id": 56129,
  "title": "Deep learning architecture?",
  "url": "/competitions/trackml-particle-identification/discussion/56129",
  "author_name": "",
  "post_date": "2018-05-06T14:20:53.696368200Z",
  "votes": 9,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I wonder what kind of deep learning architecture is the best to tackle this challenge, and I will post some approaches we can consider.\nSome of the ideas are motivated/referenced from <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/55832\">Pytorch starter kit!</a> by <a href=\"/hengck23\">@hengck23</a>.</p>\n\n<p>At first I considered the input data format &amp; its deep neural network architecture.</p>\n\n<p>1: 3D-CNN</p>\n\n<p>This is most intuitive choice that comes up first. Since the hits data is given by <code>(x, y, z)</code> coordinate, we can assign each hit in 3D voxel.\nThe input will be a \"sparse\" 3D voxel in the sense that it consists of the \"hit\" point and non-hit point.\nConcern of this approach is memory/computing efficiency, which depends on the resolution.\nIn the detector, x and y direction ranges from -1000 to 1000 mm, z direction ranges from -3000 to 3000 mm. So when we take the resolution in 1mm for 1 voxel, we need 2000 * 2000 * 6000 3D voxel, which is too big to deal with current GPU memory.</p>\n\n<p>To reduce memory/computation consumption, we may consider reducing the resolution. I think resolution can be the order of module size (which I did not investigate yet).</p>\n\n<p>2: PointNet</p>\n\n<p>To deal with the sparse hit points in 3D voxels, there is a suitable existing study called <a href=\"https://arxiv.org/abs/1612.00593\">PointNet</a>.\nThe basic idea is to process \"point\" itself rather than map the point to 3d coordinate to deal with 3d-CNN. The advantage is the memory/computation efficiency compared to 3D-CNN approach.\nSince \"hits\" data is provided as a <code>(x, y, z)</code> points, this approach seems nice for this challenge.</p>\n\n<p>I found a well summarized post about these architectures.</p>\n\n<p><a href=\"http://www.itzikbs.com/3d-point-cloud-classification-using-deep-learning\">3D POINT CLOUD CLASSIFICATION USING DEEP LEARNING – RECENT WORKS</a></p>\n\n<p>3: Graph Convolution Network</p>\n\n<p>Other possible approach is to adopt Graph Convolution Network, which applies convolution network on graphs.\nGraph Convolution Network can take any graph input consists of node and edge.\nWe can consider each hits point as \"node\", however, it is not clear that how to represent \"edge\". </p>\n\n<p>\"Edge\" should be not too dense, if we connect all the node-node pair its computation will be as heavy as fully connected layer. Thus we should only connect \"adjacent\" node, how to define \"adjacent\" node will be a key for this graph convolution approach.</p>\n\n<p>--- Other approches---</p>\n\n<p>Other than the above, these approach are studied/proposed before</p>\n\n<p>4: Extrapolation using LSTM</p>\n\n<p>Consider track finding problem. When we given a starting path, later track path can be estimated by LSTM architecture.</p>\n\n<p>Unclear point to adopt this approach to this challenge is how to define/find a starting path itself.</p>\n\n<p>5: Denoising (remove noise hit) Auto Encoder</p>\n\n<p>There exists a noise hit which does not belong to any track. To remove these noise measurement, the idea of Auto Encoder may be used.\nThis idea might become a key when we want to fine tune the score in the later stage.</p>\n\n<p>Reference</p>\n\n<ul>\n<li><a href=\"https://www.youtube.com/watch?v=QMuCSLWsks4\">The HEP.TrX Project: Deep Learning for Particle Tracking - ACAT 2017</a></li>\n<li><a href=\"https://github.com/HEPTrkX/heptrkx-ctd/blob/master/hit_classification/lstm_toy2D.ipynb\">HEPTrkX/heptrkx-ctd</a></li>\n</ul>\n\n<p>This challenge is interesting because coming up the approach itself is not easy, I am still seeking the best way!</p>",
  "messages": [
    {
      "id": "323884",
      "postDate": "05/06/2018 14:20:53",
      "content": "<p>I wonder what kind of deep learning architecture is the best to tackle this challenge, and I will post some approaches we can consider.\nSome of the ideas are motivated/referenced from <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/55832\">Pytorch starter kit!</a> by <a href=\"/hengck23\">@hengck23</a>.</p>\n\n<p>At first I considered the input data format &amp; its deep neural network architecture.</p>\n\n<p>1: 3D-CNN</p>\n\n<p>This is most intuitive choice that comes up first. Since the hits data is given by <code>(x, y, z)</code> coordinate, we can assign each hit in 3D voxel.\nThe input will be a \"sparse\" 3D voxel in the sense that it consists of the \"hit\" point and non-hit point.\nConcern of this approach is memory/computing efficiency, which depends on the resolution.\nIn the detector, x and y direction ranges from -1000 to 1000 mm, z direction ranges from -3000 to 3000 mm. So when we take the resolution in 1mm for 1 voxel, we need 2000 * 2000 * 6000 3D voxel, which is too big to deal with current GPU memory.</p>\n\n<p>To reduce memory/computation consumption, we may consider reducing the resolution. I think resolution can be the order of module size (which I did not investigate yet).</p>\n\n<p>2: PointNet</p>\n\n<p>To deal with the sparse hit points in 3D voxels, there is a suitable existing study called <a href=\"https://arxiv.org/abs/1612.00593\">PointNet</a>.\nThe basic idea is to process \"point\" itself rather than map the point to 3d coordinate to deal with 3d-CNN. The advantage is the memory/computation efficiency compared to 3D-CNN approach.\nSince \"hits\" data is provided as a <code>(x, y, z)</code> points, this approach seems nice for this challenge.</p>\n\n<p>I found a well summarized post about these architectures.</p>\n\n<p><a href=\"http://www.itzikbs.com/3d-point-cloud-classification-using-deep-learning\">3D POINT CLOUD CLASSIFICATION USING DEEP LEARNING – RECENT WORKS</a></p>\n\n<p>3: Graph Convolution Network</p>\n\n<p>Other possible approach is to adopt Graph Convolution Network, which applies convolution network on graphs.\nGraph Convolution Network can take any graph input consists of node and edge.\nWe can consider each hits point as \"node\", however, it is not clear that how to represent \"edge\". </p>\n\n<p>\"Edge\" should be not too dense, if we connect all the node-node pair its computation will be as heavy as fully connected layer. Thus we should only connect \"adjacent\" node, how to define \"adjacent\" node will be a key for this graph convolution approach.</p>\n\n<p>--- Other approches---</p>\n\n<p>Other than the above, these approach are studied/proposed before</p>\n\n<p>4: Extrapolation using LSTM</p>\n\n<p>Consider track finding problem. When we given a starting path, later track path can be estimated by LSTM architecture.</p>\n\n<p>Unclear point to adopt this approach to this challenge is how to define/find a starting path itself.</p>\n\n<p>5: Denoising (remove noise hit) Auto Encoder</p>\n\n<p>There exists a noise hit which does not belong to any track. To remove these noise measurement, the idea of Auto Encoder may be used.\nThis idea might become a key when we want to fine tune the score in the later stage.</p>\n\n<p>Reference</p>\n\n<ul>\n<li><a href=\"https://www.youtube.com/watch?v=QMuCSLWsks4\">The HEP.TrX Project: Deep Learning for Particle Tracking - ACAT 2017</a></li>\n<li><a href=\"https://github.com/HEPTrkX/heptrkx-ctd/blob/master/hit_classification/lstm_toy2D.ipynb\">HEPTrkX/heptrkx-ctd</a></li>\n</ul>\n\n<p>This challenge is interesting because coming up the approach itself is not easy, I am still seeking the best way!</p>",
      "rawMarkdown": "I wonder what kind of deep learning architecture is the best to tackle this challenge, and I will post some approaches we can consider.\nSome of the ideas are motivated/referenced from [Pytorch starter kit!][1] by @hengck23.\n\nAt first I considered the input data format &amp; its deep neural network architecture.\n\n1: 3D-CNN\n\nThis is most intuitive choice that comes up first. Since the hits data is given by `(x, y, z)` coordinate, we can assign each hit in 3D voxel.\nThe input will be a \"sparse\" 3D voxel in the sense that it consists of the \"hit\" point and non-hit point.\nConcern of this approach is memory/computing efficiency, which depends on the resolution.\nIn the detector, x and y direction ranges from -1000 to 1000 mm, z direction ranges from -3000 to 3000 mm. So when we take the resolution in 1mm for 1 voxel, we need 2000 * 2000 * 6000 3D voxel, which is too big to deal with current GPU memory.\n\nTo reduce memory/computation consumption, we may consider reducing the resolution. I think resolution can be the order of module size (which I did not investigate yet).\n\n2: PointNet\n\nTo deal with the sparse hit points in 3D voxels, there is a suitable existing study called [PointNet][2].\nThe basic idea is to process \"point\" itself rather than map the point to 3d coordinate to deal with 3d-CNN. The advantage is the memory/computation efficiency compared to 3D-CNN approach.\nSince \"hits\" data is provided as a `(x, y, z)` points, this approach seems nice for this challenge.\n\nI found a well summarized post about these architectures.\n\n[3D POINT CLOUD CLASSIFICATION USING DEEP LEARNING – RECENT WORKS][3]\n\n\n3: Graph Convolution Network\n\nOther possible approach is to adopt Graph Convolution Network, which applies convolution network on graphs.\nGraph Convolution Network can take any graph input consists of node and edge.\nWe can consider each hits point as \"node\", however, it is not clear that how to represent \"edge\". \n\n\"Edge\" should be not too dense, if we connect all the node-node pair its computation will be as heavy as fully connected layer. Thus we should only connect \"adjacent\" node, how to define \"adjacent\" node will be a key for this graph convolution approach.\n\n\n--- Other approches---\n\nOther than the above, these approach are studied/proposed before\n\n4: Extrapolation using LSTM\n\nConsider track finding problem. When we given a starting path, later track path can be estimated by LSTM architecture.\n\nUnclear point to adopt this approach to this challenge is how to define/find a starting path itself.\n\n5: Denoising (remove noise hit) Auto Encoder\n\nThere exists a noise hit which does not belong to any track. To remove these noise measurement, the idea of Auto Encoder may be used.\nThis idea might become a key when we want to fine tune the score in the later stage.\n\nReference\n\n - [The HEP.TrX Project: Deep Learning for Particle Tracking - ACAT 2017][4]\n - [HEPTrkX/heptrkx-ctd][5]\n\n\nThis challenge is interesting because coming up the approach itself is not easy, I am still seeking the best way!\n\n  [1]: https://www.kaggle.com/c/trackml-particle-identification/discussion/55832\n  [2]: https://arxiv.org/abs/1612.00593\n  [3]: http://www.itzikbs.com/3d-point-cloud-classification-using-deep-learning\n  [4]: https://www.youtube.com/watch?v=QMuCSLWsks4\n  [5]: https://github.com/HEPTrkX/heptrkx-ctd/blob/master/hit_classification/lstm_toy2D.ipynb",
      "votes": null
    },
    {
      "id": "326451",
      "postDate": "05/09/2018 18:36:42",
      "content": "<p>I think a U-net might work for this problem. We are really just looking for a contrast between a path and the background noise. I suppose the setup could be similar to the 3D-CNN you mentioned, but I don't think we need a full 3D representation. We could cut the detector space into cones originating at 0,0,0.</p>",
      "rawMarkdown": "I think a U-net might work for this problem. We are really just looking for a contrast between a path and the background noise. I suppose the setup could be similar to the 3D-CNN you mentioned, but I don't think we need a full 3D representation. We could cut the detector space into cones originating at 0,0,0.",
      "votes": null
    },
    {
      "id": "326598",
      "postDate": "05/10/2018 02:07:47",
      "content": "<p>Good idea!\nMay be each detector layer can be reduced to 2-d representation, each channel represents each layer.\nIn this idea, I feel challenging point is to combine cone type layer in the center and endcap (circle) type at the edge.</p>\n\n<p>This is related to a category called \"MultiView\" architecture.</p>",
      "rawMarkdown": "Good idea!\nMay be each detector layer can be reduced to 2-d representation, each channel represents each layer.\nIn this idea, I feel challenging point is to combine cone type layer in the center and endcap (circle) type at the edge.\n\nThis is related to a category called \"MultiView\" architecture.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 326451,
      "author_name": "jeffhebert",
      "author_url": "",
      "post_date": "05/09/2018 18:36:42",
      "content": "<p>I think a U-net might work for this problem. We are really just looking for a contrast between a path and the background noise. I suppose the setup could be similar to the 3D-CNN you mentioned, but I don't think we need a full 3D representation. We could cut the detector space into cones originating at 0,0,0.</p>",
      "votes": null,
      "replies": [
        {
          "id": 326598,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "05/10/2018 02:07:47",
          "content": "<p>Good idea!\nMay be each detector layer can be reduced to 2-d representation, each channel represents each layer.\nIn this idea, I feel challenging point is to combine cone type layer in the center and endcap (circle) type at the edge.</p>\n\n<p>This is related to a category called \"MultiView\" architecture.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "323884": "I wonder what kind of deep learning architecture is the best to tackle this challenge, and I will post some approaches we can consider.\nSome of the ideas are motivated/referenced from [Pytorch starter kit!][1] by @hengck23.\n\nAt first I considered the input data format &amp; its deep neural network architecture.\n\n1: 3D-CNN\n\nThis is most intuitive choice that comes up first. Since the hits data is given by `(x, y, z)` coordinate, we can assign each hit in 3D voxel.\nThe input will be a \"sparse\" 3D voxel in the sense that it consists of the \"hit\" point and non-hit point.\nConcern of this approach is memory/computing efficiency, which depends on the resolution.\nIn the detector, x and y direction ranges from -1000 to 1000 mm, z direction ranges from -3000 to 3000 mm. So when we take the resolution in 1mm for 1 voxel, we need 2000 * 2000 * 6000 3D voxel, which is too big to deal with current GPU memory.\n\nTo reduce memory/computation consumption, we may consider reducing the resolution. I think resolution can be the order of module size (which I did not investigate yet).\n\n2: PointNet\n\nTo deal with the sparse hit points in 3D voxels, there is a suitable existing study called [PointNet][2].\nThe basic idea is to process \"point\" itself rather than map the point to 3d coordinate to deal with 3d-CNN. The advantage is the memory/computation efficiency compared to 3D-CNN approach.\nSince \"hits\" data is provided as a `(x, y, z)` points, this approach seems nice for this challenge.\n\nI found a well summarized post about these architectures.\n\n[3D POINT CLOUD CLASSIFICATION USING DEEP LEARNING – RECENT WORKS][3]\n\n\n3: Graph Convolution Network\n\nOther possible approach is to adopt Graph Convolution Network, which applies convolution network on graphs.\nGraph Convolution Network can take any graph input consists of node and edge.\nWe can consider each hits point as \"node\", however, it is not clear that how to represent \"edge\". \n\n\"Edge\" should be not too dense, if we connect all the node-node pair its computation will be as heavy as fully connected layer. Thus we should only connect \"adjacent\" node, how to define \"adjacent\" node will be a key for this graph convolution approach.\n\n\n--- Other approches---\n\nOther than the above, these approach are studied/proposed before\n\n4: Extrapolation using LSTM\n\nConsider track finding problem. When we given a starting path, later track path can be estimated by LSTM architecture.\n\nUnclear point to adopt this approach to this challenge is how to define/find a starting path itself.\n\n5: Denoising (remove noise hit) Auto Encoder\n\nThere exists a noise hit which does not belong to any track. To remove these noise measurement, the idea of Auto Encoder may be used.\nThis idea might become a key when we want to fine tune the score in the later stage.\n\nReference\n\n - [The HEP.TrX Project: Deep Learning for Particle Tracking - ACAT 2017][4]\n - [HEPTrkX/heptrkx-ctd][5]\n\n\nThis challenge is interesting because coming up the approach itself is not easy, I am still seeking the best way!\n\n  [1]: https://www.kaggle.com/c/trackml-particle-identification/discussion/55832\n  [2]: https://arxiv.org/abs/1612.00593\n  [3]: http://www.itzikbs.com/3d-point-cloud-classification-using-deep-learning\n  [4]: https://www.youtube.com/watch?v=QMuCSLWsks4\n  [5]: https://github.com/HEPTrkX/heptrkx-ctd/blob/master/hit_classification/lstm_toy2D.ipynb",
    "326451": "I think a U-net might work for this problem. We are really just looking for a contrast between a path and the background noise. I suppose the setup could be similar to the 3D-CNN you mentioned, but I don't think we need a full 3D representation. We could cut the detector space into cones originating at 0,0,0.",
    "326598": "Good idea!\nMay be each detector layer can be reduced to 2-d representation, each channel represents each layer.\nIn this idea, I feel challenging point is to combine cone type layer in the center and endcap (circle) type at the edge.\n\nThis is related to a category called \"MultiView\" architecture."
  },
  "source": "meta"
}