{
  "id": 57503,
  "title": "Incorporating Deep Learning",
  "url": "/competitions/trackml-particle-identification/discussion/57503",
  "author_name": "maka",
  "post_date": "2018-05-24T15:23:48.193000",
  "votes": 5,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Been thinking about ways to apply deep learning to this, obviously, very big dataset. Here are my ideas so far:</p>\n\n<ul>\n<li><p>Use a siamese network to calculate the neighbour-function for clustering.  Input: data from two hits. output: Probability hits are in same cluster. Can use the output directly (like my approach below) or use the embedding found by the siamese network.\nA computationally effective way would be to first use radius search with kd_tree + helixcoordinates (Like in DBSCAN benchmark ), and then reject hits that the neural network flags. This requires implementing DBSCAN from scratch, but its not that hard(I can probably post my code, if its useful).</p></li>\n<li><p>Neural network that takes in data from N hits and scores this as a potential cluster. Needs to be used in conjunction with some other algorithm. For instance 1. Run DBSCAN with many different epsilons/unrolling angles. 2. Use this NN to score and choose which way to cluster a hit.</p></li>\n<li><p>Pure DL: From a list of candidate hits(Nearest neighbours in helix coordinates for instance), and N hits currently in cluster, predict the position of the next hit in the cluster. (N is from one to, say 20). Build up clusters from scratch using greedy search or some other computationally inexpensive search algorithm. Algorithm must also be trained to know when to stop adding hits to a cluster.</p></li>\n</ul>\n\n<p>Just my thoughts. Would love to hear peoples thoughts/experiences with these approaches, or others that I haven't thought about. I have done some experiments with the first approach, but haven't gotten any huge gains in score compared to DBSCAN benchmark</p>",
  "messages": [
    {
      "id": 333187,
      "postDate": "2018-05-24T15:23:48.193Z",
      "content": "<p>Been thinking about ways to apply deep learning to this, obviously, very big dataset. Here are my ideas so far:</p>\n\n<ul>\n<li><p>Use a siamese network to calculate the neighbour-function for clustering.  Input: data from two hits. output: Probability hits are in same cluster. Can use the output directly (like my approach below) or use the embedding found by the siamese network.\nA computationally effective way would be to first use radius search with kd_tree + helixcoordinates (Like in DBSCAN benchmark ), and then reject hits that the neural network flags. This requires implementing DBSCAN from scratch, but its not that hard(I can probably post my code, if its useful).</p></li>\n<li><p>Neural network that takes in data from N hits and scores this as a potential cluster. Needs to be used in conjunction with some other algorithm. For instance 1. Run DBSCAN with many different epsilons/unrolling angles. 2. Use this NN to score and choose which way to cluster a hit.</p></li>\n<li><p>Pure DL: From a list of candidate hits(Nearest neighbours in helix coordinates for instance), and N hits currently in cluster, predict the position of the next hit in the cluster. (N is from one to, say 20). Build up clusters from scratch using greedy search or some other computationally inexpensive search algorithm. Algorithm must also be trained to know when to stop adding hits to a cluster.</p></li>\n</ul>\n\n<p>Just my thoughts. Would love to hear peoples thoughts/experiences with these approaches, or others that I haven't thought about. I have done some experiments with the first approach, but haven't gotten any huge gains in score compared to DBSCAN benchmark</p>",
      "rawMarkdown": "Been thinking about ways to apply deep learning to this, obviously, very big dataset. Here are my ideas so far:\n\n - Use a siamese network to calculate the neighbour-function for clustering.  Input: data from two hits. output: Probability hits are in same cluster. Can use the output directly (like my approach below) or use the embedding found by the siamese network.\nA computationally effective way would be to first use radius search with kd_tree + helixcoordinates (Like in DBSCAN benchmark ), and then reject hits that the neural network flags. This requires implementing DBSCAN from scratch, but its not that hard(I can probably post my code, if its useful).\n\n - Neural network that takes in data from N hits and scores this as a potential cluster. Needs to be used in conjunction with some other algorithm. For instance 1. Run DBSCAN with many different epsilons/unrolling angles. 2. Use this NN to score and choose which way to cluster a hit.\n\n - Pure DL: From a list of candidate hits(Nearest neighbours in helix coordinates for instance), and N hits currently in cluster, predict the position of the next hit in the cluster. (N is from one to, say 20). Build up clusters from scratch using greedy search or some other computationally inexpensive search algorithm. Algorithm must also be trained to know when to stop adding hits to a cluster.\n\n\nJust my thoughts. Would love to hear peoples thoughts/experiences with these approaches, or others that I haven't thought about. I have done some experiments with the first approach, but haven't gotten any huge gains in score compared to DBSCAN benchmark\n",
      "votes": 5
    },
    {
      "id": 334464,
      "postDate": "2018-05-27T12:59:05.110Z",
      "content": "<p>@maka</p>\n\n<p>check my post</p>\n\n<p><a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/57643\">https://www.kaggle.com/c/trackml-particle-identification/discussion/57643</a></p>",
      "rawMarkdown": "@maka\n\ncheck my post\n\nhttps://www.kaggle.com/c/trackml-particle-identification/discussion/57643",
      "votes": 1,
      "replies": [
        {
          "id": 334509,
          "postDate": "2018-05-27T17:03:28.940Z",
          "content": "<p>good stuff :)</p>",
          "rawMarkdown": "good stuff :)"
        }
      ]
    },
    {
      "id": 334229,
      "postDate": "2018-05-26T18:37:22.683Z",
      "content": "<p>For the pure deep learning approach, to reduce the computational barrier you could start with labelling with a fast approach... If you're using r, Grzegorz Sionkowski's parallel script is incredibly fast and you could use the output clusters as a starting candidate list for any given cluster.</p>",
      "rawMarkdown": "For the pure deep learning approach, to reduce the computational barrier you could start with labelling with a fast approach... If you're using r, Grzegorz Sionkowski's parallel script is incredibly fast and you could use the output clusters as a starting candidate list for any given cluster.",
      "votes": 1
    },
    {
      "id": 335377,
      "postDate": "2018-05-29T16:55:08.213Z",
      "content": "<p>To carry out a deep learning approach, I expect that a new idea is needed - the work will not be incremental optimization. That’s exciting but in addition to the idea, one is working with a massive data set and, in order to even begin competing with highly optimized conventional approaches, one will probably require large computer resources at the learning stage. While eventual production throughput might be great, the learning stage can be very expensive, and that will be acceptable because learning is not real time.</p>\n\n<p>My point is: is the competition too difficult as set up? 3 months; no offer of compute assistance; data is huge simulated event files with long learning time on deep networks. These properties might prevent new ideas from emerging if the majority cant run their networks in the time provided.</p>\n\n<p>Personally, I think the problem is interesting enough to try deep learning approaches on progressively more difficult toy models of my own construction (currently 2d with mixed circular and straight tracks). But I don’t have high hopes that my efforts will be sophisticated enough in time, though, to submit anything for scoring. Perhaps if I have significant success, publication would be a route to having it used.</p>\n\n<p>Could another stage have been added? Events with a few hundred tracks in the real detector that could bring forth a range of new ideas, leaving full realism to a second stage.</p>",
      "rawMarkdown": "To carry out a deep learning approach, I expect that a new idea is needed - the work will not be incremental optimization. That’s exciting but in addition to the idea, one is working with a massive data set and, in order to even begin competing with highly optimized conventional approaches, one will probably require large computer resources at the learning stage. While eventual production throughput might be great, the learning stage can be very expensive, and that will be acceptable because learning is not real time.\n\nMy point is: is the competition too difficult as set up? 3 months; no offer of compute assistance; data is huge simulated event files with long learning time on deep networks. These properties might prevent new ideas from emerging if the majority cant run their networks in the time provided.\n\nPersonally, I think the problem is interesting enough to try deep learning approaches on progressively more difficult toy models of my own construction (currently 2d with mixed circular and straight tracks). But I don’t have high hopes that my efforts will be sophisticated enough in time, though, to submit anything for scoring. Perhaps if I have significant success, publication would be a route to having it used.\n\nCould another stage have been added? Events with a few hundred tracks in the real detector that could bring forth a range of new ideas, leaving full realism to a second stage.\n",
      "replies": [
        {
          "id": 335385,
          "postDate": "2018-05-29T17:05:29.040Z",
          "content": "<p>Keep in mind the ultimate goal of the problem requires not just a solution, but a performant solution. There is a stage 2, and it is entirely focused on making the models more simple. </p>",
          "rawMarkdown": "Keep in mind the ultimate goal of the problem requires not just a solution, but a performant solution. There is a stage 2, and it is entirely focused on making the models more simple. "
        },
        {
          "id": 335412,
          "postDate": "2018-05-29T17:47:33.257Z",
          "content": "<p>From my side I have already started constructing a big metal box in my house in order to place there all my 3  computers for concept evaluation. Hope that the metal box will be able to prevent possible fire during evaluation while I am on vacation for 3 weeks in July.</p>\n\n<p>I am writing this, because, I think, at least 3 weeks computation will be required only for concept choising / evaluation on gradually simplified data. Actual final training stage might require might longer time, and, thus even good concepts, from different participants, might not appear at Leader Board. This is because of a computational barrier.  </p>",
          "rawMarkdown": "From my side I have already started constructing a big metal box in my house in order to place there all my 3  computers for concept evaluation. Hope that the metal box will be able to prevent possible fire during evaluation while I am on vacation for 3 weeks in July.\n\nI am writing this, because, I think, at least 3 weeks computation will be required only for concept choising / evaluation on gradually simplified data. Actual final training stage might require might longer time, and, thus even good concepts, from different participants, might not appear at Leader Board. This is because of a computational barrier.  ",
          "votes": 2
        },
        {
          "id": 335454,
          "postDate": "2018-05-29T18:50:29.760Z",
          "content": "<p>A metal box is a sensible idea.  To reduce the risk of fire I recommend a good air conditioner.</p>",
          "rawMarkdown": "A metal box is a sensible idea.  To reduce the risk of fire I recommend a good air conditioner."
        },
        {
          "id": 335486,
          "postDate": "2018-05-29T20:07:46.730Z",
          "content": "<p>Develop a DL algorithm to control the thermostat inside the box though. And buy new box for the computing necessary</p>",
          "rawMarkdown": "Develop a DL algorithm to control the thermostat inside the box though. And buy new box for the computing necessary"
        }
      ]
    },
    {
      "id": 334028,
      "postDate": "2018-05-26T10:33:07.920Z",
      "content": "<p>I'm interested in a DBSCAN implementation where one can plug in a custom metric. Altenatively just the radius search bit. Right now I'm bruteforce precalculating distance metrics, which is challenging at these data sizes.</p>",
      "rawMarkdown": "I'm interested in a DBSCAN implementation where one can plug in a custom metric. Altenatively just the radius search bit. Right now I'm bruteforce precalculating distance metrics, which is challenging at these data sizes.",
      "replies": [
        {
          "id": 334089,
          "postDate": "2018-05-26T14:01:40.670Z",
          "content": "<p>Just found out that the DBSCAN implementation in mlpack has support for custom metrics, when using the C++ interface (not exposed in Python). It also looks to be well optimized in general. <a href=\"https://github.com/mlpack/mlpack/blob/master/src/mlpack/methods/dbscan/dbscan_main.cpp#L154\">https://github.com/mlpack/mlpack/blob/master/src/mlpack/methods/dbscan/dbscan_main.cpp#L154</a></p>",
          "rawMarkdown": "Just found out that the DBSCAN implementation in mlpack has support for custom metrics, when using the C++ interface (not exposed in Python). It also looks to be well optimized in general. https://github.com/mlpack/mlpack/blob/master/src/mlpack/methods/dbscan/dbscan_main.cpp#L154"
        },
        {
          "id": 334282,
          "postDate": "2018-05-26T21:48:51.410Z",
          "content": "<p>You can do this with the python sklearn DBSCAN. But it will be much slower, since one in general cannot use kd_tree</p>",
          "rawMarkdown": "You can do this with the python sklearn DBSCAN. But it will be much slower, since one in general cannot use kd_tree"
        }
      ]
    },
    {
      "id": 334011,
      "postDate": "2018-05-26T09:38:47.390Z",
      "content": "<p>Glad I am not the only one thinking about this. I tried some trivial supervised learning approaches, not necessarily using deep learning but I found everything I tried to be computationally not viable as learning the labeled data even for a single event takes hours. Hope to get some new ideas here :)</p>",
      "rawMarkdown": "Glad I am not the only one thinking about this. I tried some trivial supervised learning approaches, not necessarily using deep learning but I found everything I tried to be computationally not viable as learning the labeled data even for a single event takes hours. Hope to get some new ideas here :)",
      "replies": [
        {
          "id": 334674,
          "postDate": "2018-05-28T06:37:15.310Z",
          "content": "<p>Seems like the amount of data associated with an LHS event is not too different from the amount present in a typical image used in deep learning image classification. Why is the learning time significantly different for this problem?</p>",
          "rawMarkdown": "Seems like the amount of data associated with an LHS event is not too different from the amount present in a typical image used in deep learning image classification. Why is the learning time significantly different for this problem?"
        },
        {
          "id": 334709,
          "postDate": "2018-05-28T09:03:01.537Z",
          "content": "<p>The problem is not related to the amount of input data itself (in bytes), the problem is related to input and output data representation. If, for example, we decide to apply end-to-end network architecture similar to what is used in image classifiers, at first we would need to implement input vector transformation into 3D volume which, in turn, would make the problem solution computationally non-realistic from both memory as well as CPU perspectives. However, honestly, it would be interesting to see someone's results of end-to-end intensive neural network simulation for this problem.</p>",
          "rawMarkdown": "The problem is not related to the amount of input data itself (in bytes), the problem is related to input and output data representation. If, for example, we decide to apply end-to-end network architecture similar to what is used in image classifiers, at first we would need to implement input vector transformation into 3D volume which, in turn, would make the problem solution computationally non-realistic from both memory as well as CPU perspectives. However, honestly, it would be interesting to see someone's results of end-to-end intensive neural network simulation for this problem."
        },
        {
          "id": 334736,
          "postDate": "2018-05-28T09:41:19.540Z",
          "content": "<p>@Denis</p>\n\n<p>Refer to my post for a proposed faster-rcnn like end-to-end solution</p>\n\n<p><a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/57643\">https://www.kaggle.com/c/trackml-particle-identification/discussion/57643</a></p>",
          "rawMarkdown": "@Denis\n\nRefer to my post for a proposed faster-rcnn like end-to-end solution\n\nhttps://www.kaggle.com/c/trackml-particle-identification/discussion/57643",
          "votes": 1
        },
        {
          "id": 334737,
          "postDate": "2018-05-28T09:47:05.610Z",
          "content": "<p>@pete It's not obious exactly HOW to formulate this problem in terms of input-output in a way that will give good results(No matter how long one trains).</p>",
          "rawMarkdown": "@pete It's not obious exactly HOW to formulate this problem in terms of input-output in a way that will give good results(No matter how long one trains)."
        },
        {
          "id": 334849,
          "postDate": "2018-05-28T15:29:41.903Z",
          "content": "<p>Looks like people have been investigating some deep neural approaches for the tracking problem since 2016. A good paper can be found by the link provided by Heng CherKeng above, some other papers, and even source code can be found here:</p>\n\n<p><a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/57742\">The link</a></p>\n\n<p>The interesting thing, as I understood,  is that good results were achieved only on toy (specially simplified) sets of data, I have not found so far any score rating on deep learning algorithms trained and tested on full-scale ACTS data.</p>\n\n<p>I personally consider this as an opportunity for a deep learning approach for the TrackML Challlenge.</p>",
          "rawMarkdown": "Looks like people have been investigating some deep neural approaches for the tracking problem since 2016. A good paper can be found by the link provided by Heng CherKeng above, some other papers, and even source code can be found here:\n\n[The link][1]\n\nThe interesting thing, as I understood,  is that good results were achieved only on toy (specially simplified) sets of data, I have not found so far any score rating on deep learning algorithms trained and tested on full-scale ACTS data.\n\nI personally consider this as an opportunity for a deep learning approach for the TrackML Challlenge.\n\n\n  [1]: https://www.kaggle.com/c/trackml-particle-identification/discussion/57742"
        },
        {
          "id": 334852,
          "postDate": "2018-05-28T15:34:23.413Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 334870,
          "postDate": "2018-05-28T16:10:37.180Z",
          "content": "<p>I confirm your observation.</p>",
          "rawMarkdown": "I confirm your observation."
        }
      ]
    },
    {
      "id": 333554,
      "postDate": "2018-05-25T11:57:57.140Z",
      "content": "<p>Very good post, since deep learning approach for this problem is not a trivial task. From other side,  big amount of provided labeled train data enforces to consider deep supervised algorithms. I'am also experimenting with deep networks application for this problem, will share the results as soon as I have something interesting.</p>",
      "rawMarkdown": "Very good post, since deep learning approach for this problem is not a trivial task. From other side,  big amount of provided labeled train data enforces to consider deep supervised algorithms. I'am also experimenting with deep networks application for this problem, will share the results as soon as I have something interesting."
    }
  ],
  "comments": [
    {
      "id": 334464,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-05-27T12:59:05.110000",
      "content": "<p>@maka</p>\n\n<p>check my post</p>\n\n<p><a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/57643\">https://www.kaggle.com/c/trackml-particle-identification/discussion/57643</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 334509,
          "author_name": "maka",
          "author_url": "",
          "post_date": "2018-05-27T17:03:28.940000",
          "content": "<p>good stuff :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 334229,
      "author_name": "macfarll",
      "author_url": "",
      "post_date": "2018-05-26T18:37:22.683000",
      "content": "<p>For the pure deep learning approach, to reduce the computational barrier you could start with labelling with a fast approach... If you're using r, Grzegorz Sionkowski's parallel script is incredibly fast and you could use the output clusters as a starting candidate list for any given cluster.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 335377,
      "author_name": "pete",
      "author_url": "",
      "post_date": "2018-05-29T16:55:08.213000",
      "content": "<p>To carry out a deep learning approach, I expect that a new idea is needed - the work will not be incremental optimization. That’s exciting but in addition to the idea, one is working with a massive data set and, in order to even begin competing with highly optimized conventional approaches, one will probably require large computer resources at the learning stage. While eventual production throughput might be great, the learning stage can be very expensive, and that will be acceptable because learning is not real time.</p>\n\n<p>My point is: is the competition too difficult as set up? 3 months; no offer of compute assistance; data is huge simulated event files with long learning time on deep networks. These properties might prevent new ideas from emerging if the majority cant run their networks in the time provided.</p>\n\n<p>Personally, I think the problem is interesting enough to try deep learning approaches on progressively more difficult toy models of my own construction (currently 2d with mixed circular and straight tracks). But I don’t have high hopes that my efforts will be sophisticated enough in time, though, to submit anything for scoring. Perhaps if I have significant success, publication would be a route to having it used.</p>\n\n<p>Could another stage have been added? Events with a few hundred tracks in the real detector that could bring forth a range of new ideas, leaving full realism to a second stage.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 335385,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-05-29T17:05:29.040000",
          "content": "<p>Keep in mind the ultimate goal of the problem requires not just a solution, but a performant solution. There is a stage 2, and it is entirely focused on making the models more simple. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 335412,
          "author_name": "Denis",
          "author_url": "",
          "post_date": "2018-05-29T17:47:33.257000",
          "content": "<p>From my side I have already started constructing a big metal box in my house in order to place there all my 3  computers for concept evaluation. Hope that the metal box will be able to prevent possible fire during evaluation while I am on vacation for 3 weeks in July.</p>\n\n<p>I am writing this, because, I think, at least 3 weeks computation will be required only for concept choising / evaluation on gradually simplified data. Actual final training stage might require might longer time, and, thus even good concepts, from different participants, might not appear at Leader Board. This is because of a computational barrier.  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 335454,
          "author_name": "Robert",
          "author_url": "",
          "post_date": "2018-05-29T18:50:29.760000",
          "content": "<p>A metal box is a sensible idea.  To reduce the risk of fire I recommend a good air conditioner.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 335486,
          "author_name": "maka",
          "author_url": "",
          "post_date": "2018-05-29T20:07:46.730000",
          "content": "<p>Develop a DL algorithm to control the thermostat inside the box though. And buy new box for the computing necessary</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 334028,
      "author_name": "Jon Nordby",
      "author_url": "",
      "post_date": "2018-05-26T10:33:07.920000",
      "content": "<p>I'm interested in a DBSCAN implementation where one can plug in a custom metric. Altenatively just the radius search bit. Right now I'm bruteforce precalculating distance metrics, which is challenging at these data sizes.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 334089,
          "author_name": "Jon Nordby",
          "author_url": "",
          "post_date": "2018-05-26T14:01:40.670000",
          "content": "<p>Just found out that the DBSCAN implementation in mlpack has support for custom metrics, when using the C++ interface (not exposed in Python). It also looks to be well optimized in general. <a href=\"https://github.com/mlpack/mlpack/blob/master/src/mlpack/methods/dbscan/dbscan_main.cpp#L154\">https://github.com/mlpack/mlpack/blob/master/src/mlpack/methods/dbscan/dbscan_main.cpp#L154</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 334282,
          "author_name": "maka",
          "author_url": "",
          "post_date": "2018-05-26T21:48:51.410000",
          "content": "<p>You can do this with the python sklearn DBSCAN. But it will be much slower, since one in general cannot use kd_tree</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 334011,
      "author_name": "RichardMeyes",
      "author_url": "",
      "post_date": "2018-05-26T09:38:47.390000",
      "content": "<p>Glad I am not the only one thinking about this. I tried some trivial supervised learning approaches, not necessarily using deep learning but I found everything I tried to be computationally not viable as learning the labeled data even for a single event takes hours. Hope to get some new ideas here :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 334674,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-05-28T06:37:15.310000",
          "content": "<p>Seems like the amount of data associated with an LHS event is not too different from the amount present in a typical image used in deep learning image classification. Why is the learning time significantly different for this problem?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 334709,
          "author_name": "Denis",
          "author_url": "",
          "post_date": "2018-05-28T09:03:01.537000",
          "content": "<p>The problem is not related to the amount of input data itself (in bytes), the problem is related to input and output data representation. If, for example, we decide to apply end-to-end network architecture similar to what is used in image classifiers, at first we would need to implement input vector transformation into 3D volume which, in turn, would make the problem solution computationally non-realistic from both memory as well as CPU perspectives. However, honestly, it would be interesting to see someone's results of end-to-end intensive neural network simulation for this problem.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 334736,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-05-28T09:41:19.540000",
          "content": "<p>@Denis</p>\n\n<p>Refer to my post for a proposed faster-rcnn like end-to-end solution</p>\n\n<p><a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/57643\">https://www.kaggle.com/c/trackml-particle-identification/discussion/57643</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 334737,
          "author_name": "maka",
          "author_url": "",
          "post_date": "2018-05-28T09:47:05.610000",
          "content": "<p>@pete It's not obious exactly HOW to formulate this problem in terms of input-output in a way that will give good results(No matter how long one trains).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 334849,
          "author_name": "Denis",
          "author_url": "",
          "post_date": "2018-05-28T15:29:41.903000",
          "content": "<p>Looks like people have been investigating some deep neural approaches for the tracking problem since 2016. A good paper can be found by the link provided by Heng CherKeng above, some other papers, and even source code can be found here:</p>\n\n<p><a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/57742\">The link</a></p>\n\n<p>The interesting thing, as I understood,  is that good results were achieved only on toy (specially simplified) sets of data, I have not found so far any score rating on deep learning algorithms trained and tested on full-scale ACTS data.</p>\n\n<p>I personally consider this as an opportunity for a deep learning approach for the TrackML Challlenge.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 334852,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-28T15:34:23.413000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 334870,
          "author_name": "David Rousseau",
          "author_url": "",
          "post_date": "2018-05-28T16:10:37.180000",
          "content": "<p>I confirm your observation.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 333554,
      "author_name": "Denis",
      "author_url": "",
      "post_date": "2018-05-25T11:57:57.140000",
      "content": "<p>Very good post, since deep learning approach for this problem is not a trivial task. From other side,  big amount of provided labeled train data enforces to consider deep supervised algorithms. I'am also experimenting with deep networks application for this problem, will share the results as soon as I have something interesting.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "333187": "Been thinking about ways to apply deep learning to this, obviously, very big dataset. Here are my ideas so far:\n\n - Use a siamese network to calculate the neighbour-function for clustering.  Input: data from two hits. output: Probability hits are in same cluster. Can use the output directly (like my approach below) or use the embedding found by the siamese network.\nA computationally effective way would be to first use radius search with kd_tree + helixcoordinates (Like in DBSCAN benchmark ), and then reject hits that the neural network flags. This requires implementing DBSCAN from scratch, but its not that hard(I can probably post my code, if its useful).\n\n - Neural network that takes in data from N hits and scores this as a potential cluster. Needs to be used in conjunction with some other algorithm. For instance 1. Run DBSCAN with many different epsilons/unrolling angles. 2. Use this NN to score and choose which way to cluster a hit.\n\n - Pure DL: From a list of candidate hits(Nearest neighbours in helix coordinates for instance), and N hits currently in cluster, predict the position of the next hit in the cluster. (N is from one to, say 20). Build up clusters from scratch using greedy search or some other computationally inexpensive search algorithm. Algorithm must also be trained to know when to stop adding hits to a cluster.\n\n\nJust my thoughts. Would love to hear peoples thoughts/experiences with these approaches, or others that I haven't thought about. I have done some experiments with the first approach, but haven't gotten any huge gains in score compared to DBSCAN benchmark\n",
    "334464": "@maka\n\ncheck my post\n\nhttps://www.kaggle.com/c/trackml-particle-identification/discussion/57643",
    "334229": "For the pure deep learning approach, to reduce the computational barrier you could start with labelling with a fast approach... If you're using r, Grzegorz Sionkowski's parallel script is incredibly fast and you could use the output clusters as a starting candidate list for any given cluster.",
    "335377": "To carry out a deep learning approach, I expect that a new idea is needed - the work will not be incremental optimization. That’s exciting but in addition to the idea, one is working with a massive data set and, in order to even begin competing with highly optimized conventional approaches, one will probably require large computer resources at the learning stage. While eventual production throughput might be great, the learning stage can be very expensive, and that will be acceptable because learning is not real time.\n\nMy point is: is the competition too difficult as set up? 3 months; no offer of compute assistance; data is huge simulated event files with long learning time on deep networks. These properties might prevent new ideas from emerging if the majority cant run their networks in the time provided.\n\nPersonally, I think the problem is interesting enough to try deep learning approaches on progressively more difficult toy models of my own construction (currently 2d with mixed circular and straight tracks). But I don’t have high hopes that my efforts will be sophisticated enough in time, though, to submit anything for scoring. Perhaps if I have significant success, publication would be a route to having it used.\n\nCould another stage have been added? Events with a few hundred tracks in the real detector that could bring forth a range of new ideas, leaving full realism to a second stage.\n",
    "334028": "I'm interested in a DBSCAN implementation where one can plug in a custom metric. Altenatively just the radius search bit. Right now I'm bruteforce precalculating distance metrics, which is challenging at these data sizes.",
    "334011": "Glad I am not the only one thinking about this. I tried some trivial supervised learning approaches, not necessarily using deep learning but I found everything I tried to be computationally not viable as learning the labeled data even for a single event takes hours. Hope to get some new ideas here :)",
    "333554": "Very good post, since deep learning approach for this problem is not a trivial task. From other side,  big amount of provided labeled train data enforces to consider deep supervised algorithms. I'am also experimenting with deep networks application for this problem, will share the results as soon as I have something interesting."
  }
}